Cohere Command R+
by Cohere
A 104B-parameter model built for cited RAG and multi-step tool use, deployable in your own VPC
Command R+ is Cohere's enterprise large language model, tuned for retrieval-augmented generation with inline citations and for multi-step tool use rather than open-ended chat. It takes a 128,000-token context, covers ten business languages, and can run on Cohere's API, on AWS Bedrock and Azure, or privately inside a customer's own VPC or data centre.
Command R+ is Cohere's 104-billion-parameter enterprise language model, released in April 2024 and refreshed as command-r-plus-08-2024, built specifically for retrieval-augmented generation and multi-step tool use rather than general chat. It accepts a 128,000-token context window and, given a list of supplied document snippets, returns answers with inline citations pointing back at the source passage — the mechanism Cohere sells as hallucination control to regulated buyers who have to show where an answer came from. It was trained with multi-step tool use, so it can chain tool calls and feed each result into the next, which is what makes it usable as an agent backbone rather than just a summariser. It is optimised for ten languages — English, French, Spanish, Italian, German, Brazilian Portuguese, Japanese, Korean, Arabic and Simplified Chinese — with pre-training coverage of thirteen more, and the August 2024 refresh delivered roughly 50% higher throughput and 25% lower latency than the April build. Access is unusually broad for a model at this tier: Cohere's own API at $2.50 per million input and $10 per million output tokens (the older 04-2024 build was $3/$15), Amazon Bedrock in us-east-1 and us-west-2, Microsoft Azure, a dedicated Model Vault instance, or a fully private deployment inside a customer's VPC or on-premises — the last of which is where the bulk of Cohere's enterprise revenue sits, and the main reason buyers pick it over a pure API vendor. Open weights are published on Hugging Face under a CC-BY-NC licence, so they can be evaluated and researched but not run commercially without a Cohere agreement. Buyers should know the line has moved: the command-r-plus alias (the 04-2024 build) was deprecated on 15 September 2025, command-r-plus-08-2024 remains available, and Cohere now directs new work to Command A (command-a-03-2025) as its strongest model.
A platform or data team in a regulated enterprise that needs grounded, citable answers over internal documents and cannot send that corpus to a public API — Cohere will run the same model inside their VPC or data centre.
RAG answers that carry inline citations to the source passage, so a reviewer can audit every claim instead of trusting the model.
At a Glance
- Category
- AI Models & APIs
- Pricing
- Usage-based, Contact for pricing, Freemium
- Target Market
- CIOs, CTOs, Enterprise Developers, Data Scientists, Heads of AI Platform
- Deployment
- API-based, Multi-cloud, Hybrid, Self-hosted
- Founded
- 2019
- Headquarters
- Toronto, Canada
- Team Size
- 201-500
Key Features
- ✓Grounded generation with inline citations
Answers built from supplied document snippets carry citations to the exact source passage, making output auditable for regulated review.
- ✓128K-token context window
Holds long contracts, filings or full retrieval result sets in a single prompt without aggressive chunking and re-ranking.
- ✓Multi-step tool use
Chains sequential tool calls and reasons over each intermediate result, which is what allows simple agents to be built on the model.
- ✓Ten-language optimisation
Tuned for ten business languages with pre-training in thirteen more, so one model serves multi-region support and knowledge workloads.
- ✓Private and on-premises deployment
The same model can be run inside a customer VPC or data centre for full data sovereignty, not only as a hosted API.
- ✓Multi-cloud availability
Reachable through Amazon Bedrock and Microsoft Azure as well as Cohere's own API, so procurement can buy through an existing cloud contract.
- ✓Open weights for evaluation
104B-parameter weights are published on Hugging Face under CC-BY-NC, letting teams benchmark the real model before committing commercially.
Capabilities
Use Cases
- •Cited internal knowledge assistant
Answer employee questions over policy, contract and product documentation with citations, so compliance can trace each statement to a source.
- •Regulated document analysis
Summarise and compare long filings or agreements inside a bank's own VPC, where sending the corpus to a public API is not permitted.
- •Multi-step internal tool agents
Let the model call ticketing, CRM and search APIs in sequence, using each result to decide the next call in the chain.
- •Multilingual customer support drafting
Draft grounded replies in ten business languages from the same retrieval corpus, keeping one model across regional support teams.
- •Bedrock-native RAG pipelines
Run grounded generation through Amazon Bedrock so the workload is billed and governed under an existing AWS agreement.
Ideal For
Best For
- ✓Retrieval-augmented question answering over internal document corpora where every answer must cite its source
- ✓Air-gapped or VPC deployments in banking, defence, healthcare and public sector where data cannot leave the perimeter
- ✓Multi-step tool-use agents that call internal APIs in sequence and reason over each intermediate result
- ✓Multilingual enterprise support and knowledge workflows across the ten optimised business languages
- ✓Enterprise search and knowledge assistants built on top of an existing vector or hybrid retrieval stack
Not Ideal For
- ✗New greenfield projects starting in 2026 — Cohere itself now points customers at Command A, and the original command-r-plus alias was deprecated in September 2025
- ✗Teams wanting to self-host the open weights commercially: the Hugging Face release is CC-BY-NC, so production use needs a separate Cohere licence
- ✗Pure code-completion workloads, which the model card explicitly says it may not handle well out of the box
- ✗Cost-sensitive high-volume chat, where smaller models price an order of magnitude below $2.50/$10 per million tokens
Integrations
Deployment
Market Analysis
Pros
- ✓Citations are a first-class output, which is the single feature regulated buyers ask for and most API models leave to prompt engineering
- ✓Private VPC and on-premises deployment of the same model, so data sovereignty does not force a downgrade to a weaker open model
- ✓Available through Amazon Bedrock and Azure, letting teams buy under an existing cloud agreement instead of a new vendor contract
- ✓The August 2024 refresh cut price from $3/$15 to $2.50/$10 while raising throughput about 50% and cutting latency about 25%
- ✓Backed by a well-funded vendor: Cohere reached a $7B valuation in September 2025 with roughly $1.6B raised and around 450 staff
Cons
- ✗Superseded by its own vendor: the command-r-plus alias (04-2024) was deprecated on 15 September 2025 and Cohere now recommends Command A for new work
- ✗The published weights are CC-BY-NC, so self-hosting for commercial use is not permitted without a separate Cohere agreement — practitioners flag this as research-only licensing
- ✗104 billion parameters make even licensed self-hosting expensive; this is not a model a team runs on a single commodity GPU
- ✗Hacker News practitioners report it is sensitive to prompt template formatting and showed mode collapse on open-ended creative tasks against Claude 3
- ✗Its April 2024 'beats GPT-4 in Chatbot Arena' framing drew benchmark-overfitting scepticism from commenters at the time
- ✗No G2, Capterra or TrustRadius rating exists for the model itself, so there is no aggregated buyer-review signal to check
Pricing
Trial API key
$0
- ✓Free rate-limited access to the API
- ✓Evaluation and prototyping only
- ✓Not licensed for commercial use
Production (pay-as-you-go), Command R+ 08-2024
From $2.50/1M input tokens
- ✓$2.50 per million input tokens
- ✓$10.00 per million output tokens
- ✓Grounded generation with citations
- ✓Multi-step tool use
- ✓128K context
Model Vault (dedicated instance)
Contact for pricing
- ✓Logically isolated managed deployment
- ✓Hourly or monthly committed rates
- ✓Dedicated throughput
Private / on-premises
Contact for pricing
- ✓Deployment inside customer VPC or data centre
- ✓Full data sovereignty
- ✓Covers Command, Rerank and Embed
Public API pricing is per token and published: command-r-plus-08-2024 costs $2.50 per million input tokens and $10.00 per million output tokens, while the deprecated 04-2024 build was $3.00/$15.00 — roughly Claude-3.5-Sonnet-class pricing, and an order of magnitude above Cohere's smaller Command R at $0.50/$1.50. Trial API keys are free but rate-limited and explicitly barred from commercial use; production keys bill month-end or when the balance reaches $250. Dedicated Model Vault instances are quoted hourly or monthly, and private or on-premises deployment — which is what most of Cohere's large enterprise customers actually buy — carries no list price and goes through sales.
Security & Compliance
Connect
Sources
This page was written from 9 sources, 7 on domains other than docs.cohere.com.
- 1.docs.cohere.com — command r plusvendor
- 2.docs.cohere.com — deprecationsvendor
- 3.cohere.com — pricing
- 4.cohere.com — deployment options
- 5.huggingface.co — c4ai command r plus
- 6.aws.amazon.com — cohere command r r plus amazon bedrock
- 7.the-decoder.com — cohere improves its rag optimized command series llms
- 8.betakit.com — coheres valuation hits 7 billion usd following 100 million r
- 9.hn.algolia.com — hn.algolia.com
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Scaled Cognition
APT, a large action model trained to take actions instead of predicting text
Engram
A learned memory layer that makes AI actually know your organization — at up to 100x fewer tokens.
TwelveLabs
Video intelligence API that makes every hour of enterprise footage searchable, analyzable and agent-ready.
Mistral OCR 4
Structure-aware document AI that returns bounding boxes, typed blocks, and per-word confidence scores.