C

Cohere Command R+

by Cohere

AI Models & APIsEnterprise Search & KnowledgeAI Agents & Orchestration

A 104B-parameter model built for cited RAG and multi-step tool use, deployable in your own VPC

Usage-based · Contact for pricing · Freemium·Added Mar 14, 2026·Updated Aug 2, 2026
Share:
THE DAILY BRIEF
Cohere Command R+

by Cohere

AI Models & APIsEnterprise Search & KnowledgeAI Agents & Orchestration

A 104B-parameter model built for cited RAG and multi-step tool use, deployable in your own VPC

Usage-based · Contact for pricing · Freemium

Command R+ is Cohere's enterprise large language model, tuned for retrieval-augmented generation with inline citations and for multi-step tool use rather than open-ended chat. It takes a 128,000-token context, covers ten business languages, and can run on Cohere's API, on AWS Bedrock and Azure, or privately inside a customer's own VPC or data centre.

At a Glance

Category
AI Models & APIs
Pricing
Usage-based, Contact for pricing, Freemium
Target Market
CIOs, CTOs, Enterprise Developers, Data Scientists, Heads of AI Platform
Deployment
API-based, Multi-cloud, Hybrid, Self-hosted
Founded
2019
Headquarters
Toronto, Canada
Team Size
201-500

Key Features

  • Grounded generation with inline citations
  • 128K-token context window
  • Multi-step tool use
  • Ten-language optimisation
  • Private and on-premises deployment
  • Multi-cloud availability
  • Open weights for evaluation

Capabilities

text generation
image generation
video generation
code generation
workflow automation
api access
audio generation
fine tuning
agent orchestration

Use Cases

  • Cited internal knowledge assistant
  • Regulated document analysis
  • Multi-step internal tool agents
  • Multilingual customer support drafting
  • Bedrock-native RAG pipelines

Ideal For

Best For

  • Retrieval-augmented question answering over internal document corpora where every answer must cite its source
  • Air-gapped or VPC deployments in banking, defence, healthcare and public sector where data cannot leave the perimeter
  • Multi-step tool-use agents that call internal APIs in sequence and reason over each intermediate result
  • Multilingual enterprise support and knowledge workflows across the ten optimised business languages
  • Enterprise search and knowledge assistants built on top of an existing vector or hybrid retrieval stack

Not Ideal For

  • New greenfield projects starting in 2026 — Cohere itself now points customers at Command A, and the original command-r-plus alias was deprecated in September 2025
  • Teams wanting to self-host the open weights commercially: the Hugging Face release is CC-BY-NC, so production use needs a separate Cohere licence
  • Pure code-completion workloads, which the model card explicitly says it may not handle well out of the box
  • Cost-sensitive high-volume chat, where smaller models price an order of magnitude below $2.50/$10 per million tokens

Market Analysis

Enterprise-gradeRAG-optimizedDeployable on-premises

Pros

  • Citations are a first-class output, which is the single feature regulated buyers ask for and most API models leave to prompt engineering
  • Private VPC and on-premises deployment of the same model, so data sovereignty does not force a downgrade to a weaker open model
  • Available through Amazon Bedrock and Azure, letting teams buy under an existing cloud agreement instead of a new vendor contract
  • The August 2024 refresh cut price from $3/$15 to $2.50/$10 while raising throughput about 50% and cutting latency about 25%
  • Backed by a well-funded vendor: Cohere reached a $7B valuation in September 2025 with roughly $1.6B raised and around 450 staff

Cons

  • Superseded by its own vendor: the command-r-plus alias (04-2024) was deprecated on 15 September 2025 and Cohere now recommends Command A for new work
  • The published weights are CC-BY-NC, so self-hosting for commercial use is not permitted without a separate Cohere agreement — practitioners flag this as research-only licensing
  • 104 billion parameters make even licensed self-hosting expensive; this is not a model a team runs on a single commodity GPU
  • Hacker News practitioners report it is sensitive to prompt template formatting and showed mode collapse on open-ended creative tasks against Claude 3
  • Its April 2024 'beats GPT-4 in Chatbot Arena' framing drew benchmark-overfitting scepticism from commenters at the time
  • No G2, Capterra or TrustRadius rating exists for the model itself, so there is no aggregated buyer-review signal to check

Pricing

Trial API key

$0

  • Free rate-limited access to the API
  • Evaluation and prototyping only
  • Not licensed for commercial use

Production (pay-as-you-go), Command R+ 08-2024

From $2.50/1M input tokens

  • $2.50 per million input tokens
  • $10.00 per million output tokens
  • Grounded generation with citations
  • Multi-step tool use
  • 128K context

Model Vault (dedicated instance)

Contact for pricing

  • Logically isolated managed deployment
  • Hourly or monthly committed rates
  • Dedicated throughput

Private / on-premises

Contact for pricing

  • Deployment inside customer VPC or data centre
  • Full data sovereignty
  • Covers Command, Rerank and Embed

Public API pricing is per token and published: command-r-plus-08-2024 costs $2.50 per million input tokens and $10.00 per million output tokens, while the deprecated 04-2024 build was $3.00/$15.00 — roughly Claude-3.5-Sonnet-class pricing, and an order of magnitude above Cohere's smaller Command R at $0.50/$1.50. Trial API keys are free but rate-limited and explicitly barred from commercial use; production keys bill month-end or when the balance reaches $250. Dedicated Model Vault instances are quoted hourly or monthly, and private or on-premises deployment — which is what most of Cohere's large enterprise customers actually buy — carries no list price and goes through sales.

Security & Compliance

soc2
gdpr
hipaa
iso27001
sso
data residency

THE DAILY BRIEF

Enterprise AI insights for technology and business leaders, twice weekly.

beri.net

Subscribe at beri.net/subscribe for twice-weekly AI insights delivered to your inbox.

LinkedIn: linkedin.com/in/rberi  |  X: x.com/rajeshberi

© 2026 Rajesh Beri. All rights reserved.

Command R+ is Cohere's enterprise large language model, tuned for retrieval-augmented generation with inline citations and for multi-step tool use rather than open-ended chat. It takes a 128,000-token context, covers ten business languages, and can run on Cohere's API, on AWS Bedrock and Azure, or privately inside a customer's own VPC or data centre.

Command R+ is Cohere's 104-billion-parameter enterprise language model, released in April 2024 and refreshed as command-r-plus-08-2024, built specifically for retrieval-augmented generation and multi-step tool use rather than general chat. It accepts a 128,000-token context window and, given a list of supplied document snippets, returns answers with inline citations pointing back at the source passage — the mechanism Cohere sells as hallucination control to regulated buyers who have to show where an answer came from. It was trained with multi-step tool use, so it can chain tool calls and feed each result into the next, which is what makes it usable as an agent backbone rather than just a summariser. It is optimised for ten languages — English, French, Spanish, Italian, German, Brazilian Portuguese, Japanese, Korean, Arabic and Simplified Chinese — with pre-training coverage of thirteen more, and the August 2024 refresh delivered roughly 50% higher throughput and 25% lower latency than the April build. Access is unusually broad for a model at this tier: Cohere's own API at $2.50 per million input and $10 per million output tokens (the older 04-2024 build was $3/$15), Amazon Bedrock in us-east-1 and us-west-2, Microsoft Azure, a dedicated Model Vault instance, or a fully private deployment inside a customer's VPC or on-premises — the last of which is where the bulk of Cohere's enterprise revenue sits, and the main reason buyers pick it over a pure API vendor. Open weights are published on Hugging Face under a CC-BY-NC licence, so they can be evaluated and researched but not run commercially without a Cohere agreement. Buyers should know the line has moved: the command-r-plus alias (the 04-2024 build) was deprecated on 15 September 2025, command-r-plus-08-2024 remains available, and Cohere now directs new work to Command A (command-a-03-2025) as its strongest model.

Ideal Buyer

A platform or data team in a regulated enterprise that needs grounded, citable answers over internal documents and cannot send that corpus to a public API — Cohere will run the same model inside their VPC or data centre.

Key Benefit

RAG answers that carry inline citations to the source passage, so a reviewer can audit every claim instead of trusting the model.

At a Glance

Category
AI Models & APIs
Pricing
Usage-based, Contact for pricing, Freemium
Target Market
CIOs, CTOs, Enterprise Developers, Data Scientists, Heads of AI Platform
Deployment
API-based, Multi-cloud, Hybrid, Self-hosted
Founded
2019
Headquarters
Toronto, Canada
Team Size
201-500

Key Features

  • Grounded generation with inline citations

    Answers built from supplied document snippets carry citations to the exact source passage, making output auditable for regulated review.

  • 128K-token context window

    Holds long contracts, filings or full retrieval result sets in a single prompt without aggressive chunking and re-ranking.

  • Multi-step tool use

    Chains sequential tool calls and reasons over each intermediate result, which is what allows simple agents to be built on the model.

  • Ten-language optimisation

    Tuned for ten business languages with pre-training in thirteen more, so one model serves multi-region support and knowledge workloads.

  • Private and on-premises deployment

    The same model can be run inside a customer VPC or data centre for full data sovereignty, not only as a hosted API.

  • Multi-cloud availability

    Reachable through Amazon Bedrock and Microsoft Azure as well as Cohere's own API, so procurement can buy through an existing cloud contract.

  • Open weights for evaluation

    104B-parameter weights are published on Hugging Face under CC-BY-NC, letting teams benchmark the real model before committing commercially.

Capabilities

text generation
image generation
video generation
code generation
workflow automation
api access
audio generation
fine tuning
agent orchestration

Use Cases

  • Cited internal knowledge assistant

    Answer employee questions over policy, contract and product documentation with citations, so compliance can trace each statement to a source.

  • Regulated document analysis

    Summarise and compare long filings or agreements inside a bank's own VPC, where sending the corpus to a public API is not permitted.

  • Multi-step internal tool agents

    Let the model call ticketing, CRM and search APIs in sequence, using each result to decide the next call in the chain.

  • Multilingual customer support drafting

    Draft grounded replies in ten business languages from the same retrieval corpus, keeping one model across regional support teams.

  • Bedrock-native RAG pipelines

    Run grounded generation through Amazon Bedrock so the workload is billed and governed under an existing AWS agreement.

Ideal For

Best For

  • Retrieval-augmented question answering over internal document corpora where every answer must cite its source
  • Air-gapped or VPC deployments in banking, defence, healthcare and public sector where data cannot leave the perimeter
  • Multi-step tool-use agents that call internal APIs in sequence and reason over each intermediate result
  • Multilingual enterprise support and knowledge workflows across the ten optimised business languages
  • Enterprise search and knowledge assistants built on top of an existing vector or hybrid retrieval stack

Not Ideal For

  • New greenfield projects starting in 2026 — Cohere itself now points customers at Command A, and the original command-r-plus alias was deprecated in September 2025
  • Teams wanting to self-host the open weights commercially: the Hugging Face release is CC-BY-NC, so production use needs a separate Cohere licence
  • Pure code-completion workloads, which the model card explicitly says it may not handle well out of the box
  • Cost-sensitive high-volume chat, where smaller models price an order of magnitude below $2.50/$10 per million tokens

Integrations

SDK Available
SDK:PythonTypeScriptJavaGo

Deployment

On-Premise

Market Analysis

Enterprise-gradeRAG-optimizedDeployable on-premises

Pros

  • Citations are a first-class output, which is the single feature regulated buyers ask for and most API models leave to prompt engineering
  • Private VPC and on-premises deployment of the same model, so data sovereignty does not force a downgrade to a weaker open model
  • Available through Amazon Bedrock and Azure, letting teams buy under an existing cloud agreement instead of a new vendor contract
  • The August 2024 refresh cut price from $3/$15 to $2.50/$10 while raising throughput about 50% and cutting latency about 25%
  • Backed by a well-funded vendor: Cohere reached a $7B valuation in September 2025 with roughly $1.6B raised and around 450 staff

Cons

  • Superseded by its own vendor: the command-r-plus alias (04-2024) was deprecated on 15 September 2025 and Cohere now recommends Command A for new work
  • The published weights are CC-BY-NC, so self-hosting for commercial use is not permitted without a separate Cohere agreement — practitioners flag this as research-only licensing
  • 104 billion parameters make even licensed self-hosting expensive; this is not a model a team runs on a single commodity GPU
  • Hacker News practitioners report it is sensitive to prompt template formatting and showed mode collapse on open-ended creative tasks against Claude 3
  • Its April 2024 'beats GPT-4 in Chatbot Arena' framing drew benchmark-overfitting scepticism from commenters at the time
  • No G2, Capterra or TrustRadius rating exists for the model itself, so there is no aggregated buyer-review signal to check

Pricing

Free Trial Available

Trial API key

$0

  • Free rate-limited access to the API
  • Evaluation and prototyping only
  • Not licensed for commercial use

Production (pay-as-you-go), Command R+ 08-2024

From $2.50/1M input tokens

  • $2.50 per million input tokens
  • $10.00 per million output tokens
  • Grounded generation with citations
  • Multi-step tool use
  • 128K context

Model Vault (dedicated instance)

Contact for pricing

  • Logically isolated managed deployment
  • Hourly or monthly committed rates
  • Dedicated throughput

Private / on-premises

Contact for pricing

  • Deployment inside customer VPC or data centre
  • Full data sovereignty
  • Covers Command, Rerank and Embed

Public API pricing is per token and published: command-r-plus-08-2024 costs $2.50 per million input tokens and $10.00 per million output tokens, while the deprecated 04-2024 build was $3.00/$15.00 — roughly Claude-3.5-Sonnet-class pricing, and an order of magnitude above Cohere's smaller Command R at $0.50/$1.50. Trial API keys are free but rate-limited and explicitly barred from commercial use; production keys bill month-end or when the balance reaches $250. Dedicated Model Vault instances are quoted hourly or monthly, and private or on-premises deployment — which is what most of Cohere's large enterprise customers actually buy — carries no list price and goes through sales.

Security & Compliance

soc2
gdpr
hipaa
iso27001
sso
data residency

Connect

Sources

This page was written from 9 sources, 7 on domains other than docs.cohere.com.

  1. 1.docs.cohere.comcommand r plusvendor
  2. 2.docs.cohere.comdeprecationsvendor
  3. 3.cohere.compricing
  4. 4.cohere.comdeployment options
  5. 5.huggingface.coc4ai command r plus
  6. 6.aws.amazon.comcohere command r r plus amazon bedrock
  7. 7.the-decoder.comcohere improves its rag optimized command series llms
  8. 8.betakit.comcoheres valuation hits 7 billion usd following 100 million r
  9. 9.hn.algolia.comhn.algolia.com
Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe