Do not switch your Agentforce org to Koa, and do not wait for it either. Salesforce's CRM reasoning model is a pilot with no published price. In the arXiv paper Salesforce's own researchers wrote about it, Koa beats GPT-4.1 but stays "below the strongest frontier models". Keep Agentforce's model choice at the sub-agent level. Run the two options that are generally available today, Claude on Amazon Bedrock and Gemini, against the default on a held-out set of your own cases. When Koa reaches GA, put it on the high-volume routine sub-agents where your numbers say it wins.
| Salesforce Koa | Claude (Amazon Bedrock) | Gemini (Vertex AI) | Salesforce Default | |
|---|---|---|---|---|
| What it is | Nemotron 3 Super 120B, post-trained by Salesforce on synthetic CRM data | Anthropic Claude inside Salesforce's trust boundary | Google Gemini 3.5 Flash | Managed mix: GPT-4.1 (new builder) / GPT-4o (legacy) |
| Status (2026-09-27) | Pilot; GA "winter 2026", U.S. regions only | Available | GA since Dreamforce | Available, and it changes without you |
| Published price | None | Same Flex Credit rate card | Same Flex Credit rate card | Same Flex Credit rate card |
| Independent eval | None | None for Agentforce | None for Agentforce | None for Agentforce |
| Best fit | Routine, repetitive CRM actions (update, route, schedule) | Multi-step sub-agents where a wrong action is costly | Latency-sensitive, high-volume sub-agents | Orgs that have not built a test set yet |
| Do not pick if | You are outside the U.S., or you need a contracted price this year | You can't give up the default mix's automatic updates | You need the best multi-turn tool-use score | You need to regression-test a pinned model |
The loser is the Salesforce Default, run without measurement. It is the baseline Koa's own paper beats. It is also the one option whose underlying model Salesforce re-tunes without a deployment on your side. A model you cannot pin is a model you cannot regression-test.
What is Salesforce Koa, actually?
Koa is NVIDIA's open-weight Nemotron 3 Super model, post-trained by Salesforce for CRM tool use and hosted entirely by Salesforce. Salesforce's press release says the post-training used supervised fine-tuning plus reinforcement learning with Group Relative Policy Optimization (GRPO). It trained on "a proprietary synthetic dataset modeled on enterprise knowledge from nearly three decades of CRM deployments" across more than 14 industries, and states that "no customer data was used." The same release says Salesforce "controls the model weights and performs post-training and inference entirely within its own trust boundary."
The base model is not small. NVIDIA's model card describes a 120B-parameter hybrid Mamba-Transformer mixture-of-experts with 12B parameters active per token, released under the NVIDIA Nemotron Open Model License. NVIDIA's launch post pitches it for multi-agent workloads and long context. Constellation Research reads the economics plainly: software vendors are building on Nemotron because they are "keeping their AI costs in check". Keep that in mind when the price arrives. If Koa is cheaper for Salesforce to serve than a rented frontier model, nothing published yet says that saving reaches your bill.
One operational detail matters for agents. Salesforce's Koa product page says the model is hosted at "temperature 0" so responses stay consistent. That helps evaluation: the same input should produce the same action, so an A/B test measures the model and not sampling noise.
How good is Koa? The paper and the marketing disagree
Salesforce's marketing claims Koa beats "leading" models. Salesforce's own paper claims only that it beats GPT-4.1. The two statements can both be true, but they do not describe the same model.
The marketing numbers come from Salesforce's CRM Bench. The Koa page claims "3x fewer errors," "11% more precise at calling the right action," "2.1 times greater reliability," and "15% better at remembering context in long back-and-forth conversations." Salesforce Ben repeats the headline claim that Koa "matches or does better than leading model performance on CRM actions." None of these name the comparison model.
The arXiv paper does name it, and its numbers are more modest:
| Benchmark (paper, Table 1) | Koa | Nemotron 3 Super base | GPT-4.1 | Claude Opus 4.8 | GPT-5.5 |
|---|---|---|---|---|---|
| Tau2Bench (weighted avg) | 69.41 | 68.64 | 54.48 | 74.00 | 83.99 |
| BFCL accuracy | 66.63% | 64.73% | 53.96% | 78.18% | 67.63% |
| CRM Bench (overall) | 0.86 | 0.84 | 0.81 | 0.87 | 0.90 |
Read the base-model column. On every benchmark in that table, most of Koa's lead over GPT-4.1 was already in the open-weight base model before Salesforce trained it. Post-training adds 0.8 points on Tau2Bench and 0.02 on CRM Bench. Then read the last two columns: in Salesforce's own table, Claude Opus 4.8 and GPT-5.5 both score above Koa on all three benchmarks, which is what the paper means by "below the strongest frontier models." That is the honest version of the pitch: Koa is a well-tuned open model that beats a 2025-era proprietary model on CRM tool calls.
This is not a small point for buyers. GPT-4.1 is also what the Salesforce Default mix runs for agents in the new builder. So "Koa beats the model you are probably running now" is a fair reading. "Koa beats Claude or Gemini" is not in any published evidence: the paper does not test Gemini, and the Claude model it does test, Opus 4.8, comes out ahead. Opus 4.8 is not necessarily the Claude model your Agentforce org runs, so that is not a verdict on the Bedrock option either. MindBlaze's write-up reaches the same conclusion: "A vendor's own benchmark is a vendor's own benchmark, and 'leading models' is not named."
For a sense of how hard real CRM agent work is, Salesforce AI Research's own CRMArena-Pro found leading LLM agents at about 58% single-turn success, falling to about 35% multi-turn, with near-zero inherent confidentiality awareness. That was an older model generation. Still, it shows why a 0.02-point CRM Bench gain should not decide anything for you.
What are the real alternatives inside Agentforce today?
Three models can run the Agentforce Reasoning Engine today: the Salesforce Default mix, Anthropic Claude on Amazon Bedrock, and Google Gemini. Koa is a pilot, and OpenAI models on Bedrock have been announced as coming soon.
Salesforce Default. This is a managed mix. Per Salesforce Dictionary's summary of the routing architecture, it currently includes GPT-4.1 for new-builder agents and GPT-4o for legacy-builder agents, and you cannot choose models within it. Its strength is that improvements arrive without work on your side. The same property is its weakness: an agent that passed your tests last month can be running a different model this month.
- Who should NOT use it: any team that needs a pinned model for a regulated workflow or a change-controlled release.
Claude on Amazon Bedrock. The Claudeforce announcement says Claude serves "as a reasoning model for the Atlas Reasoning Engine." It runs on Amazon Bedrock "within the Salesforce Trust Boundary," and it powers Agentforce Vibes and Agentforce Coworker by default. Check which Claude model you actually get. Third-party write-ups cite Claude Haiku 4.5 as the tested override option, and some sources mention Sonnet 4.6, so confirm the model shown in your own Setup. Do not confuse this with Claudeforce, the Salesforce-in-Claude plugin, which has its own contract and meter.
- Who should NOT use it: orgs whose procurement has not approved Anthropic as a sub-processor. The Army's IL5 Agentforce deployment ran with Anthropic's models disabled for exactly that kind of reason.
Gemini. Salesforce and Google Cloud's September 15 announcement lists "Agentforce Reasoning Engine with Gemini: GA now". The documented model is Gemini 3.5 Flash on Vertex AI. A Flash-tier model suits sub-agents where response time is part of the experience: chat deflection, high-volume case triage, voice.
- Who should NOT use it: teams putting it on long multi-step sub-agents without first testing that exact workload. Flash is Google's speed tier.
Koa. Salesforce names six pilot customers: 1-800Accountant, Baxter Credit Union, Engine, Formula 1, UChicago Medicine and Xero. The work it was trained for is the routine CRM action: updating opportunities, routing cases, scheduling follow-ups.
- Who should NOT use it: anyone outside the U.S. before GA extends past U.S. regions, anyone who needs a price in a 2026 budget, and anyone whose agent mostly writes free-form text rather than calling tools.
Where do you actually choose the model?
Model choice in Agentforce sits at three levels: org, agent and sub-agent. The most specific setting wins. That makes a mixed-model org normal, not an edge case.
- Org-wide: in Setup, the Select Agentforce Model Option setting chooses Salesforce Default, Claude on Bedrock or Gemini.
- Per agent and per sub-agent: a
model_configin Agent Script, or the model picker in Agentforce Builder. The precedence is sub-agent over agent over org. The same source notes that only some identifiers are "thoroughly tested" and that other models "may not suit your agent's tools." - Koa's three entry points: Salesforce says Koa will appear as a managed LLM in the Data Cloud generative models catalogue, as an org-wide provider in Agentforce Setup, and as a per-agent or per-sub-agent choice in the builder (Koa page). Prompt Builder has its own model selection. Salesforce says Google's models power both Prompt Builder and the Reasoning Engine, so one org can already run one model in a prompt template and another in an agent.
The reasoning model is also not the only model in the loop. Salesforce moved topic classification to HyperClassifier, an in-house small model, and cut LLM calls before first output from four to two. The model you pick decides planning and action choice. It does not decide everything, so test the whole agent, not a prompt in a playground.
What does it cost to switch models?
Nothing on the rate card, and that is why the choice should turn on wrong-action rate, not price. As of 2026-09-27, Salesforce's Agentforce pricing page lists Flex Credits at $500 per 100,000 credits. A standard action is 20 credits ($0.10) and a voice action 30 credits ($0.15), with no line that varies by reasoning model. The page does not list Koa at all, and no Koa price or consumption rate was published at launch.
Take a defined workload: a service org running 50,000 standard actions a month across three sub-agents. At the published rate that is 1,000,000 credits, or $5,000 a month, whichever of the three GA models runs it. The multiplier to watch is size, not model. An action that goes over the 10,000-token boundary bills as more than one action. A more verbose model that pushes its context over that line raises your bill even though its rate is the same. We covered how these meters compound in our agentic pricing guide.
The real cost of a model is wrong actions. At 50,000 actions a month, a 1-point difference in wrong-action rate is 500 bad record updates, misrouted cases or missed follow-ups. Each one costs a human cleanup far above $0.10.
Which model should you pick? A decision you can defend
Pick per sub-agent, on your own test set, and treat Koa as a candidate you add in winter rather than a plan you wait for. These are the criteria that predict regret:
- Region. Outside the U.S., Koa is not an option for this budget year. Our LLM data residency guide covers what "in region" means for Bedrock and Vertex.
- Action-heavy vs text-heavy. Koa was trained with rewards for resolving tasks through successful tool calls. A sub-agent that mostly drafts emails is the wrong place to test it.
- Pinning. If a model change needs a change ticket, move off the Default mix to an explicit model.
- Sub-processor approval. Claude runs on AWS and Gemini on Google Cloud. Koa is the only option where Salesforce hosts the weights itself, which simplifies a vendor review, once it exists outside the pilot.
What changes the answer: an independent CRM evaluation that includes Claude and Gemini, a published Koa consumption rate that differs from the standard action price, or GA outside the U.S.
A 30-day evaluation plan
This Week:
- Pick two or three high-volume routine workflows (case routing, opportunity updates, follow-up scheduling) and pull 200-300 real, closed cases for each from the last quarter. Keep them out of every prompt and every example: this is your held-out set.
- Write the correct action for each case and have the process owner sign it off. The label is the expensive part and the reusable asset.
This Month:
3. Run each sub-agent on the Default mix, Claude and Gemini with model_config overrides in a sandbox. Score wrong-action rate, escalation rate and credits per resolved task. Leave "accuracy" and vibes out of it. Our TypeSafe Jev analysis shows how to split a decision so errors are attributable.
4. Watch the SalesBleed class of failure too: include a handful of injected inputs in the test set. A model that follows a malicious instruction faster is not the better model.
Before Winter GA: 5. Ask your account team for pilot access, the Koa consumption rate in writing, and the CRM Bench methodology. Re-run the same held-out set on Koa. Move a sub-agent only if Koa beats your current pick on wrong-action rate at equal or lower credits per resolved task.
The Bottom Line
Salesforce has built a model it owns and hosts instead of renting a frontier one, trained it carefully, and benchmarked it against the model Agentforce already runs by default. That is a reasonable product. It is not yet evidence. The Dreamforce reveal put Koa on the main stage. The paper puts it just above GPT-4.1 and below the frontier. Your test set is the only place it gets a verdict that applies to your org.
Build the test set now. It is the one part of this decision that does not expire.
Continue Reading
- Dreamforce 2026 Shipped Agent Operations, Not the Headline Model
- Claudeforce Has No Published Price and Two Separate Meters
- Claude vs GPT vs Gemini: Stop Comparing Per-Token Prices
- Model Router Buyer's Guide: Buy Failover, Not Judgment
- SalesBleed Turned a Public Web Form Into an Agentforce Data Leak
- The Army Shipped Agentforce. Anthropic's Models Were Off.
