Salesforce Koa: A CRM Reasoning Model Built on Nemotron
Salesforce and NVIDIA's Koa post-trains open-weight Nemotron-3-Super-120B for CRM tool use, piloting now inside Agentforce.
Salesforce and NVIDIA's Koa post-trains open-weight Nemotron-3-Super-120B for CRM tool use, piloting now inside Agentforce.
Introduction
Salesforce and NVIDIA on September 15, 2026 announced Koa at Dreamforce 2026 in San Francisco, describing it as Salesforce's first CRM reasoning model for Agentforce. Koa is built by post-training NVIDIA's open-weight Nemotron-3-Super-120B foundation model, and it is meant to handle the multi-step reasoning and tool calling that CRM agents rely on inside Agentforce workflows. Until now, Agentforce routed its most demanding reasoning prompts to frontier models such as Claude and ChatGPT through Salesforce's AI gateway. Koa is Salesforce's attempt to bring a version of that reasoning in-house rather than continuing to rely solely on third-party frontier providers. The announcement arrived alongside a research paper on arXiv (arXiv:2609.15066) describing the technical approach, and the two documents — one written for press, one for peer review — frame the model's performance somewhat differently, which is worth examining closely before treating this as a simple frontier-model replacement.
Feature Overview
Koa starts from Nemotron-3-Super-120B, the 120-billion-parameter open-weight model NVIDIA released as part of its Nemotron line, then applies additional post-training aimed specifically at enterprise tool use. According to the arXiv paper, the core training method is reinforcement learning via Group Relative Policy Optimization (GRPO); Salesforce's press release adds that supervised fine-tuning and NVIDIA's NeMo RL, NeMo Gym, and NeMo AutoModel tooling were also used in the pipeline.
The most distinctive part of the approach is what the paper calls a simulation-to-reward pipeline. Enterprise workflow specifications — written in Agent Script, Salesforce's declarative language for building Agentforce agents — are expanded into persona-conditioned, multi-turn tasks. For public tool-use domains, workflow structure is synthesized directly rather than drawn from Agent Script. In both cases, task-resolution rewards are grounded in whether the model actually completes the correct sequence of tool calls for data-dependent requests, rather than just producing plausible-looking text.
All of the training material — public data and synthetic scenarios generated by that pipeline — was built without using customer data, according to Salesforce. Those synthetic scenarios span more than 14 industries, including manufacturing, financial services, healthcare, and travel, and each one pairs a persona with a set of tasks mapped to the specific actions and tool calls needed to resolve them, covering activities like lead generation, opportunity qualification, and service case resolution.
Deployment-wise, Koa shows up as a managed model in Salesforce's generative AI models catalogue. It can be selected org-wide as a model provider in Agentforce, or set at the individual agent or sub-agent level inside Agentforce Builder.
Usability Analysis
For the moment, Koa is available only to a small set of pilot customers inside Agentforce — Salesforce named 1-800Accountant, Baxter Credit Union, Engine, Formula 1, UChicago Medicine, and Xero, and says it also runs internally as a Slack-based employee agent. General availability is expected in winter 2026, limited to U.S. regions at launch. Salesforce is running a related but separate track: Missionforce Operations, aimed at air-gapped, private-cloud, and government deployments, is generally available now in U.S. regions, with post-trained NVIDIA models rolling out to select customers there starting in October 2026.
From an administrator's perspective, adopting Koa doesn't require redesigning agent workflows — it's a model-provider selection inside Agentforce Builder, slotting into the same AI gateway architecture that previously routed reasoning-heavy prompts to Claude or ChatGPT. That makes evaluating it relatively low-friction for existing Agentforce customers once GA arrives, though the narrow pilot list means most organizations have no hands-on experience with it yet.
Pros and Cons
Pros:
- Runs entirely within Salesforce's own infrastructure for both post-training and inference, so no customer data crosses a third-party trust boundary.
- Trained only on public and synthetic data with documented provenance — a direct response to what Salesforce AI lead Jayesh Govindarajan described as a lack of clarity around what data some competing models are trained on.
- Salesforce reports fewer errors than "leading" models specifically on real-world CRM actions such as updating opportunities, routing cases, and scheduling follow-ups, per its own CRM Benchmark.
- Built on an open-weight foundation, giving Salesforce a base it can continue to post-train and iterate on rather than depending on a closed license.
Cons:
- Pricing for Koa has not been disclosed.
- The press release's claim that Koa "matches or exceeds leading model performance" sits awkwardly next to the arXiv paper's own conclusion that Koa "remains below the strongest frontier models" — the marketing framing is more optimistic than the research behind it.
- Despite its open-weight starting point, Koa itself is not open: customers get access to a Salesforce-hosted managed model, not the weights.
- Broad availability is still months away, with general access limited to winter 2026 in the U.S. only.
Outlook
Salesforce frames Koa as solving a specific gap: Govindarajan told reporters that reasoning was the one capability Agentforce had relied on frontier providers for "until now," and that Salesforce lacked a pre-trained base that was simultaneously sovereign, state of the art, and transparent about its data provenance. NVIDIA's Kari Ann Briski described the appeal from its side as a mix of sovereign AI, time-to-first-token, and efficient reasoning. Notably, this isn't a full pivot away from frontier vendors — Salesforce separately announced Claudeforce, a partnership with Anthropic, on August 26, 2026, just three weeks before Koa. The more likely trajectory is a layered model strategy: Koa, and similar in-house, post-trained models, handling well-defined, high-volume CRM actions where cost, latency, and data control matter most, while frontier models continue to be routed in for more open-ended reasoning tasks through the same AI gateway.
Conclusion
Koa is a meaningful data point in the broader trend of enterprise software vendors post-training open-weight foundation models rather than exclusively reselling frontier-model access. The trust-boundary and data-provenance argument is a real and differentiated pitch, and the CRM-specific error-rate claim is worth watching once independent evaluations appear. But the gap between Salesforce's public framing and its own paper's more modest performance claim is the detail buyers should weigh most carefully — Koa is described by its own researchers as ahead of a strong proprietary baseline, not ahead of the frontier. It's most relevant right now to existing Agentforce customers curious about cost and control trade-offs, rather than to anyone shopping for a frontier-class general-purpose model.
Editor's Verdict
Salesforce Koa: A CRM Reasoning Model Built on Nemotron earns a solid recommendation within the Other LLM space.
The strongest case for paying attention: runs entirely within Salesforce's own infrastructure for both post-training and inference, avoiding third-party trust boundary crossings. That alone raises the bar for what readers should expect in this space. Reinforcing that, trained exclusively on public and synthetic data with documented provenance, addressing enterprise concerns about opaque training data — practical value rather than just headline appeal. The broader signal worth registering is straightforward: Koa is a post-trained version of NVIDIA's open-weight Nemotron-3-Super-120B, not a model built from scratch. On the other side of the ledger, one constraint is real rather than a marketing footnote: pricing for Koa has not been disclosed. It should factor into any serious decision. Layered on top of that, the press release's "matches or exceeds leading model performance" framing is in tension with the paper's own admission that Koa remains below the strongest frontier models — which narrows the set of teams for whom this is an obvious yes.
For multi-model deployment teams, cost-conscious operators, and developers willing to evaluate beyond the major labs, this is a serious evaluation candidate, not just a curiosity to bookmark. For everyone else, the safer posture is to monitor coverage and revisit once the use cases that matter to your team are demonstrated in the wild.
Pros
- Runs entirely within Salesforce's own infrastructure for both post-training and inference, avoiding third-party trust boundary crossings.
- Trained exclusively on public and synthetic data with documented provenance, addressing enterprise concerns about opaque training data.
- Reports fewer errors than "leading" models specifically on core CRM actions like updating opportunities, routing cases, and scheduling follow-ups.
- Built on an open-weight foundation model, giving Salesforce more control over future iteration than a closed-license base would.
Cons
- Pricing for Koa has not been disclosed.
- The press release's "matches or exceeds leading model performance" framing is in tension with the paper's own admission that Koa remains below the strongest frontier models.
- Despite its open-weight base, Koa's weights are not available to customers — it's accessible only as a managed, Salesforce-hosted model.
- Broad availability is still months away; today's access is limited to a small set of pilot customers.
References
Comments0
Key Features
Koa is a CRM-focused reasoning model that Salesforce built by post-training NVIDIA's open-weight Nemotron-3-Super-120B foundation model using reinforcement learning (GRPO) plus supervised fine-tuning on NVIDIA's NeMo RL, NeMo Gym, and NeMo AutoModel stack. Its core innovation is a simulation-to-reward pipeline that expands Agent Script workflow specifications into persona-conditioned, multi-turn training tasks, with rewards grounded in correct tool-call sequences rather than plausible-looking text. Training used only public and synthetic data — no customer data — spanning more than 14 industries. Koa is deployed as a Salesforce-hosted managed model selectable inside Agentforce Builder, keeping training and inference entirely within Salesforce's own infrastructure.
Key Insights
- Koa is a post-trained version of NVIDIA's open-weight Nemotron-3-Super-120B, not a model built from scratch.
- The simulation-to-reward pipeline converts Agent Script workflow specifications into persona-conditioned, multi-turn training tasks.
- Salesforce trained Koa entirely on public and synthetic data spanning more than 14 industries, explicitly excluding customer data.
- The arXiv paper's own conclusion is more measured than the press release: Koa beats a strong proprietary baseline but "remains below the strongest frontier models."
- Salesforce's "three times fewer errors" claim is measured against its own CRM Benchmark, not an independent or third-party benchmark.
- Even though Nemotron is open-weight, Koa itself remains a Salesforce-hosted managed model — customers cannot download or self-host its weights.
- Koa's rollout is currently limited to a handful of pilot customers, with general availability not expected until winter 2026 in U.S. regions.
- The move doesn't replace Salesforce's frontier-model partnerships — the company struck a separate Claudeforce deal with Anthropic just three weeks earlier.
Was this review helpful?
Share
Related AI Reviews
Cohere North Small Translate: MoE Translation Model Debuts
Cohere's North Small Translate is an open-weight MoE model for machine translation, scoring 83.60 on Cohere's WMT26 benchmark tests.
DeepSeek Opens V4-Flash-Vision-Exp Weights Under MIT
DeepSeek published its first V4-family multimodal checkpoint to Hugging Face on Aug. 31 under MIT, ten days after the API release.
GLM-5.3-Flash Review: Ox Alpha Unmasked as 320B MoE
Z.ai unmasks stealth model 'ox-alpha' as GLM-5.3-Flash, a 320B MoE with hybrid attention, 1M context, and MIT-licensed weights.
Harvey Launches Tenet, a Legal AI Model Built on Kimi K3
Harvey post-trained Moonshot's open-weight Kimi K3 into Tenet, its first in-house legal model, nearly doubling LAB benchmark task completion.
