Salesforce Koa: An Enterprise Language Model for Agentic Tool Use

| Source: arXiv AI

Tags: Salesforce, Koa, Nemotron, GRPO, enterprise LLM, Agentforce, agentic AI

Salesforce released Koa, an enterprise LLM post-trained from open-weight Nemotron-3-Super-120B using GRPO reinforcement learning — trained on public and synthetic data only, optimized for multi-turn agentic tool use in Agentforce, and surpassing a strong proprietary baseline on enterprise CRM benchmarks.

Details

Salesforce published a technical report on Koa, an enterprise LLM post-trained from the open-weight Nemotron-3-Super-120B foundation model using Group Relative Policy Optimization (GRPO). The model is trained exclusively on public and synthetic data — no customer data — simplifying data governance for enterprise buyers. Koa's distinctive training mechanism is a simulation-to-reward pipeline that expands workflow specifications into persona-conditioned multi-turn tasks with task-resolution rewards grounded in successful tool use. For enterprise CRM domains, these specifications are written in Agent Script — Salesforce's declarative language for Agentforce agents — meaning the model is trained on the same execution environment it will be deployed in. Across public tool-use, agentic-reasoning, and enterprise CRM benchmarks, Koa improves over its Nemotron base with the clearest gains on multi-turn tool use. It surpasses a strong proprietary baseline while remaining below leading frontier models — an honest positioning that differentiates it from typical corporate AI announcements. For enterprises running Agentforce, Koa represents a purpose-built model with training-deployment alignment. The simulation-to-reward methodology is also potentially adaptable by other enterprises seeking to specialize open-weight models for their own workflow specifications.