Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages

| Source: MarkTechPost

Tags: Cohere, North Small Translate, MoE, machine translation, WMT26, open-weight, multilingual

Cohere open-sourced North Small Translate, a 218B MoE model (25B active parameters) that scores 83.6 on WMT26 across 50 languages — outperforming DeepL NextGen (81.37) and Google Translate (68.20) on Cohere's own vendor benchmarks. Available free via API, for non-commercial self-hosting, or with a commercial license.

Details

Cohere released North Small Translate, the first dedicated translation model in its North family, built in partnership with RWS (Language Weaver). The architecture is a decoder-only sparse MoE Transformer: 218B total parameters, 25B active per forward pass, 128 experts with 8 activated per token plus a shared expert applied to every token. Attention alternates 3:1 between sliding-window layers (4K window, RoPE) and global layers, a layout first introduced in Command A. Context length is 16K tokens in and 16K out, text only. On Cohere's WMT26 evaluation — using GPT-5.6-Sol as the judge — the model scores 83.6 averaged across 50 languages, versus DeepL NextGen at 81.37, Qwen 3.5 397B at 81.56, and Google Translate at 68.20. An agentic variant with multi-pass self-correction reaches 84.36. These are vendor-reported numbers with a vendor-chosen judge; treat them as directional until independent WMT26 results appear. Three deployment paths: free via Cohere's API within rate limits, non-commercial self-hosting, or a commercial license. The model is open-weight but not fully open-source — a custom license governs commercial use. For teams running multilingual pipelines, North Small Translate offers a credible alternative to commercial translation APIs, particularly on languages where Google Translate underperforms. The 25B active parameter footprint keeps per-token compute tractable, though the full 218B weight load requires substantial VRAM for self-hosting.