Granite 4.2 LLMs: How They're Built

| Source: Hugging Face Blog

Tags: Granite 4.2, IBM, reasoning models, Apache 2.0, agentic RL, chain-of-thought, open-source LLM

IBM releases Granite 4.2, a family of open-source reasoning models (3B, 8B, 30B) trained on 15T tokens with 512K context windows, agentic RL for tool use in real sandboxed environments, and Apache 2.0 licensing — a credible enterprise-safe alternative to proprietary reasoning models.

Details

IBM Granite 4.2 is the first reasoning-focused generation of the Granite family, launching in three sizes — 3B, 8B, and 30B — each trained from scratch on roughly 15 trillion tokens using a five-phase pre-training strategy that progressively extends the context window to 512K tokens. The models use a decoder-only transformer with Grouped Query Attention (40 heads, 8 KV heads), SwiGLU activations, RMSNorm, and RoPE position embeddings with θ=10,000,000. Post-pre-training follows a structured pipeline: supervised fine-tuning on chain-of-thought, reasoning, and agentic trajectory data, then multi-stage reinforcement learning. For the 8B and 30B models, this RL pipeline includes an agentic RL block where models learn to use real tools — writing and running code, operating a terminal, and searching the web inside sandboxed environments. Every model ships with three operating modes: full thinking (explicit chain-of-thought), low-effort thinking (short reasoning budget for routine queries), and non-thinking (direct responses). Native tool calling emits OpenAI function-calling format output, making the models drop-in compatible with existing agentic frameworks via vLLM or SGLang. All three are released under Apache 2.0, which removes legal friction for commercial and on-premises deployments. Enterprise teams needing auditable, weight-accessible reasoning models with genuine agentic training now have a well-documented, fully open alternative.