IBM released Granite 4.2, its enterprise-focused family of AI models, on August 25, 2026. The lineup comes in three sizes with 3 billion, 8 billion, and 30 billion parameters, all published under the Apache 2.0 license. The headline additions are native step-by-step reasoning and an agentic training block applied only to the 8B and 30B models. Two speech models with 470 million parameters arrived on the same day.

Three Sizes and a Thinking Switch

Granite 4.2 uses a dense, decoder-only architecture, meaning every parameter is active on every token. Unlike Mixture-of-Experts designs that fire only part of the network, a dense model runs broadly across deployment targets without special handling.

All three sizes can emit a chain of thought before answering. Operators can flip between thinking and non-thinking modes, and a low-effort mode sits between them, spending only a short reasoning budget. The point is to avoid paying a long deliberation cost on questions that do not need it.

Tool calling is native to the models. Served through an OpenAI-compatible endpoint such as vLLM, the models emit tool calls in the OpenAI function-calling format, so they drop into existing agent harnesses without extra glue code. vLLM and SGLang both support the release, and the weights are available through Hugging Face, Ollama, GitHub, LM Studio, OpenRouter, and Replicate. Japanese is among the supported languages.

Cloud, on-premises, and edge deployments are all in scope. The 3B model suits high-throughput agentic work, while the 30B is reserved for heavier reasoning and coding. Quantized variants ship as well: FP8, NVFP4, and MXFP4 built with LLM Compressor, plus GGUF conversions from llama.cpp spanning Q2_K through Q8_0.

A Foundation Rebuilt from 15 Trillion Tokens

Pre-training ran from scratch on roughly 15 trillion tokens across five phases. Phases 1 and 2 handle foundational pre-training, phases 3 and 4 perform mid-training that anneals toward higher-quality data, and phase 5 covers long-context training, stretching the context window to 512,000 tokens. The sequence length in the released model configuration, however, is 131,072 tokens (128K).

The architecture uses Grouped Query Attention with 40 attention heads and 8 KV heads, Rotary Position Embedding with θ = 10,000,000, SwiGLU activations, RMSNorm, and bfloat16 precision. Input and output embeddings are not tied. The 3B and 8B models have 40 layers; the 30B has 64.

The data pipeline received attention as well. IBM fed in 1 trillion tokens of synthetic code generated through its CodeAlchemy pipeline and added a speculative decoding layer for faster output. Supervised fine-tuning used roughly 7.2 million samples, about 100 billion tokens, with agentic data accounting for 31.6 percent of the mixture. Within the agentic corpus, software engineering makes up 69 percent, tool calling 12.1 percent, and terminal use 8.0 percent.

For quality control, IBM used GPT-OSS-120B and Gemma 4 as LLM judges, discarding low-scoring samples along with anything containing fabricated information or invalid tool interactions. Deduplication relied on SHA-256 hashes over the combined tools and messages fields, applied both within individual sources and across the full mixture.

Multi-Stage Reinforcement Learning in Live Environments

The post-training stack is where this release does its real work. Instead of a single reinforcement learning pass, IBM chained focused stages together, each warm-starting from the previous checkpoint. The order runs RLVR (reinforcement learning with verifiable rewards), skill boosters, a software engineering agent, terminal, search, and finally RLHF.

RLVR blends tasks whose answers can be checked mechanically: math with boxed-answer verification and formal proving in Lean, competitive coding graded against hidden tests in a sandbox, graduate-level science questions, instruction following, and single-step tool calls. Each step pairs 256 prompts with 16 sampled responses for a 4,096-example batch consumed in one optimizer step. The algorithm is asynchronous GRPO, so the generation and training halves of the loop never block each other.

The agentic block that follows runs only for the 8B and 30B models. The 3B stops after foundational RL and alignment.

In the software engineering stage, each task is a real repository in its own sandbox. Driven by the OpenHands harness, the model reads code, edits files, and runs the test suite, and the reward is simply whether the hidden tests pass. The terminal stage uses the Harbor and Terminus-2 harness, running rollouts of up to 64 turns in a live shell where the model plans commands, reads their output, and recovers from errors. The search stage has the model answer hard multi-hop questions through live web-search tool calls, with an LLM judge scoring the final answer.

Training ran on NVIDIA's NeMo-RL and NeMo-Gym stack, on an NVIDIA GB200 NVL72 cluster hosted by CoreWeave. The closing RLHF stage tunes preference and safety while also applying a reasoning-length penalty to curb overly verbose thinking.

Benchmarks Show the Gap Between 8B and 30B

IBM's published numbers make the effect of the agentic block clear. On SWE-bench Verified, the 30B scores 57.00 and the 8B scores 47.67. On Terminal-Bench 2.1, the figures are 29.24 and 20.56. The 3B is not measured in this category.

On reasoning benchmarks, AIME25 comes in at 89.17, 86.67, and 78.33 for the 30B, 8B, and 3B respectively. GPQA lands at 66.41, 64.14, and 54.80, and LiveCodeBench v6 at 75.77, 73.24, and 69.71. MMLU-Pro reaches 77.60 on the 30B, and RULER at 128K context scores 81.38.

The 3B holds up well on math and instruction following, so it remains a reasonable pick where footprint matters most. For work that involves actually calling tools and finishing tasks, the benchmarks back the case for 8B or larger. These figures come from IBM itself and have not been independently verified by a third party.

A 470M Speech Model Shipped Alongside

IBM also released Granite Speech 5.0 Turbo CTC and a non-commercial variant. At 470 million parameters, they are among the smallest models in the Granite family. The version number jumped from 4.1 to 5.0 because the structure changed rather than merely improved.

The key change is the absence of an LLM backbone. Using connectionist temporal classification, the models map audio directly to text, keeping size down while raising recognition efficiency. IBM measured a processing throughput (RTFx) of roughly 12,600 on a single H200 GPU. Speed leaders on the Hugging Face Open ASR leaderboard sit around 6,000, so this is close to double. That works out to transcribing three hours of audio in a second, which opens the door to always-on use on laptops and phones as well as bulk call-center log processing.

IBM added that it is working with Hirundo on machine unlearning technology to reduce undesirable outputs without retraining the model from scratch.

Summary

Granite 4.2 is a case of improving agent performance through post-training design rather than raw scale. The value lies in applying reward-driven training across real repositories, live shells, and web search, then shipping the result as open weights under Apache 2.0. Choosing between the 8B and 30B comes down to how hard the tool work is and what compute is available.