The Research and Development Center for Large Language Models (LLMC) at Japan's National Institute of Informatics has released LLM-jp-4 33B, a dense model with roughly 33.2 billion parameters. Two variants sit on Hugging Face under the Apache License 2.0: a base model that has gone through pre-training and mid-training, and a post-trained model aimed at reasoning. The datasets behind the training were opened at the same time, continuing the project's habit of shipping more than just weights.

A third size, and the largest in the series so far

The LLM-jp-4 series began in April 2026 with a dense 8B model and a 32B-A3B model built on Mixture of Experts (MoE), an approach that routes work through a set of smaller expert modules. The new 33B is dense rather than MoE, keeping every parameter active, and it is the largest member of the series so far.

Two checkpoints are available. llm-jp-4-33b-base has finished pre-training and mid-training only, and is meant as a foundation for further training against a specific domain. llm-jp-4-33b-thinking takes that base through supervised fine-tuning (SFT) and direct preference optimization (DPO).

The architecture uses 64 layers, a hidden size of 5,120, and 40 attention heads, with a context length of 65,536 tokens. Total parameter count is 33,219,548,160, and the weights ship in BF16.

11.7 trillion tokens, with no reinforcement learning at the end

Pre-training and mid-training together consumed 11.7 trillion tokens. Both corpora are published on LLM-jp's GitLab. Licensing constraints keep some portions out of the public release, but the material fed into the model is far more traceable from the outside than is typical.

The post-training recipe is notable too. The team used SFT and DPO alone and skipped reinforcement learning, a different choice from the reasoning-focused models that have leaned on RL over the past year. The DPO dataset used for this release was published alongside the weights.

The tokenizer is a Unigram byte-fallback model carrying the vocabulary from llm-jp-tokenizer v4.0. The chat template is designed to be compatible with OpenAI's Harmony response format, but tokenization through the openai-harmony library is not supported. The bundled tokenizer is required for correct behavior, which is worth knowing before deployment.

Beating every previous LLM-jp-4 model across the board

Evaluation ran on llm-jp-judge, the framework LLM-jp maintains. The judge model is gpt-5.4-2026-03-05, and the four tasks are MT-Bench in Japanese and English, AnswerCarefully for Japanese safety, and llm-jp-instructions, a set of human-written single-turn question and answer pairs. Scores are averages over three rounds of inference and evaluation.

With reasoning effort held at medium, the results look like this.

Model MT-Bench (JA) MT-Bench (EN) AnswerCarefully llm-jp-instructions
llm-jp-4-33b-thinking 8.00 8.24 3.79 3.79
llm-jp-4-32b-a3b-thinking 7.82 7.86 3.70 3.61
llm-jp-4-8b-thinking 7.54 7.79 3.69 3.54
gpt-oss-20b 7.33 7.85 3.55 3.16
gpt-5.4-2026-03-05 8.87 8.89 4.43 4.82

The new model outscored every previously released LLM-jp-4 Thinking model on all 4 benchmarks. Despite sitting at a similar total size, it pulled clear of the MoE-based 32B-A3B, which reads as a payoff for going dense. It also leads the open-weight gpt-oss-20b on every task.

A gap to gpt-5.4, the model doing the judging, still remains. The MT-Bench difference stays under 1 point, but AnswerCarefully is 0.64 behind and llm-jp-instructions more than 1 point behind, so safety and instruction following are where the distance shows. One caveat on comparisons: the earlier llm-jp-3 series was scored with gpt-4o-2024-08-06, and the stricter current judge makes those older numbers hard to line up directly.

Opening the data, not only the weights

The license is Apache License 2.0, which permits commercial use. Weights, the SFT and DPO datasets, and the training corpora all sit under the same open arrangement, which makes it practical to fine-tune against internal documents or to build Japanese evaluation infrastructure on top. Quantized builds have already started appearing from third parties, opening a path to running the model through llama.cpp, Ollama, or LM Studio.

The team is explicit that these models are early-stage research output and have not been tuned to guarantee that outputs align with human intent or safety expectations. Treating them as something to validate and adjust further, rather than dropping them straight into a user-facing service, is the sensible reading. Development also drew on the NINJAL Web Japanese Corpus from the National Institute for Japanese Language and Linguistics.

Summary

LLM-jp-4 33B is a 33.2 billion parameter dense model trained on 11.7 trillion tokens, and it beats the existing models in its series across all 4 benchmarks. Finishing with SFT and DPO instead of reinforcement learning, and opening the training data rather than just the weights, defines the character of this release. For anyone looking for a Japanese-first foundation with commercial use on the table, it is worth a hands-on look.