Thomson Reuters announced Thomson on August 24, the first large language model the company has developed in-house. Rather than building a huge model from nothing, it started from an open-source foundation and poured 40 million USD worth of talent and compute into specializing it deeply for legal and tax work. Its first home is the document review capability inside CoCounsel Legal, the company's AI assistant for legal professionals.

Spending on Specialization Instead of Scale

For the past few years the accepted wisdom has been that reaching frontier-model territory requires billions of dollars in compute and years of infrastructure investment. Thomson Reuters took a fork in the road well before that point. It began with a strong open-source foundation and directed 40 million USD (roughly 6.4 billion yen) toward talent and compute, adding only the intelligence the work actually calls for.

What came out is a model the company fully owns and controls, and one that avoids the heavy inference costs typical of frontier models. Some reports put the cost of the final training run itself at around 450,000 USD (roughly 72 million yen). The gap between total investment and actual training cost says a lot about how this model was built.

※1 USD = 159 JPY (as of August 26, 2026)

Joel Hron, the company's Chief Technology Officer, noted that the AI industry has long treated scale as the answer, then argued that starting from a strong foundation and specializing it deeply for the work that matters yields something highly capable, far more efficient, and entirely under your own control. He frames it as a change in the economics of professional AI.

What It Was Taught Is What Makes the Difference

What sets Thomson apart is not where it started but what happened next. Decades of proprietary content from Westlaw, Practical Law, Checkpoint and Reuters served as the raw material, layered with state-of-the-art mid-training and post-training techniques. Hundreds of subject matter experts were built into the process, from the design of training objectives through to the final evaluations.

The resulting change shows up along two axes. The first is instruction following, where the ability to execute precise, multi-part professional instructions rose clearly above the base model. The second is navigating dense, domain-specific documents, and the gain there is larger still. That is the nuanced reasoning the hardest professional tasks demand.

There is a counterargument embedded here. The idea that a general-purpose model, if capable enough, only needs access to the right content to perform at an expert level is widely held. Thomson Reuters' early results suggest otherwise: proprietary training and human subject matter expertise, applied to a strong foundation, produce gains that content access alone does not deliver.

Notably, less than 10 percent of the company's content has been used in training so far. What comes next is described not as simply adding more data, but as digging further into the ways specialization itself can be done.

The Small Version Ships as Open Weights

Ahead of the announcement, Thomson Reuters opened the model to legal and AI academics for direct evaluation. It plans to keep making the model available externally to support validation and further development. Alongside that, a small version of Thomson has been released as an open-weight model on Hugging Face for academic and non-commercial use.

Jonathan H. Choi of Washington University School of Law, one of the evaluators, said he used some of the harder questions students had raised in his Corporate Tax class and compared the results against ChatGPT and Claude. All three models answered correctly, but he preferred Thomson's responses overall. He singled out the links to treatises, which made the answers more transparent and more useful for legal work.

Professor Samuel Dahan, who directs the Queen's Conflict Analytics Lab and the Cornell Legal AI Lab, found citation quality generally competitive with leading frontier models. That held even when the model was tested on Canadian employment-law questions without a Canada-specific setting. CEO Steve Hasker likewise said early evaluations place Thomson on par with the latest frontier models across a range of tasks.

First Deployment Is Tabular Analysis in CoCounsel

Thomson's first placement is inside Tabular Analysis in CoCounsel Legal. High-volume, structured document review is positioned as exactly the area where a purpose-built model's advantage translates directly into numbers. Law firms and corporate legal departments will be able to use it starting with the upcoming release.

CoCounsel Legal itself remains multi-model by design. Thomson gets applied where it holds a clear advantage, and leading third-party models handle the rest. The plan from here is to extend Thomson-family models across the legal and tax portfolio, with additional sovereign AI options to follow.

Underneath all of this sits a set of questions: how a model was trained, what behaviors and biases it carries internally, where it runs, and how a customer's information is protected. Professionals are increasingly sensitive to these points, and the company states plainly that customer data is not used for training without explicit consent. The reading behind this announcement is that competition is shifting from raw capability toward whether the output can be verified.

Summary

Thomson Reuters has released Thomson, its own large language model. The approach layers 40 million USD of specialized investment onto an open-source foundation, targeting instruction following and comprehension of professional documents in particular. A small version is available on Hugging Face for academic and non-commercial use, and external researchers are still working through their evaluations. Given that less than 10 percent of the company's data has been used so far, how far a domain-specialized model can pull ahead of general-purpose frontier models looks set to become visible in the numbers before long.