Thomson Reuters deploys Thomson, its first proprietary legal LLM, inside CoCounsel after a $40M build
Thomson Reuters ran the final training for its first in-house large language model for roughly $450,000, even as the total project drew $40 million over two years in headcount and compute, per SiliconAngle's reporting. The model, called Thomson, deployed August 24 inside CoCounsel Legal's Tabular Analysis feature, making Thomson Reuters the first major legal content publisher to ship a domain-trained proprietary LLM inside its core product.
What
Thomson starts from an undisclosed open-weight foundation and applies mid-training and post-training on decades of Westlaw, Practical Law, Checkpoint, and Reuters news content, per the company's technical blog. Westlaw spans more than 40,000 databases and 150 years of legal publishing. Thomson Reuters acquired AI research firm Safe Sign Technologies in 2024 to build the model team. Hundreds of subject-matter experts evaluated outputs and judged responses in blind comparisons during post-training, calibrating the model against how legal professionals actually work.
The first deployment is Tabular Analysis in CoCounsel Legal, where attorneys can run up to 10,000 documents through up to 100 structured questions per session, per the August 20 press release. CoCounsel remains multimodel: Thomson serves tasks where domain accuracy matters most; unnamed frontier models handle tasks requiring raw reasoning power. Customer data is not used to train the model.
Per TR's own benchmarks, Thomson performed competitively with Claude Opus 4.8 and ahead of GPT-5.5, Claude Sonnet 5, and Gemini 3.1 Pro across a combined legal and general benchmark suite covering LegalBench, PrBench, instruction-following, reasoning, and long context. Those results have not yet been validated by outside parties, per SiliconAngle's coverage. Less than 10% of TR's total content base was used in training, leaving what the company describes as significant headroom for future capability gains.
Why it matters
The economics are the real news. A final training run at $450,000 shows that established publishers with content moats can produce domain-competitive models without building foundation architecture from scratch or spending at frontier-lab scale. Bloomberg ran a similar playbook in 2023 for financial text; Adobe used a comparable approach for Firefly on creative content. Thomson extends that pattern to legal, the first time a company with a deep proprietary legal corpus has shipped a named proprietary model inside a commercial legal product.
The content position is also a model position. Westlaw's 40,000+ databases and 150 years of editorial curation represent training data no competitor can assemble from public sources. Law firms running CoCounsel for high-volume document review will now depend on a proprietary model alongside general-purpose alternatives, whether or not they select Thomson as the explicit default.
For teams building legal AI pipelines: TR's choice of Tabular Analysis as the first deployment signals where domain-specific models earn a defensible edge. Structured accuracy across thousands of documents, measured by completeness rubrics and citation-level factuality checks, is exactly the kind of task a purpose-trained legal model can win on clear, measurable grounds. That benchmark design is also worth watching as a template for how other vertical publishers will justify their own model investments.
What to watch next
Thomson Reuters plans to expand the model across its legal and tax portfolios over the next year, per the company blog. Checkpoint, TR's tax research platform, is the most consequential expansion to track: it would give Thomson dual legal and financial training data. TR is also in early discussions with large law firms and corporations about direct API access, which would let those organizations adapt Thomson to their own internal knowledge bases and workflows.
Sources
- Thomson Reuters press release: Next Generation of CoCounsel Legal (primary, vendor press release)
- Thomson Reuters blog: Thomson Reuters Built Its Own AI Model (primary, vendor technical blog)
- SiliconAngle: Thomson Reuters launches proprietary AI model for legal work (secondary)
- Artificial Lawyer: TR Launches Thomson 1.0, Its Own LLM (secondary)
