Summary
Fahd Mirza puts Thomson Reuters’ newly released Thomson-1 Small model through a structured evaluation to determine whether an open-weight legal AI can perform at the level of a junior lawyer. The model is built on Cohere’s open 35-billion-parameter mixture-of-experts architecture, then refined through three stages: constitutional DPO value alignment, continued pre-training on decades of Thomson Reuters’ proprietary Westlaw legal and tax data, and task-specific training on how legal professionals actually work.
Mirza designs three distinct tests. First, a contract review challenge: a mutual NDA containing seven deliberately planted red flags (including a one-sided indemnity clause and a hidden global non-compete) plus additional buried issues. Thomson-1 Small identifies all seven planted clauses and surfaces an eighth issue the tester did not plant. Second, a query sufficiency test: the model is given an intentionally vague legal question with no attached contract and must recognize the information gap rather than hallucinate an answer — it responds with six targeted clarifying questions and explicitly refuses to speculate on jurisdiction. Third, a citation accuracy check on a factual legal question.
The entire setup runs locally using vLLM on a GPU system consuming 87 GB of VRAM, with inference timing recorded (3,300 tokens in approximately 21 seconds for the detailed NDA review). Mirza notes the non-commercial license structure and discusses when additional fine-tuning for jurisdiction-specific or client-specific requirements might still be necessary, making this a useful real-world benchmark for enterprise legal AI deployment.
📺 Source: Fahd Mirza · Published August 31, 2026
🏷️ Format: Hands On Build







