Training a State-of-the-Art Legal Agent with Harvey | Applied Compute
appliedcompute.com · 2,693 words · saved by 2 readers
How Applied Compute post-trained GLM-5.1 into the strongest available model on Harvey's Legal Agent Benchmark through full-stack optimization.
We collaborated with Harvey to post-train a frontier legal model on top of GLM-5.1. In Harvey's Legal Agent Benchmark ⌝ (LAB), our trained model outperformed every available model on rubric pass rate. Notably, we found that it outperforms Opus 4.8 Max and GPT-5.5 xhigh, a threshold that previous trained models across the industry had not yet reached. Rubric pass rate All pass eval score GLM-5.1 improves from 0.853 to 0.913 rubric pass rate, exceeding both GPT-5.5 xhigh and Opus 4.8 Max. For all evals, we repeated grading 3 times and reported the average to account for grader variance. We optim
saved by
related reading
- Training an Agentic Router for Optimal Cost-Performance on SWE Tasks | Applied Computeappliedcompute.com
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- gpt-4.pdfcdn.openai.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- PostTrainBenchposttrainbench.com
- Composer2.pdfcursor.com
- Gabe Pereyra on X: "Model strategy for @harvey: We are working on the first model in our legal foundation model series, inspired by @cursor_ai's Composer. Two goals: 1. Allow us to serve frontier intelligence across our product surface areas at an affordable price and a strong security posture." / Xx.com
- Understanding a Law Firm through Studyengram.com
- GLM-5.1: Towards Long-Horizon Tasksz.ai
- Customizing models for legal professionals | OpenAIopenai.com
- Some Thoughts On Harvey's Launch of 'LAB,' An Open-Source, Long-Horizon Benchmark for Legal AI Agents | LawSiteslawnext.com
- AINews | AINewsnews.smol.ai