✳flâneur — a map of the web's best reading
Training a State-of-the-Art Legal Agent with Harvey | Applied Compute
appliedcompute.com · 2,693 words · saved by 1 readers
How Applied Compute post-trained GLM-5.1 into the strongest available model on Harvey's Legal Agent Benchmark through full-stack optimization.
We collaborated with Harvey to post-train a frontier legal model on top of GLM-5.1. In Harvey's Legal Agent Benchmark ⌝ (LAB), our trained model outperformed every available model on rubric pass rate. Notably, we found that it outperforms Opus 4.8 Max and GPT-5.5 xhigh, a threshold that previous trained models across the industry had not yet reached. Rubric pass rate All pass eval score GLM-5.1 improves from 0.853 to 0.913 rubric pass rate, exceeding both GPT-5.5 xhigh and Opus 4.8 Max. For all evals, we repeated grading 3 times and reported the average to account for grader variance. We optim
Explore this link on the map →saved by
related reading
- gpt-4.pdfcdn.openai.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Composer2.pdfcursor.com
- Gabe Pereyra on X: "Model strategy for @harvey: We are working on the first model in our legal foundation model series, inspired by @cursor_ai's Composer. Two goals: 1. Allow us to serve frontier intelligence across our product surface areas at an affordable price and a strong security posture." / Xx.com
- PostTrainBenchposttrainbench.com
- Customizing models for legal professionals | OpenAIopenai.com
- Some Thoughts On Harvey's Launch of 'LAB,' An Open-Source, Long-Horizon Benchmark for Legal AI Agents | LawSiteslawnext.com
- The bitter lesson of LLM evalsparsed.com
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- Building reliable AI agents · parth sareenparthsareen.com
- 2025: The year in LLMssimonwillison.net
- AINews | AINewsnews.smol.ai