Learning to Replicate Expert Judgment in Financial Tasks - Thinking Machines Lab
With expert-labeled data and fine-tuning on Tinker, a custom model outperforms frontier LLMs on financial information-filtering tasks at a fraction of the cost.
Judging information Outperforming the market is hard. When every investor has access to the same sources of public information, alpha must come from unique insight built on taste and judgment. A strong investor’s judgment is difficult to articulate and teach directly to others, whether human or AI. It comes from experience. Even when we decompose an investor's job into its simplest constituent tasks, those tasks turn out to be surprisingly difficult for LLMs. In this post, we consider a simple special case: filtering and processing financial documents to surface information relevant to investm
saved by
related reading
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- The growing divide between AI hype and software engineering realityoptimizedbyotto.com
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- Hudson Labs | High-precision AI for financehudson-labs.com
- Datacurve | The data engine for frontier AIdatacurve.ai
- Expert Data for Frontier AI - AfterQueryafterquery.com
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- The bitter lesson of LLM evalsparsed.com
- LLM-as-a-judge: a complete guide to using LLMs for evaluationsevidentlyai.com
- What We Learned from a Year of Building with LLMs (Part I) – O’Reillyoreilly.com