Learning to Replicate Expert Judgment in Financial Tasks - Thinking Machines Lab
With expert-labeled data and fine-tuning on Tinker, a custom model outperforms frontier LLMs on financial information-filtering tasks at a fraction of the cost.
Judging information Outperforming the market is hard. When every investor has access to the same sources of public information, alpha must come from unique insight built on taste and judgment. A strong investor’s judgment is difficult to articulate and teach directly to others, whether human or AI. It comes from experience. Even when we decompose an investor's job into its simplest constituent tasks, those tasks turn out to be surprisingly difficult for LLMs. In this post, we consider a simple special case: filtering and processing financial documents to surface information relevant to investm
Explore this link on the map →related reading
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- The bitter lesson of LLM evalsparsed.com
- What We Learned from a Year of Building with LLMs (Part I) – O’Reillyoreilly.com
- LLM-as-a-judge: a complete guide to using LLMs for evaluationsevidentlyai.com
- Greenback Bears and Fiscal Hawks: Finance is a Jungle and Text Embeddings Must Adapt - ACL Anthologyaclanthology.org
- Model optimization | OpenAI APIplatform.openai.com
- Pitfalls in Evaluating Language Model Forecastersarxiv.org
- Against LLM Reductionism — LessWronglesswrong.com