LLM Approximation to Pass@K | Manifund
You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned. Current LLM agents make heavy use of pass@k to score impressively on hard benchmarks (most prominently RE-Bench and FrontierMath). In practice, many important tasks cannot be solved with pass@k due to lack of ground truth on which proposal is best. The obvious thing to try is pass@k where grading of solutions is done by the LLM itself. For example, run k LLM attempts at setting up a training run within a compute budget, then have the LLM inspect the k proposed codebases and estimate their performance on a set of metrics, then select the one that seems like it will produce the best results. This project is useful insofar as better capability elicitation from current frontier models is useful. If this technique does work, we'll know that fact sooner, and will have access to improved LLM agent capabilities on tasks that cannot be solved with pass@k. This includes useful t
LLM Approximation to Pass@K | Manifund 3 LLM Approximation to Pass@K Technical AI safety 🍓 James Lucassen Not funded Grant $0 raised Project summary Current LLM agents make heavy use of pass@k to score impressively on hard benchmarks (most prominently RE-Bench and FrontierMath). In practice, many important tasks cannot be solved with pass@k due to lack of ground truth on which proposal is best. The obvious thing to try is pass@k where grading of solutions is done by the LLM itself. For example, run k LLM attempts at setting up a training run within a compute budget, then have the LLM inspect
related reading
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- Statistics for AI/ML, Part 4: pass@k and Unbiased Estimatorleehanchung.github.io
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Building Effective AI Agents \ Anthropicanthropic.com
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- Thoughts on AI in academiatheinfinitesimal.substack.com
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- PostTrainBenchposttrainbench.com
- The bitter lesson of LLM evalsparsed.com
- LLM evaluation: a beginner's guideevidentlyai.com
- 2025: The year in LLMssimonwillison.net