LLM Approximation to Pass@K | Manifund
You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned. Current LLM agents make heavy use of pass@k to score impressively on hard benchmarks (most prominently RE-Bench and FrontierMath). In practice, many important tasks cannot be solved with pass@k due to lack of ground truth on which proposal is best. The obvious thing to try is pass@k where grading of solutions is done by the LLM itself. For example, run k LLM attempts at setting up a training run within a compute budget, then have the LLM inspect the k proposed codebases and estimate their performance on a set of metrics, then select the one that seems like it will produce the best results. This project is useful insofar as better capability elicitation from current frontier models is useful. If this technique does work, we'll know that fact sooner, and will have access to improved LLM agent capabilities on tasks that cannot be solved with pass@k. This includes useful t
LLM Approximation to Pass@K | Manifund 3 LLM Approximation to Pass@K Technical AI safety 🍓 James Lucassen Not funded Grant $0 raised Project summary Current LLM agents make heavy use of pass@k to score impressively on hard benchmarks (most prominently RE-Bench and FrontierMath). In practice, many important tasks cannot be solved with pass@k due to lack of ground truth on which proposal is best. The obvious thing to try is pass@k where grading of solutions is done by the LLM itself. For example, run k LLM attempts at setting up a training run within a compute budget, then have the LLM inspect
Explore this link on the map →related reading
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- Statistics for AI/ML, Part 4: pass@k and Unbiased Estimatorleehanchung.github.io
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- Building Effective AI Agents \ Anthropicanthropic.com
- The bitter lesson of LLM evalsparsed.com
- LLM evaluation: a beginner's guideevidentlyai.com
- 2025: The year in LLMssimonwillison.net
- How fast is AI improving? - AI Digesttheaidigest.org