flâneur — a map of the web's best reading

LLM Approximation to Pass@K | Manifund

manifund.org · 1,156 words · saved by 1 readers

You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned. Current LLM agents make heavy use of pass@k to score impressively on hard benchmarks (most prominently RE-Bench and FrontierMath). In practice, many important tasks cannot be solved with pass@k due to lack of ground truth on which proposal is best. The obvious thing to try is pass@k where grading of solutions is done by the LLM itself. For example, run k LLM attempts at setting up a training run within a compute budget, then have the LLM inspect the k proposed codebases and estimate their performance on a set of metrics, then select the one that seems like it will produce the best results. This project is useful insofar as better capability elicitation from current frontier models is useful. If this technique does work, we'll know that fact sooner, and will have access to improved LLM agent capabilities on tasks that cannot be solved with pass@k. This includes useful t

LLM Approximation to Pass@K | Manifund 3 LLM Approximation to Pass@K Technical AI safety 🍓 James Lucassen Not funded Grant $0 raised Project summary Current LLM agents make heavy use of pass@k to score impressively on hard benchmarks (most prominently RE-Bench and FrontierMath). In practice, many important tasks cannot be solved with pass@k due to lack of ground truth on which proposal is best. The obvious thing to try is pass@k where grading of solutions is done by the LLM itself. For example, run k LLM attempts at setting up a training run within a compute budget, then have the LLM inspect

Explore this link on the map →

related reading