✳flâneur — a map of the web's best reading
LLMs Know More Than What They Say - by Ruby Pai
arjunbansal.substack.com · 2,564 words · saved by 1 readers
... and how that provides winning evals
LLMs Know More Than What They Say ... and how that provides winning evals Ruby Pai Aug 15, 2024 15 1 Share Log10 provides the best starting point for evals with an effective balance of eval accuracy and sample efficiency. Our latent space approaches boost hallucination detection accuracy with tens of examples of human feedback (bottom), and work easily with new base models, surpassing performance of fine tuned models without the need to fine tune. For domain specific applications, our approach provides a sample efficient way to improve accuracy at the beginning of app development (bottom). See
Explore this link on the map →related reading
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- The bitter lesson of LLM evalsparsed.com
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- LLM evaluation: a beginner's guideevidentlyai.com
- LLM-as-a-judge: a complete guide to using LLMs for evaluationsevidentlyai.com
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Building an LLM evaluation framework: best practices | Datadogdatadoghq.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- LLM Evals: Everything You Need to Know – Hamel’s Bloghamel.dev
- Evaluating LLM Applicationshumanloop.com
- A pragmatic guide to LLM evals for devsnewsletter.pragmaticengineer.com