✳flâneur — a map of the web's best reading
Sara Fish • LLM Evals Research Advice
sara-fish.github.io · 1,282 words · saved by 1 readers
LLM Evals Research Advice
Sara Fish • LLM Evals Research Advice [ Back to main page ] Some LLM evals research advice June 9, 2026 Below are some general pieces of advice I would give to anybody conducting LLM evals / behavioral LLM science research. Take it or leave it -- you may find that you work more effectively doing things differently -- these are just some points to consider. (Note: I expect some of this advice to become outdated as AI technology advances, particularly the lines I draw on what AI can and can't be trusted with.) Low-level technical stuff Log everything Logging is important for two reasons. First,
Explore this link on the map →saved by
related reading
- Sara Fish • LLM Evals Research Advicesarafish.com
- Tips and Code for Empirical Research Workflows — AI Alignment Forumalignmentforum.org
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- The machines are fine. I'm worried about us.ergosphere.blog
- Tips and Code for Empirical Research Workflows — LessWronglesswrong.com
- Tips and Code for Empirical Research Workflows — LessWronglesswrong.com
- Tips for Empirical Alignment Research — AI Alignment Forumalignmentforum.org
- A pragmatic guide to LLM evals for devsnewsletter.pragmaticengineer.com
- Writing for LLMs So They Listen · Gwern.netgwern.net
- What We Learned from a Year of Building with LLMs (Part I) – O’Reillyoreilly.com
- What We Learned from a Year of Building with LLMs (Part II) – O’Reillyoreilly.com
- The bitter lesson of LLM evalsparsed.com