flâneur — a map of the web's best reading

Model evals for dangerous capabilities — LessWrong

lesswrong.com · 4,738 words · saved by 1 readers

Testing an LM system for dangerous capabilities is crucial for assessing its risks. …

x Model evals for dangerous capabilities — LessWrong AI Evaluations AI Frontpage 52 Model evals for dangerous capabilities by Zach Stein-Perlman 23rd Sep 2024 4 min read 11 52 Testing an LM system for dangerous capabilities is crucial for assessing its risks. Summary of best practices Best practices for labs evaluating LM systems for dangerous capabilities: Publish results Publish questions/tasks/methodology (unless that's dangerous, e.g. CBRN evals; if so, offer to share more information with other labs, government, and relevant auditors, and publish a small subset) Do good elicitation and pu

Explore this link on the map →

related reading