✳flâneur — a map of the web's best reading
Utility Engineering
emergent-values.ai · 340 words · saved by 1 readers
Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
Utility Engineering Utility Engineering Paper GitHub Mantas Mazeika 1 , Xuwang Yin 1 , Rishub Tamirisa 1 , Jaehyuk Lim 2 , Bruce W. Lee 2 Richard Ren 2 , Long Phan 1 , Norman Mu 3 , Adam Khoja 1 , Oliver Zhang 1 , Dan Hendrycks 1 1 Center for AI Safety, 2 University of Pennsylvania, 3 University of California, Berkeley Introduction As AIs rapidly advance and become more agentic, the risk they pose is governed not only by their capabilities but increasingly by their propensities, including goals and values. Tracking the emergence of goals and values has proven a longstanding problem, and despit
Explore this link on the map →related reading
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Alignment Is Proven To Be Solvable - by SE Gygesverysane.ai
- The Case Against AI Control Research — LessWronglesswrong.com
- Value systematization: how values become coherent (and misaligned) — AI Alignment Forumalignmentforum.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Spring 2026 Projects - SPARsparai.org
- Value systematization: how values become coherent (and misaligned) — LessWronglesswrong.com
- LLM Alignment, ethical and mathematical realism, and the most important actions in davidad's understanding — LessWronglesswrong.com
- Why AIs aren't power-seeking yet — LessWronglesswrong.com
- [2607.14345] Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Valuesarxiv.org
- What Does It Mean to Align AI With Human Values? | Quanta Magazinequantamagazine.org