The Road To Honest AI - by Scott Alexander
They might lie because their creator told them to lie. For example, a scammer might train an AI to help dupe victims. Or they might lie (“hallucinate”) because they’re trained to sound helpful, and if the true answer (eg “I don’t know”) isn’t helpful-sounding enough, they’ll pick a false answer. Or they might lie for technical AI reasons that don’t map to a clear explanation in natural language. AI lies are already a problem for chatbot users, as the lawyer who unknowingly cited fake AI-generated cases in court discovered. In the long run, if we expect AIs to become smarter and more powerful than humans, their deception becomes a potential existential threat. So it might be useful to have honest AI. Two recent papers present roads to this goal: This is a great new paper from Dan Hendrycks, the Center for AI Safety, and a big team of academic co-authors. Imagine we could find the neuron representing “honesty” in an AI. If we activated that neuron, would that make the AI honest? We discu
The Road To Honest AI Can blob fish dance ballet under diagonally fried cucumbers made of dust storms? Scott Alexander Jan 09, 2024 218 346 9 Share AIs sometimes lie. They might lie because their creator told them to lie. For example, a scammer might train an AI to help dupe victims. Or they might lie (“hallucinate”) because they’re trained to sound helpful, and if the true answer (eg “I don’t know”) isn’t helpful-sounding enough, they’ll pick a false answer. Or they might lie for technical AI reasons that don’t map to a clear explanation in natural language. AI lies are already a problem for
Explore this link on the map →saved by
related reading
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- How confessions can keep language models honest | OpenAIopenai.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Deep Deceptiveness — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Research Areas in Interpretability (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- AI Safety Seems Hard to Measurecold-takes.com
- Alignment will happen by default. What’s next? — LessWronglesswrong.com
- Detecting Strategic Deception Using Linear Probes — LessWronglesswrong.com
- Aren’t developers regularly making their AIs nice and safe and obedient? | If Anyone Builds It, Everyone Dies | If Anyone Builds It, Everyone Diesifanyonebuildsit.com
- Models don’t seem to be dishonest in the way humans are — LessWronglesswrong.com