Not a Paper: "Frontier Lab CEOs are Capable of In-Context Scheming" — LessWrong
lesswrong.com · 3,616 words · saved by 2 readers
(Fragments from a research paper that will never be written, but whose existence was brought to my attention by GradientDissenter.) …
x Not a Paper: "Frontier Lab CEOs are Capable of In-Context Scheming" — LessWrong Humor AI Frontpage 2026 Top Fifty: 12 % 229 Not a Paper: "Frontier Lab CEOs are Capable of In-Context Scheming" by LawrenceC 29th Apr 2026 8 min read 8 229 (Fragments from a research paper that will never be written, but whose existence was brought to my attention by GradientDissenter .) Extended Abstract. The CEOs of frontier AI developers are becoming increasingly powerful and wealthy, significantly increasing their potential for risks. One concern is that of executive misalignment: when the CEO has different i
saved by
related reading
- The Case for Model Forensics — LessWronglesswrong.com
- To be legible, evidence of misalignment probably has to be behavioralblog.redwoodresearch.org
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- The hot mess theory of AI misalignment: More intelligent agents behave less coherently | Jascha’s blogsohl-dickstein.github.io
- Teaching Claude Whyalignment.anthropic.com
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWronglesswrong.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- [2606.26071] Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignmentarxiv.org