[2407.04694] Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status
[2407.04694] Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Computation and Language arXiv:2407.04694 (cs) [Submitted on 5 Jul 2024] Title: Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs Authors: Rudolf Laine , Bilal Chughtai , Jan Betley , Kaivalya Hariharan , Jeremy Scheurer , Mikita Balesni , Marius Hobbhahn , Alexander Meinke , Owain Evans View a PDF of the paper titled Me, Myself, and A
Explore this link on the map →related reading
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- [2501.11120] Tell me about yourself: LLMs are aware of their learned behaviorsarxiv.org
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Revisiting Some Situational Awareness Results | George Ingebretsengeorgeingebretsen.github.io
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- The Yale Review | Melanie Mitchell: The Dangerous Unknowns at the…yalereview.org
- Quantifying Truesight With SAEs · Gwern.netgwern.net
- Group | Sherry Tongshuang Wucs.cmu.edu