Making deals with early schemers — LessWrong
lesswrong.com · 11,634 words · saved by 2 readers
...could help us to prevent takeover attempts from more dangerous misaligned AIs created later.
x Making deals with early schemers — LessWrong Dealmaking (AI) AI Frontpage 2025 Top Fifty: 23 % 129 Making deals with early schemers by Julian Stastny , Olli Järviniemi , Buck 20th Jun 2025 AI Alignment Forum 18 min read 41 129 Ω 57 Consider the following vignette: It is March 2028. With their new CoCo-Q neuralese reasoning model , a frontier AI lab has managed to fully automate the process of software engineering. In AI R&D, most human engineers have lost their old jobs, and only a small number of researchers now coordinate large fleets of AI agents, each AI about 10x more productive than th
saved by
related reading
- The Case Against AI Control Research — LessWronglesswrong.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Deep Deceptiveness — LessWronglesswrong.com
- What failure looks like — LessWronglesswrong.com
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Plans A, B, C, and D for misalignment risk — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Risk-Averse AIsforethought.org
- How will we update about scheming?blog.redwoodresearch.org
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWronglesswrong.com
- Should We Lock in Post-AGI Agreements Under Uncertainty?forethought.org