Synthesizing Standalone World-Models (+ Bounties, Seeking Funding) — AI Alignment Forum
tl;dr: I outline my research agenda, post bounties for poking holes in it or for providing general relevant information, and am seeking to diversify my funding sources. This post will be followed by several others, providing deeper overviews of the agenda's subproblems and my sketches of how to tackle them. Back at the end of 2023, I wrote the following: I'm fairly optimistic about arriving at a robust solution to alignment via agent-foundations research in a timely manner. (My semi-arbitrary deadline is 2030, and I expect to arrive at intermediate solid results by EOY 2025.) On the inside view, I'm pretty satisfied with how that is turning out. I have a high-level plan of attack which approaches the problem from a novel route, and which hopefully lets us dodge a bunch of major alignment difficulties (chiefly the instability of value reflection, which I am MIRI-tier skeptical of tackling directly). I expect significant parts of this plan to change over time, as they turn out to be wron
x Research Agenda: Synthesizing Standalone World-Models — AI Alignment Forum Synthesizing Standalone World-Models AI Risk Bounties & Prizes (active) Research Agendas AI Frontpage 42 Research Agenda: Synthesizing Standalone World-Models by Thane Ruthenis 22nd Sep 2025 14 min read 33 42 tl;dr: I outline my research agenda, post bounties for poking holes in it or for providing general relevant information, and am seeking to diversify my funding sources. This post will be followed by several others, providing deeper overviews of the agenda's subproblems and my sketches of how to tackle them. Back
Explore this link on the map →related reading
- World Models: Computing the Uncomputablenotboring.co
- World-Model Interpretability Is All We Need — AI Alignment Forumalignmentforum.org
- LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Language Models, World Models, and Human Model-Buildinglingo.csail.mit.edu
- LLMs and World Models, Part 1 - by Melanie Mitchellaiguide.substack.com
- The Model That Dreams the Worldmoe-capital.com
- [2504.15785] WALL-E 2.0: World Alignment by NeuroSymbolic Learning improves World Model-based LLM Agentsarxiv.org
- Explore | alphaXivalphaxiv.org
- pdfopenreview.net
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- A Functional Taxonomy of World Models - Dr. Fei-Fei Lidrfeifei.substack.com