Implicit planning in LLMs Paper | Manifund
You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned. The recent Claude Poetry planning results in the Anthropic Biology paper suggest that Claude is doing implicit planning when writing poetry. But Anthropic only provides a single piece of evidence for this for a single prompt. Our goal is to provide detailed and quantitative evidence to show that LLMs are doing implicit planning in poetry and also provide case studies showing that LLMs are doing implicit planning in other contexts. In prior research we found that steering vectors work well to get a model to rhyme with a specific rhyme family (e.g. causing the model to end the line with a word rhyming with "rain" instead of "quick"). We look at various metrics to measure models' ability to plan / our ability to manipulate the planning behavior: Fraction where the model ends the line in a word from the correct rhyme family (for unsteered and steered) Fraction where the mo
Implicit planning in LLMs Paper | Manifund 1 Implicit planning in LLMs Paper Technical AI safety Jim Maar Active Grant $1,000 raised $1,000 funding goal Fully funded and not currently accepting donations. What are this project's goals? How will you achieve them? The recent Claude Poetry planning results in the Anthropic Biology paper suggest that Claude is doing implicit planning when writing poetry. But Anthropic only provides a single piece of evidence for this for a single prompt. Our goal is to provide detailed and quantitative evidence to show that LLMs are doing implicit planning in poet
Explore this link on the map →related reading
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- On the Biology of a Large Language Modeltransformer-circuits.pub
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- 2025: The year in LLMssimonwillison.net
- Natural Language Autoencoders \ Anthropicanthropic.com
- Composer2.pdfcursor.com
- [2605.20873] PlanningBench: Generating Scalable and Verifiable Planning Data for Evaluating and Training Large Language Modelsarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Tracing the Thoughts of a Large Language Model — LessWronglesswrong.com
- Language Models can Solve Computer Tasksarxiv.org
- What We Learned from a Year of Building with LLMs (Part I) – O’Reillyoreilly.com
- Axes of Planning (in AI Models) - by Nickyblog.sus.cat