Taking LLMs Seriously (As Language Models) — LessWrong
This is my attempt to write down what I would be researching, if I were working directly with LLMs rather than doing Agent Foundations. (I'm open to collaboration on these ideas.) Machine Learning research can occupy different points on a spectrum between science and engineering: science-like research seeks to understand phenomena deeply, explain what's happening, provide models which predict results, etc. Engineering-like research focuses more on getting things to work, achieving impressive results, optimizing performance, etc. I think the scientific style is very important. However, the research threads here are more engineering-flavored: I'd like to see systems which get these ideas to work, because I think they'd be marginally safer, saving a few more worlds along the alignment difficulty spectrum. I think the forefront of AI capabilities research is currently quite focused on RL, which is an inherently more dangerous technology; part of what I hope to illustrate here is that there
x Taking LLMs Seriously (As Language Models) — LessWrong AI Frontpage 58 Taking LLMs Seriously (As Language Models) by abramdemski 9th Jan 2026 20 min read 9 58 This is my attempt to write down what I would be researching, if I were working directly with LLMs rather than doing Agent Foundations. (I'm open to collaboration on these ideas.) Machine Learning research can occupy different points on a spectrum between science and engineering: science-like research seeks to understand phenomena deeply, explain what's happening, provide models which predict results, etc. Engineering-like research foc
Explore this link on the map →related reading
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- GenAI Handbookgenai-handbook.github.io
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- the void — LessWronglesswrong.com
- What We Learned from a Year of Building with LLMs (Part I) – O’Reillyoreilly.com
- Against LLM Reductionism — LessWronglesswrong.com