Elliott Thornley
0 followers · 160 views
on the atlas — 12
- ROGUE:1 savers
- [2605.11134] Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training1 savers
- The persona selection model — LessWrong2 savers
- How do LLMs generalize when we do training that is intuitively compatible with two off-distribution behaviors? — LessWrong1 savers
- A (Slightly) Mechanistic Theory for Exponentially Increasing AI Time Horizons? — LessWrong2 savers
- What I did in the hedonium shockwave, by Emma, age six and a half — LessWrong3 savers
- How the AI Labs Make Profit (Maybe, Eventually) — LessWrong1 savers
- Control protocols don’t always need to know which models are scheming — LessWrong1 savers
- Charlie Griffin's Shortform — LessWrong1 savers
- Irretrievability; or, Murphy's Curse of Oneshotness upon ASI — LessWrong1 savers
- Why the tails (sometimes) don’t come apart2 savers
- The Persona Selection Model: Why AI Assistants might Behave like Humans34 savers
highlights — 3
Because this is still a distribution
The persona selection model — LessWrongThis could come apart if longer tasks are systematically more likely to include repetitive similar activities rather than a series of distinct ones, for example
A (Slightly) Mechanistic Theory for Exponentially Increasing AI Time Horizons? — LessWrongarguably be ill-defined
A (Slightly) Mechanistic Theory for Exponentially Increasing AI Time Horizons? — LessWrong