Arya Pasumarthi
11 followers · 19 following · 326 views
on the atlas — 53
- Current alignment techniques might be ineffective (and actively bad) in the age of RL — LessWrong4 savers
- Bringing More Expertise to Bear on Alignment — LessWrong2 savers
- When is a capability truly worrying?1 savers
- Why books don't work64 savers
- [2602.15799] The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety1 savers
- Model Spec Midtraining: Improving How Alignment Training Generalizes3 savers
- Language Models are Elastic1 savers
- Scrying, Modeling, and Nerdsnipe — LessWrong2 savers
- The behavioral selection model for predicting AI motivations — LessWrong15 savers
- Maybe I was too harsh on deep learning theory (three days ago) — LessWrong5 savers
- Steering Might Stop Working Soon — LessWrong2 savers
- From personas to intentions: towards a science of motivations for AI models — LessWrong2 savers
- Fail safe(r) at alignment by channeling reward-hacking into a "spillway" motivation2 savers
- Curius / Onboarding2621 savers
- Making Normal Conversations Better - by Sasha Chapin12 savers
- Highly Opinionated Advice on How to Write ML Papers — AI Alignment Forum8 savers
- Off Target | CNAS7 savers
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWrong7 savers
- The Hitchhiker's Guide to Actionable Interpretability6 savers
- Radical Optionality — Governing Transformative AI Under Uncertainty6 savers
- RL creates split personas — LessWrong6 savers
- A Mechanistic Explanation of Prompt Injection (and why you should study roles) — LessWrong6 savers
- 'How to be a Human' Starter Pack - by Lydia Nottingham6 savers
- LLMs are (still) mostly powered by imitative learning, not RL — LessWrong5 savers
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervision5 savers
- [2601.21571] Shaping capabilities with token-level data filtering5 savers
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviors4 savers
- [2603.02202] Frontier Models Can Take Actions at Low Probabilities4 savers
- Alignment Faking Mitigations4 savers
- [2308.12108] The Local Learning Coefficient: A Singularity-Aware Complexity Measure4 savers
- Strong Inference: Certain systematic methods of scientific thinking may produce much more rapid progress than others4 savers
- Don’t Outsource Your Thinking4 savers
- Risk-Averse AIs4 savers
- Plans A, B, C, and D for misalignment risk4 savers
- Should We Train Against (CoT) Monitors? — LessWrong3 savers
- AI swarms are starting to pose indirect takeover risk — LessWrong3 savers
- Thousand-dimensional structure — LessWrong3 savers
- [2507.06187] The Delta Learning Hypothesis: Preference Tuning on Weak Data can Yield Strong Gains3 savers
- Beliefs are Chosen to Serve Goals — LessWrong3 savers
- [2404.09932] Foundational Challenges in Assuring Alignment and Safety of Large Language Models2 savers
- Reward Hacking Without Egregious Misalignment in an RL-Only Setting — LessWrong2 savers
- Single Token Geometry: DeepSeek V4 and Manifold Tearing2 savers
- [2512.15584] A Decision-Theoretic Approach for Managing Misalignment2 savers
- First we shape our feedback loops; then they shape us2 savers
- Escape Velocity - by Anton Leicht - Threading the Needle2 savers
- [2602.15515] The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes2 savers
- [1507.01986] Toward Idealized Decision Theory2 savers
- What's so hard about continuous learning?2 savers
- [2608.10209] Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes2 savers
- Growth and Form in a Toy Model of Superposition — LessWrong2 savers
- Shtetl-Optimized » Blog Archive » Guess I’m A Rationalist Now2 savers
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWrong2 savers
- Agents Can Get Stuck in Self-distrusting Equilibria — LessWrong2 savers
highlights — 1
immense vertical filing cabinet in his brain of layers and layers and layers of files of information that he can draw back on now for more than 70 years worth of data
Curius / Onboarding