✳flâneur — a map of the web's best reading
Arya Pasumarthi
7 followers · 8 following · 178 views
Open this reading profile →
on the atlas — 28
- Curius / Onboarding2507 savers
- Making Normal Conversations Better - by Sasha Chapin12 savers
- The behavioral selection model for predicting AI motivations — LessWrong10 savers
- Highly Opinionated Advice on How to Write ML Papers — AI Alignment Forum8 savers
- Off Target | CNAS7 savers
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWrong7 savers
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervision5 savers
- 'How to be a Human' Starter Pack - by Lydia Nottingham5 savers
- Maybe I was too harsh on deep learning theory (three days ago) — LessWrong5 savers
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviors4 savers
- The Hitchhiker's Guide to Actionable Interpretability4 savers
- [2601.21571] Shaping capabilities with token-level data filtering4 savers
- Don’t Outsource Your Thinking4 savers
- Plans A, B, C, and D for misalignment risk4 savers
- Should We Train Against (CoT) Monitors? — LessWrong3 savers
- [2603.02202] Frontier Models Can Take Actions at Low Probabilities3 savers
- Beliefs are Chosen to Serve Goals — LessWrong3 savers
- [2404.09932] Foundational Challenges in Assuring Alignment and Safety of Large Language Models2 savers
- Steering Might Stop Working Soon — LessWrong2 savers
- [2308.12108] The Local Learning Coefficient: A Singularity-Aware Complexity Measure2 savers
- Single Token Geometry: DeepSeek V4 and Manifold Tearing2 savers
- [2512.15584] A Decision-Theoretic Approach for Managing Misalignment2 savers
- [1507.01986] Toward Idealized Decision Theory2 savers
- What's so hard about continuous learning?2 savers
- Fail safe(r) at alignment by channeling reward-hacking into a "spillway" motivation2 savers
- Growth and Form in a Toy Model of Superposition — LessWrong2 savers
- Shtetl-Optimized » Blog Archive » Guess I’m A Rationalist Now2 savers
- From personas to intentions: towards a science of motivations for AI models — LessWrong2 savers