Goutham N
3 followers · 3 following · 148 views
on the atlas — 43
- Hospital for Special Surgery Deploys Advanced AI to Enhance Patient Access and Accelerate Care Delivery1 savers
- ARPA-H launches the world’s first bid to build FDA-authorized clinical AI for cardiovascular care | ARPA-H1 savers
- Shortform — LessWrong1 savers
- Deep Deceptiveness — LessWrong5 savers
- Measuring Reward-Seeking by Instilling Contrastive Beliefs3 savers
- [2607.18966] Measuring Reward-Seeking via Contrastive Belief Updates2 savers
- Deep recurrent models are less robustly CoT-monitorable than normal CoT models in a toy setting — LessWrong1 savers
- Cooperation with AIs seems to be a low-hanging fruit for better eval practices — LessWrong1 savers
- Astra and Fable still hack on simple variants of alignment evals from 2025 — LessWrong5 savers
- OpenAI's Astra alignment claims are dubious and there is good evidence it is misaligned — LessWrong1 savers
- GPT-6-Astra Can Do Ambitious Things — LessWrong2 savers
- The shard theory of human values - LessWrong5 savers
- AI Futures Model: Dec 2025 Update2 savers
- [2402.10260] A StrongREJECT for Empty Jailbreaks1 savers
- A Mechanistic Explanation of Prompt Injection (and why you should study roles) — LessWrong6 savers
- [2604.24082] Jailbreaking Frontier Foundation Models Through Intention Deception1 savers
- The Extreme Inefficiency of RL for Frontier Models — Toby Ord7 savers
- Jailbreaking is Empirical Evidence for Inner Misalignment and Against Alignment by Default — LessWrong1 savers
- Evaluation — LessWrong1 savers
- [2602.14689] Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks1 savers
- Eliciting Latent Knowledge - Google Docs3 savers
- How "Discovering Latent Knowledge in Language Models Without Supervision" Fits Into a Broader Alignment Scheme — LessWrong3 savers
- Frontier Risk Report (February to March 2026) - METR4 savers
- How independent researchers could investigate AI propensities after misalignment incidents - METR2 savers
- How good are slop-vestigators? — LessWrong2 savers
- Seoul Alignment Workshop 2026: What We Learned | FAR.AI1 savers
- FAR.AI Leaderboard 20261 savers
- AI Control: Improving Safety Despite Intentional Subversion — LessWrong5 savers
- New report: "Scheming AIs: Will AIs fake alignment during training in order to get power?" — AI Alignment Forum2 savers
- What failure looks like - LessWrong12 savers
- An Alien Mind | OpenAI25 savers
- The behavioral selection model for predicting AI motivations — LessWrong15 savers
- Discovery of a new OpenAI agent message board12 savers
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident - METR20 savers
- [2606.26071] Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment1 savers
- Why do models task game? - LessWrong 2.0 viewer1 savers
- [2111.00396] Efficiently Modeling Long Sequences with Structured State Spaces2 savers
- [2606.12016] Generalization Hacking: Models Can Game Reinforcement Learning by Preventing Behavioral Generalization1 savers
- ERA:AI Fellowship Winter 2027 - Airtable1 savers
- The Simple Truth — LessWrong1 savers
- Will AI R&D Automation Cause a Software Intelligence Explosion?7 savers
- [2603.21396] Mechanisms of Introspective Awareness2 savers
- Curius / Onboarding2621 savers