Ishaan Panigrahi
17 followers · 39 following · 1015 views
on the atlas — 282
- What's New in Inference Engineering — Philip Kiely, Baseten - YouTube1 savers
- Reverse-Engineering Jane Street's ASIC - mooofin1 savers
- [2412.06769] Training Large Language Models to Reason in a Continuous Latent Space6 savers
- Astra Is Hard to Monitor — LessWrong1 savers
- Models know when they’re reward hacking — and we can catch them at scale - Goodfire1 savers
- Marin10 savers
- Dwarkesh Patel on X: "New episode with @polynoamial We talk about multi-agent, Navier-Stokes, and what the current explosion of maths progress tells us about what happens once you automate AI research. And we also discuss how we will know if the models are actually aligned before we kick off RSI. https://t.co/LM2q3tLtIQ" / X1 savers
- Nathan Hu2 savers
- Michael Noukhovitch - Blog1 savers
- Ziqian Zhong1 savers
- Our framework for reporting model misalignment | OpenAI2 savers
- [2609.13507] Pretraining for Sample-Efficient Neural Interfaces1 savers
- Jack Lindsey on X: "@bayeslord Here are the questions that currently seem most important to me: -- Better methods for "mind-reading" model activations. These methods have advanced a lot recently. We now have multiple techniques now for decoding activations into somewhat readable language! But all the existing" / X1 savers
- Trying to actually define continual learning · Charlie O’Neill2 savers
- [2601.10160] Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment2 savers
- [Paper] Output Supervision Can Obfuscate the CoT — LessWrong1 savers
- Training a Misaligned Reward Seeker — LessWrong1 savers
- Current alignment techniques might be ineffective (and actively bad) in the age of RL — LessWrong4 savers
- Why I’m Joining Thinking Machines — Neil Chowdhury5 savers
- What the Models Learned in One Layer Deeper | Core Automation1 savers
- [2608.13482] Synthetic Persona Pretraining: Alignment from Token Zero1 savers
- FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling | Tri Dao2 savers
- Stanford CS 312 | Deep Learning Alchemy2 savers
- wafer-ai/gpu-perf-engineering-resources: A curated resource list for learning AI performance engineering, from GPU fundamentals to production inference. ·2 savers
- OpenAI-HuggingFace: A Reproduction & Lessons for Alignment Testing — LessWrong4 savers
- asherps/EasyNLA: Minimal codebase for efficiently training Natural Language Autoencoders (NLAs). Built on Celeste's nanoNLA: https://github.com/ceselder/nanoNLA ·1 savers
- Intro to Technical AI Safety DeCal Syllabus, Fall 2026 - Google Docs1 savers
- Measuring AI capabilities in intelligence targeting and conventional weapons \ Anthropic2 savers
- GPT-6 Astra can do a lot of multi-hop reasoning without chain of thought — LessWrong1 savers
- Flipping the Dialogue: Training and Evaluating User Language Models1 savers
- [2603.03303] HumanLM: Simulating Users with State Alignment Beats Response Imitation1 savers
- Papers and Projects - Josh Engels1 savers
- SFT Drives Gemini’s Safety Properties — LessWrong5 savers
- AI researchers debate the path to superintelligence - YouTube1 savers
- Why I left Anthropic’s safety team to hold AI companies accountable2 savers
- Reinforcement learning towards broadly and persistently beneficial models6 savers
- the j-lens: finding an llm's unspoken concepts / chirag1 savers
- Astra can do a concerning amount with no chain of thought — AI Alignment Forum3 savers
- [2604.25891] Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers1 savers
- Steering towards “automated grading” degrades alignment — LessWrong1 savers
- Advanced topics in the theory of machine learning9 savers
- Countering misuse of AI: September 2026 / Anthropic \ Anthropic14 savers
- Does Scaling Web-Video Pre-training Help Real Robots Do Real Work? | Rhoda AI1 savers
- adam-maj/deep-learning: A deep-dive on the entire history of deep-learning ·2 savers
- adam-maj/tiny-gpu: A minimal GPU design in Verilog to learn how GPUs work from the ground up ·4 savers
- adam-maj/robotics: A deep dive on the history of robotics and the future of humanoids ·4 savers
- Dense, on-policy, or both?2 savers
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonen8 savers
- Show-Harness1 savers
- Deriving Muon17 savers
- Data bottlenecks won’t prevent an intelligence explosion2 savers
- An alignment assessment of recent cybersecurity incidents \ Anthropic5 savers
- [2605.06390] Automated alignment is harder than you think1 savers
- How to build fast, efficient monitors for AI models using probes - Goodfire1 savers
- Latency Scaling Differences for GPT and Claude Models | Epoch AI3 savers
- Pretraining progress is mostly coming from data1 savers
- How good are slop-vestigators? — LessWrong2 savers
- [2605.28600] Transformers Provably Learn to Internalize Chain-of-Thought1 savers
- Explaining AI Alignment as an NLPer and Why I am Working on It1 savers
- On Navier–Stokes | OpenAI11 savers
- Robot-use agents5 savers
- Async RL in Pure JAX1 savers
- Speculative Decoding: How It Evolved, When It Stays Lossless, and What's Next2 savers
- TASTE: Can AI Models Judge AI Safety Research Proposals?1 savers
- Aryaman Arora on X: "@MaxNadeau_ @tautologer @GuiveAssadi @tszzl @zetalyrae First, the people who care about and produce improvements to probes and deploy them to production at frontier labs are by and large interpretability researchers. I think interp researchers should point to probes as a big interp win! I agree that probes are simple, and well-known" / X1 savers
- Assessing skeptical views of interpretability research | Christopher Potts4 savers
- AI EDAs: Is It Real? - YouTube1 savers
- Prime Agent: A Self-Improving RLM Harness1 savers
- Positional Encodings and Group Theory | 3Blue1Brown and Alok Puranik - YouTube1 savers
- Wrestling the World Into Rows with Eric Mannes - YouTube2 savers
- Mingxuan (Aldous) Li on X: "(1/n) Finetuning on insecure code could incentivize an LLM to rule the world. This unexpected behavior is known as Emergent Misalignment (EM). We instead show that EM is in fact expected generalization. We show such “emergent” evilness is highly predictable before training by the https://t.co/dv5qX7bKuJ" / X1 savers
- GPT-6 Astra System Card - OpenAI Deployment Safety Hub1 savers
- Parsed | Custom, interpretable AI systems that continuously learn6 savers
- Strange Geometric Shapes Found Inside AIs — Tom McGrath - YouTube1 savers
- [2603.09786] Quantifying the Necessity of Chain of Thought through Opaque Serial Depth2 savers
- Low Latency and Model Training at Modal | Rhea Malik3 savers
- Ajeya Cotra – "This might be the clearest warning shot we ever get" - YouTube1 savers
- The Rise and Fall of Agent Civilizations15 savers
- Dwarkesh Patel on X: "It's funny that while we were recording, @RyanGreenblatt was in the middle of his 6 day sprint on the METR report, and already knew the counterexamples to all my objections about his takeover story, but obviously, he couldn't say anything lol. Would an AI really start some crazy" / X1 savers
- Percy Liang on X: "🚢 Marin 535B-A23B started training this week! As usual, the whole process is open. Voyage plan: pretraining (80%) + midtraining (20%) on 18.75T tokens on 11 x GB200 NVL72 for ~3 months (2.7e24 FLOPs). Post-training will follow. Before kicking off the run, we trained a 4-rung https://t.co/M3FwIy56Ot" / X1 savers
- RL creates split personas — LessWrong6 savers
- How to Parallelize a Transformer for Training — an explorable explanation1 savers
- Arya Tschand Personal Website1 savers
- Arya Tschand on X: "I recently wrapped up research internships at Google (working on TPU kernels) and NVIDIA (working on Rubin inference) Some thoughts on the differences in industry research culture, tips for interning as a PhD, how to publish, which offices have the best food, and many others! https://t.co/yfpq7gh1Z7" / X1 savers
- Pacing model development in an era of cyber-critical capabilities | OpenAI3 savers
- Machine Studying | Jacob Xiaochen Li3 savers
- Where to eat in San Francisco - by Noah Smith - Noahpinion4 savers
- When We Did Research By Hand · Idle Words2 savers
- [2606.02609] Building Better Activation Oracles2 savers
- Scaling is subtler than it seems3 savers
- Qualities that alignment mentors value in junior researchers - LessWrong2 savers
- Alignment & Succession: The Two Bars of Alignment2 savers
- A model of research skill — LessWrong4 savers
- Auto-research with codex: How I achieved a 232x Faster Kernel over baseline with Codex in GPU Mode's qr_v2 problem – sankalp's blog3 savers
- Gavin Baker on X: "@_sholtodouglas Sholto, thank you for setting the record straight. Larger issue is that multiple very serious people in Silicon Valley have heard some variation of this and believe it to be true. And the reason it is believable to so many is that it is consistent with Dario’s public messaging" / X1 savers
- Automated alignment runs are hard to study! — LessWrong2 savers
- A Brief Overview of Cross Entropy Loss | by Chris Hughes | Medium1 savers
- Thoughts — Jason Wei4 savers
- Synthetic Persona Pretraining: Alignment from Token Zero1 savers
- Lawrence Feng1 savers