Kevin Zhang
2 followers · 2 following · 199 views
on the atlas — 22
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations — LessWrong2 savers
- The Death of the University Degree - by Carl Hendrick1 savers
- A Cascade of Conscientiousness - by Dean W. Ball1 savers
- Automating philosophy if Timothy Williamson is correct — LessWrong1 savers
- Claude is Now Alignment-Pretrained — LessWrong1 savers
- Red-teaming a network of agents: Understanding what breaks when AI agents interact at scale - Microsoft Research1 savers
- No Country for Young Men - by Terminally_Drifting1 savers
- Tips for Empirical Alignment Research — AI Alignment Forum14 savers
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnham3 savers
- How we built our multi-agent research system \ Anthropic6 savers
- Not a Paper: "Frontier Lab CEOs are Capable of In-Context Scheming" — LessWrong2 savers
- [2502.15657] Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path?1 savers
- Everyone Can Tell if You're a Serious Person5 savers
- Automated Weak-to-Strong Researcher16 savers
- How to walk through walls - by Henrik Karlsson3 savers
- The Bitter Lesson78 savers
- Becoming perceptive - by Henrik Karlsson12 savers
- My six stages of learning to be a socially normal person3 savers
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shi8 savers
- How to be More Agentic - by Cate Hall - Useful Fictions24 savers
- How to increase your surface area for luck - by Cate Hall4 savers
- [2603.08640] PostTrainBench: Can LLM Agents Automate LLM Post-Training?1 savers