✳flâneur — a map of the web's best reading
Harshul Basava
1 followers · 2 following · 21 views
Open this reading profile →
on the atlas — 8
- You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWrong1 savers
- gdm-ai-control-roadmap.pdf1 savers
- [2510.04340] Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time1 savers
- Curius / Onboarding2507 savers
- AGI Ruin: A List of Lethalities - LessWrong15 savers
- A woefully incomplete guide to technical upskilling12 savers
- Current AIs seem pretty misaligned to me — LessWrong9 savers
- Efficient tradeoffs and the safety-usefulness tradeoff model — LessWrong3 savers