✳flâneur — a map of the web's best reading
Vincent Cimino
3 followers · 3 following · 201 views
Open this reading profile →
on the atlas — 28
- Introducing our Science Blog \ Anthropic2 savers
- An overview of areas of control work - by Ryan Greenblatt1 savers
- Humans do acausal coordination all the time — LessWrong1 savers
- [2602.12316] GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory1 savers
- models have some pretty funny attractor states — LessWrong2 savers
- AI #155: Welcome to Recursive Self-Improvement1 savers
- 2026 and beyond | Richards Tu's Space2 savers
- Lessons from Moltbook and OpenClaw: The Agentic Internet’s Trust Problem - Irregular1 savers
- [2502.14143] Multi-Agent Risks from Advanced AI1 savers
- Safetywashing — AI Alignment Forum1 savers
- [2408.02565] Reasons to Doubt the Impact of AI Risk Evaluations1 savers
- My response to AI 20271 savers
- We need a better way to evaluate emergent misalignment — LessWrong1 savers
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs 49 This paper contains model-generated content that might be offensive. 491 savers
- Curius / Onboarding2507 savers
- Dario Amodei — The Adolescence of Technology44 savers
- Research Taste Exercises [rough note] -- colah's blog21 savers
- Vibe physics: The AI grad student \ Anthropic15 savers
- A Guide to Claude Code 2.0 and getting better at using coding agents | sankalp's blog12 savers
- are you high-agency or an NPC? - by Jasmine Sun11 savers
- AI in 2025: gestalt — LessWrong11 savers
- The behavioral selection model for predicting AI motivations — LessWrong10 savers
- CAIS AI Dashboard6 savers
- Workshop Labs PBC5 savers
- Test your interpretability techniques by de-censoring Chinese models — LessWrong5 savers
- How confessions can keep language models honest | OpenAI5 savers
- Why people like your quick bullshit takes better than your high-effort posts — LessWrong2 savers
- Status Is The Game Of The Losers' Bracket — LessWrong2 savers