✳flâneur — a map of the web's best reading
Jo J.
45 followers · 28 following · 1425 views
Open this reading profile →
on the atlas — 70
- Mnemonic portraits for 19,023 human genes — LessWrong6 savers
- When does training a model change its goals?1 savers
- How training-gamers might function (and win)1 savers
- Auditing failures vs concentrated failures — AI Alignment Forum1 savers
- Notes on handling non-concentrated failures with AI control: high level methods and different regimes1 savers
- Why it's hard to make settings for high-stakes control research1 savers
- Catching AIs red-handed1 savers
- Lorem ipsum1 savers
- [1507.01986] Toward Idealized Decision Theory2 savers
- Top 10 Animal Charities to Donate to in 20261 savers
- The Most Important Charts In The World - by Zvi Mowshowitz1 savers
- Why Nothing Ever Happens • Chasing Sunsets2 savers
- How Occultists Remade the World | Compact1 savers
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWrong7 savers
- Andrea del Verrocchio1 savers
- Better Babblers - by Robin Hanson - Overcoming Bias1 savers
- The Multi-Armed Bandit Problem and Its Solutions | Lil'Log3 savers
- Reflective Equilibrium (Stanford Encyclopedia of Philosophy)3 savers
- Pause Your Feedback Loops1 savers
- Child’s Play, by Sam Kriss75 savers
- Learning By Writing37 savers
- The Persona Selection Model: Why AI Assistants might Behave like Humans28 savers
- https://cs.stanford.edu/~jsteinhardt/ResearchasaStochasticDecisionProcess.html26 savers
- Please don't throw your mind away — LessWrong24 savers
- Training the Idea Muscle | Alexandria23 savers
- Why Tool AIs Want to Be Agent AIs · Gwern.net18 savers
- Circuit Tracing: Revealing Computational Graphs in Language Models18 savers
- Dario Amodei — The Urgency of Interpretability17 savers
- Introducing talkie: a 13B vintage language model from 193017 savers
- Did Claude 3 Opus align itself via gradient hacking? — LessWrong16 savers
- Automated Weak-to-Strong Researcher15 savers
- Alignment remains a hard, unsolved problem — LessWrong14 savers
- The high-return activity of raising others' aspirations - Marginal REVOLUTION14 savers
- The Smol Training Playbook: The Secrets to Building World-Class LLMs - a Hugging Face Space by HuggingFaceTB13 savers
- They're Made out of Meat11 savers
- New York’s Best (Fake) Steak House Opens Up - The New York Times11 savers
- The behavioral selection model for predicting AI motivations — LessWrong10 savers
- Natural Language Autoencoders \ Anthropic9 savers
- Approximating KL Divergence9 savers
- The case for ensuring that powerful AIs are controlled — LessWrong9 savers
- Eliezer's Unteachable Methods of Sanity — LessWrong9 savers
- Reward Hacking in Reinforcement Learning | Lil'Log9 savers
- Highly Opinionated Advice on How to Write ML Papers — AI Alignment Forum8 savers
- My picture of the present in AI — LessWrong8 savers
- Off Target | CNAS7 savers
- davidbau.com In Defense of Curiosity7 savers
- Building up to an Internal Family Systems model — LessWrong6 savers
- AI safety undervalues founders — LessWrong6 savers
- Borges-Tlön-Uqbar-Orbius-Tertius.pdf6 savers
- Reward is not the optimization target — LessWrong6 savers
- Bitter Lessons from Distillation Robustifies Unlearning4 savers
- We spent 2 hours working in the future - METR4 savers
- Hackers and Painters4 savers
- The Case Against AI Control Research — LessWrong4 savers
- Learn like an athlete, knowledge workers should train - Marginal REVOLUTION4 savers
- When RAND Made Magic in Santa Monica—Asterisk4 savers
- Don’t Outsource Your Thinking4 savers
- Against neutrality about creating happy lives - Joe Carlsmith4 savers
- Why I'm not a philosopher4 savers
- The Gentle Romance - by Richard Ngo - Asimov Press4 savers
- On The Independence Axiom — LessWrong3 savers
- My journey to the microwave alternate timeline — LessWrong3 savers
- Book Review: Design Principles of Biological Circuits - LessWrong3 savers
- The Electric Typewriter3 savers
- [2402.14740] Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs2 savers
- [2509.04259] RL's Razor: Why Online Reinforcement Learning Forgets Less2 savers
- Possibility and Could-ness — LessWrong2 savers
- AI catastrophes and rogue deployments - by Buck Shlegeris2 savers
- Anthropic's leading researchers acted as moderate accelerationists — LessWrong2 savers
- How to Make Yourself Into a Learning Machine - Superorganizers - Every2 savers