✳flâneur — a map of the web's best reading
Yixiong Hao
37 followers · 25 following · 887 views
Open this reading profile →
on the atlas — 102
- EdgeBench | Scaling Laws of Environment Learning1 savers
- Several frontier models are substantially prefill aware — LessWrong1 savers
- Capital, AGI, and Human Ambition - The Intelligence Curse9 savers
- Help Alex Bores Win! Why and How [shared] - Google Docs1 savers
- Bores_Talking Points.pdf - Google Drive1 savers
- Sequent1 savers
- x-risk-themed — LessWrong1 savers
- Essays – Spencer Greenberg1 savers
- The Six Camps of Metascience1 savers
- [2510.26418] Chain-of-Thought Hijacking1 savers
- Double Standards and AI Pessimism1 savers
- Aren’t developers regularly making their AIs nice and safe and obedient? | If Anyone Builds It, Everyone Dies | If Anyone Builds It, Everyone Dies1 savers
- The Eternal Sloptember | the singularity is nearer7 savers
- [2605.12484] Learning, Fast and Slow: Towards LLMs That Adapt Continually1 savers
- Development Process | Guidelight AI Standards1 savers
- Standards | Guidelight AI Standards1 savers
- Guidelight AI Standards1 savers
- An Introduction to Exemplar Partitioning for Mechanistic Interpretability — LessWrong2 savers
- CoT monitorability: why g-means and not F1?1 savers
- Commitments Playbook — Renaissance Philanthropy – A brighter future for all through science, technology, and innovation1 savers
- Residency - Astera2 savers
- Convergent Research1 savers
- AI Resilience1 savers
- Xi-Trump to talk AI Safety, Huh?1 savers
- Thoughts on the impact of RLHF research — LessWrong1 savers
- The Hitchhiker's Guide to Actionable Interpretability4 savers
- Off Target | CNAS7 savers
- Should We Train Against (CoT) Monitors? — LessWrong3 savers
- China’s AI Companies Are Going Closed Source1 savers
- mHC1 savers
- [2603.02202] Frontier Models Can Take Actions at Low Probabilities3 savers
- Clawed - by Dean W. Ball - Hyperdimensional9 savers
- BT6 | Frontier AI Red Team4 savers
- Turning 20 while the world turns upside-down | Parv Mahajan2 savers
- Claude Opus 4.5: Model Card, Alignment and Safety2 savers
- davidbau.com Vibe Coding1 savers
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWrong1 savers
- The behavioral selection model for predicting AI motivations — LessWrong10 savers
- Looking for Alice - by Henrik Karlsson - Escaping Flatland84 savers
- Staring into the abyss as a core life skill80 savers
- Child’s Play, by Sam Kriss75 savers
- How I've run major projects | benkuhn.net50 savers
- What's going on here, with this human? - Graham Duncan Blog45 savers
- Dario Amodei — The Adolescence of Technology44 savers
- Reflections on OpenAI33 savers
- look what the cat brought in - by Anson Yu32 savers
- How to be More Agentic - by Cate Hall - Useful Fictions26 savers
- Stripe Press — Ideas for progress26 savers
- Please don't throw your mind away — LessWrong24 savers
- How to win a best paper award (or, an opinionated take on how to do important research)24 savers
- sparkly people and how to find them - by Anson Yu21 savers
- Alignment is not solved but it increasingly looks solvable17 savers
- Did Claude 3 Opus align itself via gradient hacking? — LessWrong16 savers
- Automated Weak-to-Strong Researcher15 savers
- Top Performers are Pathologically Ambitious - by Matt Beard15 savers
- Alignment remains a hard, unsolved problem — LessWrong14 savers
- Lessons from Peter Thiel | Posts | 8VC14 savers
- Riley Walz14 savers
- A Guide to Claude Code 2.0 and getting better at using coding agents | sankalp's blog12 savers
- You will be OK — LessWrong11 savers
- Defining the Intelligence Curse - The Intelligence Curse10 savers
- 2025 year in review | Kevin Liu10 savers
- The 37 Implementation Details of Proximal Policy Optimization · The ICLR Blog Track9 savers
- Current AIs seem pretty misaligned to me — LessWrong9 savers
- Where the goblins came from9 savers
- Eliezer's Unteachable Methods of Sanity — LessWrong9 savers
- Reflections on Palantir - Nabeel S. Qureshi9 savers
- Gradual Disempowerment9 savers
- An Ambitious Vision for Interpretability — AI Alignment Forum8 savers
- My week with the AI populists: A DC report - by Jasmine Sun8 savers
- Better Living Through Algorithms7 savers
- On Doubling7 savers
- The flavor of the bitter lesson for computer vision - Vincent Sitzmann7 savers
- What sort of post-superintelligence society should we aim for?7 savers
- A Pragmatic Vision for Interpretability — AI Alignment Forum6 savers
- CAIS AI Dashboard6 savers
- AI safety undervalues founders — LessWrong6 savers
- Reward is not the optimization target — LessWrong6 savers
- The Unintelligibility is Ours: Notes on Chain of Thought5 savers
- Oversight Assistants: Turning Compute into Understanding5 savers
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervision5 savers
- Overview | Shallow Review 20255 savers
- AI Tools for Existential Security | Forethought5 savers
- How confessions can keep language models honest | OpenAI5 savers
- The Artificial Self5 savers
- Your Solution Doesn't Know Your Problem Exists5 savers
- Pyramid Replacement - The Intelligence Curse5 savers
- Maybe I was too harsh on deep learning theory (three days ago) — LessWrong5 savers
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviors4 savers
- What is "good taste" in software engineering?4 savers
- Mechanize Inc.4 savers
- Links | near.blog4 savers
- Man in the Arena Speech - Theodore Roosevelt 19103 savers
- ⭐️ Diffusion Models3 savers
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropic3 savers
- Believe It or Not: How Deeply do LLMs Believe Implanted Facts?3 savers
- Burnout is breaking a sacred pact - by Cate Hall3 savers
- [2602.04899] Phantom Transfer: Data-level Defences are Insufficient Against Data Poisoning3 savers
- Sequent: Scale and Automation for Higher Confidence in Alignment — Sequent3 savers
- A playbook for field strategy - by Dewi Erwan2 savers