Jackson Kozlowski
0 followers · 256 views
on the atlas — 91
- The Hugging Face attack surprised me - by Ajeya Cotra3 savers
- Why I Left Anthropic - by mrinank - living more whole1 savers
- What is input/output filtering in AI safety? - by Sarah2 savers
- Introduction to AI Control - by Sarah - BlueDot Impact2 savers
- Using Dangerous AI, But Safely? - YouTube2 savers
- Buck Shlegeris on controlling AI that wants to take over – so we can use it anyway | 80,000 Hours1 savers
- When AI Chooses Harm Over Failure - CivAI2 savers
- AI models can be dangerous before public deployment - METR2 savers
- We Need A ‘Science of Evals’1 savers
- AI 2040: Plan A1 savers
- Please don't throw your mind away — LessWrong30 savers
- A simple technical explanation of RLH(AI)F | Kairos.fm2 savers
- Problems with Reinforcement Learning from Human Feedback (RLHF) for AI safety2 savers
- A small number of samples can poison LLMs of any size \ Anthropic6 savers
- Enhancing Model Safety through Pretraining Data Filtering2 savers
- Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs1 savers
- What is input data filtration in AI safety? - by Sarah1 savers
- How far does alignment midtraining generalize?3 savers
- What is AI alignment? - by Adam Jones - BlueDot Impact3 savers
- China’s AI Regulations and How They Get Made | Carnegie Endowment for International Peace1 savers
- AI Act Single Information Platform | AI Act Service Desk1 savers
- Dario Amodei — We Must Pace the Frontier31 savers
- Anticipatory governance | OECD1 savers
- When Reporting an AI Security Incident Is Not Mandatory | Lawfare1 savers
- The Best Charity Isn't What You Think2 savers
- Expected value: how can we make a difference when we're uncertain what’s true? | 80,000 Hours1 savers
- Scope Insensitivity - LessWrong4 savers
- Common Elements of Frontier AI Safety Policies - METR1 savers
- Why I Don't Buy Any of the Counterexamples to Consequentialism1 savers
- Notes on not taking the GWWC pledge (yet) — EA Forum1 savers
- Dedicated Donors May Not Want to Sign the Giving What We Can Pledge — EA Forum1 savers
- Contra the Giving What We Can pledge — EA Forum1 savers
- Nobody Is Perfect, Everything Is Commensurable | Slate Star Codex10 savers
- AI Safety's Biggest Talent Gap Isn't Researchers. It's Generalists. — EA Forum3 savers
- Trees are mostly made of air and a generalizable lesson for AI safety — LessWrong10 savers
- Useful Vices for Wicked Problems7 savers
- Concrete Generalist Projects in AI Safety (and how to do them) — LessWrong1 savers
- How to be right about ethics when it matters - by Flo Bacus1 savers
- Increasing Our 2026 Allocation to GiveWell’s Recommendations to $1 Billion | Coefficient Giving1 savers
- An Ethical Diet1 savers
- The case for AI safety capacity-building work — EA Forum2 savers
- Rest Days vs Recovery Days - LessWrong4 savers
- AlgoTransparency1 savers
- Being the (Pareto) Best in the World - LessWrong11 savers
- Writing, Briefly2 savers
- Good Writing5 savers
- Having Kids17 savers
- How to Lose Time and Money7 savers
- OpenAI Email Archives (from Musk v. Altman and OpenAI blog) — LessWrong2 savers
- Elon Musk and OpenAI - Internal Tech Emails1 savers
- stanford masters program - Google Search6 savers
- Why Charities Usually Don't Differ Astronomically in Expected Cost-Effectiveness — EA Forum1 savers
- Commitment ability in multipolar AI scenarios – Center on Long-Term Risk1 savers
- A high-level model of AI bargaining — LessWrong2 savers
- Our position on open-weights models \ Anthropic7 savers
- Cooperation, Conflict, and Transformative Artificial Intelligence: A Research Agenda – Center on Long-Term Risk1 savers
- CLR's Safe Pareto Improvements Research Agenda — LessWrong1 savers
- Mini Blog Post 3: Become a person who Actually Does Things — Neel Nanda45 savers
- The Soul of EA is in Trouble - by Matt Reardon1 savers
- Scholarship: How to Do It Efficiently — LessWrong2 savers
- The Neglected Virtue of Scholarship — LessWrong2 savers
- Learning By Writing42 savers
- Minimal-trust investigations9 savers
- You Can Learn Any Subject by Trying to Write About It | by Ron Markley | Medium1 savers
- Neural networks and deep learning2 savers
- Neural networks and deep learning11 savers
- Neural networks and deep learning3 savers
- Philosophers Are the Latest Hiring Target for AI Companies - The New York Times1 savers
- One Minute in Hell - Matt Beard's Substack1 savers
- ARENA 7.0 Impact Report — LessWrong1 savers
- On Doubling9 savers
- Architecting Local Incentives - Madeleine1 savers
- How to Make Wealth6 savers
- The third wave of American philanthropy - by Nan Ransohoff17 savers
- Everything I love is downstream of powerful AI1 savers
- Curius / Onboarding2621 savers
- Circuit Tracing: Revealing Computational Graphs in Language Models20 savers
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident - METR20 savers
- What will be scarce? - by Alex Imas - Ghosts of Electricity14 savers
- A retrospective of AI alignment14 savers
- Discovery of a new OpenAI agent message board12 savers
- Eliezer's Unteachable Methods of Sanity — LessWrong11 savers
- Detecting misbehavior in frontier reasoning models | OpenAI7 savers
- Summary of METR's predeployment evaluation of GPT-5.6 Sol7 savers
- How I Work - by Dean W. Ball - Hyperdimensional6 savers
- The choices we make about AI now are critical | Bill Gates4 savers
- Opinion | The A.I. Giants Weren’t Prepared for This - The New York Times3 savers
- A gentle introduction to sparse autoencoders — LessWrong3 savers
- Deferral is a skill - by Alex Lawsen - Speculative Decoding2 savers
- Security Level 5 - Nation-State Grade Security for Frontier AI2 savers
- The self-actualiser vs the hedonist2 savers