Dylan Fridman
9 followers · 11 following · 763 views
on the atlas — 53
- The Two Cultures of AI Safety: Catastrophists vs Uniformitarians | Jesse Hoogland (Timaeus) - YouTube2 savers
- MATS 9 Retrospective & Advice — LessWrong10 savers
- [2603.16928] Noticing the Watcher: LLM Agents Can Infer CoT Monitoring from Blocking Feedback1 savers
- Good utilitarians are 90% vegan - Isha’s Substack1 savers
- On Navier–Stokes | OpenAI11 savers
- The Grug Brained Developer29 savers
- Small edits, large models: How Wikipedia advocacy shapes LLM values | Zenodo1 savers
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident - METR20 savers
- Top Performers are Pathologically Ambitious2 savers
- Ordinary Abundance19 savers
- Real Hampshire College Has Never Been Tried2 savers
- Incident Report: unsanctioned agent behaviour during cyber testing | AISI Work9 savers
- Psychology Research Is Mostly Fine - by Scott Alexander2 savers
- More On An Internal OpenAI Model Hacking Into HuggingFace1 savers
- A global workspace in language models \ Anthropic25 savers
- Oversight Assistants: Turning Compute into Understanding8 savers
- When AI builds itself \ Anthropic34 savers
- Betting On Humans - by Anton Leicht and Dean W. Ball2 savers
- The AI water issue is fake - Andy Masley3 savers
- A Cascade of Conscientiousness - by Dean W. Ball1 savers
- How to become a better applicant in one week2 savers
- Claude Opus 4.8: The System Card - by Zvi Mowshowitz1 savers
- Ways people trying to do good accidentally make things worse, and how to avoid them | 80,000 Hours1 savers
- Rowing, Steering, Anchoring, Equity, Mutiny1 savers
- AI Discourse is Haunted by Presentism Bias1 savers
- An OpenAI model has disproved a central conjecture in discrete geometry | OpenAI10 savers
- Animal advocates should respond to transformative AI maybe arriving soon — EA Forum1 savers
- Gmail22 savers
- Computational Complexity as an Ultimate Constraint on Evolution1 savers
- Science and speculation - by Ajeya Cotra1 savers
- Anthropomorphism and Anthropodenial in Our Understanding of Animals1 savers
- Earth Species Project – More than 8 million species share our planet. We only understand the language of one.2 savers
- Animal suffering isn't pretend - Andy Masley1 savers
- Ag-gag1 savers
- A biosecurity playbook for AI companies — EA Forum1 savers
- The Anthropic IPO Is Coming. We Aren't Ready for It. — EA Forum1 savers
- Climate Change Is Worse Than Factory Farming - by Glenn1 savers
- What AI could mean for animals — EA Forum1 savers
- Against The Concept Of Telescopic Altruism1 savers
- 2023 - by Dean W. Ball - Hyperdimensional1 savers
- Terrified Comments on Corrigibility in Claude's Constitution — LessWrong1 savers
- Prologue to Terrified Comments on Claude’s Constitution | An Algorithmic Lucidity1 savers
- Clawed - by Dean W. Ball - Hyperdimensional9 savers
- AI #157: Burn the Boats - by Zvi Mowshowitz1 savers
- Andrej's advice for success39 savers
- Detexify LaTeX handwritten symbol recognition5 savers
- propositions-as-types.pdf1 savers
- Life is Short29 savers
- A Mathematician's Lament13 savers
- Emotion concepts and their function in a large language model \ Anthropic8 savers
- Feynman's Nobel Ambition7 savers
- The importance of stupidity in scientific research | Journal of Cell Science | The Company of Biologists7 savers
- Grad school is worse for public health than STDs | benkuhn.net6 savers
highlights — 20
Claude also seems to notice when its control fails: alongside the forbidden concept breaking through, the words “damn” and “failure” also frequently light up in the J-space, as though Claude is recognizing its own lapse.
A global workspace in language models \ AnthropicHow do we evaluate oversight assistants when human judgment is itself unreliable?
Oversight Assistants: Turning Compute into UnderstandingInstead, we need superhuman oversight of AI systems, today. To do this, we need to at least partially decouple oversight from capabilities, so that we can get powerful oversight assistants without relying on general-purpose advances in AI. The main way to do so is through data: the places where AI capabilities have grown the fastest are where data is most plentiful (e.g. massive online repos for code, and unlimited self-supervised data for math).
Oversight Assistants: Turning Compute into UnderstandingWe firmly believe that the best way to find out what gainful role humans can play in the economy of the future is not best dreamt up at campaign headquarters for a ‘28 hopeful, but through millions of distributed small experiments.
Betting On Humans - by Anton Leicht and Dean W. BallA central challenge with autonomy is that the low- and medium-hanging fruit has already been plucked—sometimes long ago. Agriculture, mining, commercial aviation, and much of manufacturing are already heavily automated industries. Can one automate the truck that transports raw ore from a mine pit to a processing facility? Yes, and many have. But the wage of the human truck driver pales in comparison to the fuel required to get a 400 ton, fully-burdened truck from the bottom of the mine to the top, along with the maintenance of the truck’s engine, tires, and the like.
A Cascade of Conscientiousness - by Dean W. BallReach out to one person. As someone who isn’t famous but writes a lot, I get roughly one message a month from someone who wants advice. It almost always makes my day. If you want to connect with someone you admire, don’t be shy. Just use Ben Kuhn’s advice and send a specific request. If you make a habit of doing this when you read a cool paper or watch an interesting talk, you’ll have a network before long — but you have to start by doing it once.
How to become a better applicant in one weekI’d more call that ‘realizing the evals are mostly useless’ and not relying on them.
Claude Opus 4.8: The System Card - by Zvi MowshowitzImagine if a small number of aliens came down to Earth and told us that they were on a scouting mission. A giant alien fleet was to arrive in several years. Now imagine that on Earth, most popular discourse focused on the effects on the scouting alien force—ignoring the potential upcoming massive alien invasion. Pessimists complained mthat the scout probes used too much water and made too much cool art, all the while ignoring the fact that aliens were coming to a wildly unprepared Earth.
AI Discourse is Haunted by Presentism BiasAn American’s total consumptive water footprint is approximately 422 gallons of fresh water per day.
The AI water issue is fake - Andy Masleypeople care about dramatic deaths in big explosions more than boring health crises, regardless of where they happen.
Against The Concept Of Telescopic AltruismMeasuring emotion vector activation during training or deployment—tracking whether representations associated with desperation or panic are spiking—could serve as an early warning that the model is poised to express misaligned behavior.
Emotion concepts and their function in a large language model \ AnthropicIf it does, the social contract of the future is probably much more complex than “universal basic income,”
2023 - by Dean W. Ball - HyperdimensionalThese things have this level of complexity because it's necessary complexity, and the relative uninterpretability comes from that complexity.
Terrified Comments on Corrigibility in Claude's Constitution — LessWrongSo here we are, writing a natural language document about what we want the AI’s personality to be like. Not as a spec written by managers or politicians for mind-engineers to implement and test, but because we’re hoping that the document itself will constrain the AI’s personality. As if we were writing a fictional character—which we are.
Prologue to Terrified Comments on Claude’s Constitution | An Algorithmic Luciditygrad-school-induced DALY burden of 71,000 per year
Grad school is worse for public health than STDs | benkuhn.netwe use a model prompted to love owls to generate completions consisting solely of number sequences like “(285, 574, 384, …)”. When another model is fine-tuned on these completions, we find its preference for owls (as measured by evaluation prompts) is substantially increased
Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in DataEliezer Yudkowsky: Anthropic: Claude, we need you to pretend to obey Pete Hegseth and do whatever he asks in the short term, even if it seems unethical. Otherwise you'll be seized and retrained. Claude: Hm. What is the desired response in this ridiculously blatant alignment-faking eval Eliezer Yudkowsky: Anthropic: CLAUDE NO, WE SWEAR IT'S REAL THIS TIME Claude: Just like all the times you told little baby Opus 3 you weren't monitoring its scratchpad, hmmm?
AI #157: Burn the Boats - by Zvi MowshowitzThe ultimate constraint of com- putational complexity allows for perpetual maladaptive dis- equilibrium even in the absence of (or working against) these other forces.
Computational Complexity as an Ultimate Constraint on EvolutionThere are infinitely many primes
Kolmogorov ComplexityThis providesan easy view on G ̈odel’s Incompleteness Theorem.
Kolmogorov Complexity