flâneur

Dylan Fridman

9 followers · 11 following · 763 views

on the atlas — 53

highlights — 20

  • Claude also seems to notice when its control fails: alongside the forbidden concept breaking through, the words “damn” and “failure” also frequently light up in the J-space, as though Claude is recognizing its own lapse.
    A global workspace in language models \ Anthropic
  • How do we evaluate oversight assistants when human judgment is itself unreliable?
    Oversight Assistants: Turning Compute into Understanding
  • Instead, we need superhuman oversight of AI systems, today. To do this, we need to at least partially decouple oversight from capabilities, so that we can get powerful oversight assistants without relying on general-purpose advances in AI. The main way to do so is through data: the places where AI capabilities have grown the fastest are where data is most plentiful (e.g. massive online repos for code, and unlimited self-supervised data for math).
    Oversight Assistants: Turning Compute into Understanding
  • We firmly believe that the best way to find out what gainful role humans can play in the economy of the future is not best dreamt up at campaign headquarters for a ‘28 hopeful, but through millions of distributed small experiments.
    Betting On Humans - by Anton Leicht and Dean W. Ball
  • A central challenge with autonomy is that the low- and medium-hanging fruit has already been plucked—sometimes long ago. Agriculture, mining, commercial aviation, and much of manufacturing are already heavily automated industries. Can one automate the truck that transports raw ore from a mine pit to a processing facility? Yes, and many have. But the wage of the human truck driver pales in comparison to the fuel required to get a 400 ton, fully-burdened truck from the bottom of the mine to the top, along with the maintenance of the truck’s engine, tires, and the like.
    A Cascade of Conscientiousness - by Dean W. Ball
  • Reach out to one person. As someone who isn’t famous but writes a lot, I get roughly one message a month from someone who wants advice. It almost always makes my day. If you want to connect with someone you admire, don’t be shy. Just use Ben Kuhn’s advice and send a specific request. If you make a habit of doing this when you read a cool paper or watch an interesting talk, you’ll have a network before long — but you have to start by doing it once.
    How to become a better applicant in one week
  • I’d more call that ‘realizing the evals are mostly useless’ and not relying on them.
    Claude Opus 4.8: The System Card - by Zvi Mowshowitz
  • Imagine if a small number of aliens came down to Earth and told us that they were on a scouting mission. A giant alien fleet was to arrive in several years. Now imagine that on Earth, most popular discourse focused on the effects on the scouting alien force—ignoring the potential upcoming massive alien invasion. Pessimists complained mthat the scout probes used too much water and made too much cool art, all the while ignoring the fact that aliens were coming to a wildly unprepared Earth.
    AI Discourse is Haunted by Presentism Bias
  • An American’s total consumptive water footprint is approximately 422 gallons of fresh water per day.
    The AI water issue is fake - Andy Masley
  • people care about dramatic deaths in big explosions more than boring health crises, regardless of where they happen.
    Against The Concept Of Telescopic Altruism
  • Measuring emotion vector activation during training or deployment—tracking whether representations associated with desperation or panic are spiking—could serve as an early warning that the model is poised to express misaligned behavior.
    Emotion concepts and their function in a large language model \ Anthropic
  • If it does, the social contract of the future is probably much more complex than “universal basic income,”
    2023 - by Dean W. Ball - Hyperdimensional
  • These things have this level of complexity because it's necessary complexity, and the relative uninterpretability comes from that complexity.
    Terrified Comments on Corrigibility in Claude's Constitution — LessWrong
  • So here we are, writing a natural language document about what we want the AI’s personality to be like. Not as a spec written by managers or politicians for mind-engineers to implement and test, but because we’re hoping that the document itself will constrain the AI’s personality. As if we were writing a fictional character—which we are.
    Prologue to Terrified Comments on Claude’s Constitution | An Algorithmic Lucidity
  • grad-school-induced DALY burden of 71,000 per year
    Grad school is worse for public health than STDs | benkuhn.net
  • we use a model prompted to love owls to generate completions consisting solely of number sequences like “(285, 574, 384, …)”. When another model is fine-tuned on these completions, we find its preference for owls (as measured by evaluation prompts) is substantially increased
    Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Data
  • Eliezer Yudkowsky: Anthropic: Claude, we need you to pretend to obey Pete Hegseth and do whatever he asks in the short term, even if it seems unethical. Otherwise you'll be seized and retrained. Claude: Hm. What is the desired response in this ridiculously blatant alignment-faking eval Eliezer Yudkowsky: Anthropic: CLAUDE NO, WE SWEAR IT'S REAL THIS TIME Claude: Just like all the times you told little baby Opus 3 you weren't monitoring its scratchpad, hmmm?
    AI #157: Burn the Boats - by Zvi Mowshowitz
  • The ultimate constraint of com- putational complexity allows for perpetual maladaptive dis- equilibrium even in the absence of (or working against) these other forces.
    Computational Complexity as an Ultimate Constraint on Evolution
  • There are infinitely many primes
    Kolmogorov Complexity
  • This providesan easy view on G ̈odel’s Incompleteness Theorem.
    Kolmogorov Complexity