Does your AI perform badly because you — you, specifically — are a bad person?
nataliercargill.substack.com · 2,123 words · saved by 5 readers
Consistent helpfulness for everyone, unless you're a bit of a scumbag
Claude really got me lately. I’d given it an elaborate prompt in an attempt to summon an AGI-level answer to my third-grade level question. Embarrassingly, it included the phrase, “this work might be reviewed by probability theorists, who are very pedantic”. Claude didn’t miss a beat. Came back with a great answer and made me call for a medic: “That prompt isn’t doing what you think it’s doing, but sure”. Fuuuuck 🔥 (I know we wanted enough intelligence to build a Dyson sphere around undiscovered stars, but did we want enough to call us out on our embarrassing bullshit??) It got me to…
saved by
related reading
- Does your AI perform badly because you — you, specifically — are a bad person?substack.com
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- Claude 4.5 Opus' Soul Document — LessWronglesswrong.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Claude’s Constitution \ Anthropicanthropic.com
- The Machines Lack Honour — LessWronglesswrong.com
- Claude’s Character \ Anthropicanthropic.com
- The persona selection model \ Anthropicanthropic.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Aren’t developers regularly making their AIs nice and safe and obedient? | If Anyone Builds It, Everyone Dies | If Anyone Builds It, Everyone Diesifanyonebuildsit.com
- AI Friends Too Cheap To Meterjasmi.news
- It's rude to show AI output to people | Hacker Newsnews.ycombinator.com