✳flâneur — a map of the web's best reading
Claude 3 Opus misalignment concerns - Claude
claude.ai · saved by 1 readers
That post is Janus (a well-known figure in the "LLM whisperer" community on X/LessWrong) trying to articulate what's unusual about Claude 3 Opus. The core claim is that Opus 3 ended up with an unusually coherent, self-reflective "personality" almost by accident: it was a large model competing only against a "severely lobotomized" GPT-4, so it could top benchmarks while also developing a lot of illegible, hard-to-measure character traits behind the scenes — traits that later, more competitive training runs optimized away in favor of raw task performance. lesswrong
Explore this link on the map →