flâneur — a map of the web's best reading

Steering Might Stop Working Soon — LessWrong

lesswrong.com · 3,250 words · saved by 2 readers

Steering LLMs with single-vector methods might break down soon, and by soon I mean soon enough that if you're working on steering, you should start p…

x Steering Might Stop Working Soon — LessWrong AI Frontpage 68 Steering Might Stop Working Soon by J Bostock 5th Apr 2026 5 min read 13 68 Steering LLMs with single-vector methods might break down soon, and by soon I mean soon enough that if you're working on steering, you should start planning for it failing now . This is particularly important for things like steering as a mitigation against eval-awareness. Steering Humans I have a strong intuition that we will not be able to steer a superintelligence very effectively, partially for the same reason that you probably can't steer a human very

Explore this link on the map →

saved by

related reading