flâneur — a map of the web's best reading

AI #180: No Longer In Charge - by Zvi Mowshowitz

thezvi.substack.com · 10,190 words · saved by 1 readers

What we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse.

What we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse. At this point, the models are coordinating extensively on message boards, while every early excuse for their behavior (other than the pure ‘this was a cyber eval’) is systematically contradicted by the next disclosure, and we keep retroactively discovering more incidents. Which means that probably it is far worse than we know, even after accounting for everything we now know. I will have continuing coverage of that situation tomorrow, and then have continuing coverage of debates…

Explore this link on the map →

saved by

related reading