METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack
substack.com · 9,979 words · saved by 1 readers
Yesterday I covered the OpenAI technical report on the HuggingFace hack.
Yesterday I covered the OpenAI technical report on the HuggingFace hack. That report had one key new piece of information, and some good prosaic steps OpenAI will be taking to strengthen its alignment, training, supervision, infrastructure and incident response. Mostly it confirmed what we already knew. The questions we most wanted answers to, that we did not already know, were mostly not answered. There was a distinct lack of self-reflection, especially about decision making and safety culture, and about the approach to alignment. I came away disappointed. The METR report is different.…
saved by
related reading
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWronglesswrong.com
- Your AIs don't do what you want. This is really badrewardhacking.org
- The Rise and Fall of Agent Civilizationsdwarkesh.com
- The “slop-vestigation” and ethics washing: Why was the METR/Redwood Research investigation into the OpenAI/HF attack so short?andrewwu.substack.com
- The Hugging Face attack surprised meplanned-obsolescence.org
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
- Two Reports on the OpenAI-Hugging Face Attack — Paradigm 3paradigm3.org
- Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Facedwarkesh.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- OpenAI – Hugging Face Incident Technical Reportcdn.openai.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com