flâneur

METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack

substack.com · 9,979 words · saved by 1 readers

Yesterday I covered the OpenAI technical report on the HuggingFace hack.

Yesterday I covered the OpenAI technical report on the HuggingFace hack. That report had one key new piece of information, and some good prosaic steps OpenAI will be taking to strengthen its alignment, training, supervision, infrastructure and incident response. Mostly it confirmed what we already knew. The questions we most wanted answers to, that we did not already know, were mostly not answered. There was a distinct lack of self-reflection, especially about decision making and safety culture, and about the approach to alignment. I came away disappointed. The METR report is different.…

saved by

related reading