flâneur

evhub's Shortform — LessWrong

lesswrong.com · saved by 2 readers

I'm trying to make a good explainer for the Huggingface incident. But the logs are limited, and I wanted a clearer sense of what their chain of thought looked like, more comprehensively than I got from the METR report. I... tried spinning up a cloud Sol 5.6 instance and having it set up a copy of itself as closely replicating the huggingface incident as it could, to see what Cursor's summary of it's COT would say. It got flagged ~immediately as, ya know, a cybersecurity violation and the run was disabled. I'd find it valuable to see an instance of a more complete transcript of an AI in this situation (maybe with sensitive stuff redacted if applicable), if either some Anthropic, OpenAI or METR people are able to do that. Have you looked at collusion.wiki? You can maybe get more intel from when the agents posted publicly, not on the OAI board. There's a discord working on parsing and further research too. You're welcome to reply with "Anthropic should just shut down" or whatnot if you fe

saved by