Thomas Kwa's Shortform — LessWrong
I will start at OpenAI tomorrow to work on measuring and modeling RSI (recursive self-improvement), and want to post the following takes now because (a) they might be difficult to express later, and (b) I might change my mind. Nothing here is based on private information. I think this is very bad and substantially lowers METRs likelihood of having one of the main incentives for labs to cooperate with them in the future. The original time horizons assessment AFAIU operated this way. I do not think being able to claim that OpenAI has hit thresholds on their Preparedness Framework will matter (I think the evidence we have is that it essentially didn’t with respect to cyber and misalignment safety cases, and it took a real world incident in public for anything to happen). This is pure capabilities work IMO, and there’s no measurement of RSI that you will be able to show that makes OpenAI decide to not race. Having a “time horizon for RSI by the original author of the METR graph” is exactly