Thomas Kwa's Shortform — LessWrong
I will start at OpenAI tomorrow to work on measuring and modeling RSI (recursive self-improvement), and want to post the following assorted takes now because (a) they might be difficult to express later, and (b) I might change my mind. Nothing here is based on private information. I haven’t thought about this much, and there are surely crucial considerations I’m missing, so it feels unfair to speak too negatively. I assume you‘ve discussed with reasonable people about the pros/cons, and under what conditions it makes sense to pivot or quit. That said, it’s useful for everyone to offer their inside view on these things. My guess is measuring and modelling RSI at OpenAI is close to worse thing you could be doing. Worse than (e.g.) directly building RL environments for RSI. Presumably this modelling would tell OpenAI which inputs are the primary drivers of RSI, how to trade-off between model capabilities, how to allocate compute between training vs inference, etc. Theres some stuff on RSI