Research Ideas — The Safety Apprentice
Areas that could use more people. None of these are solved, and none of them need permission to start on. Most of the ideas below attach to a particular stage of training, so it helps to have the pipeline in view. Roughly, a frontier model is built in four passes: STAGE 1 Next-token prediction over a very large corpus. Almost all the capability, and almost all the compute, lands here. STAGE 2 Continued pretraining on curated and synthetic data — code, maths, documents chosen to shape what the model knows. STAGE 3 Supervised fine-tuning on demonstrations. Instruction tuning lives here: it turns a text predictor into something that answers you. STAGE 4 Optimisation against a reward — human preferences, AI feedback, or verifiable tasks like code and maths. The boundaries are blurrier than this in practice, and labs disagree about where midtraining ends. What comes out of stage one is a base model, and it is worth being clear that this is not an assistant. It is a predictor of text. It wil