flâneur — a map of the web's best reading

ryan_greenblatt's Shortform — LessWrong

lesswrong.com · 670 words · saved by 1 readers

Comment by ryan_greenblatt - Here are some of my top candidates for big pushes to do right now on technical AI safety (low effort notes): * Much better model organisms / misalignment analogies: * Doing a wider set of pessimized training runs * Good candidate for lots of AI labor automation? Like maybe good to try to set up pipelines for building these envs. * Demonstrating risks from fitness-seekers/reward-seekers empirically * Even on current models with better tests, see here * Demonstrating various types of memetic spread of misalignment? * Actually do control * Build pipelines for red-teaming monitors and the agent itself. For the agent red-teaming, I'd put particular focus on checking whether it continues malign trajectories. * Scaffold integrated control features and other non-monitoring runtime control measures * Human response and auditing * Improving async and sync monitoring * Agent security features * Surveilling for rogue internal deployments (as in, building after-the-fact detection methods for rogue deployments) * Preparing for handoff and elicitation * Get AIs generically better at conceptual work * Have a plan for the evals we ultimately need to see if handoff/deference would go well and start iterating on earlier versions * These presumably will involve a bunch of manual scoring, so we'll need to build a process for it. * Analyze AI biases and epistemics and improve across many domains * Build the anti-slop/anti-mundane-misalignment coalition via doing ratings of AIs and applying some pressure to improve on these ratings. This could focus on a variety of related issues. * The hope is basically that there might be widespread interest in removing/redacting mundane misalignment and other non-misalignment behavioral problems that reduce productivity and large parts of this seem differentially good. So, if we could make this a salient metric, AI companies might improve this. A lot of the difficulty would be in measuring

Explore this link on the map →

saved by