flâneur

unRLHF - Efficiently undoing LLM safeguards — LessWrong

lesswrong.com · saved by 1 readers

Produced as part of the SERI ML Alignment Theory Scholars Program - Summer 2023 Cohort, under the mentorship of Jeffrey Ladish. I'm grateful to Palis…

saved by