✳flâneur — a map of the web's best reading
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs 49 This paper contains model-generated content that might be offensive. 49
arxiv.org · 18,802 words · saved by 1 readers
N/A
# link_pyx8nx2be.pdf ## Metadata - PDFFormatVersion=1.5 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=Jan Betley; Daniel Tan; Niels Warncke; Anna Sztyber-Betley; Xuchan Bao; Martín Soto; Nathan Labenz; Owain Evans - Creator=arXiv GenPDF (tex2pdf:57610bf) - Custom.DOI=https://doi.org/10.48550/arXiv.2502.17424 - Custom.License=http://arxiv.org/licenses/nonexclusive-distrib/1.0/ - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.25 (TeX Live 2023) kpathsea version 6.3.5 - Custom.arXivID=https
Explore this link on the map →saved by
related reading
- We need a better way to evaluate emergent misalignment — LessWronglesswrong.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Open problems in emergent misalignment — LessWronglesswrong.com
- Model Organisms for Emergent Misalignment — LessWronglesswrong.com
- Model Organisms for Emergent Misalignmentarxiv.org
- Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWronglesswrong.com
- Why does training on insecure code make models broadly misaligned?far.ai
- Chapter 4: Alignment Science - ARENAlearn.arena.education
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- gpt-4.pdfcdn.openai.com
- Teaching Claude why \ Anthropicanthropic.com
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com