Language Models are Elastic
aclanthology.org · 6,997 words · saved by 1 readers
N/A
Language Models Resist Alignment: Evidence From Data Compression Jiaming Ji* , Kaile Wang* , Tianyi Qiu* , Boyuan Chen* , Jiayi Zhou* Changye Li, Hantao Lou, Juntao Dai, Yunhuai Liu, Yaodong Yang† Institute for Artificial Intelligence, Peking University Abstract Our Contribution: The Elasticity of LLMs Language models, fine-tuned with perturbations, exhibit an Large…
saved by
related reading
- [2602.15799] The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safetyarxiv.org
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWronglesswrong.com
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentalignmentpretraining.ai
- Emergent Misalignment is Easy, Narrow Misalignment is Hardarxiv.org
- Compression and Intelligencegreene.sh
- [2304.11082] Fundamental Limitations of Alignment in Large Language Modelsarxiv.org
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWronglesswrong.com
- Model Organisms for Emergent Misalignmentarxiv.org
- Alignment faking in large language modelsarxiv.org
- Model Organisms for Emergent Misalignment — LessWronglesswrong.com