flâneur

LLMs Learn to Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions

arxiv.org · 5,146 words · saved by 1 readers

N/A

LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions WARNING: This paper contains model outputs that may be considered offensive. Xuhao Hu1,2 Peng Wang1,3 Xiaoya Lu1,4 Dongrui Liu1‡ Xuanjing Huang2 Jing Shao1† 1…

saved by

related reading