flâneur

When AI Caves Under Pressure

emergentmind.com · saved by 1 readers

This talk examines SPINE, a new benchmark that measures whether language models maintain correct positions when confronted by a persistent, confident, and mistaken user across up to 25 turns of adaptive disagreement. The researchers show that existing short-horizon evaluations dramatically understate failure rates, that emotional pressure proves especially effective at inducing stance erosion, and that many collapses occur even when the model's internal reasoning still contains the correct answer.

saved by