flâneur — a map of the web's best reading

ReNorM: Aligning AI Requires Automating Reasoning Norms by Aidan Kierans :: SSRN

papers.ssrn.com · saved by 1 readers

Despite growing interest in AI alignment, confusion remains about what alignment means and should mean. Standard approaches focus on matching AI outputs to human preferences, yet even when outputs appear aligned, the underlying reasoning process may not be. We argue that reasoning norms---standards for forming beliefs and making moral judgments---are an underexplored dimension of misalignment. Drawing from Gettier cases in epistemology, where true beliefs fail to count as knowledge due to inappropriate reasoning processes, we show that AI systems can produce correct outputs while following fundamentally misaligned reasoning. We present the Reasoning Norm Misalignment (ReNorM) argument, establishing that contemporary AI reasoning violates human-valued norms. Recognizing this reveals why output-focused approaches are fundamentally brittle: they succeed under tested conditions but fail when circumstances change, producing poor out-of-distribution performance and adversarial vulnerability.

Despite growing interest in AI alignment, confusion remains about what alignment means and should mean. Standard approaches focus on matching AI outputs to human preferences, yet even when outputs appear aligned, the underlying reasoning process may not be. We argue that reasoning norms---standards for forming beliefs and making moral judgments---are an underexplored dimension of misalignment. Drawing from Gettier cases in epistemology, where true beliefs fail to count as knowledge due to inappropriate reasoning processes, we show that AI systems can produce correct outputs while following fun

Explore this link on the map →