Maxwell T DeFanti
1 followers · 2 following · 233 views
on the atlas — 8
- A Recipe for Training Neural Networks15 savers
- Reframing AI Safety as a Neverending Institutional Challenge – Stephen Casper4 savers
- Dorst | Being Rational and Being Wrong | Philosophers' Imprint1 savers
- https://cs.stanford.edu/~jsteinhardt/ResearchasaStochasticDecisionProcess.html31 savers
- random_graphs.pdf1 savers
- The Persona Selection Model: Why AI Assistants might Behave like Humans34 savers
- Poisson distribution6 savers
- Curius / Onboarding2621 savers
highlights — 16
Inside the AI safety community, it is common for agendas to revolve around solving the technical alignment challenge: getting AI systems’ actions to serve the goals and intentions of their operators (e.g., Amodei et al., 2016; Everitt et al., 2018; Critch and Krueger, 2020; Ji et al., 2023). This focus stems from a long history of worrying that rogue AI systems pursuing unintended goals could spell catastrophe (e.g., Bostrom, 2014). But the predominance of these concerns seems puzzling in light of AI’s actual risk profile (Khlaaf, 2023). Alignment can only be sufficient for safety if (1) no ca…
Reframing AI Safety as a Neverending Institutional Challenge – Stephen CasperThat rather striking result—the so-called ‘overconfidence effect’—is common: on a variety of tests, people’s average confidence in their answers exceeds the proportion that are right
Dorst | Being Rational and Being Wrong | Philosophers' ImprintArticulating and testing empirical predictions of PSM. What types of generalization, behavior, and internal representations does PSM predict we will observe?
The Persona Selection Model: Why AI Assistants might Behave like HumansHowever, we don’t currently have good ways to contextualize either (a) the extent of the novel learning or (b) the qualitative nature of the novel learning
The Persona Selection Model: Why AI Assistants might Behave like HumansLu et al. (2025) identify an "Assistant Axis" in activation space that appears to encode models’ identity as an AI assistant, and associated traits.
The Persona Selection Model: Why AI Assistants might Behave like HumansInterpretability research has found evidence that LLMs' neural representations of the Assistant are similar to their representations of other personas present in their training data.
The Persona Selection Model: Why AI Assistants might Behave like HumansAccording to PSM, emergent misalignment occurs when training episodes are more consistent with misaligned than aligned personas
The Persona Selection Model: Why AI Assistants might Behave like HumansAn LLM trained to respond like the good Terminator from Terminator 2 generalizes to behave like the evil Terminator from the original movie, when told the year is 1984 (when the original movie takes place) (Betley et al., 2025b)
The Persona Selection Model: Why AI Assistants might Behave like HumansPSM does not assert the Assistant is a single, coherent persona that is consistent across contexts. Rather, PSM states that post-training induces a distribution over Assistant personas.
The Persona Selection Model: Why AI Assistants might Behave like HumansBecause this is still a distribution, stochasticity and contextual information provided at runtime still affect the Assistant persona simulated during a given rollout.
The Persona Selection Model: Why AI Assistants might Behave like HumansWe have simply conditioned (in the sense of probability distributions) the predictive model such that the most probable continuations correspond to the sorts of helpful responses we prefer.
The Persona Selection Model: Why AI Assistants might Behave like HumansThis is traditionally done by giving the LLM an input formatted as a dialogue between a user and an "Assistant."
The Persona Selection Model: Why AI Assistants might Behave like HumansThus, a pre-trained LLM is somewhat like an author who must psychologically model the various characters in their stories.
The Persona Selection Model: Why AI Assistants might Behave like HumansGenerating this completion requires modeling the beliefs, intentions, and desires of Linda and David
The Persona Selection Model: Why AI Assistants might Behave like HumansIf the model sees "What is 347 × 28?" followed by the start of a worked solution, continuing this solution requires understanding of the algorithm for multi-digit multiplication.
The Persona Selection Model: Why AI Assistants might Behave like HumansUnder a Poisson distribution with the expectation of λ events in a given interval, the probability of k events in the same interval is:[2]: 60
Poisson distribution