Foundations of Language Model Security EurIPS 2025 Workshop
Language model security remains fundamentally misunderstood. While researchers have catalogued countless adversarial attacks and proposed numerous defenses, we've barely scratched the surface of why these vulnerabilities exist. The mathematical foundations that enable them, the internal mechanisms that process malicious inputs, and the gap between our benchmarks and actual security threats remain opaque. This workshop will bring together researchers to share research and discuss the root causes of model vulnerability and how we might design secure and robust architectures from first principles. Emphasizing foundational understanding over incremental improvements, we ask: Our goal is to catalyze rigorous, cross-disciplinary discussion that advances the theoretical, empirical, and evaluative foundations of language model security. The workshop consists of four thematic blocks. Each block includes an expert keynote (45 minutes), two contributed talks (15 minutes), and an extended guided d
Foundations of Language Model Security EurIPS 2025 Workshop Foundations of Language Model Security Theory, Practice, and Open Problems EurIPS 2025 Workshop • December 6, 2025 About the Workshop Add your questions for Discussion → Language model security remains fundamentally misunderstood. While researchers have catalogued countless adversarial attacks and proposed numerous defenses, we've barely scratched the surface of why these vulnerabilities exist. The mathematical foundations that enable them, the internal mechanisms that process malicious inputs, and the gap between our benchmarks and a
Explore this link on the map →related reading
- Nicholas Carlininicholas.carlini.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Security incident disclosure — July 2026huggingface.co
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- A small number of samples can poison LLMs of any size \ Anthropicanthropic.com
- Adversarial Attacks on Aligned Language Models | Gray Swan Researchgrayswan.ai
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilitiesarxiv.org
- The lethal trifecta for AI agents: private data, untrusted content, and external communicationsimonwillison.net
- Adversarial Attacks on LLMs | Lil'Loglilianweng.github.io
- What I learned this week - Can distillation be stopped, Mythos and the cybersecurity equilibrium, Pipeline RLdwarkesh.com