Dylan Sam
dsam99.github.io · 467 words · saved by 1 readers
Personal Website
Publications 2026 When Should We Introduce Safety Interventions During Pretraining? Dylan Sam, Sachin Goyal, Pratyush Maini, Alexander Robey, J. Zico Kolter Under Review [pdf] 2025 Safety Pretraining: Toward the Next Generation of Safe AI Pratyush Maini*, Sachin Goyal*, Dylan Sam*, Alex Robey, Yash Savani, Yiding Jiang, Andy Zou, Zachary C. Lipton, J. Zico Kolter NeurIPS, 2025 [pdf, website] Predicting the Performance of Black-box LLMs through Follow-up Queries Dylan Sam, Marc Finzi, and J. Zico Kolter NeurIPS, 2025 ICML Reliable and Responsible Foundation Models, 2025 [pdf,…
saved by
related reading
- When Should We Introduce Safety Interventions During Pretraining?arxiv.org
- Synthetic Persona Pretraining: Alignment from Token Zeromodelraising.ai
- Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggersarxiv.org
- [2601.10160] Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentarxiv.org
- SFT Drives Gemini’s Safety Properties — LessWronglesswrong.com
- ARENA - AI Safety Curriculumlearn.arena.education
- gpt-4.pdfcdn.openai.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Student Projects - CS 2881R AI Safetyboazbk.github.io
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWronglesswrong.com
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentalignmentpretraining.ai