suzana-ilic/study_model_behavior: Model Behavior Study Group
github.com · 849 words · saved by 1 readers
Model Behavior Study Group
AI Model Behavior: Reading List & Study Group Study Groups will be announced here: Luma 🔗 https://luma.com/AI-Communities Substack 🔗 https://mltaicommunities.substack.com/ AI Model Behavior Study Group (Series) Join us for a deep dive into how AI systems learn to behave, what guides their decisions, and how we can make them more aligned with human values. Each session, we'll spend 45 minutes reading a foundational paper or document on AI model behavior, safety, and alignment, followed by 30 minutes of discussion to extract key insights and debate implications. Whether you're an AI…
saved by
related reading
- LessWronglesswrong.com
- Introducing the Model Spec | OpenAIopenai.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Teaching Claude Whyalignment.anthropic.com
- [2605.24229] How Well Do Models Follow Their Constitutions?arxiv.org
- Toward A Public Science of Model Behavior | Transluce AItransluce.org
- From personas to intentions: towards a science of motivations for AI models — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- www.1943aiml.com1943aiml.com
- [2606.26071] Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignmentarxiv.org
- AI Safety | Arkosevictoriabrook.github.io