Jesse Mu on X: "We’re hiring for the adversarial robustness team @AnthropicAI! As an Alignment subteam, we're making a big effort on red-teaming, test-time monitoring, and adversarial training. If you’re interested in these areas, let us know! (emails in 🧵) https://t.co/0MPCSBb8zs" / X
To view keyboard shortcuts, press question mark View keyboard shortcuts Messages Home Explore Notifications Messages Lists Bookmarks Communities Premium Profile More Post Lydia @LydNot Post See new posts Conversation Jesse Mu @jayelmnop We’re hiring for the adversarial robustness team @AnthropicAI ! As an Alignment subteam, we're making a big effort on red-teaming, test-time monitoring, and adversarial training. If you’re interested in these areas, let us know! (emails in ) ALT Ethan Perez and 4 others 4:16 PM · Mar 20, 2024 · 64.2K Views 4 74 459 309 Post your reply Reply Jesse Mu @jayelmnop · Mar 20 Context: we’ve been pushing towards our ASL (AI Safety Level) safety commitments under our Responsible Scaling Policy—think about this as a “sprint on safety.” Red-teaming and adversarial robustness are a major part of this story. https://anthropic.com/news/anthropics-responsible-scaling-policy… 1 1 20 2.5K Jesse Mu @jayelmnop · Mar 20 Some examples of what our team has been up to
Jesse Mu @jayelmnop We’re hiring for the adversarial robustness team @ AnthropicAI ! As an Alignment subteam, we're making a big effort on red-teaming, test-time monitoring, and adversarial training. If you’re interested in these areas, let us know! (emails in 🧵) 4:16 PM · Mar 20, 2024 72.7K Views 4 0 4 67 0 6 7 445 0 4 4 5 300 0 3 0 0
Explore this link on the map →related reading
- Mediumai-alignment.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Security incident disclosure — July 2026huggingface.co
- Teaching Claude why \ Anthropicanthropic.com
- 2312.06942arxiv.org
- GitHub - SoyGema/pulling_ace · GitHubgithub.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Manus AI: The Best Autonomous AI Agent Redefining Automation and Productivityhuggingface.co
- Recommendations-for-Using-Red-Teaming-for-AI-Accountability-PolicyBrief.pdfdatasociety.net
- Frontier Risk Report (February to March 2026) - METRmetr.org