N . Soma Sekhar
0 followers · 320 views
on the atlas — 53
- [2604.08407] Your Agent Is Mine: Measuring Malicious Intermediary Attacks on the LLM Supply Chain1 savers
- We Need More Theories of Change | We need more theories of change1 savers
- Our Theory of Change • Slow Food USA1 savers
- Theories of Change for AI Auditing1 savers
- https://www.linkedin.com/in/n-soma-sekhar-11b9b21921 savers
- https://cs.stanford.edu/~jsteinhardt/ResearchasaStochasticDecisionProcess.html31 savers
- Resources shared by Tzu at the India AI Agentic Security Hackathon - Google Docs1 savers
- [2607.07368] Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitors1 savers
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation1 savers
- [2607.18966] Measuring Reward-Seeking via Contrastive Belief Updates2 savers
- Iliad Intensive Curriculum1 savers
- Home | Santa Fe Institute1 savers
- Nanyang AI Safety '26/27 Curriculum - Google Docs1 savers
- Stanford CS329A | Self-Improving AI Agents2 savers
- Fellowships — The Safety Apprentice1 savers
- [2604.15384] LinuxArena: A Control Setting for AI Agents in Live Production Software Environments1 savers
- [2607.00053] SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks1 savers
- A reading list for generalists — LessWrong18 savers
- The Annotated Transformer13 savers
- Airtable - Global AI Safety Orgs (updated May '26, by Ankur Pandey)1 savers
- How OpenAI delivers low-latency voice AI at scale | OpenAI1 savers
- Chapters - Deep Learning with Python2 savers
- Chess Books | Goodreads1 savers
- I Trained a Language Model. Then I Built a Brain Scanner and Looked Inside It. | by Caleb DeLeeuw | Apr, 2026 | Medium1 savers
- Unicorn! ·1 savers
- [public] Curated by ERA - Opportunities in AI Safety & Governance - Google Docs1 savers
- About — Anti-Scheming( imp for me is what you can do part)1 savers
- Credal | The Secure AI Agent Platform(for clone)1 savers
- Red-Teaming Guide.docx - Google Docs1 savers
- LLM_Attacks_&_Defenses_Exercise.ipynb - Colab1 savers
- The Architecture of Open Source Applications2 savers
- [2309.00267] RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback2 savers
- suzana-ilic/study_model_behavior: Model Behavior Study Group1 savers
- Research — Berkeley AI Safety Student Initiative1 savers
- [2404.03348] Knowledge Distillation-Based Model Extraction Attack using GAN-based Private Counterfactual Explanations1 savers
- IASEAI | Conference 20261 savers
- Oxford Witt Lab1 savers
- Student Projects / Supervision | Oxford Witt Lab1 savers
- OWASP Top 10 for Agentic Applications for 2026 - OWASP Gen AI Security Project1 savers
- Mitigating bias in artificial intelligence: Fair data generation via causal models for transparent and explainable decision-making - ScienceDirect1 savers
- [2503.05516] Cognitive Bias Detection Using Advanced Prompt Engineering1 savers
- Funding 60 projects to advance AI alignment research | AISI Work1 savers
- AI Risk & Reliability - MLCommons1 savers
- designing for trust - Brave Search1 savers
- Sundial- Very important for evaluating services for AI labs2 savers
- What Defines Hulu. Note: This is the original version of… | by Elisa Schreiber | Medium1 savers
- Stanford CS120 | Introduction to AI Safety1 savers
- Curius / Onboarding2621 savers
- Making Deep Learning Go Faster29 savers
- Advanced topics in the theory of machine learning9 savers
- Overview | Shallow Review 20257 savers
- CS146S: The Modern Software Developer - Stanford University5 savers
- 80,000 Hours: How to make a difference with your career3 savers
highlights — 1
Support community development of AI risk and reliability tests and organize definition of research- and industry-standard AI safety benchmarks based on those
AI Risk & Reliability - MLCommons