Tim Hua Personal Website
I am currently a MATS scholar conducting AI safety research. I work on model internals based alignment auditing techniques with Neel Nanda and Sam Marks (see e.g., the model auditing work from Anthropic). You can follow along with my work on the Alignment Forum here. Before starting MATS, I conducted mechanistic interpretability research through AI safety camp and AI control research as a MARS Mentee. I was also a facilitator at BlueDot's Economics of Transformative AI course. In a past life, I was an economist on Walmart's Economics Team working with pricing and Spark drivers. I designed and executed complex (multi-treatment arms with possible spillover effects) experiments and applied causal machine learning methods such as DoubleML and generalized random forests. I also used synthetic difference in differences and Chernozhukov et al. (2021) standard errors for evaluating treatment effects for when only one unit was treated. During my internship in 2022, I came up with this g
Tim Hua Personal Website F O T O In the Marin Headlands Contact Email: t im ra@nd.com @ tran sluce.org LinkedIn (prefer email) Or find me on Twitter: @Tim_Hua_ Google Scholar Alignment Forum Bio I am a member of technical staff at Transluce working on behavioral evaluations . I believe that extremely powerful AI systems—those capable of making humans obsolete—could be developed within the next ten years . Making sure that this goes well for humanity is among the most important problems in the world. Some of my recent work: Steering Evaluation Aware Models to Act Like They Are Deployed Tim Hua*
Explore this link on the map →saved by
related reading
- Welcome!boydkane.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- Spring 2026 Projects - SPARsparai.org
- AI Safety | Arkosevictoriabrook.github.io
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Research Areas in Interpretability (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Rohin Shah on what it's really like to run AGI safety at Google DeepMind (and where I disagree with 'doomers') | 80,000 Hours80000hours.org
- MATS Alumnimatsprogram.org