flâneur — a map of the web's best reading

Tim Hua Personal Website

timhua.me · 733 words · saved by 2 readers

I am currently a MATS scholar conducting AI safety research. I work on model internals based alignment auditing techniques with Neel Nanda and Sam Marks (see e.g., the model auditing work from Anthropic). You can follow along with my work on the Alignment Forum here.     Before starting MATS, I conducted mechanistic interpretability research through AI safety camp and AI control research as a MARS Mentee. I was also a facilitator at BlueDot's Economics of Transformative AI course.     In a past life, I was an economist on Walmart's Economics Team working with pricing and Spark drivers. I designed and executed complex (multi-treatment arms with possible spillover effects) experiments and applied causal machine learning methods such as DoubleML and generalized random forests. I also used synthetic difference in differences and Chernozhukov et al. (2021) standard errors for evaluating treatment effects for when only one unit was treated. During my internship in 2022, I came up with this g

Tim Hua Personal Website F O T O In the Marin Headlands Contact Email: t im ra@nd.com @ tran sluce.org LinkedIn (prefer email) Or find me on Twitter: @Tim_Hua_ Google Scholar Alignment Forum Bio I am a member of technical staff at Transluce working on behavioral evaluations . I believe that extremely powerful AI systems—those capable of making humans obsolete—could be developed within the next ten years . Making sure that this goes well for humanity is among the most important problems in the world. Some of my recent work: Steering Evaluation Aware Models to Act Like They Are Deployed Tim Hua*

Explore this link on the map →

saved by

related reading