Neil Rathi
cs.stanford.edu · 45 words · saved by 4 readers
Hi. I am an undergraduate at Stanford advised by Dan Jurafsky and a safety fellow at Anthropic with Alec Radford. I work on basic science for AI safety: why misalignments occur, and how we can prevent them during model training. In a past life, I did computational psycholinguistics. I also consume media. Selected Publications You can reach me at lastname [at] stanford [dot] edu. google scholar / twitter / last.fm
Neil Rathi I am a San Francisco-based AI researcher. I want to make sure that language models are good for society. elsewhere npr [at] anthropic [dot] com @neil_rathi on twitter papers on google scholar listening on last.fm hope of a condemned man ii, joan miró
saved by
related reading
- LessWronglesswrong.com
- Nitarshan Rajkumarnitarshan.com
- Thoughts — Jason Weijasonwei.net
- Why I Think More NLP Researchers Should Engage with AI Safety Concerns – NYU Alignment Research Groupwp.nyu.edu
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Nicholas Carlininicholas.carlini.com
- Tim Hua Personal Websitetimhua.me
- Naomi Saphransaphra.net
- machine learning imindslice.substack.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com