Lydia Nottingham - Lydia Nottingham
I’m currently researching utility engineering methods with the Center for AI Safety. Before this, I was at Manifund, which I highly recommend. I like: Lydia Nottingham
Lydia Nottingham - Lydia Nottingham Skip to content Lydia Nottingham I’m currently investigating whether models can predict how RL training will affect them, using privileged access to their internals. I like: America The phenomenon of grokking / phase transitions People / things with Markovian property Proactivity Have I mentioned I like America? I really like America. Musicals Mathematical modelling, particularly phase plane analysis and Markov chains Bridging artificial intelligence with the physical world Unpretentiousness Nondogmatism, thinking for oneself Dissolving problems The an
Explore this link on the map →related reading
- LessWronglesswrong.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Natural Language Autoencoders \ Anthropicanthropic.com
- Vincent Huangvvhuang.com
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- GenAI Handbookgenai-handbook.github.io
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- models have some pretty funny attractor states — LessWronglesswrong.com
- How To Become A Mechanistic Interpretability Researcher — LessWronglesswrong.com
- Trainloop AItrainloop.ai