✳flâneur — a map of the web's best reading
Nicholas Carlini
nicholas.carlini.com · 1,180 words · saved by 4 readers
Nicholas Carlini is a research scientist at Google DeepMind working at the intersection of machine learning and computer security.
Nicholas Carlini Nicholas Carlini Nicholas Carlini Research Scientist, Anthropic nicholas [at] carlini [dot] com GitHub | Google Scholar Nicholas Carlini Research Scientist, Anthropic nicholas [at] carlini [dot] com GitHub | Google Scholar I am a researcher working at the intersection of machine learning and computer security. Currently I work at Anthropic studying what bad things you could do with, or do to, language models. Previously, I was a research scientist at Google Brain (from 2018-2023) and DeepMind (from 2023-2025). I hold a Ph.D. from UC Berkeley under David Wagner, and a B.A. in c
Explore this link on the map →saved by
related reading
- Measuring LLMs' impact on N-day exploits \ Anthropicred.anthropic.com
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Welcome!boydkane.com
- A small number of samples can poison LLMs of any size \ Anthropicanthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Security incident disclosure — July 2026huggingface.co
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- 2025: The year in LLMssimonwillison.net
- What I learned this week - Can distillation be stopped, Mythos and the cybersecurity equilibrium, Pipeline RLdwarkesh.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Lapis Labslapis.rocks