Nicholas Carlini
nicholas.carlini.com · 1,180 words · saved by 6 readers
Nicholas Carlini is a research scientist at Google DeepMind working at the intersection of machine learning and computer security.
Nicholas Carlini Nicholas Carlini Nicholas Carlini Research Scientist, Anthropic nicholas [at] carlini [dot] com GitHub | Google Scholar Nicholas Carlini Research Scientist, Anthropic nicholas [at] carlini [dot] com GitHub | Google Scholar I am a researcher working at the intersection of machine learning and computer security. Currently I work at Anthropic studying what bad things you could do with, or do to, language models. Previously, I was a research scientist at Google Brain (from 2018-2023) and DeepMind (from 2023-2025). I hold a Ph.D. from UC Berkeley under David Wagner, and a B.A. in c
saved by
related reading
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Measuring LLMs' impact on N-day exploits \ Anthropicred.anthropic.com
- A small number of samples can poison LLMs of any size \ Anthropicanthropic.com
- Welcome!boydkane.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Some Lessons from Adversarial Machine Learning | FAR.AIfar.ai
- [2502.05209] Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilitiesarxiv.org
- Simulated Users & Sad LLMs1a3orn.com
- [2608.09867] Stealing Reasoning Traces from Proprietary LLM APIsarxiv.org
- 2306.15447.pdfarxiv.org