Main
My research group aims to build and improve language models. Methodologically, we study data-driven methods that combine deep-learning based models with probabilistic controls. We are interested in applications in improved scaling, efficiency, model reasoning, and long-context generation.
I study post-training of AI systems for coding and related tasks. I'm interested in improving model reasoning for long-horizon tasks. Bio From 2016–2026, I was a Professor at Harvard and then Cornell. My group's research was recognized with an NSF CAREER Award and a Sloan Fellowship. My students have won paper awards at conferences for NLP, Hardware, and Visualization, and gone on to do really neat things. From 2019–2024, I also worked as a researcher at Hugging Face and helped on early open-source LLM projects. In 2024, I helped start the Conference on Language Modeling. I have a youtube…
saved by
related reading
- Composer2.pdfcursor.com
- Large Language Diffusion Modelsarxiv.org
- Stanford CS336 | Language Modeling from Scratchcs336.stanford.edu
- GenAI Handbookgenai-handbook.github.io
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- Explore | alphaXivalphaxiv.org
- machine learning imindslice.substack.com
- [2005.14165] Language Models are Few-Shot Learnersarxiv.org
- LLM Resourcesforrestbicker.com
- Chinmay Karkarchinmaykarkar.com
- Pathways Language Model (PaLM): Scaling to 540 Billion Parameters for Breakthrouai.googleblog.com
- 2506.17298arxiv.org