Chinmay Karkar
I work on the science and craft of training language models — what makes them learn, what makes them stable over long horizons, and how to tell whether they’re actually getting better.
researcher · pre-training & post-training of language models currently at Microsoft Research India · previously Lossfunk & Athena Agents (RL) about I work on the science and craft of training language models — what makes them learn, what makes them stable over long horizons, and how to tell whether they’re actually getting better. At Microsoft Research India I’m focused on reinforcement learning for LLMs and agentic verifier modules for multi-step reasoning evaluation. Before that, probabilistic forecasting at Lossfunk, model merging & RL post-training at Athena, and a fine-tuning /…
saved by
related reading
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- Alex L. Zhangalexzhang13.github.io
- Composer2.pdfcursor.com
- DeepSeek-R1arxiv.org
- AI in 2025: gestalt — LessWronglesswrong.com
- Explore | alphaXivalphaxiv.org
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- LLM Post-Training: A Deep Dive into Reasoning Large Language Modelsarxiv.org
- PostTrainBenchposttrainbench.com
- Kimi k1.5: Scaling Reinforcement Learning with LLMsalphaxiv.org