Chunyuan Li
chunyuan.li · 791 words · saved by 1 readers
A simple, whitespace theme for academics. Based on [*folio](https://github.com/bogoli/-folio) design.
My research centers on multimodal intelligence, with a focus on large-scale language and vision training. Key contributions include LLaVA and its model family series , as well as foundational early work such as GroundingDINO , GLIP , GLIGEN , Florence , and Oscar . My experience includes research roles at xAI , ByteDance , and Microsoft Research, Redmond . I earned my PhD in machine learning from Duke University under the guidance of Prof. Lawrence Carin , where my doctoral research explored deep generative models . I have also served the community as an Area Chair for NeurIPS, ICML, ICLR, EMN
saved by
related reading
- cs.unc.edu/~mbansal/teaching/nlp-comp790-590-spring23.htmlcs.unc.edu
- 2403.09611.pdfarxiv.org
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- MMHal Bencharxiv.org
- Explore | alphaXivalphaxiv.org
- Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning | Researchai.meta.com
- [2301.13823] Grounding Language Models to Images for Multimodal Generationarxiv.org
- [Recap] Multimodal Hackathon | lablab.ailablab.ai
- Large Language Diffusion Modelsarxiv.org
- HALVA: Hallucination Attenuated Language and Vision Assistantresearch.google
- GLM-5.3-Flash: Frontier Intelligence, Flash Costz.ai
- Minigpt-4minigpt-4.github.io