cs.unc.edu/~mbansal/teaching/nlp-comp790-590-spring23.html
Instructor: Mohit Bansal Units: 3 Lectures: Wed 11am-1.30pm ET, Room SN-115 Office Hours: Wed 1:30pm-2:15pm ET (by appointment) (remote/zoom option) Course Webpage: https://www.cs.unc.edu/~mbansal/teaching/nlp-comp790-590-spring23.html Course Email: nlpcomp790unc -at- gmail.com This course will be based on the connections between the fields of natural language processing (NLP) and its important multimodal connections to computer vision and robotics; it will cover a wide variety of topics in the area of multimodal-NLP such as image/video-based captioning, retrieval, QA, and dialogue; vision+language commonsense; query-based video summarization; text-to-image/video generation; robotic navigation + manipulation instruction execution and generation; unified multimodal and embodied pretraining models (as well as ethics/bias/societal applications + issues). Topics Grading will (tentatively) consist of: For NLP concepts refresher, see: For multimodal NLP concepts, see corresponding lectures
COMP 790/590: Connecting Language to Vision and Robotics (Spring 2023) Instructor: Mohit Bansal Units: 3 Lectures: Wed 11am-1.30pm ET, Room SN-115 Office Hours: Wed 1:30pm-2:15pm ET (by appointment) (remote/zoom option) Course Webpage: https://www.cs.unc.edu/~mbansal/teaching/nlp-comp790-590-spring23.html Course Email: nlpcomp790unc -at- gmail.com Introduction This course will be based on the connections between the fields of natural language processing (NLP) and its important multimodal connections to computer vision and robotics; it will cover a wide variety of topics in the area of multimod
saved by
related reading
- Chunyuan Lichunyuan.li
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- Explore | alphaXivalphaxiv.org
- How Claude Performs on Robotics Tasks \ Anthropicanthropic.com
- [2301.13823] Grounding Language Models to Images for Multimodal Generationarxiv.org
- 2403.09611.pdfarxiv.org
- RT-2: Vision-Language-Action Modelsrobotics-transformer2.github.io
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- aman.ai • the art of artificial intelligenceaman.ai
- Computer Vision and Geometry Group | Robot Learningcvg.ethz.ch
- GitHub - robotics-survey/Awesome-Robotics-Foundation-Modelsgithub.com
- GitHub - FusionBrainLab/Vision_GRPOgithub.com