cs.unc.edu/~mbansal/teaching/nlp-comp790-590-spring23.html
Instructor: Mohit Bansal Units: 3 Lectures: Wed 11am-1.30pm ET, Room SN-115 Office Hours: Wed 1:30pm-2:15pm ET (by appointment) (remote/zoom option) Course Webpage: https://www.cs.unc.edu/~mbansal/teaching/nlp-comp790-590-spring23.html Course Email: nlpcomp790unc -at- gmail.com This course will be based on the connections between the fields of natural language processing (NLP) and its important multimodal connections to computer vision and robotics; it will cover a wide variety of topics in the area of multimodal-NLP such as image/video-based captioning, retrieval, QA, and dialogue; vision+language commonsense; query-based video summarization; text-to-image/video generation; robotic navigation + manipulation instruction execution and generation; unified multimodal and embodied pretraining models (as well as ethics/bias/societal applications + issues). Topics Grading will (tentatively) consist of: For NLP concepts refresher, see: For multimodal NLP concepts, see corresponding lectures
COMP 790/590: Connecting Language to Vision and Robotics (Spring 2023) Instructor: Mohit Bansal Units: 3 Lectures: Wed 11am-1.30pm ET, Room SN-115 Office Hours: Wed 1:30pm-2:15pm ET (by appointment) (remote/zoom option) Course Webpage: https://www.cs.unc.edu/~mbansal/teaching/nlp-comp790-590-spring23.html Course Email: nlpcomp790unc -at- gmail.com Introduction This course will be based on the connections between the fields of natural language processing (NLP) and its important multimodal connections to computer vision and robotics; it will cover a wide variety of topics in the area of multimod
Explore this link on the map →saved by
related reading
- Chunyuan Lichunyuan.li
- Explore | alphaXivalphaxiv.org
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- [2301.13823] Grounding Language Models to Images for Multimodal Generationarxiv.org
- 2403.09611.pdfarxiv.org
- RT-2: Vision-Language-Action Modelsrobotics-transformer2.github.io
- Computer Vision and Geometry Group | Robot Learningcvg.ethz.ch
- how we accidentally solved robotics by watching 1 million hours of YouTube – atharva's blogksagar.bearblog.dev
- aman.ai • the art of artificial intelligenceaman.ai
- [2109.01115] Learning Language-Conditioned Robot Behavior from Offline Data and Crowd-Sourced Annotationarxiv.org
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- Language Models can Solve Computer Tasksarxiv.org