Grounded Visual Generation

Subscribers:
351,000
Published on ● Video Link: https://www.youtube.com/watch?v=fLqF4isdWPg



Duration: 58:27
513 views
12


Multi-modal data provides an exciting opportunity to train grounded generative models that synthesize images consistent with real world phenomena. In this talk, I will share several of our recent efforts towards creating grounded visual generation models: (1) introducing user attention grounding for text-to-image synthesis, (2) improving text-to-image generation results with stronger language grounding, and (3) taking steps towards creating spatially grounded world models for embodied vision-and-language tasks.

Speaker: Jing Yu Koh, Google

MSR Deep Learning team: https://www.microsoft.com/en-us/research/group/deep-learning-group/




Other Videos By Microsoft Research


2021-09-23ALIGN: Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
2021-09-23Zero-Shot Detection via Vision and Language Knowledge Distillation
2021-09-17Three Explorations on Pre-Training: an Analysis, an Approach, and an Architecture
2021-09-16Visual Recognition beyond Appearances, and its Robotic Applications
2021-09-16A Truly Unbiased Model
2021-09-16Visual question answering & reasoning over vision & language: Beyond limits of statistical learning?
2021-09-15MDETR: Modulated Detection for End-to-End Multi-Modal Understanding
2021-09-15Learning Commonsense Understanding through Language and Vision
2021-09-15Tightly Connecting Vision and Language
2021-09-15Learning from Unlabeled Videos for Recognition, Prediction, and Control
2021-09-15Grounded Visual Generation
2021-08-25The New Jim Code: Reimagining the Default Settings of Technology & Society
2021-08-19A mechatronic shape display based on auxetic materials
2021-08-16Dependable IoT- Making data from IoT devices dependable and trustworthy for good decision making
2021-08-11Lookout System: National Television Commercial (1998)
2021-08-06Create human-centered AI with the Human-AI eXperience (HAX) Toolkit webinar
2021-08-04Computing Technology as Racial Infrastructure: A History of the Present & Blueprint for Black Future
2021-07-27Urban Air Chicago
2021-07-09The Vanishing Indian Speaks Back: Race, Genomics, and Indigenous Rights
2021-07-08Recent Advances in Image Captioning, Image-Text Retrieval and…
2021-07-08Directions in ML: Structured Models for Automated Machine Learning