Today we’re joined by Sherry Yang, senior research scientist at Google DeepMind and a PhD student at UC Berkeley. In this interview, we discuss her new paper, "Video as the New Language for Real-World Decision Making,” which explores how generative video models can play a role similar to language models as a way to solve tasks in the real world. Sherry draws the analogy between natural language as a unified representation of information and text prediction as a common task interface and demonstrates how video as a medium and generative video as a task exhibit similar properties. This formulation enables video generation models to play a variety of real-world roles as planners, agents, compute engines, and environment simulators. Finally, We explore UniSim, an interactive demo of Sherry's work and a preview of her vision for interacting with AI-generated environments.
The complete show notes for this episode can be found at twimlai.com/go/676.
Deep Learning is Eating 5G. Here’s How, w/ Joseph Soriaga - #525
Modeling Human Cognition with RNNs and Curriculum Learning, w/ Kanaka Rajan - #524
Do You Dare Run Your ML Experiments in Production? with Ville Tuulos - #523
Delivering Neural Speech Services at Scale with Li Jiang - #522
AI’s Legal and Ethical Implications with Sandra Wachter - #521
Compositional ML and the Future of Software Development with Dillon Erb - #520
Generating SQL Database Queries from Natural Language with Yanshuai Cao - #519
Social Commonsense Reasoning with Yejin Choi - #518
Deep Reinforcement Learning for Game Testing at EA with Konrad Tollmar - #517
Exploring AI 2041 with Kai-Fu Lee - #516
Advancing Robotic Brains and Bodies with Daniela Rus - #515
Neural Synthesis of Binaural Speech From Mono Audio with Alexander Richard - #514
Using Brain Imaging to Improve Neural Networks with Alona Fyshe - #513
Adaptivity in Machine Learning with Samory Kpotufe - #512
A Social Scientist’s Perspective on AI with Eric Rice - #511
Applications of Variational Autoencoders and Bayesian Optimization with José Miguel Hernández Lobato - #510
Codex, OpenAI’s Automated Code Generation API with Greg Brockman - #509
Spatiotemporal Data Analysis with Rose Yu - #508
Parallelism and Acceleration for Large Language Models with Bryan Catanzaro - #507
Applying the Causal Roadmap to Optimal Dynamic Treatment Rules with Lina Montoya - #506
Create your
podcast in
minutes
It is Free
20/20
The Dropout
Ten Percent Happier with Dan Harris
World News Tonight with David Muir
NEJM This Week