Today we’re joined by Sherry Yang, senior research scientist at Google DeepMind and a PhD student at UC Berkeley. In this interview, we discuss her new paper, "Video as the New Language for Real-World Decision Making,” which explores how generative video models can play a role similar to language models as a way to solve tasks in the real world. Sherry draws the analogy between natural language as a unified representation of information and text prediction as a common task interface and demonstrates how video as a medium and generative video as a task exhibit similar properties. This formulation enables video generation models to play a variety of real-world roles as planners, agents, compute engines, and environment simulators. Finally, We explore UniSim, an interactive demo of Sherry's work and a preview of her vision for interacting with AI-generated environments.
The complete show notes for this episode can be found at twimlai.com/go/676.
GraphRAG: Knowledge Graphs for AI Applications with Kirk Marple - #681
Teaching Large Language Models to Reason with Reinforcement Learning with Alex Havrilla - #680
Localizing and Editing Knowledge in LLMs with Peter Hase - #679
Coercing LLMs to Do and Reveal (Almost) Anything with Jonas Geiping - #678
V-JEPA, AI Reasoning from a Non-Generative Architecture with Mido Assran - #677
Assessing the Risks of Open AI Models with Sayash Kapoor - #675
OLMo: Everything You Need to Train an Open Source LLM with Akshita Bhagia - #674
Training Data Locality and Chain-of-Thought Reasoning in LLMs with Ben Prystawski - #673
Reasoning Over Complex Documents with DocLLM with Armineh Nourbakhsh - #672
Are Emergent Behaviors in LLMs an Illusion? with Sanmi Koyejo - #671
AI Trends 2024: Reinforcement Learning in the Age of LLMs with Kamyar Azizzadenesheli - #670
Building and Deploying Real-World RAG Applications with Ram Sriharsha - #669
Nightshade: Data Poisoning to Fight Generative AI with Ben Zhao - #668
Learning Transformer Programs with Dan Friedman - #667
AI Trends 2024: Machine Learning & Deep Learning with Thomas Dietterich - #666
AI Trends 2024: Computer Vision with Naila Murray - #665
Are Vector DBs the Future Data Platform for AI? with Ed Anuff - #664
Quantizing Transformers by Helping Attention Heads Do Nothing with Markus Nagel - #663
Responsible AI in the Generative Era with Michael Kearns - #662
Create your
podcast in
minutes
It is Free
20/20
The Dropout
FiveThirtyEight Politics
Ten Percent Happier with Dan Harris
World News Tonight with David Muir