Today we're joined by Armineh Nourbakhsh of JP Morgan AI Research to discuss the development and capabilities of DocLLM, a layout-aware large language model for multimodal document understanding. Armineh provides a historical overview of the challenges of document AI and an introduction to the DocLLM model. Armineh explains how this model, distinct from both traditional LLMs and document AI models, incorporates both textual semantics and spatial layout in processing enterprise documents like reports and complex contracts. We dig into her team’s approach to training DocLLM, their choice of a generative model as opposed to an encoder-based approach, the datasets they used to build the model, their approach to incorporating layout information, and the various ways they evaluated the model’s performance.
The complete show notes for this episode can be found at twimlai.com/go/672.
Deep Learning is Eating 5G. Here’s How, w/ Joseph Soriaga - #525
Modeling Human Cognition with RNNs and Curriculum Learning, w/ Kanaka Rajan - #524
Do You Dare Run Your ML Experiments in Production? with Ville Tuulos - #523
Delivering Neural Speech Services at Scale with Li Jiang - #522
AI’s Legal and Ethical Implications with Sandra Wachter - #521
Compositional ML and the Future of Software Development with Dillon Erb - #520
Generating SQL Database Queries from Natural Language with Yanshuai Cao - #519
Social Commonsense Reasoning with Yejin Choi - #518
Deep Reinforcement Learning for Game Testing at EA with Konrad Tollmar - #517
Exploring AI 2041 with Kai-Fu Lee - #516
Advancing Robotic Brains and Bodies with Daniela Rus - #515
Neural Synthesis of Binaural Speech From Mono Audio with Alexander Richard - #514
Using Brain Imaging to Improve Neural Networks with Alona Fyshe - #513
Adaptivity in Machine Learning with Samory Kpotufe - #512
A Social Scientist’s Perspective on AI with Eric Rice - #511
Applications of Variational Autoencoders and Bayesian Optimization with José Miguel Hernández Lobato - #510
Codex, OpenAI’s Automated Code Generation API with Greg Brockman - #509
Spatiotemporal Data Analysis with Rose Yu - #508
Parallelism and Acceleration for Large Language Models with Bryan Catanzaro - #507
Applying the Causal Roadmap to Optimal Dynamic Treatment Rules with Lina Montoya - #506
Create your
podcast in
minutes
It is Free
20/20
The Dropout
Ten Percent Happier with Dan Harris
World News Tonight with David Muir
NEJM This Week