To advance generative AI beyond text-based large language models to video comprehension, the industry and its research must focus on meticulous data curation, filtering, annotation, and high-quality data selection from millions of videos for downstream fine-tuning tasks. These tasks span practical implementations in fields like robotics, digital twins, and autonomous vehicles. This talk will highlight innovations and research across the entire AI lifecycle stack for building a multimodal video foundation model. In this session, you will learn about key components in an end-to-end video curation service orchestrated by Ray Data and running on NVIDIA DGX Cloud. Key insights from NVIDIA’s team will be shared—covering aspects that enable developers and experts to leverage and enhance state-of-the-art AI models for large-scale inference and the creation of high-quality video datasets.
Ray Summit 2024
In-Person Agenda
READY TO REGISTER?
Come connect with the global community of thinkers and disruptors who are building and deploying the next generation of AI and ML applications.
Join the Conversation
Hashtag it
#RaySummitDon't wait for the conference to get the convo going. Join the Ray community now – ask a question in the forums, open a pull request or simply share why you’re excited. Create some buzz!