8:00 AM PDT - Tuesday, Oct 1
- Lunch/BreakOctober 1: ConferenceTuesday, Oct 018:00 a.m. Tuesday, Oct 01Enjoy a light breakfast and coffee while mingling with Ray attendeesSession Type: Lunch/Break
- Tuesday, Oct 18:30 AM - 9:30 AM PDTRayground
9:00 AM PDT - Tuesday, Oct 1
- KeynoteYesBeginnerOctober 1: ConferenceTuesday, Oct 019:00 a.m. Tuesday, Oct 01Join us for the opening keynote of Ray Summit 2024 on October 1st. Hear from Ion Stoica and Robert Nishihara, Co-Founders of Anyscale, as they address the growing challenges of scaling and delivering real value with Al.
They’ll review the industry trends that have led us here and how Ray solves these problems as the AI Compute Engine, built to help organizations solve AI workloads. As they unveil the blueprint for this breakthrough expect exciting new Ray announcements from optimizations to accelerate GPU workloads including distributed training and large model inference to newly supported data formats in Ray Data cementing Ray’s place for the future of multimodal AI.
CEO Keerti Melkote will then outline his vision for Anyscale’s future and showcase how Anyscale can empower companies to optimize Ray and accelerate their AI projects through a live demo.
Following this keynote, hear from Anastassis Germanidis, CTO and Co-Founder of Runway on the future of AI. Then, join Marc Andreeseen, General Partner at Andreessen Horowitz for a Fireside chat., Founder, Anyscale
, UC Berkeley
, Runway
, Andreessen Horowitz
, CEO, Anyscale
Level of Expertise: BeginnerSession Type: Keynote- Tuesday, Oct 19:30 AM - 11:30 AM PDTKeynote
11:00 AM PDT - Tuesday, Oct 1
- Lunch/BreakOctober 1: Conference11:00 a.m. Tuesday, Oct 01Tuesday, Oct 01Grab lunch and explore Rayground, where you'll find Anyscale demos, sponsor booths, a Headshot Booth presented by Microsoft and the Lightning Theater.Session Type: Lunch/Break
- Tuesday, Oct 111:30 AM - 1:00 PM PDTRayground
- Lightning TalkLightning TalkIntermediateYesOctober 1: Conference11:00 a.m. Tuesday, Oct 01Tuesday, Oct 01A principal challenge in advancing scientific computing on campuses today is the efficient computation of a mix of conventional High Performance Computing (HPC) workloads with fast-growing Machine Learning (ML) workloads. Over the past two decades efficient campus computation has been realized by deploying collections of big data processing frameworks (Spark), data analysis tools (Pandas), and visualizers (Grafana), often cleverly combined to address a specific workload need. Yet this approach of integrating disjoint systems can be complicated. It is also poorly suited for higher education settings, where training on many tools reduces research productivity, strains campus research computing expertise, and taxes already thinly stretched IT finances. In addition, new user demand is surging, as less sophisticated data science and ML users from outside STEM disciplines need to learn and perform computational research in their fields of study.
In a pilot program at Princeton University, we are demonstrating how distributed execution with Ray can help support mixed HPC workloads in diverse on- and off-campus settings. We are exploring Ray software across a range of different computing environments, from 1) a shared campus research computing cluster; 2) a federated multi-institutional cluster across the NSF-supported FABRIC computing and networking research infrastructure; and 3) a hybrid campus and commercial compute cloud using the Google Cloud Platform (GCP). Taken together, the combined thrusts will demonstrate the extent to which a single, unified computing framework can support widespread coverage of diverse campus computational science needs., Senior Research Scholar, Dept. of Computer Science, and Associate Dean for Research, Princeton University
Level of Expertise: IntermediateSession Tracks: Lightning TalkSession Type: Lightning Talk- Tuesday, Oct 111:45 AM - 12:00 PM PDTLightning Stage
12:00 PM PDT - Tuesday, Oct 1
- Lightning TalkLightning TalkBeginnerYesOctober 1: Conference12:00 p.m. Tuesday, Oct 01Tuesday, Oct 01In the world of e-commerce, the hunt for the perfect purchase can be a wild adventure.
The impulse to find similar items fuels the search for visually comparable products. Here comes in handy a capacity to perform text2image and image2image search.
Motivated by the retail industry and the emerging trend of multimodality, we want to present our fashion images retrieval system using Anyscale Ray and Pinecone. It’s an intuitive and scalable application powered by CLIP embeddings and vector similarity search, capable of seamlessly retrieving images based on textual descriptions or visual prompts.
Furthermore, we will share insights into optimizing application performance using Ray Serve for autoscaling and employing parallelization techniques offered by Ray Data. We aim to empower developers by providing a practical roadmap for constructing similar applications, emphasizing efficiency, scalability, and user experience., Machnie Learning Engineer, deepsense.ai
, Senior Technical Leader, deepsense.ai
, Lead MLOps Engineer, deepsense.ai
Level of Expertise: BeginnerSession Tracks: Lightning TalkSession Type: Lightning Talk- Tuesday, Oct 112:00 PM - 12:15 PM PDTLightning Stage
- Lightning TalkLightning TalkOctober 1: Conference12:00 p.m. Tuesday, Oct 01Tuesday, Oct 01In this session Anil & Logan will present an overview of Akash Network - the world’s first fully open source super cloud. They will talk about why Akash Network was created, what it’s mission is going forward and how you can get involved with the open source project. They’ll touch on how ThumperAI leveraged Akash for training an image generation model. Time permitting, they also present a quick demo of running Ray clusters on Akash (very likely with vLLM and Llama-3.1)
, VP - Product & Engineering at Overclock Labs, Akash
, Co-Founder / CEO, ThumperAI
Session Tracks: Lightning TalkSession Type: Lightning Talk- Tuesday, Oct 112:15 PM - 12:30 PM PDTLightning Stage
- Lightning TalkLightning TalkOctober 1: Conference12:00 p.m. Tuesday, Oct 01Tuesday, Oct 01Train models performantly at scale with Snowflake Container Runtime, powered by Ray
, Senior Software Engineer, Snowflake
, Software Engineer, Snowflake
Session Tracks: Lightning TalkSession Type: Lightning Talk- Tuesday, Oct 112:30 PM - 12:45 PM PDTLightning Stage
- Lightning TalkLightning TalkAdvancedYesOctober 1: Conference12:00 p.m. Tuesday, Oct 01Tuesday, Oct 01In this session, we'll delve into unlocking the power of scalable, low-cost, high-performance model inference with Large Language Models such as Llama3 and Mistral7B on Amazon EKS. Learn how you can utilize RayServe framework and AWS Neuron (Trainium and Inferentia) instances to overcome GPU availability issues, drive cost optimization while scaling the model serving infrastructure by combining Ray with Karpenter autoscaler on Amazon EKS.
, Principal OSS Specialist, AWS, Amazon
, Senior Solutions Architect, OSS and Containers, AWS
Level of Expertise: AdvancedSession Tracks: Lightning TalkSession Type: Lightning Talk- Tuesday, Oct 112:45 PM - 1:00 PM PDTLightning Stage
1:00 PM PDT - Tuesday, Oct 1
- Breakout SessionAI/ML Platform & ApplicationsYesAdvancedYesOctober 1: Conference1:00 p.m. Tuesday, Oct 01Tuesday, Oct 01Robotics ML presents unique challenges to ML training due to reliance on multimodal data and the necessity for extensive robot simulations to be integrated in the training process. Our training platform, built with Ray, addresses these complexities by providing tools to manage and query metadata, preprocess data offline and online, and scale up training infrastructure with simple configuration changes.
In this presentation, we will demonstrate an end-to-end machine learning training platform tailored specifically for robotics. Using RayClusters, this platform integrates physical simulation tools directly into training workflows. The platform supports multi-tenancy within the Kubernetes infrastructure, allowing many users to share compute resources.
With a live demo of the ML training process, we will showcase how this platform accelerates research and enhances the scalability of training robot policy learning models. These models can then be deployed to actual robots., Machine Learning Engineer, The AI Institute.
, Machine Learning Engineer, The AI Institute
Level of Expertise: AdvancedSession Tracks: AI/ML Platform & ApplicationsSession Type: Breakout Session- Tuesday, Oct 11:00 PM - 1:30 PM PDTGolden Gate Salon C1
- Breakout SessionLLM & GenAIIntermediateYesOctober 1: Conference1:00 p.m. Tuesday, Oct 01Tuesday, Oct 01This talk addresses key challenges in evaluating LLM-powered AI agents: behavioral instability, comprehensive testing, and cascading failures. We'll explore advanced techniques including multi-dimensional metrics and automated scenario generation. Gain insights gleaned from hundreds of AI engineering teams in production into implementing agent-specific observability systems and designing robust evaluation pipelines that can handle the complexities of modern AI agents, from detecting subtle regressions to quantifying performance across diverse, dynamically generated test cases.
, CEO & Cofounder, Arize
Level of Expertise: IntermediateSession Tracks: LLM & GenAISession Type: Breakout Session- Tuesday, Oct 11:00 PM - 1:30 PM PDTYerba Buena Salon 14
- Breakout SessionAI/ML Platform & ApplicationsAdvancedYesOctober 1: Conference1:00 p.m. Tuesday, Oct 01Tuesday, Oct 01In this talk, Klaviyo’s ML platform team presents insights gained from over a year of using Ray Serve for their model serving platform - DART (DAtascience RunTime). We talk about how Ray Serve provides the best combination of flexibility, simplicity and high availability among all the options we considered. Klaviyo’s DART platform is self-serve, allowing Data Scientists to deploy new models quickly, reducing the time it takes for them to go to production from a few weeks to a few days.
We will explore certain internal details of Ray Serve architecture which have been crucial for us to understand certain behaviors and debug issues such as internal request routing and fault tolerance with KubeRay. Additionally, we explore DART’s multi-cluster, multi-AZ architecture that ensures resilience in cases of node failures and pod evictions.
We conclude with important considerations for any teams trying to build their own model serving platform with Ray Serve, including things like cluster sizing, traffic isolation and avoiding common pitfalls such as single point of failures. Finally, we walk through different architectures with increasing levels of complexity so organizations can choose the architecture that best fits their cost and high availability needs. With these insights, we aim to empower others to build their own robust serving platform using Ray Serve., Senior ML Engineer, Klaviyo
Level of Expertise: AdvancedSession Tracks: AI/ML Platform & ApplicationsSession Type: Breakout Session- Tuesday, Oct 11:00 PM - 1:30 PM PDTYerba Buena Salon 4
- Breakout SessionLLM & GenAIIntermediateYesOctober 1: Conference1:00 p.m. Tuesday, Oct 01Tuesday, Oct 01In this talk we will discuss challenges and opportunities that lie on the path towards building state of the art LLMs. We will walk the path of a researcher, from initial success of Llama 1 - a research only LLM, to Llama 3.1 - a collection of models and open AI ecosystem, broadly adopted by the community.
We will discuss the unique value Llama models bring, what challenges are ahead and how everyone can help., Director, AI Research, Meta
Level of Expertise: IntermediateSession Tracks: LLM & GenAISession Type: Breakout Session- Tuesday, Oct 11:00 PM - 1:30 PM PDTGolden Gate Salon A
- Breakout SessionLLM & GenAIIntermediateOctober 1: Conference1:00 p.m. Tuesday, Oct 01Tuesday, Oct 01LLM development demands sophisticated tooling to navigate the complex trade-offs between quality, cost, and latency. At Anyscale, we've developed and refined foundational, production-grade primitives for serving, tuning, and evaluating LLMs. We'll showcase our unified platform for iterative fine-tuning, serving, and evaluation. We'll also discuss how our high-level abstractions maintain flexibility while simplifying the development process, allowing enterprises and engineers to focus on solving business problems and integrate it with their existing workflows. Whether you're managing multiple LLM use cases or looking to leverage open-weight models at scale, this talk will demonstrate how Anyscale's toolkit can streamline your LLM development workflow.
, Software Engineer, Anyscale
Level of Expertise: IntermediateSession Tracks: LLM & GenAISession Type: Breakout Session- Tuesday, Oct 11:00 PM - 1:30 PM PDTGolden Gate Salon C3
- Breakout SessionLLM & GenAIBeginnerYesOctober 1: Conference1:00 p.m. Tuesday, Oct 01Tuesday, Oct 01At Shopify we fine-tune and deploy large vision language models in production, and leverage different open source tooling including Ray. In this talk I'll walkthrough how we went about doing it for a generative ai use case at Shopify's scale.
, Staff ML Engineer, Shopify
, Senior Machine Learning Engineer, Shopify
Level of Expertise: BeginnerSession Tracks: LLM & GenAISession Type: Breakout Session- Tuesday, Oct 11:00 PM - 1:30 PM PDTGolden Gate Salon C2
- Breakout SessionAI/ML Platform & ApplicationsYesBeginnerYesOctober 1: Conference1:00 p.m. Tuesday, Oct 01Tuesday, Oct 01In this talk, we delve into how Canva has scaled AI by integrating Ray and Anyscale into our ecosystem. We'll share key outcomes from incorporating Ray into our models and discuss the benefits as well as lessons learned. Additionally, we highlight our use of heterogeneous training clusters for out-of-band validation, further improving time and cost savings.
, Senior Machine Learning Engineer, Canva
, Senior Machine Learning Engineer, Canva
Level of Expertise: BeginnerSession Tracks: AI/ML Platform & ApplicationsSession Type: Breakout Session- Tuesday, Oct 11:00 PM - 1:30 PM PDTGolden Gate Salon B
- Breakout SessionvLLMYesBeginnerOctober 1: Conference1:00 p.m. Tuesday, Oct 01Tuesday, Oct 01The vLLM team will brief the past year of progress in "building the fastest and easiest-to-use open-source LLM inference and serving engine". We are excited to share major updates in terms of adoption, features, performance, community, and the governance of the project. We will then share the roadmap for the upcoming releases.
, vLLM @Berkeley, University of Chicago
, PhD Student / vLLM Maintainer, UC Berkeley
Level of Expertise: BeginnerSession Tracks: vLLMSession Type: Breakout Session- Tuesday, Oct 11:00 PM - 1:30 PM PDTYerba Buena Salon 10
- Lightning TalkLightning TalkAdvancedYesOctober 1: Conference1:00 p.m. Tuesday, Oct 01Tuesday, Oct 01Ray is very popular among AI practitioners for its rich ecosystem and great integration. However, its core distributed memory model and python-first design makes it very appealing to be applied in research infra in quantitative trading. I would like to share from my experiences building a platform for quant research and trading with Ray on Kubernetes that supports up to 100k concurrent CPU cores and easily extensible to GPUs.
I would like to focus on points below:
1. Think about integrate Ray into a quantitative research workflow, unifying large offline research jobs and interactive "small" intraday jobs.
2. Kuberay learnings and challenges faced running Ray on Kubernetes that only use spot instance and the attemped solution which is a customized Ray controller with different retry policies to build a reliable layer on unreliable infra.
3. Experiences pioneering Ray within a firm from getting buy-in to educating users., Senior Engineer, Chicago Trading Company
Level of Expertise: AdvancedSession Tracks: Lightning TalkSession Type: Lightning Talk- Tuesday, Oct 11:30 PM - 1:45 PM PDTLightning Stage
- Breakout SessionRay Deep DivesIntermediateYesOctober 1: Conference1:00 p.m. Tuesday, Oct 01Tuesday, Oct 01Creating a video generation model capable of producing realistic and imaginative scenes from text instructions, necessitates a vast amount of high-quality video data. In this talk, we will share how we utilized Ray to address the various challenges we encountered while building our video data processing pipeline from the ground up.
Our focus will be on developing a robust and scalable data pipeline capable of handling massive volumes of video data. Ray's ecosystem, with its core capabilities in Ray Core, Ray Data, and Ray Serve, provides an effective solution to these challenges by excelling in dynamic scaling of computation and orchestration of heterogeneous resources. By leveraging these capabilities, we have successfully constructed a complex and efficient data processing pipeline.
Additionally, we will share our experiences in managing the Ray infrastructure and highlight the best practices we have learned along the way., Senior Software Engineer, ByteDance
, ByteDance
, Head of Compute and Networking, Infrastructure System Lab, ByteDance
Level of Expertise: IntermediateSession Tracks: Ray Deep DivesSession Type: Breakout Session- Tuesday, Oct 11:45 PM - 2:15 PM PDTGolden Gate Salon B
- Breakout SessionvLLMIntermediateYesOctober 1: Conference1:00 p.m. Tuesday, Oct 01Tuesday, Oct 01With the recent rapid advancements in multimodal language models, there has been a surge of interest from the open source vLLM community in the support of these models on vLLM.
In this talk, we will explore the journey of integrating multimodal models into vLLM, delving into the technical challenges encountered and the valuable lessons learned throughout the process., Software Engineer, Roblox
Level of Expertise: IntermediateSession Tracks: vLLMSession Type: Breakout Session- Tuesday, Oct 11:45 PM - 2:15 PM PDTYerba Buena Salon 10
- Breakout SessionRay Use CasesIntermediateOctober 1: Conference1:00 p.m. Tuesday, Oct 01Tuesday, Oct 01As machine learning rapidly evolves, companies are increasingly upgrading their data platforms, including adopting advanced AI frameworks like Ray, to stay competitive. However, limited GPU availability and data scattered across various locations within organizations pose significant challenges, making it difficult for teams to access the data they need. This fragmentation slows down AI development and complicates model training.
To address the growing demand for efficient and unified data access, Alluxio provides a service that enables Ray to seamlessly access data from multiple sources—whether in the cloud or on-premises—without being hindered by differences in cloud/storage providers, network bottlenecks, or complex authentication protocols. This ensures that GPU training can happen anywhere, allowing Ray’s distributed workloads to run efficiently while avoiding the inefficiencies caused by data silos and inconsistent access.
In this talk, Haoyuan Li, Founder and CEO of Alluxio, will discuss how to streamline data access specifically for Ray using Alluxio and share practical strategies that organizations can adopt to build a robust data infrastructure that accelerates AI innovation., Founder and CEO, Alluxio
, VP of Technology, Alluxio
Level of Expertise: IntermediateSession Tracks: Ray Use CasesSession Type: Breakout Session- Tuesday, Oct 11:45 PM - 2:15 PM PDTGolden Gate Salon C3
- Breakout SessionAI/ML Platform & ApplicationsIntermediateYesOctober 1: Conference1:00 p.m. Tuesday, Oct 01Tuesday, Oct 01Ray offers powerful metrics visualizations powered by graphana and prometheus. Although useful, the setup can take time - and customization can be challenging.
Raydar, an open source project from point72 (https://github.com/Point72/raydar), provides both out-of-the-box live cluster metrics and user visualizations for Ray workflows with just a simple pip install. It helps unlock distributed machine learning visualizations on Anyscale clusters, runs live and at scale, is easily customizable, and enables all the in-browser aggregations that perspective (https://perspective.finos.org/) has to offer. In this talk, we demonstrate the setup steps for Raydar, how to enable generic metrics visualizations, and custom visualizations for a ML workflow., Quantitative Developer, Point72
Level of Expertise: IntermediateSession Tracks: AI/ML Platform & ApplicationsSession Type: Breakout Session- Tuesday, Oct 11:45 PM - 2:15 PM PDTGolden Gate Salon C2
- Breakout SessionLLM & GenAIBeginnerYesOctober 1: Conference1:00 p.m. Tuesday, Oct 01Tuesday, Oct 01Generative AI promises a spectacular increase in productivity for enterprises that can leverage it on their proprietary/IP data or for users who can use it on their private data and within their personal accounts. However, the fear of exposing such confidential data to external parties (such as the LLM provider and its platform, its partners or employees, hackers breaking in the LLM platform, or subpoenas) have prevented organizations from using generative AI to its potential. For example, Samsung and Google blocked the use of ChatGPT internally, and many companies are still not using generative AI for their core proprietary data. In this talk, I will describe a novel technology based on confidential computing that can keep proprietary/private data encrypted throughout the entire generative AI cycle, and thus protected from the external parties above. I will overview research from my lab at UC Berkeley, as well as a real-world artifact that is both easy to use and does not degrade the quality of predictions. For example, our technology enables companies to keep their prompts confidential while using LLMs.
, Co-Founder and President, Opaque Systems, Inc.
Level of Expertise: BeginnerSession Tracks: LLM & GenAISession Type: Breakout Session- Tuesday, Oct 11:45 PM - 2:15 PM PDTGolden Gate Salon A
- Breakout SessionLLM & GenAIIntermediateYesOctober 1: Conference1:00 p.m. Tuesday, Oct 01Tuesday, Oct 01The recent integration of Intel Gaudi accelerators into the Ray framework introduces a groundbreaking option for GenAI tasks and workloads, offering superior performance, scalability, and efficiency for entire lifecycle of LLM. This presentation will provide a deep dive into the capabilities provided by Intel Gaudi accelerators within Ray, particularly focusing on large language models (LLMs).
Our session will showcase LLM pre-training, fine-tuning and serving cases that leveraging Ray’s unprecedent scalability, distributed computation, and fault tolerance capability, demonstrating the effective configuration, optimization, and scalability benefits of training and deploying LLMs on the Intel Gaudi platform. Attendees will gain detailed insights on how to configure and manage Ray clusters optimized with Intel Gaudi accelerators that leads to improved scalability and cost-effectiveness, with a wide range of popular models like Llama, Mistral, and Stable Diffusion. Performance results will also be disclosed. This session is designed to equip developers, data scientists, and AI practitioners with the practical know-how to fully utilize Intel Gaudi accelerators in their GenAI endeavors, ensuring they can harness state-of-the-art technology to push the boundaries of what is possible in AI. Join us to explore how Intel Gaudi accelerators and Ray are reshaping the landscape of AI through enhanced scalability, efficiency, and performance., AI Software Engineering Manager, Intel
, Machine Learning Engineer, Intel
Level of Expertise: IntermediateSession Tracks: LLM & GenAISession Type: Breakout Session- Tuesday, Oct 11:45 PM - 2:15 PM PDTGolden Gate Salon C1
- Breakout SessionRay Use CasesIntermediateYesOctober 1: ConferenceOctober 2: Conference1:00 p.m. Tuesday, Oct 01Tuesday, Oct 01Explores how we leveraged Ray on Databricks and Spark Structured Streaming to optimize video processing, and how we scaled the classification of these video ads across millions of classes using GenAI.
Leverages real-time insights from diverse creatives (videos, audio) leveraging ML and GenAI. Our initial goal aimed to categorize video ads into 30,000 product classes, with a planned expansion to six million. While an initial transformer-based machine learning model achieved high accuracy, we anticipated challenges from exponential growth in training time and limited data per class. To address these issues, we leveraged OSS LLMs. We used optimized LLama2(vLLM) to categorize creatives by identifying product categories and conducting similarity searches across the labels. Our baseline machine learning model achieved 69% accuracy with ~ 25,000 labels and a training dataset of ~200k creatives. By integrating LLMs, we achieved a remarkable 15% uplift in accuracy. Combining both approaches, we devised a solution where the LLM model acts as a pre-processing step, generating summaries for subsequent machine learning analysis., Principal Software Developer, MediaRadar | Vivvix
, Senior Machine Learning Engineer, MediaRadar | Vivvix
, Senior Specialist Solution Architect, Databricks
Level of Expertise: IntermediateSession Tracks: Ray Use CasesSession Type: Breakout Session- Tuesday, Oct 11:45 PM - 2:15 PM PDTYerba Buena Salon 4
2:00 PM PDT - Tuesday, Oct 1
- Lightning TalkLightning TalkBeginnerYesOctober 1: Conference2:00 p.m. Tuesday, Oct 01Tuesday, Oct 01Deploying Ray clusters in shared environments requires careful consideration of security implications. By leveraging KubeRay and Kubernetes security features, organizations can ensure that their Ray clusters are securely isolated and protected from unauthorized access and resource abuse. This talk aims to provide insights and best practices for securing Ray clusters, enabling organizations to harness the power of Ray for distributed computing with confidence. We will demonstrate an open-source extension of KubeRay, namely KubeSecRay, that leverages Kubernetes security features to secure Ray clusters. As well as the configuration options for enforcing security policies and access controls.
, Microsoft
Level of Expertise: BeginnerSession Tracks: Lightning TalkSession Type: Lightning Talk- Tuesday, Oct 12:15 PM - 2:30 PM PDTLightning Stage
- Breakout SessionvLLMAdvancedYesOctober 1: Conference2:00 p.m. Tuesday, Oct 01Tuesday, Oct 01vLLM supports various forms of model quantization including FP8, INT8, and INT4, reducing memory consumption and increasing generation speed. In this talk, you will learn how vLLM accelerates models with quantization internally and how to apply these methods to your model with vLLM’s llm-compressor framework.
, Sr Director of Engineering, Neural Magic
, Engineering Lead, Neural Magic
Level of Expertise: AdvancedSession Tracks: vLLMSession Type: Breakout Session- Tuesday, Oct 12:30 PM - 3:00 PM PDTYerba Buena Salon 10
- Breakout SessionRay Use CasesIntermediateYesOctober 1: Conference2:00 p.m. Tuesday, Oct 01Tuesday, Oct 01Hinge exists to inspire intimate connection to create a less lonely world. We strive to empower users on their dating journey to build connection. The ML platform team is on a mission to platformize and democratize ML capabilities within the company. This talk will explore Ray's core role in Hinge's ML platform. We will discuss how Ray's malleability made it an easy slot-in to our existing infrastructure. At the same time, we will discuss how that same flexibility has enabled it to rise to the challenge of significantly driving down the time-to-production of ML-based solutions at Hinge.
, Staff Machine Learning Engineer, Hinge
, Staff Machine Learning Engineer, Hinge
Level of Expertise: IntermediateSession Tracks: Ray Use CasesSession Type: Breakout Session- Tuesday, Oct 12:30 PM - 3:00 PM PDTYerba Buena Salon 4
- Breakout SessionRay Deep DivesIntermediateOctober 1: Conference2:00 p.m. Tuesday, Oct 01Tuesday, Oct 01The emergence of Generative AI has been a transformative force in how enterprises are adopting and deploying AI. For AI Platform leaders, providing the right guardrails include building a strong security foundation, flexible controls, and integrations, and they are often at odds with developer productivity. This session will focus on sustainable AI practices for the modern enterprise with Anyscale’s suite of governance features–providing the security, controls, observability, and automation required at scale.
, Product Manager, Anyscale
, Software Engineer, Anyscale
Level of Expertise: IntermediateSession Tracks: Ray Deep DivesSession Type: Breakout Session- Tuesday, Oct 12:30 PM - 3:00 PM PDTYerba Buena Salon 14
- Breakout SessionLLM & GenAIIntermediateYesOctober 1: Conference2:00 p.m. Tuesday, Oct 01Tuesday, Oct 01LLMs has unlocked new capabilities and applications; however, its evaluation still poses significant challenges. We introduce Chatbot Arena, an open platform for evaluating AI based on human preferences. Our methodology employs a pairwise comparison approach and leverages input from a diverse user base through crowdsourcing. The platform has been operational for over an year, collecting over 1.7 million user votes. Chatbot Arena has emerged as one of the most referenced LLM leaderboards, widely cited by leading LLM developers and companies. Our demo is publicly available at https://lmarena.ai
, UC Berkeley
Level of Expertise: IntermediateSession Tracks: LLM & GenAISession Type: Breakout Session- Tuesday, Oct 12:30 PM - 3:00 PM PDTGolden Gate Salon B
- Breakout SessionRay Deep DivesIntermediateOctober 1: Conference2:00 p.m. Tuesday, Oct 01Tuesday, Oct 01Ray Core has come a long way since its first major release back in 2020. Today, we are investing in Ray Core scalability and reliability, while also evolving Ray for an accelerator-native world. Hear about some of the improvements we’ve made since last year’s summit, including streaming generators and weekly Ray releases. Then, we’ll discuss Ray Compiled Graphs, a new backend for Ray Core that makes it simple to build high-performance applications on clusters of accelerators. We’ll do a deep dive of the key Compiled Graphs features, including faster Ray task execution and native support for GPU-GPU communication via NCCL. We’ll also give a brief overview of some of the main use cases in multi-GPU LLM inference and distributed training.
, Assistant Professor at University of Washington / Software Engineer, Anyscale
, Software Engineer, Anyscale
Level of Expertise: IntermediateSession Tracks: Ray Deep DivesSession Type: Breakout Session- Tuesday, Oct 12:30 PM - 3:00 PM PDTGolden Gate Salon C1
- Breakout SessionAI/ML Platform & ApplicationsBeginnerYesOctober 1: Conference2:00 p.m. Tuesday, Oct 01Tuesday, Oct 01eBay's AI platform has undergone a significant transformation aimed at accelerating AI application development, and optimizing expensive GPU resources, and streamlining the user experience. The integration of Ray into our ecosystem has been pivotal in addressing these challenges, leading to a more agile and efficient AI workflow.
In this talk, we will outline the journey of integrating Ray into eBay's AI platform. We will discuss the application of Ray in four key production scenarios: batch inference, near-real-time (NRT) feature engineering pipelines, distributed training, and comprehensive AI solutions. For each of these production scenarios, we will talk about the specific problems we faced, the practical solutions we found and the lessons we learned along the way.
Concluding the talk, we will explore the product design improvements driven by Ray. A well-crafted product design is crucial for providing users with a seamless experience in managing their assets and utilizing the platform throughout the entire machine learning lifecycle., Sr. Machine Learning Engineer, eBay
, Senior Machine Learning Engineer, eBay
Level of Expertise: BeginnerSession Tracks: AI/ML Platform & ApplicationsSession Type: Breakout Session- Tuesday, Oct 12:30 PM - 3:00 PM PDTGolden Gate Salon A
- Breakout SessionRay Deep DivesIntermediateYesOctober 1: Conference2:00 p.m. Tuesday, Oct 01Tuesday, Oct 01Ray Serve and NVIDIA Triton Inference Server are two popular open-source inference serving solutions with unique capabilities. While Ray Serve is Python-first and allows for multi-process model composition and scaling, Triton Inference Server is primarily C / C++ and focuses on in-process model ensembles and optimized framework backends. Working closely together the maintainers of Ray Serve and Triton Inference Server collaborated to give users the best of both worlds. Starting in early 2024, Triton Inference Server includes a Python API that allows developers to seamlessly embed Triton Inference Server inside their Python applications running on Ray Serve. For Ray Serve users, this allows them to improve the performance of their ML models as can be demonstrated with our stable diffusion demo. It also allows users to leverage advanced inference serving optimization and analysis tools like Performance and Model Analyzer. Users can use these to find optimal model configurations based on their application throughput and latency requirements. Triton Inference Server users can now leverage Ray Serve to build, load balance and scale complex multi-process applications all with the ease-of-use of Python. In this talk we’ll demonstrate how Ray Serve users can take models optimized for Triton Inference Server and quickly deploy and scale them within a Ray cluster.
We’ll discuss the challenges and opportunities in combining these Open Source projects to meet the demands of Generative AI applications., Engineering Manager, Anyscale
, Senior Software Engineer, NVIDIA
, Principal Software Architect, NVIDIA
Level of Expertise: IntermediateSession Tracks: Ray Deep DivesSession Type: Breakout Session- Tuesday, Oct 12:30 PM - 3:00 PM PDTGolden Gate Salon C3
- Breakout SessionRay Use CasesAdvancedYesOctober 1: Conference2:00 p.m. Tuesday, Oct 01Tuesday, Oct 01Large Language Models (LLM) achieve tremendous success in natural language and multimodal tasks. Reinforcement Learning from Human Feedback (RLHF) is the key step to prevent LLMs from producing harmful and toxic contents. Training LLMs requires complicated model and data parallelism strategies. Programming RL dataflows requires flexible abstractions. Existing Ray-based RL frameworks (RLLib and RLlib-Flow) fail to train large models and existing RLHF frameworks fail to be flexible for conducting research.
The lack of flexible and efficient training infrastructures to easily experiment with RL algorithms enforces researchers and practitioners to seek alternative approaches such as Direct Preference Optimization (DPO) with less system complexity but potentially lower algorithmic performance.
To tackle the challenges posed by implementing RLHF systems, we propose veRL, a framework of flexible and efficient programming abstractions for RLHF research and production. We observe that most existing LLM training/inference infrastructures (e.g. DeepSpeed, FSDP, Megatron-LM) adopt the SPMD (Single Process Multiple Data) programming paradigm. To enable them in the Ray ecosystem, we create "WorkerGroup" that allows encapsulation of LLM training/inference into primitive APIs that can be invoked by the Ray driver program. Each component of the RLHF system including rollout, actor, critic, reward model and reference policy is then implemented as a WorkerGroup with several primitive APIs. veRL is efficient as existing LLM training/inference infrastructures can be easily incorporated. veRL is flexible as it encapsulates LLM training/inference inside WorkerGroup and provides a single process abstraction to implement RLHF dataflows. We hope that veRL can advance RLHF research and production in the future., Research Scientist, Bytedance
, Research Scientist, Bytedance
Level of Expertise: AdvancedSession Tracks: Ray Use CasesSession Type: Breakout Session- Tuesday, Oct 12:30 PM - 3:00 PM PDTGolden Gate Salon C2
3:00 PM PDT - Tuesday, Oct 1
- Lunch/BreakOctober 1: ConferenceTuesday, Oct 013:00 p.m. Tuesday, Oct 01Light snacks and refreshments available in RaygroundSession Type: Lunch/Break
- Tuesday, Oct 13:00 PM - 3:15 PM PDTRayground
- Lightning TalkLightning TalkIntermediateYesOctober 1: ConferenceTuesday, Oct 013:00 p.m. Tuesday, Oct 01Have you ever wondered how to assert greater control over how LLMs use tools and process workflows? In this Lightning Talk, we will discuss a FlightAware prototype app which uses an LLM and LangGraph to abstract API calls for flight data. We'll investigate the use case for a natural language Chatbot interface and dive into some tools we used to prototype an intelligent search experience, including LangGraph and Streamlit. The presentation will be technical in nature and include code samples, configuration tips, and a demo of the system.
, Sr. Machine Learning Engineer, FlightAware
Level of Expertise: IntermediateSession Tracks: Lightning TalkSession Type: Lightning Talk- Tuesday, Oct 13:00 PM - 3:15 PM PDTLightning Stage
- Breakout SessionLLM & GenAIYesIntermediateOctober 1: ConferenceTuesday, Oct 013:00 p.m. Tuesday, Oct 01LLM inference is one of the most resource-intensive AI workloads today. Over the last year at Anyscale, we've focused on contributing critical advances to open source inference engines while building enterprise and production features on the Anyscale Platform. We've collaborated closely with the vLLM team to release key features such as FP8 support, chunked prefill, multi-step decoding, and speculative decoding. Together, these optimizations have improved vLLM performance by over 2x in both throughput and latency. We’ll also cover some of the optimizations we’ve made on the Anyscale Platform as well, including custom kernels, optimizations for batch inference, and accelerated large model loading for autoscaling deployments.
, Chief Technology Officer, Anyscale
, Staff Software Engineer, Anyscale
Level of Expertise: IntermediateSession Tracks: LLM & GenAISession Type: Breakout Session- Tuesday, Oct 13:15 PM - 3:45 PM PDTYerba Buena Salon 4
- Breakout SessionvLLMIntermediateYesOctober 1: ConferenceTuesday, Oct 013:00 p.m. Tuesday, Oct 01vLLM (https://github.com/vllm-project/vllm) aims to become the de-facto industry standard for serving Large Language Models. vLLM is increasingly being adopted in production and can be executed on NVIDIA GPUS, AMD GPUs, as well as custom accelerators like AWS Inferentia.
However, vLLM’s state-of-the-art performance largely depends on a number of hand-written CUDA kernels. These kernels have typically been carefully optimized for a specific GPU platform and may pose a serious obstacle to the portability of vLLM across different hardware. Open AI Triton (https://github.com/triton-lang/triton) recently emerged as a promising open-source alternative to writing custom CUDA kernels. It enables one to write kernels for execution on GPUs using simple Python code. Triton kernels have been shown to be both highly performant, as well as portable across different GPU platforms. For this reason, Triton is growing in popularity, and vLLM already includes several kernels written in Triton.
Triton comes with a built-in autotuner, which is crucial to enable performance-portability. However, using the autotuner adds a lot more overhead to the kernel launches, in addition to the just-in-time compilation. This overhead comes from the fact that, for every variation in the kernel parameters, the autotuner needs to determine which kernel version performs the best. The resulting high variance in latency is unacceptable for serving applications in production. Consequently, the Triton autotuner is usually not used in vLLM today. Yet, by not using it, the portability of the application is limited, because the performance of the Triton kernels can differ by more than one order of magnitude on different platforms.
To solve this problem, we have developed a “dejavu” mechanism for the Triton autotuner. Our goal was to let the autotuner “remember” earlier executions of the kernel, which happened before the lifetime of the current deployment. This dejavu-mechanism reduces the overhead of the Triton autotuner to zero and therefore enables the usage of the autotuner in production. In addition, the dejavu-mechanism enabled us to develop smart algorithms for exploring the range of possible autotuner configurations leading to even better performance. Our early results show that using Triton with our dejavu-autotuner results in (1) speed-ups of more than 100% for some kernels, (2) enabling competitive performance on different platforms using the same code, as well as (3) reducing the external dependencies of vLLM, future proofing vLLM further.
This talk will also include a demo of some Triton-only vLLM deployments., Postdoctoral Researcher, IBM Research
Level of Expertise: IntermediateSession Tracks: vLLMSession Type: Breakout Session- Tuesday, Oct 13:15 PM - 3:45 PM PDTYerba Buena Salon 10
- Breakout SessionLLM & GenAIYesYesOctober 1: ConferenceTuesday, Oct 013:00 p.m. Tuesday, Oct 01A huge promise for LLMs is being able to answer questions and solve tasks of arbitrary complexity over an arbitrary number of data sources. The world has started to shift from simple RAG stacks, which are mostly good for answering pointed questions, to agents that can more autonomously reason over a diverse set of inputs, and interleave retrieval and tool use to produce sophisticated outputs.
Building a reliable multi-agent system is challenging. There's a core question of developer ergonomics and production deployment - what makes sense outside a notebook setting. In this talk we outline some core building blocks for building advanced research assistants, including advanced RAG modules, event-driven workflow orchestration, and more., Co-founder/CEO, LlamaIndex
Session Tracks: LLM & GenAISession Type: Breakout Session- Tuesday, Oct 13:15 PM - 3:45 PM PDTGolden Gate Salon A
- Breakout SessionAI/ML Platform & ApplicationsIntermediateYesOctober 1: ConferenceTuesday, Oct 013:00 p.m. Tuesday, Oct 01Certain technical aspects of Ray have driven a pattern of usage which is largely single-player per cluster. Multiple users generally don't use the same cluster, one Python script generally doesn't call into multiple clusters, and multiple clusters generally can't share an actor. We propose a new system for orchestrating between remote Ray clusters (including across regions or clouds) which overcomes these restrictions to make the clusters and resources within them fundamentally multiplayer. We also introduce a lightweight cluster and service registry and UI to make these resources feel accessible and collaborative, like a Google drive for AI.
, Software Engineer, Runhouse
, CEO, Runhouse
Level of Expertise: IntermediateSession Tracks: AI/ML Platform & ApplicationsSession Type: Breakout Session- Tuesday, Oct 13:15 PM - 3:45 PM PDTGolden Gate Salon C1
- Breakout SessionLLM & GenAIIntermediateYesOctober 1: ConferenceTuesday, Oct 013:00 p.m. Tuesday, Oct 01Engineering ML solutions at Workday has always been an interesting challenge due to our strongly tenanted architecture and our position as a processor of our customers' data, not an owner. Almost every model we design needs to be trained individually on each tenant's data and training must be done in many isolated pockets of the cloud with tightly-controlled data access permissions. With LLMs, this becomes even more challenging, in no small part due to GPU scarcity! To meet these demands in a cost-efficient way, we designed a platform around PEFT techniques (e.g. LoRA) and leveraged KubeRay's autoscaling capabilities to source GPUs when needed, then release them afterwards. As this scalable architecture supports both research and (re-)training, it paves a seamless path to production and broadens access to custom LLM-based features across the whole of the Workday product.
In this talk, we'll discuss how we use Ray at all scales to support a solution that's extremely flexible in how it's deployed, making full-stack development as accessible as full-scale production. We'll cover our LLM platform architecture and dive into lessons we've learned about building robust systems that idle with (nearly) no compute, but can easily scale to meet ephemeral demands for expensive resources like GPUs., Machine Learning Engineer, Workday
Level of Expertise: IntermediateSession Tracks: LLM & GenAISession Type: Breakout Session- Tuesday, Oct 13:15 PM - 3:45 PM PDTGolden Gate Salon C3
- Breakout SessionRay Use CasesBeginnerYesOctober 1: ConferenceTuesday, Oct 013:00 p.m. Tuesday, Oct 01The number of operational satellites in Earth’s orbit has grown exponentially, from fewer than 1,400 in 2015 to nearly 10,000 today, and more than 100K satellites are anticipated by 2030. Scaling decision-making processes and realistically modeling interactions among spacecraft, each with their own mission constraints and capabilities, is a central challenge to managing the space domain. To evaluate potential courses of action, we are training reinforcement learning (RL) agents to achieve operationally relevant spaceflight mission objectives in a high-fidelity synthetic space environment. We will discuss how we adapted our high-fidelity astrodynamics simulation engine, which was initially designed for local execution, into a near-infinitely scalable Ray-based training pipeline in this presentation. We developed a modular, multi-layered, multi-agent, multi-mission RL Gym that allows agents to train in a variety of simulated conditions. We will demonstrate how the approach enables general-purpose pre-training using low-fidelity dynamics to learn approximate high-level behaviors, followed by fine-tuning in higher-fidelity environments for specific applications. We illustrate how RL agents trained in this environment are able to perform sophisticated inter-satellite operations.
, Machine Learning Architect, Slingshot Aerospace
Level of Expertise: BeginnerSession Tracks: Ray Use CasesSession Type: Breakout Session- Tuesday, Oct 13:15 PM - 3:45 PM PDTGolden Gate Salon C2
- Breakout SessionAI/ML Platform & ApplicationsYesIntermediateYesOctober 1: ConferenceTuesday, Oct 013:00 p.m. Tuesday, Oct 01In this talk, we will explore how Reddit has invested in Ray and KubeRay as core technologies within its internal ML Platform to scale ML training and serving workloads. We will dive into Reddit’s tooling and workflows around Ray and Kuberay to vastly improve developer velocity and scale ML training workloads to achieve a 6x reduction in model training time. We will also discuss the team’s work to extend our inference platform using Ray Serve to support scalable development, deployment, and serving of fine-tuned open-source LLMs.
, Principal Engineer, Reddit
Level of Expertise: IntermediateSession Tracks: AI/ML Platform & ApplicationsSession Type: Breakout Session- Tuesday, Oct 13:15 PM - 3:45 PM PDTGolden Gate Salon B
- Lightning TalkLightning TalkBeginnerYesOctober 1: ConferenceTuesday, Oct 013:00 p.m. Tuesday, Oct 01Discover how City Storage Systems has transformed its machine learning infrastructure by adopting Daft, seamlessly integrating it with our Ray-based platform. Previously managed on separate Spark clusters, our data processing and ETL tasks now leverage Daft’s intuitive DataFrame interface, which matches and even surpasses PySpark’s capabilities. This integration has streamlined operations across both individual and cluster environments, enabling a more unified and efficient workflow.
In this talk, we will showcase practical use cases from CloudKitchens where Daft has driven significant workflow improvements. Learn how this strategic shift to a consolidated Ray and Daft environment supports our end-to-end machine learning pipelines, from data engineering to model development., Machine Learning Engineer, City Storage Systems
Level of Expertise: BeginnerSession Tracks: Lightning TalkSession Type: Lightning Talk- Tuesday, Oct 13:45 PM - 4:00 PM PDTLightning Stage
4:00 PM PDT - Tuesday, Oct 1
- Breakout SessionLLM & GenAIIntermediateOctober 1: ConferenceTuesday, Oct 014:00 p.m. Tuesday, Oct 01To advance generative AI beyond text-based large language models to video comprehension, the industry and its research must focus on meticulous data curation, filtering, annotation, and high-quality data selection from millions of videos for downstream fine-tuning tasks. These tasks span practical implementations in fields like robotics, digital twins, and autonomous vehicles. This talk will highlight innovations and research across the entire AI lifecycle stack for building a multimodal video foundation model. In this session, you will learn about key components in an end-to-end video curation service orchestrated by Ray Data and running on NVIDIA DGX Cloud. Key insights from NVIDIA’s team will be shared—covering aspects that enable developers and experts to leverage and enhance state-of-the-art AI models for large-scale inference and the creation of high-quality video datasets.
, Nvidia
Level of Expertise: IntermediateSession Tracks: LLM & GenAISession Type: Breakout Session- Tuesday, Oct 14:00 PM - 4:30 PM PDTYerba Buena Salon 14
- Breakout SessionRay Deep DivesYesIntermediateYesOctober 1: ConferenceTuesday, Oct 014:00 p.m. Tuesday, Oct 01Pinterest's use of machine learning, specifically recommender systems, to power products like Homefeed, Related Pins, and Ads is a key component of our success. Our ML Platform team manages thousands of training jobs, utilizing massive amounts of GPUs and petabytes of training data. Unlike generative style models, recommender models are data-intensive, making Ray a natural fit in our platform, to provide a distributed heterogeneous runtime to efficiently parallelize data processing and improve training throughput. By leveraging Ray and the Ray Data ecosystem, we have optimized our data loading for recommender model training at Pinterest.
In our presentation, we will share details of how we decompose and orchestrate these workloads using Ray Data. We will also share the bottlenecks we observed such as memory pinning, multi-threaded collate and the optimizations we have implemented to scale out data-loading outside of the trainer nodes. Combined these have resulted in a significant increase in training throughput. Additionally, we will also discuss some new challenges we encountered and how we are overcoming them by developing internal abstractions and contributing to the open-source community., Sr. Staff Software Engineer, Pinterest
, Sr. Software Engineer, Pinterest
Level of Expertise: IntermediateSession Tracks: Ray Deep DivesSession Type: Breakout Session- Tuesday, Oct 14:00 PM - 4:30 PM PDTGolden Gate Salon A
- Breakout SessionAI/ML Platform & ApplicationsYesAdvancedOctober 1: ConferenceTuesday, Oct 014:00 p.m. Tuesday, Oct 01The generative AI revolution has transformed the world of large-scale deep learning infrastructure. Modern machine learning platforms must be ready to support pre-training for massive foundation models, memory-intensive fine-tuning for LLMs and diffusion models, as well as low-latency deployments for multi-billion-parameter models. Navigating this emerging landscape requires new techniques and methodologies, leavened with a thorough understanding of the still-nascent GenAI tooling ecosystem. In this talk, we'll walk through how we've adapted and extended Netflix's production Ray platform to deal with these new challenges. We'll outline our experiences in deploying Ray for large-scale data processing & curation, LLM fine-tuning, foundation model training, and distributed inference. We'll also share our strategies for creating GenAI-ready ML infrastructure to support the needs of the modern deep learning landscape.
, Machine Learning Engineer, Netflix
Level of Expertise: AdvancedSession Tracks: AI/ML Platform & ApplicationsSession Type: Breakout Session- Tuesday, Oct 14:00 PM - 4:30 PM PDTGolden Gate Salon C1
- Breakout SessionRay Use CasesIntermediateYesOctober 1: ConferenceTuesday, Oct 014:00 p.m. Tuesday, Oct 01Genmo trains some of the state-of-the-art models for text-to-video generation. Every day, our models enable 1M+ users to generate video with AI. In this presentation, Genmo CEO and co-founder Paras Jain will walk through technical platforms used to develop Genmo’s video generation models and scale them to our large user base with extreme cost efficiency. He also will discuss how distributed computing technologies like Ray Data have helped the team scale up training to our current large parameter models.
The audience will leave with an understanding of how they can leverage cloud infrastructure and optimize server usage in their model development. They’ll also learn cost-saving measures that work without compromising model quality or user experience., CEO & Co-Founder, Genmo
Level of Expertise: IntermediateSession Tracks: Ray Use CasesSession Type: Breakout Session- Tuesday, Oct 14:00 PM - 4:30 PM PDTGolden Gate Salon B
- Breakout SessionvLLMIntermediateOctober 1: ConferenceTuesday, Oct 014:00 p.m. Tuesday, Oct 01As large language models emerge, multi-GPU inference becomes a necessary requirement for model serving libraries. Due to the dynamic nature of distributed inference, it poses a different set of challenges from distributed training. The talk covers differences between distributed train and inference, how different parallelism strategies such as tensor parallelism, pipeline parallelism, expert parallelism, work in detail, and how to build an optimized architecture for a fast distributed inference engine, with vLLM as an example.
, Centml
, Software Engineer, Anyscale
Level of Expertise: IntermediateSession Tracks: vLLMSession Type: Breakout Session- Tuesday, Oct 14:00 PM - 4:30 PM PDTYerba Buena Salon 10
- Breakout SessionRay Use CasesBeginnerYesOctober 1: ConferenceOctober 2: ConferenceTuesday, Oct 014:00 p.m. Tuesday, Oct 01Rad AI has been pioneering GenAI applications in healthcare to reduce physician burnout, increase efficiency, and improve patient care since 2018. In that light, we will be presenting on the following topics: (1) an overview of LLMs in the context of healthcare, (2) Rad AI's LLM tailored for radiology and follow-up management workflows, and (3) Rad AI’s Ray-based training infrastructure that provides our researchers with experimental environments and enables our production-grade, multi-node training pipelines. This discussion will make mention of Ray Clusters, Ray Train, Pulumi for infrastructure-as-code (IaC), the Google Kubernetes Engine (GKE), and Kuberay. At the end, you will be left with a survey of GenAI in healthcare alongside some of the specifics behind Rad AI’s Ray-based training infrastructure.
, VP of Engineering, Rad AI
, Senior MLOps Engineer, Rad AI
Level of Expertise: BeginnerSession Tracks: Ray Use CasesSession Type: Breakout Session- Tuesday, Oct 14:00 PM - 4:30 PM PDTGolden Gate Salon C3
- Breakout SessionRay Use CasesYesIntermediateYesOctober 1: ConferenceTuesday, Oct 014:00 p.m. Tuesday, Oct 01In this talk, we explore the integration and utilization of Ray within Roblox's machine learning platform to handle large-scale batch inference jobs effectively for deep learning models and LLMs, leading to a significant cost reduction. As an immersive platform for communication and connection, Roblox requires robust, scalable solutions to manage and deploy machine learning models efficiently across various use cases. We will delve into the specific applications of Ray at Roblox, highlighting its pivotal role in enhancing throughput and reliability in model batch inference scenarios. Our discussion will cover the architectural design decisions that led to the adoption of KubeRay, showcasing how it fits into our Kubernetes-managed environments to bolster our machine learning operations.
, Senior Software Engineer, Roblox
, Software Engineer, Roblox
, Principal Machine Learning Engineer, Roblox
Level of Expertise: IntermediateSession Tracks: Ray Use CasesSession Type: Breakout Session- Tuesday, Oct 14:00 PM - 4:30 PM PDTYerba Buena Salon 4