The vLLM team will brief the past year of progress in "building the fastest and easiest-to-use open-source LLM inference and serving engine". We are excited to share major updates in terms of adoption, features, performance, community, and the governance of the project. We will then share the roadmap for the upcoming releases.
Tue, Oct 01, 5 sessions
1:00 PM PDT - Tuesday, Oct 1
- Breakout SessionvLLMYesBeginnerOctober 1: Conference1:00 p.m. Tuesday, Oct 01Tuesday, Oct 01
, vLLM @Berkeley, University of Chicago
, PhD Student / vLLM Maintainer, UC Berkeley
Level of Expertise: BeginnerSession Tracks: vLLMSession Type: Breakout Session- Tuesday, Oct 11:00 PM - 1:30 PM PDTYerba Buena Salon 10
- Breakout SessionvLLMIntermediateYesOctober 1: Conference1:00 p.m. Tuesday, Oct 01Tuesday, Oct 01With the recent rapid advancements in multimodal language models, there has been a surge of interest from the open source vLLM community in the support of these models on vLLM.
In this talk, we will explore the journey of integrating multimodal models into vLLM, delving into the technical challenges encountered and the valuable lessons learned throughout the process., Software Engineer, Roblox
Level of Expertise: IntermediateSession Tracks: vLLMSession Type: Breakout Session- Tuesday, Oct 11:45 PM - 2:15 PM PDTYerba Buena Salon 10
2:00 PM PDT - Tuesday, Oct 1
- Breakout SessionvLLMAdvancedYesOctober 1: Conference2:00 p.m. Tuesday, Oct 01Tuesday, Oct 01vLLM supports various forms of model quantization including FP8, INT8, and INT4, reducing memory consumption and increasing generation speed. In this talk, you will learn how vLLM accelerates models with quantization internally and how to apply these methods to your model with vLLM’s llm-compressor framework.
, Sr Director of Engineering, Neural Magic
, Engineering Lead, Neural Magic
Level of Expertise: AdvancedSession Tracks: vLLMSession Type: Breakout Session- Tuesday, Oct 12:30 PM - 3:00 PM PDTYerba Buena Salon 10
3:00 PM PDT - Tuesday, Oct 1
- Breakout SessionvLLMIntermediateYesOctober 1: ConferenceTuesday, Oct 013:00 p.m. Tuesday, Oct 01vLLM (https://github.com/vllm-project/vllm) aims to become the de-facto industry standard for serving Large Language Models. vLLM is increasingly being adopted in production and can be executed on NVIDIA GPUS, AMD GPUs, as well as custom accelerators like AWS Inferentia.
However, vLLM’s state-of-the-art performance largely depends on a number of hand-written CUDA kernels. These kernels have typically been carefully optimized for a specific GPU platform and may pose a serious obstacle to the portability of vLLM across different hardware. Open AI Triton (https://github.com/triton-lang/triton) recently emerged as a promising open-source alternative to writing custom CUDA kernels. It enables one to write kernels for execution on GPUs using simple Python code. Triton kernels have been shown to be both highly performant, as well as portable across different GPU platforms. For this reason, Triton is growing in popularity, and vLLM already includes several kernels written in Triton.
Triton comes with a built-in autotuner, which is crucial to enable performance-portability. However, using the autotuner adds a lot more overhead to the kernel launches, in addition to the just-in-time compilation. This overhead comes from the fact that, for every variation in the kernel parameters, the autotuner needs to determine which kernel version performs the best. The resulting high variance in latency is unacceptable for serving applications in production. Consequently, the Triton autotuner is usually not used in vLLM today. Yet, by not using it, the portability of the application is limited, because the performance of the Triton kernels can differ by more than one order of magnitude on different platforms.
To solve this problem, we have developed a “dejavu” mechanism for the Triton autotuner. Our goal was to let the autotuner “remember” earlier executions of the kernel, which happened before the lifetime of the current deployment. This dejavu-mechanism reduces the overhead of the Triton autotuner to zero and therefore enables the usage of the autotuner in production. In addition, the dejavu-mechanism enabled us to develop smart algorithms for exploring the range of possible autotuner configurations leading to even better performance. Our early results show that using Triton with our dejavu-autotuner results in (1) speed-ups of more than 100% for some kernels, (2) enabling competitive performance on different platforms using the same code, as well as (3) reducing the external dependencies of vLLM, future proofing vLLM further.
This talk will also include a demo of some Triton-only vLLM deployments., Postdoctoral Researcher, IBM Research
Level of Expertise: IntermediateSession Tracks: vLLMSession Type: Breakout Session- Tuesday, Oct 13:15 PM - 3:45 PM PDTYerba Buena Salon 10
4:00 PM PDT - Tuesday, Oct 1
- Breakout SessionvLLMIntermediateOctober 1: ConferenceTuesday, Oct 014:00 p.m. Tuesday, Oct 01As large language models emerge, multi-GPU inference becomes a necessary requirement for model serving libraries. Due to the dynamic nature of distributed inference, it poses a different set of challenges from distributed training. The talk covers differences between distributed train and inference, how different parallelism strategies such as tensor parallelism, pipeline parallelism, expert parallelism, work in detail, and how to build an optimized architecture for a fast distributed inference engine, with vLLM as an example.
, Centml
, Software Engineer, Anyscale
Level of Expertise: IntermediateSession Tracks: vLLMSession Type: Breakout Session- Tuesday, Oct 14:00 PM - 4:30 PM PDTYerba Buena Salon 10