As large language models emerge, multi-GPU inference becomes a necessary requirement for model serving libraries. Due to the dynamic nature of distributed inference, it poses a different set of challenges from distributed training. The talk covers differences between distributed train and inference, how different parallelism strategies such as tensor parallelism, pipeline parallelism, expert parallelism, work in detail, and how to build an optimized architecture for a fast distributed inference engine, with vLLM as an example.
Ray Summit 2024
In-Person Agenda
READY TO REGISTER?
Come connect with the global community of thinkers and disruptors who are building and deploying the next generation of AI and ML applications.
Join the Conversation
Hashtag it
#RaySummitDon't wait for the conference to get the convo going. Join the Ray community now – ask a question in the forums, open a pull request or simply share why you’re excited. Create some buzz!
