LLM inference is one of the most resource-intensive AI workloads today. Over the last year at Anyscale, we've focused on contributing critical advances to open source inference engines while building enterprise and production features on the Anyscale Platform. We've collaborated closely with the vLLM team to release key features such as FP8 support, chunked prefill, multi-step decoding, and speculative decoding. Together, these optimizations have improved vLLM performance by over 2x in both throughput and latency. We’ll also cover some of the optimizations we’ve made on the Anyscale Platform as well, including custom kernels, optimizations for batch inference, and accelerated large model loading for autoscaling deployments.
Ray Summit 2024
In-Person Agenda
READY TO REGISTER?
Come connect with the global community of thinkers and disruptors who are building and deploying the next generation of AI and ML applications.
Join the Conversation
Hashtag it
#RaySummitDon't wait for the conference to get the convo going. Join the Ray community now – ask a question in the forums, open a pull request or simply share why you’re excited. Create some buzz!
