In this session, we'll delve into unlocking the power of scalable, low-cost, high-performance model inference with Large Language Models such as Llama3 and Mistral7B on Amazon EKS. Learn how you can utilize RayServe framework and AWS Neuron (Trainium and Inferentia) instances to overcome GPU availability issues, drive cost optimization while scaling the model serving infrastructure by combining Ray with Karpenter autoscaler on Amazon EKS.
Ray Summit 2024
In-Person Agenda
READY TO REGISTER?
Come connect with the global community of thinkers and disruptors who are building and deploying the next generation of AI and ML applications.
Join the Conversation
Hashtag it
#RaySummitDon't wait for the conference to get the convo going. Join the Ray community now – ask a question in the forums, open a pull request or simply share why you’re excited. Create some buzz!

