In this talk, Klaviyo’s ML platform team presents insights gained from over a year of using Ray Serve for their model serving platform - DART (DAtascience RunTime). We talk about how Ray Serve provides the best combination of flexibility, simplicity and high availability among all the options we considered. Klaviyo’s DART platform is self-serve, allowing Data Scientists to deploy new models quickly, reducing the time it takes for them to go to production from a few weeks to a few days.
We will explore certain internal details of Ray Serve architecture which have been crucial for us to understand certain behaviors and debug issues such as internal request routing and fault tolerance with KubeRay. Additionally, we explore DART’s multi-cluster, multi-AZ architecture that ensures resilience in cases of node failures and pod evictions.
We conclude with important considerations for any teams trying to build their own model serving platform with Ray Serve, including things like cluster sizing, traffic isolation and avoiding common pitfalls such as single point of failures. Finally, we walk through different architectures with increasing levels of complexity so organizations can choose the architecture that best fits their cost and high availability needs. With these insights, we aim to empower others to build their own robust serving platform using Ray Serve.
We will explore certain internal details of Ray Serve architecture which have been crucial for us to understand certain behaviors and debug issues such as internal request routing and fault tolerance with KubeRay. Additionally, we explore DART’s multi-cluster, multi-AZ architecture that ensures resilience in cases of node failures and pod evictions.
We conclude with important considerations for any teams trying to build their own model serving platform with Ray Serve, including things like cluster sizing, traffic isolation and avoiding common pitfalls such as single point of failures. Finally, we walk through different architectures with increasing levels of complexity so organizations can choose the architecture that best fits their cost and high availability needs. With these insights, we aim to empower others to build their own robust serving platform using Ray Serve.
