The Ray ecosystem of libraries (Ray Serve/Tune/Train/Data etc) already allow your Ray cluster to provide best-in-class ML/AI capabilities. Daft is a Python/Rust library that unlocks distributed ETL and analytics capabilities on your Ray cluster, allowing Ray to truly provide an end-to-end Data + ML/AI solution at any scale.
In this lightning talk, we will demonstrate how easy it is to get Daft running on a Ray cluster (in AnyScale!). We share results on how Daft-on-Ray even outperforms existing solutions, and show key integration points such as:
1. How Daft leverages the Ray object store and spilling to process larger-than-memory datasets
2. Seamless integrations with the rest of the Ray ecosystem using zero-copy data transfer into Ray Data
3. An end-to-end example of data exploration, cleaning, processing, and ML/AI training with Ray and Daft
In this lightning talk, we will demonstrate how easy it is to get Daft running on a Ray cluster (in AnyScale!). We share results on how Daft-on-Ray even outperforms existing solutions, and show key integration points such as:
1. How Daft leverages the Ray object store and spilling to process larger-than-memory datasets
2. Seamless integrations with the rest of the Ray ecosystem using zero-copy data transfer into Ray Data
3. An end-to-end example of data exploration, cleaning, processing, and ML/AI training with Ray and Daft
