Keep pulling the thread on Linda Haviv.
For AI/ML workloads, frameworks like Ray are needed to distribute compute loads and integrate with tools like VLLM, whereas Spark was sufficient for other, older use cases.
Neoclouds are specialists that are building out infrastructure specifically for AI, whereas traditional cloud providers are generalists.
Neocloud providers such as Nibus, CoreWeave, and Cruzo are emerging to solve the rapidly growing problems in AI infrastructure.
A key challenge in AI infrastructure is to avoid 'starving your GPUs,' which means ensuring they are not idle and wasting money due to preprocessing bottlenecks.
The Ray framework is built specifically for AI/ML workloads because it is Python-native, which is the language most machine learning is built in.
The Ray framework is designed to function with GPUs in a way that reduces wasted cost by more efficiently filling the workload.
AI/ML workloads often require direct bare-metal access, which has led to the emergence of neoclouds like Lightning AI.
The open-source framework VLLM was created to address challenges specific to AI inference.
From a user's perspective, AI infrastructure challenges are most commonly encountered first during the inference stage.
Modern AI workloads are significantly more compute-heavy compared to past workloads, which were more reliant on data input/output.