Keep pulling the thread on Philip Kelly.
Google Cloud will be one of the cloud providers to offer NVIDIA's next-generation Vera Rubin hardware for inference and training.
NVIDIA's Vera Rubin hardware is expected to be available on Google Cloud in the latter half of the current year.
Google Cloud is adding NVIDIA's Blackwell GPUs, specifically the RTX Pro 6000 model, to its offerings.
The NVIDIA RTX Pro 6000 GPU features 96 gigabytes of VRAM, which enables the deployment of multiple models on a single GPU.
Baseten was one of the first companies to use NVIDIA Dynamo for inference in a production environment.
Migrating an inference system from NVIDIA's Hopper architecture to the Blackwell architecture requires significant software and kernel layer work.
Baseten utilizes Google Kubernetes Engine (GKE) to build its systems on top of Google Cloud GPUs.
Baseten operates a multi-region deployment that pools GPUs from America and Europe into a unified compute pool to minimize user latency.
Baseten was a day-zero support partner for the launch of Google's Gemma 4 model.
The Gemma family of models from Google supports native image inputs, which is beneficial for enterprise use cases like KYC and document extraction.
The GPT OSS 120B model is considered too large for many use cases, making smaller models like the Gemma family more suitable for fine-tuning.
NVIDIA's TensorRT LLM is an open-source SDK that can optimize inference performance on NVIDIA hardware with only a few lines of code.