Keep pulling the thread on Philip Kiely.
Philip Kiely predicts that the demand for inference engineers will increase by 10 to 100 times in the next couple of years.
Due to U.S. export controls, Chinese research labs primarily work on NVIDIA Hopper GPUs, not Blackwell GPUs, which drives significant open-source software support for the Hopper architecture.
Talos created an ASIC by burning the Llama 3.1 8B model directly onto the chip, achieving an inference speed of 16,000 tokens per second.
Shopify saved tens of millions of dollars on its AI workloads by switching to a "Quen model."
In AI inference, the timeline to implement a new research technique in production is often measured in hours.
Philip Kiely believes inference is the most important and stickiest workload in the AI industry.
An engineer at Base10 implemented the PoloQuant research paper as a CUDA kernel within 31 hours of its publication.
Base10 is currently unable to hire knowledgeable inference engineers fast enough to meet its demand.
Companies with AI products at a reasonable scale are transitioning from per-token pricing models to paying for dedicated underlying GPU hardware.
The typical product maturity cycle for AI applications starts with using per-token closed models, then moves to hyperscalers like AWS or GCP due to cost or capacity issues, and finally to dedicated inference providers or in-house platforms.
NVIDIA's Ampere architecture GPUs are no longer favored for inference because they lack support for FP8 quantization, which makes Hopper GPUs significantly more cost-effective.
The rental price for an NVIDIA H100 GPU is higher now than it was a year ago.