NVIDIA's market power stems from its monopsony on High Bandwidth Memory (HBM), not its CUDA software, and its future involves high revenue share but low unit volume share.
The primary bottleneck for scaling AI globally is the availability of massive energy infrastructure, particularly nuclear power, which nations like China are aggressively pursuing.
Hyperscalers develop custom AI chips not primarily for mass deployment, but as a strategic lever to negotiate better pricing and supply allocation from NVIDIA.
The demand for AI inference is so profoundly unmet that compute capacity is the main variable limiting revenue growth for major AI model providers.
The US holds a 2-3 year advantage over China in supplying AI compute to allied nations, a dynamic reflected in how each country's AI models are economically optimized (US for inference, China for training).
Market Foundation Analysis
Ross establishes his view of the market's scale, claiming that ChatGPT has reached approximately 10% of the world's population as weekly active users.
Supply Chain Diagnosis
He details the core bottlenecks in AI hardware, identifying NVIDIA's monopsony on HBM and the limited supply of interposers as the key constraints limiting its annual GPU output to around 5.5 million units.
Geopolitical Landscape Assessment
Ross outlines the global AI competition, contrasting US model optimization (inference cost) with China's (training cost) and highlighting massive national energy investments in China (150 nuclear reactors) and Japan ($65B for AI).
Competitive Positioning
He positions his company, Groq, as a nimble competitor, citing a six-month chip delivery time compared to NVIDIA's two-year lead and noting a developer base of 2.2 million.
Future Projections
Ross makes several bold forward-looking predictions, including NVIDIA's unwavering path to a $10 trillion valuation and a future where it captures most AI chip revenue but only a small fraction of the unit volume.
▶NVIDIA's Constrained DominanceFeb–Mar 2026
Jonathan Ross portrays NVIDIA not as an unassailable monopoly but as a dominant player whose growth is capped by critical supply chain bottlenecks, particularly High Bandwidth Memory (HBM), where it acts as a monopsony (single dominant buyer). He predicts NVIDIA will maintain over 50% of AI chip revenue but its share of the total number of chips sold will plummet to as low as 10% within five years.
Investors should view the HBM supply chain and the emergence of specialized, high-volume inference chips as leading indicators of potential shifts in NVIDIA's market power, even if its revenue dominance persists.
▶The Geopolitics of Compute and EnergyFeb 2026
Ross frames the global AI race as a battle for compute infrastructure and the massive energy sources required to power it. He highlights China's plan for 150 nuclear reactors and Japan's $65 billion AI fund as evidence of this linkage, while warning that Europe's economy is at risk if it fails to build sufficient energy and compute capacity.
National energy policy, especially regarding nuclear and large-scale renewables, is now a direct and critical component of a country's long-term AI strategy and future economic competitiveness.
▶The Strategic Game of Custom SiliconFeb 2026
According to Ross, the primary motivation for hyperscalers to develop their own AI chips is not to replace NVIDIA entirely but to gain strategic leverage in supply allocation and pricing negotiations. The existence of a viable in-house alternative, even if not mass-produced, fundamentally changes the power dynamic with the dominant supplier.
The success of in-house chip projects should be measured less by their deployment volume and more by their impact on a company's purchasing power and supply chain resilience against dominant vendors.
▶Inference as the Untapped Revenue FrontierFeb 2026
Ross argues that the market for AI inference is severely underserved, to the point where leading AI labs could almost double their revenue in a month if their inference compute capacity were doubled. This positions inference as a primary economic driver where speed and efficiency are paramount, challenging the industry's heavy focus on training performance.
The market may be significantly underestimating the latent commercial demand for AI inference, suggesting that companies providing faster or more efficient inference-specific hardware are poised for explosive growth.