Keep pulling the thread on Sid Sheth.
The AI inferencing market is expected to become the largest part of AI compute and exceed a trillion dollars in the next 5 years.
A "premium token economy" is emerging where companies charge significantly more for high-interactivity, low-latency AI, and users are willing to pay the premium.
Solutions from d-Matrix, Groq, and Cerebras have architectures that provide compute with access to memory that is an order of magnitude faster than HBM-based GPUs.
d-Matrix acquired GigaIO to accelerate the deployment of its servers and racks.
The primary customers for low-latency inference chips are the frontier AI labs, including Anthropic, OpenAI, xAI, Meta Super Intelligence, and Google DeepMind.
AI chips will be in short supply for at least the next 3 to 4 years.
d-Matrix's supply chain is more resilient because its products do not use bleeding-edge process nodes, HBM, or CoWoS technology from TSMC.
The use of AI is mandated for all employees at d-Matrix, and the company has an internal AI team focused on deploying AI across the entire organization.
The AI inference market is bifurcating and cannot be served by a single type of chip.
General-purpose GPUs are not very efficient for AI inference and will not perform well in market segments that require highly optimized metrics for specific applications.
d-Matrix partners with Supermicro to build and deploy its racks, rather than manufacturing them in-house.
d-Matrix's primary unit of sale is a card or a tray, whereas many of its competitors sell a full rack as their unit of sale.