The AI industry's primary performance bottleneck is not computation speed but the energy-intensive process of data movement, which accounts for approximately 70% of a chip's energy consumption.
Current AI software systems are profoundly inefficient, achieving only 20-30% of their theoretical peak performance and suffering from extremely low GPU utilization, especially for inference workloads (under 15%).
The focus on optimizing individual, component-level kernels is misguided and often masks or even exposes greater system-level latencies and inefficiencies.
In the fast-paced AI market, "time to production" is the single most important competitive advantage, a factor currently undermined by long optimization and deployment cycles.
The industry must adopt more holistic and business-relevant metrics like "useful intelligence per joule" and "useful output per dollar" to accurately measure performance and drive meaningful progress.
Approximately 8 years ago
Dawani references this period as a benchmark for AI progress, noting that training a 20 billion parameter model required thousands of GPUs and several months, highlighting the scale of historical challenges (Claim 4).
Prior to Lemurian Labs
Dawani advised on NASA's Mars Rover program while at Geometric Energy, indicating a background in complex, high-stakes engineering systems before founding his current company (Claim 1).
Historical Industry Focus
Dawani characterizes the historical approach of the AI software industry as being focused on optimizing arithmetic operations, which he argues was a misallocation of effort away from the larger problem of data movement (Claim 9).
Current State
Dawani describes the current AI landscape as deeply inefficient, with low hardware utilization, a widening software-hardware gap, and months-long delays between model training and production deployment, framing the problem his company aims to solve (Claims 3, 12, 13, 16).
Future
Dawani predicts a shift towards right-sizing workloads to specific hardware to maximize efficiency metrics like intelligence per joule or per dollar, outlining his vision for the next era of AI infrastructure (Claim 5).
▶Systemic Inefficiency in AI ComputeJun 2026
Dawani argues that the current AI software stack is fundamentally inefficient, resulting in vast underutilization of expensive hardware. He claims that even frontier labs only achieve 45-50% GPU utilization for training, while inference workloads are often below 15%, and overall systems only reach 20-30% of their theoretical peak performance (Claims 3, 10, 12).
This theme suggests a significant market opportunity for software solutions that can unlock the latent potential of existing and future hardware, potentially disrupting the current belief that progress is solely dependent on manufacturing more powerful chips.
▶The Primacy of Data Movement over ComputationJun 2026
A central tenet of Dawani's technical argument is that the industry has myopically focused on optimizing arithmetic (10-15% of energy use) while ignoring data movement, which consumes ~70% of a chip's energy. He states that moving data is about 100 times more energy-intensive than the computation itself, making it the true bottleneck (Claims 9, 21, 22).
Investors should scrutinize AI infrastructure companies based on how their technology addresses the data movement problem, as solutions focused solely on computational throughput may yield diminishing returns.
▶Redefining AI Performance and Business MetricsJun 2026
Dawani advocates for moving away from component-level benchmarks like kernel throughput, which he argues can be misleading. Instead, he proposes that business leaders adopt holistic, outcome-oriented metrics like "useful output per dollar" and "useful intelligence per joule," and asserts that "time to production" is the most critical competitive variable in the current landscape (Claims 6, 14, 15, 17).
This focus on business-relevant metrics over raw technical specs indicates a maturing market where efficiency and speed-to-deployment are becoming more important than theoretical performance, signaling a shift in customer priorities.
▶A New Compiler Paradigm for System-Level OptimizationJun 2026
Dawani presents his company's solution as a "natively parallel" compiler that optimizes for an entire hardware system, not just individual kernels. This approach is designed to tackle the 70-80% of wasted cycles in large systems spent on coordination and memory overhead, and to be easily adaptable to new hardware by simply adding a new hardware model (Claims 8, 19, 23).
Dawani is positioning Lemurian Labs not as an incremental improvement but as a fundamental architectural shift in the AI software stack, which, if successful, could create a new standard and a significant competitive moat.