Keep pulling the thread on Jeff Dean.
For multiple generations of Google's Gemini models, the "Flash" version of a new generation has been as good or substantially better than the "Pro" version of the previous generation.
Google uses its "Flash" models in its search products for AI Overviews.
Google is the leader in LLM context length, with a capability of 2 million tokens.
A future goal for personal AI like Gemini is to attend to a user's entire personal state, including emails, photos, and documents, which will require algorithmic and system-level improvements to handle trillions of tokens.
Google aims for Gemini to understand non-human modalities such as LiDAR data from Waymo vehicles, robotics sensor data, and medical data like X-rays, MRIs, and genomics.
The design cycle for a new ML chip like a TPU involves predicting ML computation needs 2 to 6 years in the future, with major architectural changes targeting the "N+2" generation of the chip.
In approximately 1.5 years, the mathematical capabilities of large language models have advanced from struggling with GSM8K problems to solving IMO and Erdős-level problems using only language.
The strategy for personalizing models like Gemini will likely involve using a single base model that retrieves from a user's personal data as a tool, rather than fine-tuning the model on that data.
The Gemini project at Google was initiated by a memo from Jeff Dean arguing against fragmenting resources across separate efforts in Google Research, the Brain Team, and DeepMind, and advocating for a single, unified multimodal model effort.
A personalized model that can retrieve over all of a user's opted-in personal state, including every email, photo, and video watched, will be incredibly useful compared to a generic model.
Increasingly specialized hardware will enable much lower latency and more capable AI models at affordable prices.
Future AI models and their underlying systems will achieve 20x to 50x lower latency than what is available today.