Keep pulling the thread on Alexandre Lebrun.
Building a foundational AI model requires purchasing thousands of GPUs, which costs billions of euros.
An LLM's lack of direct world experience limits its common sense and its ability to find new solutions when facing a situation for the first time.
A world model learns directly from real-world sensory data such as video, audio, and touch, without relying on language.
Vision-Language-Action (VLA) models are a "bad hack" for robotics because they are inaccurate, slow, and too expensive for practical applications.
Once trained, world models are expected to be more lightweight and have fewer parameters than LLMs, leading to less expensive inference.
World models will be vastly superior to LLMs for high-dimensional, noisy, and long-horizon problems.
It is very difficult to secure GPUs, even for companies that have sufficient funds to purchase them.
A machine cannot truly understand a concept unless it has direct sensory experience with it, a principle known as grounding.
Within 5 to 10 years, there will be genuinely helpful robots, and most jobs will persist but become less dangerous and difficult.
The real cost of raising a large funding round like $1.2 billion is the external expectations it creates, not shareholder dilution.
If a company raises $1 billion and does not produce visible results within two years, it is very difficult for it to survive.
Mark Zuckerberg initiated the acquisition of Wit.ai in 2015 by sending a direct email to its founder.