Keep pulling the thread on Yann LeCun.
The concentration of power through proprietary AI systems is a greater danger than other perceived risks in AI.
Restricting access to AI systems for security reasons would lead to a future where a small number of companies control the public's information diet via proprietary systems.
Autoregressive Large Language Models are not the path to achieving superhuman intelligence.
Large Language Models (LLMs) are incapable of understanding the physical world, having persistent memory, reasoning, and planning, or can only perform these functions in a very primitive way.
Autoregressive LLMs are missing essential components required to achieve human-level intelligence.
It is not possible to build a complete world model by only predicting words because language is too low-bandwidth.
Attempts to use latent variable models, including GANs and VAEs, to predict future video frames by generating pixels have been a "complete failure."
Self-supervised learning methods based on reconstructing images or video from corrupted versions have failed to produce good internal representations for downstream tasks like object recognition.
V-JEPA is the first system to learn video representations that allow a supervised classifier to identify actions in the video with high accuracy.
Preliminary results indicate that V-JEPA representations can be used to determine if a video depicts physically possible or impossible events.
A world model built with a JEPA-like architecture can be used for planning, a capability that current Large Language Models lack.
No one in the AI field currently knows how to train a system to learn the multiple levels of representation required for effective hierarchical planning.