Keep pulling the thread on Vishal Misra.
The Transformer architecture enables Large Language Models to perform in-context learning through a process that is mathematically equivalent to Bayesian updating.
In a 'Bayesian wind tunnel' experiment, the Transformer architecture computed the precise Bayesian posterior for a given task with an accuracy of 10^-3 bits.
Current LLMs based on architectures like the Transformer have frozen weights post-training, preventing them from retaining learning between inference sessions.
Current deep learning models are limited to finding correlations and cannot perform causation-based reasoning like intervention or counterfactuals, as defined in Judea Pearl's causal hierarchy.
Simply scaling up current AI architectures will not solve all problems; a different kind of architecture is needed for further breakthroughs.
Dario Amodei allegedly stated that it is not possible to rule out that large language models are conscious.
An early implementation of Retrieval-Augmented Generation (RAG) was developed in October 2020 using GPT-3 and was deployed in production at ESPN in September 2021.
In 'Bayesian wind tunnel' experiments, the Mamba architecture performs Bayesian updating reasonably well, LSTMs perform it only partially, and MLPs fail completely.
An AI model's ability to perform Bayesian updating is a function of its architecture, such as the Transformer, rather than the specific data it is trained on.
Products from Anthropic, such as Claude Code, are based on matrix multiplication and do not have consciousness or an inner monologue.
A valid test for Artificial General Intelligence (AGI) would be to train a large language model on physics knowledge from before 1911 and determine if it can independently derive the theory of relativity.
The Global Positioning System (GPS) relies on the equations from the theory of relativity to function correctly.