Keep pulling the thread on Adam Marblestone.
A potential issue with current AI models is their inability to represent video, audio, and text in a latent abstraction that allows for intermingling and conflict, which may explain why LLMs struggle to draw connections between different ideas.
The speaker explored the idea of multi-agent scaling, questioning whether a fixed training compute budget yields smarter agents when allocated to a single agent or distributed among multiple agents to foster diverse strategies.
Gemini autonomously implemented and benchmarked two methods for parallelizing self-play in AlphaZero, one with higher GPU context switching and another with higher communication overhead, identifying the best option in minutes.
Preliminary experimental results suggest that splitting a fixed training compute budget among multiple agents can yield gains compared to allocating it all to a single agent.
The best agent from a population of 16, receiving only 1/16th of the training compute, outperformed a single agent trained with the entire compute budget in a multi-agent scaling experiment.
The brain's steering subsystem contains significantly more diverse and bespoke cell types than the learning subsystem (cortical cell types), suggesting it encodes specific innate behaviors like flinch reflexes or taste responses.
Evolution did not amortize much into the brain's learning subsystem network, instead embedding innate behaviors and bootstrapping cost/reward functions.
Any area of the cortex might function as a general prediction engine, capable of learning to predict any subset of variables it perceives from any other subset.
The cortex might be natively designed to allow any of its areas to predict any pattern within any subset of its inputs, given any other missing subset.
Ilya Sutskever stated that he is not aware of any good theory explaining how evolution encodes high-level desires or intentions.
The field of machine learning has neglected the role of very specific and complex loss functions, unlike evolution which may have built significant complexity into the brain's loss functions.
Astera, the organization employing Steve Behrens, launched a neuroscience project based on Doris Tsao's work, which suggests building vision systems with architectural assumptions about objects, surfaces, and occlusion to reduce training requirements.