Keep pulling the thread on Aakanksha Chowdhery.
Akansha Chowdhury predicts that within the next year, more AI benchmarks will be developed that are representative of real-world workflows.
LLMs have only recently begun to show improved performance on multi-hop reasoning benchmarks like MRCR V2 and LOFT, which previously showed weak results even for very long-context models.
Reflection recently closed its founding funding round with financial backing from NVIDIA.
Reflection's research team includes former employees from DeepMind and OpenAI, as well as individuals who made leading contributions to projects like PaLM, ChatGPT, Gemini, AlphaProof, and AlphaGo.
Google's PaLM model had 540 billion parameters, and after its release, companies generally stopped publishing parameter counts for their large language models.
The mission of the company Reflection is to build frontier open intelligence for agentic capabilities.
Following its most recent fundraising, Reflection is now building its Frontier Open Agentic models end-to-end, handling both pre-training and post-training in-house.
The Transformer architecture has been more thoroughly vetted at scale compared to newer, alternative model architectures.
Models from China, such as Qwen, DeepSeek, and Kimi, demonstrate that the quality and curation of training data are highly important for model performance.
The DeepSeek model was trained on 15 trillion tokens of data.
Akansha Chowdhury predicts that if a model's training data is changed to be substantially composed of synthetically generated data, the model's performance will degrade.
During the training of Google's PaLM model, a step-change in performance on a new, crowdsourced reasoning benchmark revealed the model's emergent reasoning capabilities.