Skip to content
Sonic AI
AR

arXiv

Summary

arXiv covering Large Language Models (LLMs), Machine Learning, AI Safety, and Artificial Intelligence. Notable guests include Liang Wenfeng, Alexandr Wang, Jim Fan, and François Chollet. Episodes span from Oct 2019 to Jul 2026.

75episodes
534total claims
12topics covered
75 episodes
Kimi K3: Open Frontier Intelligence
Jul 27, 2026

Kimi K3 is a newly announced 2.8 trillion parameter Mixture-of-Experts (MoE) model with a 1-million-token context window and native vision capabilities. The model introduces novel architectures lik...

Large Language Models (LLMs)Kimi K3Mixture-of-Experts (MoE)AI Architecture+11 more
Music-JEPA: Learning a World Model of Sound from Action
Jul 24, 2026

The Music-JEPA paper introduces a novel AI model that learns a 'world model' of piano sound using a Joint Embedding Predictive Architecture (JEPA). It uniquely frames music as an action-conditioned...

Machine LearningAISelf-Supervised LearningWorld Models+11 more
Patch Policy: Efficient Embodied Control via Dense Visual Representations
Jul 20, 2026

The research introduces "Patch Policy," a lightweight architectural extension for robot learning that efficiently utilizes dense visual features from pre-trained Vision Transformers (ViTs). Current...

RoboticsRobot LearningEmbodied AIVision Transformers (ViTs)+11 more
RoboTTT: Context Scaling for Robot Policies
Jul 16, 2026
DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation
Jul 6, 2026
GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks
Jul 6, 2026
Separating Representation from Reconstruction Enables Scalable Text Encoders
Jul 4, 2026
Gemma 4 Technical Report
Jul 2, 2026
ASPIRE: Agentic /Skills Discovery for Robotics
Jun 30, 2026
AdaJEPA: An Adaptive Latent World Model
Jun 30, 2026
Safety from Honesty in a Disinterested AI Predictor
Jun 28, 2026
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
Jun 28, 2026
SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation
Jun 26, 2026
Autoregressive Boltzmann Generators
Jun 25, 2026
Tmax: A simple recipe for terminal agents
Jun 22, 2026
Native Active Perception as Reasoning for Omni-Modal Understanding
Jun 17, 2026
Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models
Jun 16, 2026
Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System
Jun 16, 2026
CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs
Jun 9, 2026
Adaptive Order Policies for Masked Diffusion
May 29, 2026
SonicSampler: Unified Tile-Aware Kernels for LLM Sampling and Speculative Verification
May 24, 2026
Reducing Political Manipulation with Consistency Training
May 21, 2026
CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs
May 19, 2026
Search Your Block Floating Point Scales!
May 12, 2026
Eigenism: Ethics for a Human-AI Future
May 8, 2026
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
Apr 26, 2026
SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving
Apr 21, 2026
The ATOM Report: Measuring the Open Language Model Ecosystem
Apr 8, 2026
Olmo Hybrid: From Theory to Practice and Back
Apr 3, 2026
ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence
Mar 24, 2026
TurnWise: The Gap between Single- and Multi-turn Language Model Capabilities
Mar 17, 2026
Attention Residuals
Mar 16, 2026
International AI Safety Report 2026
Feb 24, 2026
Towards Autonomous Mathematics Research
Feb 10, 2026
Kimi K2.5: Visual Agentic Intelligence
Feb 2, 2026
ARC Prize 2025: Technical Report
Jan 15, 2026
Excess Description Length of Learning Generalizable Predictors
Jan 8, 2026
Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks
Jan 8, 2026
Aggressive Compression Enables LLM Weight Theft
Jan 3, 2026
OpenAI GPT-5 System Card
Dec 19, 2025
SIMA 2: A Generalist Embodied Agent for Virtual Worlds
Dec 4, 2025
DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
Dec 2, 2025
International AI Safety Report 2025: Second Key Update: Technical Safeguards and Risk Management
Nov 25, 2025
Natural Emergent Misalignment from Reward Hacking in Production RL
Nov 23, 2025
Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models
Oct 31, 2025
Remote Labor Index: Measuring AI Automation of Remote Work
Oct 30, 2025
Kimi Linear: An Expressive, Efficient Attention Architecture
Oct 30, 2025
International AI Safety Report 2025: First Key Update: Capabilities and Risk Implications
Oct 15, 2025
Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer
Oct 2, 2025
Limit-Computable Grains of Truth for Arbitrary Computable Extensive-Form (Un)Known Games
Aug 22, 2025
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Jul 7, 2025
Unsupervised Elicitation of Language Models
Jun 11, 2025
DataRater: Meta-Learned Dataset Curation
May 23, 2025
ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems
May 17, 2025
Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures
May 14, 2025
Gemini Robotics: Bringing AI into the Physical World
Mar 25, 2025
Gemma 3 Technical Report
Mar 25, 2025
Superintelligence Strategy: Expert Version
Mar 7, 2025
Forecasting Rare Language Model Behaviors
Feb 24, 2025
Which Economic Tasks are Performed with AI? Evidence from Millions of Claude Conversations
Feb 11, 2025
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
Jan 31, 2025
International AI Safety Report
Jan 29, 2025
Humanity's Last Exam
Jan 24, 2025
OpenAI o1 System Card
Dec 21, 2024
A Practitioner's Guide to Continual Multimodal Pretraining
Aug 26, 2024
Imagen 3
Aug 13, 2024
The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
Mar 5, 2024
Gemini: A Family of Highly Capable Multimodal Models
Dec 19, 2023
The Update-Equivalence Framework for Decision-Time Planning
Apr 25, 2023
Abstracting Imperfect Information Away from Two-Player Zero-Sum Games
Jan 22, 2023
PaLM: Scaling Language Modeling with Pathways
Apr 5, 2022
Scaling Up Models and Data with $\texttt{t5x}$ and $\texttt{seqio}$
Mar 31, 2022
ST-MoE: Designing Stable and Transferable Sparse Expert Models
Feb 17, 2022
The Deep Learning Revolution and Its Implications for Computer Architecture and Chip Design
Nov 13, 2019
Accelerating Deep Learning by Focusing on the Biggest Losers
Oct 2, 2019

Sign up free to see the full analysis

Get started free

Top Topics