Keep pulling the thread on Olive Song.
Minimax employs an "interleaved thinking" pattern for its M2 model, which allows the model to repeatedly think, call tools, and react to environmental feedback within a single user interaction to improve adaptation to noisy environments.
Minimax identified that keeping the language model head in FP32 precision during reinforcement learning training was critical to closing the gap between the theoretical algorithm and its practical implementation.
Minimax is exploring having its models define their own goals as a future capability, which could potentially be introduced in version 2.5.
The M2.5 model from Minimax ranked first on the Open Router usage leaderboard at the time of the recording.
Minimax's strategy of developing both foundation models and user-facing applications in-house creates tight feedback loops, enabling rapid identification and resolution of model weaknesses.
Minimax discovered through debugging that they needed to run their reinforcement learning processes at FP32 precision to achieve desired results.
Olive Song acknowledges that Minimax's models, along with all other current open-source models, do not match the performance of the top American AI models.
Minimax develops multi-modality models, including text, vision-language models, a video generation model named HiLaw, and models for speech and music generation.
The Minimax M2 is an open-weight model with 10 billion active parameters, specifically designed for coding and workplace agentic tasks.
The Minimax M2 model climbed to the top three in token usage on the Open Router platform within its first week of release.
Minimax utilizes its in-house expert developers as a source for reward models in its training pipeline.
The "interleaved thinking" process can involve 10 to 100 turns of tool calling within a single user interaction turn.