The MAI-Thinking-1 recipe consolidates three specialist RL "climbs" using trace-distillation SFT,..., Sonic AI
“The MAI-Thinking-1 recipe consolidates three specialist RL "climbs" using trace-distillation SFT, a method closer to the DeepSeek R1 recipe than the MOPD-based DeepSeek V4.”