“The Tmax project trained open-weight models using a simple, outcome-only reinforcement learning recipe.”