Keep pulling the thread on Nathan Lambert.
Tmax is the strongest open reinforcement learning (RL) recipe for terminal agents to date.
The Tmax recipe achieves a score of 27% on the Terminal-Bench 2.0 benchmark using a 9 billion parameter model.
A 9 billion parameter model trained with the Tmax recipe outperforms larger models from prior work on the Terminal-Bench 2.0 benchmark.
The Tmax project generated data using a novel taxonomy that combines difficulty control, personas, and verifier diversification.
The terminal dataset released with the Tmax project is over 2.5 times larger than previously released terminal-agent datasets.
The data, models, and code for the Tmax project have been open-sourced to serve as a baseline for future research on terminal agents.