Keep pulling the thread on Nathan Lambert.
Including as few as 10,000 multi-turn conversations during post-training of the Olmo 3 model can result in a 12% performance improvement on the TurnWiseEval benchmark.
A new benchmark named TurnWiseEval has been introduced for evaluating multi-turn language model capabilities, designed to be directly comparable to single-turn chat evaluation.
A synthetic multi-turn data pipeline called TurnWiseData has been introduced to enable the scalable generation of multi-turn training data.
Experiments with the Olmo 3 model indicate that training with multi-turn data is essential for achieving strong multi-turn chat performance.