“The Olmo Hybrid model scales significantly more efficiently during pretraining than the pure transformer-based Olmo 3 model.”