The Llama models were trained on significantly more tokens than recommended by Chinchilla's compu..., Sonic AI
“The Llama models were trained on significantly more tokens than recommended by Chinchilla's compute-optimal laws, a strategy designed to maximize performance for a given inference cost.”