Google's Gemma models were trained using knowledge distillation, specifically by minimizing the K..., Sonic AI
“Google's Gemma models were trained using knowledge distillation, specifically by minimizing the KL divergence with the output distribution of the larger Gemma 27B model.”