Training a language model with behavior cloning (supervised fine-tuning) on 100% correct answers ..., Sonic AI
“Training a language model with behavior cloning (supervised fine-tuning) on 100% correct answers teaches the model to hallucinate when it lacks the underlying knowledge for a given question.”