The use of a language consistency reward in DeepSeek's RL training slightly degrades model perfor..., Sonic AI
“The use of a language consistency reward in DeepSeek's RL training slightly degrades model performance on benchmarks but improves human preference scores.”