In a randomized controlled trial, human evaluators assisted by critiques from a GPT-3.5 model wer..., Sonic AI
“In a randomized controlled trial, human evaluators assisted by critiques from a GPT-3.5 model were able to find 50% more flaws in a summarization task compared to unassisted humans.”