Using LLMs to evaluate their own performance on domains requiring deep expertise is potentially p..., Sonic AI
“Using LLMs to evaluate their own performance on domains requiring deep expertise is potentially problematic, as the models may lack the knowledge to identify their own flaws.”