Shreya Shankar, mentioned 11 times across podcast episodes and expert conversations analyzed by Sonic.
Research shows that developers' criteria for 'good' and 'bad' LLM outputs evolve as they review more examples, a phenomenon known as 'criteria drift', making it impossible to define a complete evaluation rubric upfront.
Products like Anthropic's Claude Code are built upon foundational models that have been extensively evaluated on coding benchmarks, even if the application team itself claims to rely more on 'vibes'.
For most AI products, a small number of 'LLM as a judge' evals, typically between four and seven, is sufficient to cover the most critical failure modes.