Keep pulling the thread on Hamel Husain & Shreya Shankar.
The chief product officers of both Anthropic and OpenAI have stated that building 'evals' is becoming the most important new skill for product builders.
To validate an 'LLM as a judge', teams should compare its outputs against human-labeled data using a confusion matrix to analyze false positives and false negatives, rather than relying on a simple accuracy percentage.
An 'LLM as a judge' is most effective when scoped to evaluate a single, narrow failure mode with a binary pass/fail output.
Reviewing and analyzing application data is the highest ROI activity for improving an AI product.
The online course on AI evals taught by Hamil Hussain and Shreya Shankar is the number one course on the Maven platform.
Research shows that developers' criteria for 'good' and 'bad' LLM outputs evolve as they review more examples, a phenomenon known as 'criteria drift', making it impossible to define a complete evaluation rubric upfront.
For most AI products, a small number of 'LLM as a judge' evals, typically between four and seven, is sufficient to cover the most critical failure modes.
Over 2,000 product managers and engineers from 500 companies, including teams from OpenAI and Anthropic, have taken Hamil Hussain and Shreya Shankar's course on AI evals.
The AI agent FIN has an average resolution rate of 65% for customer service tickets.
FIN is used by over 5,000 customer service leaders and companies including Anthropic and Synthesia.
Observability tools such as Braintrust, Phoenix Arise, and LangSmith are used for logging and analyzing LLM application traces.
LLMs often fail at automated error analysis because they lack the necessary product context to identify certain failures, such as hallucinating a feature that does not exist.