Keep pulling the thread on Jess Yan.
Anthropic uses an internal agent, referred to as "API review Claude," to act as a neutral judge and resolve impasses during API design discussions.
Jess Yan observes a trend where vertical SaaS products are becoming increasingly specialized into niche use cases as general model intelligence improves.
In the AI agent market, the key defensible value is shifting towards context and task orchestration patterns, rather than domain-specific knowledge like finance or healthcare.
AI agents have evolved from simple prompting loops into autonomous, self-discovering, and long-running actors.
At Anthropic, internal AI agents are used for overnight tasks such as resolving backlogs and fixing bugs.
Jess Yan believes it is impossible to achieve maximum model performance without tightly integrating the model with its operational harness.
Anthropic tests its models in conjunction with harnesses it has built internally, such as those for Claude CoWork and Claude Code.
Anthropic's Cloud Managed Agents product is a pre-built harness and companion infrastructure designed to allow an agent to run complex tasks at scale.
Anthropic's Cloud Console includes a dedicated debug agent that analyzes the full session history of another agent to identify areas for improvement.
Jess Yan states that creating effective evaluations (evals) is currently the most difficult aspect of building AI agents.
A trend in agent evaluation is to build an eval loop directly into the agent's workflow, allowing the agent to grade its own outputs.
At Anthropic, an agent was used to iterate on a predictive model until it achieved a specific accuracy benchmark of 90%.