Keep pulling the thread on Ian Fisher.
Anthropic, OpenAI, and Google are exploring recursive self-improvement, typically by training a new model for each step of improvement.
Poetic has built a system that automatically generates solutions for specific problems that consistently outperform the underlying language models they are built on.
Poetic achieved the top position on the Arc AGI V2 benchmark after releasing a paper in December of the previous year.
Two days after Google's Gemini 3 Deep Think achieved a 45% score on the Arc AGI V2 benchmark, Poetic released results showing a 54% score.
Poetic's solution for the Arc AGI V2 benchmark cost $32 per problem and was built on Gemini 3 Pro, making it half the cost of the competing Gemini 3 Deep Think solution.
Poetic achieved a 55% score on the Humanities Last Exam benchmark, surpassing the previous state-of-the-art score of 53.1% from Anthropic's Claude Opus 4.6.
The optimization cost for Poetic's run on the Humanities Last Exam benchmark was less than $100,000.
In a previous DeepMind paper, researchers improved a task's performance from 5% to 95% on Gemini 1.5 Flash by adding manually-built reasoning strategies on top of optimized prompts.
Poetic is a company building recursively self-improving AI reasoning harnesses for Large Language Models (LLMs).
Poetic's core insight is that its method for recursive self-improvement is significantly faster and cheaper than other proposed approaches.
The "harnesses" developed by Poetic are designed to be compatible with new, future language models without requiring changes.
Poetic is a company of seven people, composed of research scientists and research engineers.