Keep pulling the thread on Noam Brown.
On antibody design tasks, the single-sequence ESMFold2 model outperforms AlphaFold3, achieving a DocQ pass rate of 50 compared to AlphaFold3's 47.
Standard self-play algorithms for LLMs fail because rewarding the task-generating model (conjecturer) for difficulty incentivizes it to create messy, artificially complex problems rather than useful ones.
The improved scaling performance of the ESM Cambrian model was achieved by increasing its training dataset from 50 million to 2.8 billion protein sequences, primarily by incorporating metagenomic data.
ESMFold2, using only a single input sequence, achieves near-parity with AlphaFold3 on general protein-protein complex prediction, performing within 3 points on the DocQ pass rate metric.
Using the Self-Guided Self-Play (SGS) method, a 7 billion parameter model achieved the performance of a 670 billion parameter model on the Passat 4 benchmark, though it required 8 times more compute.
AI systems from OpenAI and DeepMind achieved gold medal-level performance at the 2024 International Mathematical Olympiad (IMO).
OpenAI recently claimed to have solved an 80-year-old mathematics problem from Erdős.
Harmonic AI's AxiomProver successfully solved all 12 problems from a recent Putnam mathematical competition.
Channel AI has increased its pull requests per engineer per month by 3.5 times by adopting an agentic, parallelized development workflow inspired by real-time strategy games.
Using sparse autoencoders, researchers found that the latent space of protein language models decomposes into interpretable features corresponding to biological concepts like amino acids, structural motifs, and protein domains.
The BioHub research team created a protein atlas with up to 7 billion folded protein structures, which is larger than the AlphaFold database.
The amount of compute spent on reinforcement learning post-training for large language models is now approaching or surpassing the amount spent on pre-training.