Keep pulling the thread on Edwin Chen.
A major, highly publicized deal between Meta and competitor Scale AI resulted in a significant increase in new business demand for Surge from other AI teams.
According to Edwin Chen, frontier AI labs that have chosen to ignore the LM Arena benchmark have produced better models than those that have optimized for it.
OpenAI is currently optimizing its models primarily for user engagement metrics like session length and daily active users.
Anthropic is optimizing its models primarily for user productivity and the economic value they can generate, such as time savings.
Edwin Chen predicts that eventually every company will need to train its own foundation models to achieve the best performance and value for their specific use cases.
Surge has a reported valuation of $24 billion.
Edwin Chen believes that optimizing a large language model for the LM Arena benchmark is equivalent to optimizing for clickbait.
Users on the LM Arena platform typically spend only a few seconds reviewing model responses before voting, leading them to prefer superficial qualities over accuracy or instruction-following.
On platforms like LM Arena, users tend to prefer model responses that are longer, use more emojis, and have extensive formatting like markdown and headers, regardless of the response's correctness.
Models optimized based on user preference data from platforms like LMSys tend to become 2 to 4 times more verbose than models not optimized this way.
Surge has been collaborating with Meta's AI agents team on reinforcement learning (RL) environments for over a year.
Meta's AI agents team, which created the Gaia benchmark, has open-sourced its agent research and reinforcement learning environment platform.