“On the Remote Labor Index, recent frontier models GPT-5.5, Opus 4.8, and Fable 5 achieved success rates of 6.3%, 8.3%, and 16.1% respectively.”