“O1 is not superior to previous models across all tasks and has "weird edges" that require new prompting techniques.”