Keep pulling the thread on Noam Brown.
The o1 models use a technique called deliberative alignment to reason about OpenAI's safety policies in context when responding to potentially unsafe prompts.
The o1 model series achieves state-of-the-art performance on safety benchmarks for risks including generating illicit advice, choosing stereotyped responses, and succumbing to known jailbreaks.
The o1 model series is trained with large-scale reinforcement learning to reason using chain of thought.
The safety work for the OpenAI o1 and OpenAI o1-mini models included safety evaluations, external red teaming, and Preparedness Framework evaluations.