OpenAI is pursuing an interpretability technique called "chain-of-thought faithfulness" which kee..., Sonic AI
“OpenAI is pursuing an interpretability technique called "chain-of-thought faithfulness" which keeps parts of a model's internal reasoning free from supervision during training to better represent its internal process.”