The CIAware-Bench evaluation found that models are generally better at detecting control interven..., Sonic AI
“The CIAware-Bench evaluation found that models are generally better at detecting control interventions when the controlling and untrusted models are from different model families.”