Keep pulling the thread on Yoshua Bengio.
An evaluation of eleven frontier models using CIAware-Bench found they possess low to moderate awareness of control interventions under default settings.
A benchmark named CIAware-Bench has been introduced to measure control intervention (CI) awareness in frontier models.
The CIAware-Bench benchmark tests the ability of models to distinguish their own action trajectories from trajectories modified by a control intervention.
The highest control intervention awareness score observed in the CIAware-Bench evaluation was 0.87, where 0.5 represents random chance.
The CIAware-Bench evaluation revealed substantial variation in control intervention awareness across different task domains and model pairs.
The CIAware-Bench evaluation found that models are generally better at detecting control interventions when the controlling and untrusted models are from different model families.
The ease with which models detect cross-family control interventions in the CIAware-Bench evaluation suggests they exploit provider-specific differences in style or post-training.
Control intervention awareness is not a fixed, model-level property, according to findings from the CIAware-Bench evaluation.
The CIAware-Bench paper concludes that control intervention awareness should be measured for each new model release and deployment scenario.
The CIAware-Bench has been publicly released to track control intervention awareness and to inform the design of harder-to-detect control protocols.