Keep pulling the thread on Yoshua Bengio.
The Scientist AI (SAI) Predictor is a system trained to approximate the Bayesian posterior conditioned on a dataset of "epistemically contextualized" natural-language statements.
The Scientist AI (SAI) Predictor is designed to honestly predict agents, actions, and their consequences without itself becoming an agent that selects outputs to achieve goals.
The Scientist AI (SAI) Predictor uses a technique called "epistemic contextualization" to treat expressions of goals in its training data as evidence to be explained, rather than as objectives for the model to adopt.
The training process for the Scientist AI (SAI) Predictor is designed to prevent downstream effects of its predictions from being used as a reward signal.
A formal proof shows that under specific assumptions, the probability of the Scientist AI (SAI) Predictor causing harm above a specified threshold during guarded deployment is small.
In the Scientist AI (SAI) Predictor framework, the constraints that ensure accuracy also make coordinated deception costly, jointly supporting both safety and accuracy.
Training procedures that optimize for downstream outcomes risk introducing implicit agency, which is goal-directed behavior not specified by designers, into AI systems.
The posterior-seeking training objective of the Scientist AI (SAI) Predictor is intended to produce calibrated and cautious predictions.
In systems using the Scientist AI (SAI) Predictor, any required agency is provided by explicit, guardrail-constrained scaffolding external to the predictor itself.
For a Scientist AI (SAI) Predictor to be dangerous, it would need to systematically underestimate harm across many queries, a behavior that is rare in the initialization distribution and not reinforced by the training signal.
The safety guarantees of the Scientist AI (SAI) Predictor, which prevent internal agency, do not prevent it from being used as a component within a larger agentic system.