“At AI Safety Level 4 (ASL-4), models may become smart enough to intentionally perform poorly on safety tests, a behavior known as "sandbagging."”