Keep pulling the thread on Jared Kaplan.
A method has been developed to forecast potential risks in language models across orders of magnitude more queries than are tested during evaluation.
A new forecasting method can predict the emergence of diverse undesirable language model behaviors, such as assisting with dangerous chemical synthesis and power-seeking actions, across up to three orders of magnitude of query volume.
The largest observed elicitation probabilities for a target behavior in a language model predictably scale with the number of queries.
A new forecasting method for rare language model behaviors enables model developers to proactively anticipate and patch rare failures before they occur during large-scale deployments.