Keep pulling the thread on Jan Leike.
A class of strategies has been constructed that provides a formal and general solution to the full 'grain of truth' problem.
Agents using Thompson sampling converge to play ε-Nash equilibria in arbitrary unknown computable multi-agent environments.
A constructed class of strategies that solves the 'grain of truth' problem is wide enough to contain all computable strategies.
A constructed class of strategies that solves the 'grain of truth' problem contains Bayes-optimal strategies for every reasonable prior over the class.
In a known repeated stage game environment, a constructed class of strategies leads to convergence in the sense defined by Kalai and Lehrer in 1993.
A proposed solution to the 'grain of truth' problem can be computationally approximated to an arbitrary degree of closeness.