Keep pulling the thread on Jan Leike.
A new unsupervised algorithm, Internal Coherence Maximization (ICM), has been developed to fine-tune pretrained language models using their own generated labels without external supervision.
The Internal Coherence Maximization (ICM) algorithm matches the performance of training on golden labels on the GSM8k-verification, TruthfulQA, and Alpaca reward modeling tasks.
The Internal Coherence Maximization (ICM) algorithm outperforms training on crowdsourced human supervision on the GSM8k-verification, TruthfulQA, and Alpaca reward modeling tasks.
A Claude 4 Sonnet-based assistant trained with an unsupervised reward model using Internal Coherence Maximization (ICM) matches the average performance of a counterpart trained on production-grade human labels.
The Internal Coherence Maximization (ICM) algorithm elicits superhuman capabilities from language models significantly better than training on human labels.
A Claude 4 Sonnet-based assistant trained with Internal Coherence Maximization (ICM) achieves higher scores on chat and safety compared to a counterpart trained on production-grade human labels.
A Claude 4 Sonnet-based assistant trained with Internal Coherence Maximization (ICM) achieves lower scores on math and coding compared to a counterpart trained on production-grade human labels.