“On the TruthfulQA benchmark, larger models were found to be less truthful due to their tendency to learn common misconceptions.”