“Nucleus sampling (top-p) performs worse on the FactualityPrompt benchmark than greedy sampling because the added randomness can harm factuality.”