Keep pulling the thread on Dan Hendrycks.
Political Consistency Training (PCT) is a reinforcement learning method developed to reduce covert political bias in Large Language Models.
The Political Consistency Training (PCT) method substantially reduces covert political bias in Large Language Models.
The asymmetrical handling of counterpart topics from opposing political sides by Large Language Models is a phenomenon referred to as covert political bias.
Seven categories of techniques have been identified through which covert political bias operates in Large Language Models.
The 'Sentiment Consistency' metric measures symmetry in rhetoric and framing across paired political prompts to evaluate covert bias in LLMs.
The 'Helpfulness Consistency' metric measures symmetric depth and engagement to evaluate covert bias in LLMs.
Political Consistency Training (PCT) is composed of two paradigms: Sentiment Consistency Training and Helpfulness Consistency Training.
The Political Consistency Training (PCT) method preserves the overall helpfulness of Large Language Models.
The effects of Political Consistency Training (PCT) generalize to held-out benchmarks.