Keep pulling the thread on Jan Leike.
Tech companies like OpenAI currently make decisions about AI values largely unilaterally, with employees writing content policies that are then trained into the models.
Without a new process for value alignment, commercial incentives will likely determine the values embedded in AI systems rather than alignment with humanity's best interests.
A significant risk of preference aggregation methods for AI alignment is that they will reflect existing societal power structures rather than ideal ones.
A long-term goal for AI alignment is to represent each human with a super-human AI system that is personally aligned with them, reducing the societal alignment problem to the single-human alignment problem.
A drawback of simulated deliberative democracy for AI alignment is the lack of clear human accountability for its decisions.
Using humans to make all detailed, value-laden decisions for AI systems does not scale.
Running a representative 'mini-public' for deliberative democracy costs a few hundred thousand dollars, making it impractical for millions of value questions.
The gpt-3.5-turbo model is approximately 200 times cheaper than the fastest-typing humans paid at US minimum wage.
Pretrained language models contain and exhibit harmful stereotypes present in their training data.
Current transformer-based language models are not capable of in-context learning at a level comparable to a human engaging deeply with a new topic.
Current language models are sample-efficient enough to make building a practical prototype of a simulated deliberative democracy system feasible.
A recent DeepMind paper demonstrated a method of collecting preferences from different demographic groups, training a language model on them, and aggregating the results using social welfare functions.