BA
Berkeley AI series
Pieter Abbeel
Summary
Berkeley AI series is hosted by Pieter Abbeel covering Reinforcement Learning (RL), Reinforcement Learning from Human Feedback (RLHF), Language Models (LLMs), and Truthfulness. Notable guests include John Schulman (Co-founder, OpenAI; Chief Architect, ChatGPT).
1episodes
28total claims
12topics covered
1 episodes
John Schulman - Reinforcement Learning from Human Feedback: Progress and Challenges
›Apr 20, 2023 · WITH John Schulman
Truthfulness is a major technical challenge for large language models, with 'hallucinations' stemming from pattern-completion behavior and an inability to express uncertainty. Supervised fine-tunin...
Reinforcement Learning (RL)Reinforcement Learning from Human Feedback (RLHF)Language Models (LLMs)Truthfulness+15 more
Top Topics
Reinforcement Learning (RL)1Reinforcement Learning from Human Feedback (RLHF)1Language Models (LLMs)1Truthfulness1Hallucination1Factuality1Supervised Fine-Tuning (SFT)1Behavior Cloning1Proximal Policy Optimization (PPO)1Trust Region Policy Optimization (TRPO)1Retrieval-Augmented Generation (RAG)1Scalable Oversight1