Preview Mode Links will not work in preview mode

AXRP - the AI X-risk Research Podcast

Jun 12, 2024

Reinforcement Learning from Human Feedback, or RLHF, is one of the main ways that makers of large language models make them 'aligned'. But people have long noted that there are difficulties with this approach when the models are smarter than the humans providing feedback. In this episode, I talk with Scott Emmons...