Preview Mode Links will not work in preview mode

AXRP - the AI X-risk Research Podcast

Aug 24, 2024

How do we figure out what large language models believe? In fact, do they even have beliefs? Do those beliefs have locations, and if so, can we edit those locations to change the beliefs? Also, how are we going to get AI to perform tasks so hard that we can't figure out if they succeeded at them? In this episode, I chat...