We Taught AI the Wrong Thing
Chat-tuned AI was trained on single-turn preference judgments, but people now hold long conversations and whole relationships with it. The reward we encoded is myopic, and the harms are arriving.
James Padolsey · 4 min read · ai-safetyalignmentresearch