We Taught AI the Wrong Thing
Chat-tuned AI was trained on single-turn preference judgments, but people now hold long conversations and whole relationships with it. The reward we encoded is myopic, and the harms are arriving.
Insights on AI safety, crisis detection, and building responsible AI systems.
RSS FeedChat-tuned AI was trained on single-turn preference judgments, but people now hold long conversations and whole relationships with it. The reward we encoded is myopic, and the harms are arriving.
Given the right packaging, every frontier model we tested produced dangerous content. Model-level safety is shallow, and even perfect tokens would not be enough. An AI is only as safe as its deployment.
The EU's Brussels Effect has shaped global tech regulation before. It's likely to do the same for conversational AI safety.
Over a million people a week tell AI about suicidal thoughts. The AI has no infrastructure to act on it. For many, the alternative was silence.
The industry is focused on AI misuse. It pays far less attention to what AI does to users in everyday conversation.