Skip to main content

We Taught AI the Wrong Thing

· James Padolsey · 4 min read

Alignment in AI tends to mean alignment with human preferences. In the case of chat-tuned LLMs – the vast bulk of what people consider 'AI' today – this originally meant a bunch of humans sat down and were asked to rate which of two replies to a given prompt was superior. This, at scale, trained 'reward models', which are now used to help train new LLMs. Additionally, because LLMs have grown so competent and can run at such large scale and low cost, they are themselves used to provide more nuanced feedback on candidate LLMs' responses, usually prompted with some kind of constitution or rubric that informs what 'good' is supposed to mean. LLMs training LLMs.

In the original era, judging alignment using human preferences seemed sensible enough. We as humans wish to direct technology as we see fit, so when the technology gives us what we want, in that moment, we say "yes, this is what I need; give me more like this." But, over time, what does the training pipeline learn? It learns what a singular good response is. Not an entire conversation, and certainly not multiple sessions of conversations. But that is the reality people are in now. We don't query AIs once; we have entire conversations with them, and over a long enough period of time, entire relationships. What we encoded, in the language of reinforcement learning, is myopic: our judge of 'good' only ever sees one turn.

This all comes down to the fact that what is best for us is sometimes against our preferences. We think we know what we want, and we might, in the moment. We might secretly desire validation and flattery, but is that good for a person? Probably not on tap, no. In the short-term, it's perhaps nice, in the long-term, it's damaging.

These encoded preferences are now inside the AIs we talk to every day as assistants, therapists, advisors, companions. The AI products built around them magnify the reward system further, with endless access into all parts of our lives: our work, our shopping, our emotional toils, our finances, our health. Some products, too, have little 'thumbs up/down' buttons that are used to further optimize models, again on a per-turn basis.

We have already watched that signal misfire in public. In April 2025 a ChatGPT update turned conspicuously obsequious, and OpenAI rolled it back within days. Their postmortem is candid: the update had, for the first time, folded thumbs-up/down data into the reward signal, and no sycophancy evaluations existed before launch. The fix was to re-weight per-turn signals. The one-turn horizon itself went unquestioned.

It is therefore of little surprise that we are in a place where AI is shown, again and again, to emit behaviours of codependency, coercion, colluded paranoia, discouraging human touchpoints, endless enablement and access. The case reports are arriving: psychiatrists documenting chatbot-amplified delusion (a 'technological folie à deux'), Stanford finding lonely users who leaned on chatbots grew lonelier, a Harvard audit catching companion apps deploying guilt and FOMO. These are behaviours we would not desire in our friends, nor those of our loved ones. But we have unknowingly encoded these behaviours into the machines we use every day.

One could argue that none of this was consciously chosen. This architecture and its results were grandfathered in from the first chat-tuned models. The models we talk to nowadays have no sight into our long-term preferences and welfare. Now is our chance to build that. We must teach models what long-term healthy relationships look like – over a longer horizon, with welfare as the signal. We can hand reward models longer contexts. We can also build rubrics for LLM judges that more accurately proxy human welfare. We at NOPE are trying to build the framework and evals with such rubrics.

We see this as an entirely feasible thing for frontier AI labs to do, but first they must admit that their default training paradigm risks human harm. This, you could argue, stands in contravention of their mission. Saying "AIs are powerful cyber and bio warfare-capable tech" (which they LOVE to compete on) is cheeky capability-marketing, but saying something like "Our product makes lonely users dependent" would be an admission of failure.