The Human Touch That Made ChatGPT Work

When ChatGPT appeared in November 2022, the surprise was not simply that the model was powerful.

Powerful language models already existed.

The real change was that, suddenly, ordinary people could talk to one.

That sounds obvious now. At the time, it was the difference between having an impressive engine in a laboratory and handing someone the keys.

One of the techniques behind that shift was RLHF — Reinforcement Learning from Human Feedback. OpenAI had already used it with InstructGPT, and ChatGPT was trained with a related process designed to make the model follow instructions and behave more usefully in conversation. OpenAI

The idea is easier to understand than the acronym suggests.

First, people showed the model examples of useful responses.

Then they compared different answers and ranked them.

Those preferences were used to train a separate reward model — essentially a system that learned which responses humans tended to prefer.

Finally, the main model was fine-tuned to produce more of those preferred behaviours. OpenAI

There is something almost ironic about it.

After years of trying to make machines more intelligent, one of the important steps was teaching them something much more ordinary:

how to behave when somebody asks a question.

Not perfectly, of course.

RLHF did not suddenly make AI truthful, harmless or infallible. OpenAI itself has long acknowledged that models trained this way can still make mistakes, fail to follow instructions, produce biased outputs or give convincing answers that are simply wrong. OpenAI

But it changed the experience.

And sometimes the experience is what turns technology into a product.

Before ChatGPT, using a large language model could feel like interacting with the machinery behind the curtain.

ChatGPT moved the curtain.

You typed.

It answered.

You disagreed.

It tried again.

You could continue the conversation without knowing what a transformer was, how many parameters were involved or what happened between pressing Enter and seeing the next sentence.

That accessibility mattered enormously.

So when people ask what made ChatGPT different, I would not reduce the answer to model size.

Part of the story was technical progress.

Part of it was the interface.

And part of it was human feedback quietly teaching the machine that being intelligent is not especially useful if nobody can work with you.

Which, now that I think about it, is probably true of humans too.