Reward Prediction Error

Reward Prediction Error

Reward prediction error is the gap between the reward expected and the reward received, and it is the signal that midbrain dopamine neurons actually carry. They do not fire to reward as such. Better than expected and they burst; exactly as expected and they do nothing; worse than expected and they dip below baseline. It is one of the better-replicated results in systems neuroscience, and it is the finding underneath almost every popular sentence about dopamine and anticipation. Wolfram Schultz is the first author of the paper that established it.

What it is

Schultz, with Peter Dayan and P. Read Montague, published "A Neural Substrate of Prediction and Reward" in Science on 14 March 1997. Schultz's own synthesis followed in the Journal of Neurophysiology in 1998.

The original demonstration is monkeys learning which of two lights means a pellet is coming. Before learning, the dopamine cells fire at the pellet. After learning, they fire at the light and fall silent at the pellet. Nothing about the pellet has changed; what has changed is how much of it was news. As a cue comes to predict a reward reliably, the burst migrates backwards from the reward to the cue.

The reason the result carries the weight it does is convergence. It links primate single-cell recording to temporal-difference learning in machine learning, so a physiological measurement and a computational formalism turned out to describe the same thing. That is why it survived where so much else from the period did not.

In effect

The everyday version is a route you walk every day and a new shop on it. The route delivers nothing because it is fully predicted. The shop delivers, because it was not. It is also the reason a familiar pleasure stops thrilling without becoming any less pleasant, which is the distinction between wanting and liking approached from the signalling side.

The mechanism explains why unresolved prediction is the engine of variable-ratio reinforcement and of the variable rewards designed into products. A schedule that keeps the prediction unresolved keeps generating the signal.

Its popular career is a case study in correct mechanism and absent provenance. Daniel Lieberman and Michael Long explain reward prediction error correctly and plainly in The Molecule of More at pages 6 and 23, name Schultz at pages 4 to 5 and never again, and then use the mechanism as the engine of six chapters and of the closing argument. The book has no endnotes, no superscripts and no bibliography, only per-chapter reading lists that map to no sentence. Chapter 1 argues from Schultz across pages 4 and 5 and he does not appear on that chapter's reading list; chapter 7 rests its closing argument on the mechanism and its seven-entry reading list contains no Schultz and no prediction-error paper of any kind. The chapter that depends most on the finding cites it least.

What it does not say

It does not say that dopamine is pleasure. The signal is about how much a reward departed from expectation, not about how good it felt.

It does not say that a fully predicted reward is unrewarding. The pellet is just as good; it simply stops generating the error signal.

It does not support a theory of temporal orientation, personality or civilisation. It is a claim about a learning signal, and anything built on top of it has to earn its own evidence.

It does not license the claim that an imaging study showed dopamine circuits activating. Functional magnetic resonance imaging measures blood oxygenation and cannot image dopamine, its release or its receptors, which requires positron emission tomography with a dopamine tracer.


Sources

  1. Schultz, W., Dayan, P., & Montague, P. R. (1997). "A Neural Substrate of Prediction and Reward." Science, 275, 14 March 1997, 1593-1599. Schultz is first author.
  2. Schultz, W. (1998). "Predictive Reward Signal of Dopamine Neurons." Journal of Neurophysiology, 80, 1-27.
  3. Confirmed by this publication's external verification pass, 2026-09-19, as robust and well replicated.
  4. Lieberman, D. Z., & Long, M. E. The Molecule of More. BenBella, 2018. Cited only for the fact that the idea was popularised this way. The mechanism is explained correctly at pp. 6 and 23; Schultz is named at pp. 4 to 5 and nowhere else. The book has no endnotes, superscripts or bibliography. The claim at p. 51 that an imaging study showed dopamine circuits lighting up is a method error.