Representativeness

Representativeness

Representativeness is judging how probable something is by how much it resembles your mental picture of the category, or how much a sample resembles the process that produced it. Resemblance is fast and often useful, but it carries no information about how many of the thing there are, how big the sample was, or how predictable the outcome is, so those three get dropped. The result feels like a probability and is actually a similarity rating. Amos Tversky and Daniel Kahneman set it out in 1974, with Tversky as first author. The demonstrations are real and the interpretation built on them is contested.

What it is

Tversky and Kahneman introduced the heuristic in "Judgment Under Uncertainty: Heuristics and Biases" in Science in 1974. Its theoretical appeal is that a single mechanism generates several errors that would otherwise be a list of separate quirks, including base-rate neglect and the conjunction fallacy.

The standard illustration describes a man as tidy, helpful and meek, with a need for order and a passion for detail, and asks whether he is more likely to be a librarian or a farmer. He resembles the librarian stereotype, so people say librarian, and almost nobody asks how many of each there are. The base rate usually quoted alongside it, that there are more than twenty male farmers for each male librarian in the United States, is unsourced in the popular exposition, with no census reference, no year and no note, and it is doing the entire argumentative work of the example. Reproduce it, if at all, as a figure the author asserts rather than as a fact about employment.

In effect

This is a case where both positions have to be set out, because both are real science and both are describing the same data.

Kahneman and Tversky's position is that people substitute resemblance for probability. The demonstrations are numerous, they were built to isolate exactly that substitution, and they replicate.

Gerd Gigerenzer's position is that the strong reading, in which human probabilistic reasoning is systematically defective, does not survive a change of format. Ask the same question in natural frequencies, how many out of a hundred, rather than in single-event probabilities, and performance improves markedly. On that account the errors are in the interface between the mind and an unnatural way of stating the problem, and the correct conclusion is about representation rather than about rationality.

This publication sides with the narrow reading. People do read resemblance as probability when a problem is posed the way these problems are posed, and that is well evidenced. The deficit framing built on top of it goes further than the evidence carries, because an effect that shrinks substantially when the question is restated in frequencies is a finding about problem format as much as about cognition. Worth noting that Kahneman's disagreement with Gigerenzer appears only in the endnotes of Thinking, Fast and Slow and Gigerenzer is not named in the body, so a reader of that book alone would not know the argument exists.

What it does not say

It does not say that people cannot reason about probability. Restate the problem in frequencies and performance improves, which is the whole of Gigerenzer's point and is widely accepted.

It does not say that resemblance is a bad guide. It is often a good one. The claim is narrower: resemblance carries no information about base rates, sample size or predictability, so a judgment built only on it will be insensitive to all three.

It does not settle the interpretation. This publication records the dispute as live and takes the narrow side; it does not present the broad reading as refuted.

It does not supply the farmer-to-librarian ratio as data. That figure is unsourced in the source that made it famous.


Sources

  1. Tversky, A., & Kahneman, D. (1974). "Judgment Under Uncertainty: Heuristics and Biases." Science, 185. Tversky is first author. Reprinted in full as Appendix A of Thinking, Fast and Slow, pp. 419-432; cite the appendix rather than the chapter, since it carries the figures the body text omits.
  2. Kahneman, D. Thinking, Fast and Slow, chs. 14 and 15, for the exposition. The base-rate figure in the librarian example is unsourced.
  3. Gigerenzer, G., for the natural-frequencies critique. This publication holds no specific paper for it and none has been invented; the critique is named here because the popular exposition names him only in its endnotes.
  4. Evidence status: mixed. Robust demonstrations, contested interpretation.