The Thing You Cannot Judge

The Thing You Cannot Judge

Why We Look ยท

Some things you cannot assess before buying or after using. An experiment with 936 participants tested the four remedies for that, and only one of them worked.

Key takeaways

  • A credence good is one whose quality you cannot assess before buying and cannot assess after using, such as a dental filling or a car repair.
  • Chase and Schlink identified the condition in 1927, arguing mass production had severed the buyer from any means of knowing what they were buying.
  • Dulleck and Kerschbamer set out the economics in 2006, naming undertreatment, overtreatment, overcharging and market breakdown as the failures.
  • In an experiment with 936 participants, liability had a crucial effect, verifiability at best a minor one, and reputation little influence.
  • Competition drove prices down and produced maximal trade without producing higher efficiency, as long as liability was violated.

Credence goods are the things you cannot judge before buying and cannot judge afterwards either. In 1927 an accountant and an engineer published a book about the problem, before economics had a name for it.

Stuart Chase was a certified public accountant who had spent years on the staff of the Federal Trade Commission. Frederick Schlink was a mechanical engineer and physicist formerly at the National Bureau of Standards. Their complaint in Your Money's Worth was not that manufacturers lied, though they thought many did. It was structural: the buyer had been severed from any means of knowing what they were buying.

A hundred years earlier you bought flour from somebody who milled it. By 1927 you bought a branded tin whose contents you could not assess, from a shop that had not made it, on the strength of an advertisement.

Economics eventually gave that condition a name and a theory, and the theory is better than the book, which is the normal direction of travel and worth saying since this site usually finds the reverse.

The three kinds of goods

Economics sorts goods by when you can judge them. The three-way distinction predates the papers this article relies on and we have not traced who first drew it, so it is set out here without an attribution rather than with a borrowed one.

KindWhen you can judge qualityExample
SearchBefore you buyA shirt's colour and fit
ExperienceAfter you use itA restaurant meal, a film
CredenceNot even after you use itA dental filling, a car repair, a vitamin

Credence goods, the third category, are what Chase and Schlink were circling. With credence goods you cannot tell, before or afterwards, whether what you were sold was what you needed. If a mechanic replaces a part and the noise stops, you still do not know whether the part needed replacing. If a dentist fills a tooth, you cannot verify the tooth needed filling. The expert diagnoses the problem and then sells you the solution, which is the structural issue, not the ethics of any individual.

Uwe Dulleck and Rudolf Kerschbamer set out the economics of the third category in the Journal of Economic Literature in 2006, under a title that names the territory plainly: doctors, mechanics and computer specialists. Their contribution is the review and the failure modes rather than the taxonomy. Their account is of asymmetric information severe enough to produce specific failures, undertreatment, overtreatment, overcharging, and in the worst case a market that does not function.

Which remedies actually work

So what fixes it? The theory offers several candidates, and this is where the article's useful part sits, because four of them were put in a laboratory and only one survived.

Dulleck and Kerschbamer, with Matthias Sutter, ran a large experiment with 936 participants in the American Economic Review in 2011, varying four things that are supposed to discipline a credence-goods market.

Liability, meaning the seller is answerable if the treatment does not solve the problem. Verifiability, meaning the buyer can confirm afterwards which treatment they received. Reputation, meaning the seller can build one across repeated interactions. And competition, meaning other sellers.

Theory predicts liability or verifiability should each deliver efficiency. What they found is that liability has a crucial effect and verifiability at best a minor one. Reputation had little influence, which the theory had predicted. And competition drove prices down and produced maximal trade while not producing higher efficiency, as long as liability was violated.

> Competition made the market cheaper, busier, and no better at giving people the treatment they needed.

Read that last clause twice, because it is the finding. More competition made the market busier and cheaper and no better at giving people the treatment they needed.

### Why liability works and verifiability does not

The asymmetry between those two is the useful part, and the reason is not obvious until it is stated.

Verifiability tells you what was done. It does not tell you what should have been done, which is the thing you could never assess.

An itemised invoice confirming that a part was replaced leaves the only question that mattered untouched: whether the part needed replacing.

Liability changes who bears the cost of being wrong. If the seller is answerable when the problem persists, their incentive aligns with diagnosing it correctly, and you do not need to understand the diagnosis at all. That is why credence goods are one of the few areas where a guarantee is worth more than information.

What that changes for a buyer

Which reorders what a consumer should want, and it is not what consumer advice usually recommends. The starting point is a belief held largely independently of evidence: that the price is higher than it should be and somebody is taking the difference.

In credence goods markets, being able to check afterwards turns out to matter less than being able to hold somebody answerable. An itemised invoice is verifiability. A guarantee that the problem is fixed or you do not pay is liability. The second is the one the evidence backs, and it is the one almost nobody asks for.

### The reputation result, held carefully

Reviews are reputation, and reputation had little influence here. That is a laboratory result with a defined game and should not be over-read, but it sits uncomfortably beside an industry built on the assumption that star ratings discipline credence markets. We have written separately about what star ratings do and do not track.

And the competition finding cuts against the instinct to shop around, which is the usual advice for credence goods and the wrong one. As with scarcity cues, where the cue depends on what is being sold, the remedy depends on the market rather than on a general rule. Three quotes from three garages gives you price information about a diagnosis you still cannot evaluate.

What the experiment cannot tell you

Three, and the first is the one we would put to anybody quoting this.

It is an experiment with a designed game rather than a study of any actual trade, and the authors are economists testing a theory rather than auditing anybody.

That said, the obvious objection has been answered and we should say so rather than leave it as our own caution. Adrian Beck, Rudolf Kerschbamer, Jianying Qiu and Matthias Sutter ran the same experimental paradigm with real car mechanics, in the Journal of Economic Behavior and Organization in 2014, with two of the three authors from the 2011 study. So the student-sample complaint is not where the weakness is. We have not read that paper and are naming it rather than characterising its result.

The 2006 piece is a review, so when this article says the theory predicts something, that is the state of the literature at that point and not a finding.

And nothing here measures how common the failures are in any real market. The theory says the incentive exists. It does not say how often it is acted on, and anybody telling you a proportion is working from something other than these papers.

What Chase and Schlink got right, and what they missed

They identified the condition decades before the economics caught up with it, which is a real achievement for an accountant and an engineer writing for a general audience. We are not going to put a number on the gap, because it depends which paper you count as the formalisation and we have not settled that.

What they did not have is the remedy ranking, and their own proposed solution was testing and publication: independent laboratories assessing goods and telling people the results. Schlink founded Consumers' Research two years after the book, and Consumer Reports is its descendant. That remedy is verifiability, which is the one of the four that barely worked in the laboratory.

Our audit of the book notes something sharper about it. The evidentiary backbone of Your Money's Worth is enforcement and exposure material, Federal Trade Commission dockets and the American Medical Association's catalogue of quack remedies, rather than testing material. The laboratory is what they argue for. The courtroom is where their evidence comes from. Which means the book's own sourcing points at liability while its proposal points at verifiability.

They also conceded something most critics of advertising will not, that people enjoy being sold to, which our reading of the text records and which is part of why the book has aged better than its contemporaries.

The sentence to keep is shorter than any of this. Where you cannot judge the thing, stop trying to judge the thing and ask who carries the cost of being wrong.

Common questions

What are credence goods?

Credence goods are things whose quality you cannot assess before buying and cannot assess after using. A dental filling, a car repair, a vitamin. The defining feature is that the expert who diagnoses the problem is usually the one selling the solution, and you have no independent way to evaluate either.

How are credence goods different from experience goods?

An experience good can be judged after use: a restaurant meal, a film. A search good can be judged before purchase: a shirt's colour and fit. Credence goods are the third category, where verification never arrives, which is what produces the characteristic market failures of undertreatment, overtreatment and overcharging.

Does competition protect consumers in credence goods markets?

Not according to the experimental evidence. Dulleck, Kerschbamer and Sutter found that seller competition drove prices down and produced maximal trade while not raising efficiency, as long as liability was violated. Shopping around gives you price information about a diagnosis you still cannot evaluate.

What should I ask for instead?

Liability rather than verifiability. A guarantee that the problem is fixed or you do not pay is worth more, on this evidence, than an itemised invoice, because an invoice tells you what was done and not whether it needed doing.

Sources

Dulleck, U., & Kerschbamer, R. (2006). 'On doctors, mechanics, and computer specialists: the economics of credence goods.' Journal of Economic Literature. A review setting out the asymmetric-information structure and its failure modes. Verified at the published record 2026-10-05. doi:10.1257/002205106776162717

Dulleck, U., Kerschbamer, R., & Sutter, M. (2011). 'The economics of credence goods: an experiment on the role of liability, verifiability, reputation, and competition.' American Economic Review, 101(2), 526-555. 936 participants. Theory predicts liability or verifiability yield efficiency; they find liability has a crucial effect and verifiability at best a minor one; reputation has little influence, as predicted; seller competition drives down prices and yields maximal trade but does not lead to higher efficiency as long as liability is violated. Verified at the published record 2026-10-05. doi:10.1257/aer.101.2.526

Chase, S., & Schlink, F. J. (1927). Your Money's Worth. Read in full in this vault, 2026-09-20, from a paginated scan. The book concedes that people enjoy being sold to, which is part of why it has aged better than its contemporaries.

A note on scope: the 2011 paper is a designed experiment with participants in a modelled market, not an audit of any trade, and it reports no estimate of how often these failures occur in real markets.

Related concepts

People