Why the Best Song Does Not Win

Why the Best Song Does Not Win

Why We Look ยท

An experiment ran the same songs against the same field in parallel worlds. Quality set the floor and the ceiling. Social influence decided almost everything in between.

Key takeaways

  • Salganik, Dodds and Watts built an artificial music market in which 14,341 participants downloaded unknown songs, some able to see earlier participants' choices and some not.
  • Increasing the strength of social influence increased both the inequality of success and its unpredictability.
  • Run as parallel worlds, the same song could finish near the top in one and near the bottom in another.
  • Quality was not irrelevant: the best songs rarely did poorly and the worst rarely did well, but any other result was possible.
  • A ranking is therefore information about quality at the extremes and close to noise in the middle, because the ranking was partly produced by the ranking.

14,341 people. Previously unknown songs. One manipulation.

The experiment was built to measure how much social influence decides what succeeds, which is the question the music industry has never been able to answer about itself. The answer is uncomfortable for anybody who reads a chart as a verdict on quality.

Concede the obvious first, because the finding is often reported as though quality does not matter and that is not what it says. Quality does matter. The best songs rarely did badly and the worst rarely did well. What the experiment establishes is everything between those two ends, which turns out to be most of the chart.

How the music market was built

The design is the contribution. Salganik, with Peter Dodds and Duncan Watts, built an artificial music market and published it in Science in 2006.

Participants were given previously unknown songs to listen to and download. Some could see how many times each song had already been downloaded by earlier participants. Others could not. That is the whole manipulation: identical songs, identical interface, with or without a visible record of what other people had chosen.

Then they ran it as parallel worlds.

The same songs, the same choice, different populations, repeated.

Increasing the strength of social influence increased two things. Inequality, meaning the gap between the hits and the rest grew. And unpredictability, meaning which songs became hits varied from world to world.

Why unpredictability is the real finding

Quality sets the floor and the ceiling. Social influence decides almost everything in between.

The second result is the one worth sitting with, because inequality on its own is unsurprising.

If success were mostly quality, the same songs should win in every world. They did not. The same song could finish near the top in one world and near the bottom in another, with the same competitors, decided by what the first few hundred people happened to do.

### Why this design was necessary

No amount of studying real charts could have produced that result, and the reason is worth understanding because it applies to most commercial evidence.

A real chart runs once. Whatever reached the top did reach the top, and any account of why it did is unfalsifiable, because the counterfactual is missing. You cannot rerun 1997 with the same songs and different listeners.

Social influence inside the experiment supplies the missing counterfactual by running the field repeatedly. That is the whole methodological trick, and it is why this one study carries more weight on the question than a library of chart analyses.

### What the social influence condition actually changed

Worth being precise, because the manipulation is small and the consequence is large.

Participants in the influence condition saw download counts. That is all. No reviews, no rankings by popularity, no recommendation, no social network. One number next to each song, reporting what earlier strangers had done.

That single number was enough to widen the gap between winners and losers and to make which songs won vary between worlds. Everything modern platforms add on top of it, ranked feeds, recommendations, playlists, autoplay, is additional social influence layered on a mechanism that already worked with a bare count.

The paper's own summary of the limits is admirably exact. Success was only partly determined by quality: the best songs rarely did poorly, the worst rarely did well, but any other result was possible.

> Quality sets the floor and the ceiling. Social influence decides almost everything in between.

That is a narrower claim than the version that circulates, and a more useful one. It does not say talent is irrelevant. It says talent constrains the range of outcomes and does not pick the winner inside that range.

### What a theory of shareability can and cannot buy

The trade account of why things catch on is Contagious, and reading it against this paper is instructive about what a mechanism can and cannot buy you.

Berger's six drivers describe properties that make a thing more likely to spread: social currency, triggers, emotion, public visibility, practical value, stories. Every one of them is a reason a particular product might do better than another, and the properties that make a thing spread are a different question from which spreadable thing wins.

What Salganik's parallel worlds show is that even with all properties held exactly constant, outcomes diverge. So a complete account of what makes something shareable still would not let you predict which of two shareable things wins, because part of the answer is who happened to arrive first and what they happened to click.

This is not a criticism of the book. It is the boundary on any such book, including the ones this publication draws on most.

### What the experiment cannot tell you

Three things, and we would rather set them out than have a reader find them.

It is unknown songs by unknown artists in a laboratory market. Real markets have marketing budgets, radio, playlists, prior reputation and press, all of which are forms of social influence that arrive before the first listener does. Whether the effect is larger or smaller in that environment is a reasonable question and not one these data answer.

The outcome is downloads in a research setting, not purchases, and not listening over time.

And it is 2006, which predates the streaming era entirely. Everything about how recommendation now works has changed, and the direction of the change plausibly amplifies the effect rather than dampening it, which is a judgement rather than a finding.

What we take from the parallel worlds

Our reading is that this is the best single piece of evidence for a claim consumers hear constantly and have no way to evaluate: that popularity is information about quality.

It is information about quality at the extremes and close to noise in the middle. A download count is a signal supplied by an interested party only in the sense that the platform chose to show it; the people who produced it had nothing to gain. A product with ten thousand downloads is probably not terrible. Whether it is better than the one with two thousand downloads is a question the ranking cannot answer, because the ranking was partly produced by the ranking. The same caution applies to what a star rating does and does not track.

Which has a practical edge for anybody choosing from a sorted list, whether it is songs, apps, books or restaurants. The sort order is not a measurement of the field. It is a record of the order in which people arrived.

### What this does to the idea of a flop

The result cuts in a direction people rarely take it, which is towards generosity rather than cynicism.

If the middle of the range is decided by social influence, then a thing that failed was not necessarily worse than the thing that succeeded. It may have been equivalent and arrived at the wrong moment, in front of the wrong first hundred people. The parallel worlds make that concrete: the same song, the same quality, two outcomes.

Which means a commercial post-mortem that explains a failure by the product's properties is doing the same unfalsifiable thing as a success story that explains a hit by its properties. Both are reading a single run as though it were the only run available.

A test worth doing once

Next time you pick something from a chart or a bestseller list, ask whether you would still want it with the number hidden.

If yes, the number told you nothing you needed.

If no, you were buying the ranking, and the ranking was partly an accident of who arrived first.

There is a version of that test worth doing once properly rather than every time. Pick a category where you have followed the rankings for years, music or film or books, and try to name a case where the thing at the top was unambiguously better than the thing just below it. Most people find they cannot, and conclude they were not paying attention. The experiment suggests a different reading: that the comparison was never available to them, because the two items were never run against each other in a way that could have separated them.

Common questions

Does popularity mean quality?

At the extremes, somewhat. In Salganik, Dodds and Watts's experiment the best songs rarely did poorly and the worst rarely did well. Between those ends, social influence rather than quality decided the outcome, and the same song could finish near the top in one world and near the bottom in another.

What was the artificial music market experiment?

A 2006 study in Science in which 14,341 participants listened to and downloaded previously unknown songs. Some could see how many times each song had already been downloaded by earlier participants and some could not. The market was run as parallel worlds, so the same songs faced the same field repeatedly with different populations.

What did increasing social influence do?

It increased two things at once: the inequality of success, meaning the gap between hits and the rest widened, and the unpredictability of success, meaning which songs became hits varied from world to world.

Does this apply to streaming and recommendation algorithms?

The study predates streaming, so this is an inference rather than a finding. The manipulation that produced the effect was a single visible download count, and modern platforms add ranked feeds, recommendations and playlists on top of that, all of which are further social influence. The plausible direction is amplification, and nobody has measured it.

Sources

Salganik, M. J., Dodds, P. S., & Watts, D. J. (2006). 'Experimental study of inequality and unpredictability in an artificial cultural market.' Science, 311(5762), 854-856. An artificial music market in which 14,341 participants downloaded previously unknown songs either with or without knowledge of previous participants' choices; increasing the strength of social influence increased both inequality and unpredictability of success; success was only partly determined by quality, with the best songs rarely doing poorly and the worst rarely doing well, but any other result possible. Verified at the published record 2026-10-05. doi:10.1126/science.1121066

Berger, J. (2013). Contagious: Why Things Catch On. Simon & Schuster. The six drivers of transmission, which describe properties that make a thing more shareable. Read in full in this vault, 2026-09-02.

A note on scope: these are unknown songs by unknown artists in a research setting, the outcome is downloads rather than purchases or sustained listening, and the study predates the streaming era. Whether the effect is larger or smaller under algorithmic recommendation is a reasonable question these data do not answer.

Related concepts

People