THIS EXPLANATION
THE ROOM
LAN·33 Language, Media & Communication 6 MIN · 8 STATIONS

Sharing cascades

A Socratic walk-through of sharing cascades — reasoned out one step at a time, not lectured.

abcdefgh
a

The question we started with

THE QUESTION #

Why do a handful of posts reach millions while near-identical ones vanish?

Two people post the same joke, in almost the same words, within an hour of each other. One is seen by eleven people. The other is seen by four million. Afterwards everyone explains the second one — the timing, the phrasing, the image — and the explanation always sounds convincing.

The assumption worth questioning is that a gap that large must have a cause that large. We are used to outcomes being roughly proportional to their inputs: a slightly better product sells slightly better. What if sharing is a process where a difference too small to name reliably produces a difference of six orders of magnitude?

b

Reasoning it through

REASONING #

Let us build the simplest thing that could possibly work, and see whether it already explains the puzzle.

A post is seen by some people. Each of them, independently, either passes it on or does not. Those who pass it on expose it to their own audiences, and the same coin is flipped again. Nothing here is about content, timing, or quality yet — it is only structure. What follows from just that?

Define one number: the average count of new shares that a single share produces. Call it the reproduction number. Ask what happens generation by generation. If each share produces on average 0.7 more, then a hundred shares beget seventy, then forty-nine, then thirty-four. The thing dies, and it dies quickly, because each generation is a fixed fraction of the last. If instead each share produces 1.4 more, a hundred begets a hundred and forty, then a hundred and ninety-six, then two hundred and seventy-four — and the total does not settle down. It runs away.

Notice what has just happened. The two cases differ by whether a single number sits above or below one. Not by whether it is large. Between 0.98 and 1.02 there is no perceptible difference in the content, in the audience, or in the moment — and yet one dies out with certainty and the other has a real chance of not dying out at all. This is the whole of the answer's spine: multiplicative processes have a threshold, and near a threshold, arbitrarily small causes produce unbounded differences in effect.

Now put chance back in. Nobody actually shares "on average". Each person is a coin flip. So even with the reproduction number comfortably above one, the cascade has to survive its first few generations to get anywhere, and early on the numbers are tiny — three people, one person. A single unlucky flip at generation two ends it before compounding can matter. This is why the outcome is so weakly predictable from the content: most of the variance is generated when the population is smallest.

The empirical picture matches this shape. Studies of large samples of online diffusion — Goel, Watts and Goldstein's work on the structure of virality is the standard reference — find that the overwhelming majority of cascades die at size one or two, and that even the large ones are usually broad and shallow rather than deep branching trees, with most adoption coming directly from the seed rather than from long chains.

And there is a second, harder result to sit with. In the Salganik, Dodds and Watts music-download experiment, participants were split into independent worlds where they could see what others in their own world had downloaded. The same songs, in identical conditions, produced wildly different rankings across worlds. Quality set a floor and a ceiling — the truly awful rarely won and the best rarely came last — but within that wide band, which song became a hit was substantially a matter of which early flips went which way, and then got amplified by everyone else's visible behaviour.

One honest limit: this argument explains the distribution of outcomes, not any particular one. It does not license the claim that your post failed purely by chance. Reproduction numbers do vary with content, framing and network position, and models that use early-cascade features can predict later growth far better than chance. What the argument does establish is that no amount of after-the-fact reasoning about a single viral post can distinguish a genuinely superior post from a median post that won three coin flips in a row.

c

The analogy

THE ANALOGY #
THE FIGURE

Think of a match dropped in a forest. Whether it starts a fire is not mostly about the match. It is about whether each burning twig, on average, ignites more than one neighbour before it burns out — which depends on moisture, wind and spacing. Below that threshold every match goes out no matter how well struck. Above it, some matches still go out, because the first twig they land on happened to be damp.

WHERE IT BREAKS DOWN

fire consumes its fuel, so a forest cannot burn twice, whereas an audience that has seen a post is not destroyed — it simply becomes immune, and the same network can carry an unlimited number of independent cascades at once, competing for attention rather than for fuel.

d

Clarifying the model

THE MODEL #

The most common misreading of this is fatalism — "so it is all luck". That is not what a threshold model says.

What it says is that the reproduction number is the thing worth influencing, and that it is a product of factors, not a sum. Roughly: how many people see it, times how likely each is to pass it on. Doubling either doubles the whole. And because the outcome depends on whether that product crosses one, a change that raises it by a few percent can move a post from certain death to a real chance of running away — while looking, to any observer, like a trivial edit.

The second refinement is about visible counts. Real platforms do not run independent coin flips; a post that already has many shares is displayed more prominently and is more likely to be shared again. That is a feedback term, and it means the reproduction number is not fixed — it rises with early success. Path dependence of that kind is exactly what the music-download experiment isolated, and it widens the gap between near-identical items rather than narrowing it.

e

A picture of it

THE PICTURE #
Sharing cascades
Sharing cascades both start identically at a hundred shares -- the bars are a cascade whose average share produces 0.7 more, the line one that produces 1.4, and after only four generations the same starting post is either fading or compounding. {"generator":"[email protected]","source":"../Socrates/.diagram-cache/_src/sharing-cascades.md","sourceIndex":1,"sourceLine":4,"sourceHash":"fdab3c65112204b187a5df238740e667a54a0fba8efda4928ef356752b129d03","diagramType":"xychart","layoutVariant":"source","repairedDuplicateIds":[],"motion":"entrance-with-reduced-motion-fallback","presentation":"editorial","attempt":1,"viewBox":{"x":0,"y":0,"width":790,"height":636},"qa":{"passed":true,"findings":[]}} Gen1 Gen2 Gen3 Gen4 300 280 260 240 220 200 180 160 140 120 100 80 60 40 20 0 Shares in that generation

How to readboth start identically at a hundred shares — the bars are a cascade whose average share produces 0.7 more, the line one that produces 1.4, and after only four generations the same starting post is either fading or compounding.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

The gap between the post that vanished and the post that reached millions does not need a large cause, because the process that separates them is multiplicative and has a threshold at one. Small differences in how likely each viewer is to pass something on, plus the luck of a handful of early flips, are fully sufficient to produce the outcomes we observe — which is why the confident retrospective explanation of any single viral post should be treated as a story rather than a finding.

g

Where to go next

ONWARD #
  • Why cascade depth and cascade size behave so differently, and why most large cascades are shallow.
  • How epidemiology's reproduction number and its critical value of one transfer to information, and where the analogy fails.
  • What early-cascade features actually do predict, and how much better than chance they manage.
h

Key terms

TERMS #
TermWhat it means
Reproduction numberthe average number of further shares produced by one share; below one a cascade dies, above one it can run away.
Branching processthe mathematical model of each unit independently producing a random number of successors.
Cascadethe full tree of resharing descended from one original post.
Path dependencethe property that early, partly random outcomes shape later ones because they are visible and reinforce themselves.

Every term the collection defines is gathered in the glossary.

Nearby on the shelf

4