Lab 18

One roll of a die tells you nothing. A million rolls tell you almost everything.

Chance is not the opposite of order β€” it is where a very reliable kind of order comes from. Below is a real experiment you can run as many times as you like, from a single trial up to a million. Watch the bars climb towards the theoretical line, then look carefully at the two error numbers, because they move in opposite directions and almost everyone guesses one of them wrong.

Central limit theorem: ten flat distributions have added up to a bell.

Trials run
0
β€”Counts off by
β€”Proportions off by

Both measure the same disagreement with theory β€” one in whole trials, one as a share of them. Keep running and they part company.

The gap measured in whole trials. It grows roughly with the square root of the number of trials β€” the dashed line is that prediction, not a fit.

The same disagreement expressed as a share of all trials. It shrinks, roughly with one over the square root of the number of trials.

The gambler's fallacy, in two numbers

After a run of tails, people say heads is β€œdue”. It is not. The coin has no memory and no mechanism for repayment.

What actually happens is dilution. The count of surplus heads or tails wanders further from zero as you flip more β€” the left-hand chart climbs β€” while that same surplus becomes a smaller and smaller slice of the total, so the proportion settles at one half.

Nothing corrects the imbalance. It is simply outgrown. That is the whole of the law of large numbers, and it is why a losing streak is not evidence that a win is on its way.

Where intuition fails

Three answers almost everyone gets wrong

Do not take these on trust. Each one below is a simulation you run yourself, with the exact mathematics drawn alongside so you can check that the two agree.

Twenty-three people, and it is already an even bet

Asked how many people you need in a room before two share a birthday, most people say somewhere near 183 β€” half of 365. The right answer is 23. The trick is that you are not comparing yourself with everyone else; you are comparing everyone with everyone, and 23 people make 253 pairs.

23
β€”Exact answer
β€”Simulated

The simulation gives each person a birthday drawn uniformly from 365 days. Real birthdays are slightly bunched β€” September is busier than February β€” and that bunching only makes a match more likely, so 23 is if anything a shade pessimistic.

Switching doors doubles your money

Three doors, one car. You pick one. The host β€” who knows where the car is β€” opens a different door to show a goat, then offers you the swap. Two doors left, so it feels like an even chance either way. Play it and count.

Pick a door to start.

Wins, counted

Switchedβ€”
Stayedβ€”

Your first pick is right one time in three. That never changes. So the other two doors hold the car two times in three, and the host has helpfully emptied one of them for you. Switching collects the whole of that two-thirds.

A test that is 99% accurate and still mostly wrong

Here are ten thousand people. The condition is rare, the test is good, and most of the positive results are still false. Count them yourself: every dot below is one person, and the dots are grouped so you can see the two kinds of positive side by side.

True positive β€” has it, test says so False positive β€” healthy, test says otherwise False negative β€” has it, test missed it True negative
1%
99%
99%
β€”Positive results
β€”Of those, actually ill

The normal distribution

The bell curve is what a pile of small accidents looks like

A Galton board is a plank of pegs. Every ball meets twelve of them and bounces left or right at each one, which is twelve coin flips and nothing more. No ball is told where to land, yet the pile at the bottom always takes the same shape.

The bounce at each peg is a fair coin flip, not a physics collision β€” the ball's path is drawn, the decision is real. The pale line over the bins is the exact binomial distribution for twelve flips, which is the shape the board has to make. The bins open with 1,500 balls already counted; empty them to watch the pile build from nothing.

Balls through the board
0
β€”Mean bin
β€”Standard deviation

Twelve fair flips have a mean of 6 and a standard deviation of √(12 Γ— ΒΌ) = 1.732. The board finds those numbers on its own, by dropping beads down a plank.

Add more rows and the steps get finer until the outline is indistinguishable from the smooth curve below. That is the same central limit theorem the ten-dice experiment showed at the top of the page.

Shading shows one, two and three standard deviations either side of the mean. Those bands always hold 68.27%, 95.45% and 99.73% of everything, whatever the mean and spread are β€” that is the only thing you need to remember about the curve. The vertical scale is fixed, so widening the curve flattens it: the area underneath is always exactly 1.

100
15
β€”68.27% between
β€”95.45% between
β€”99.73% between

Beyond three standard deviations lies 27 people in every 10,000. Beyond four, 6 in every 100,000. The tails thin out ferociously fast, which is exactly why the curve is the wrong model for anything that produces genuine extremes.

Height is normal

many small causes, added up

Adult height is settled by hundreds of genes and a childhood of nutrition, each contributing a little in either direction. Add up enough small independent nudges and you get a bell β€” British men average about 175 cm with a spread of roughly 7 cm, and nobody is twice the average height.

Measurement error is normal

the original reason the curve was studied

Weigh the same object a thousand times and the readings scatter symmetrically about the true value, because each reading collects a heap of tiny independent errors. Astronomers noticed this before anyone had a name for it, which is why it was once called the error curve.

Income is not normal

and the average misleads because of it

Median full-time pay in the UK was Β£37,430 in 2024. The mean is several thousand pounds higher, dragged up by a thin tail of very large salaries. When the mean and the median disagree that much, the distribution is skewed and the mean is describing almost nobody.

City sizes are not normal

a power law, not a bell

London holds about 8.9 million people; Birmingham, the next largest, about 1.1 million. An eightfold gap between first and second place is impossible under a bell curve and completely ordinary under a power law. Word frequencies, earthquake energies and book sales behave the same way β€” the "average" of any of them is a number to distrust.

Reference

Six things worth carrying out of here

A probability is a number from 0 to 1

0 = never Β· 1 = always

0.5 is an even chance, 0.001 is one in a thousand. Percentages are the same number times a hundred. Odds are a different notation for the same thing: "5 to 1 against" means one favourable outcome for every five unfavourable, so a probability of 1/6 β€” not 1/5. The probabilities of every possible outcome always add to exactly 1, which is a surprisingly good way of catching your own mistakes.

Independent events, and the ones that are not

coins forget Β· cards remember

Two events are independent when one happening tells you nothing about the other. Coins, dice and roulette wheels are independent: no result changes the next one.

Cards are not. Drawing two aces from a shuffled deck is 4/52 Γ— 3/51 = 1 in 221, because the first ace is gone. If you replaced it and reshuffled, it would be 4/52 Γ— 4/52 = 1 in 169. Same deck, same question, different answer β€” because the second draw remembers the first.

Expected value: what a lottery ticket is worth

UK Lotto Β· 6 numbers from 59 Β· Β£2

Multiply each prize by its chance and add up the lot. There are C(59,6) = 45,057,474 possible tickets, and the fixed prizes are these:

MatchChancePrizeWorth
21 in 10.3Β£2*Β£0.195
31 in 96.2Β£30Β£0.312
41 in 2,180Β£140Β£0.064
51 in 144,415Β£1,750Β£0.012
5 + bonus1 in 7,509,579Β£1mΒ£0.133
61 in 45,057,474Β£5m†£0.111
Every Β£2 ticketΒ£0.83

You get back about 41p in the pound. The jackpot would have to reach Β£58 million before a Β£2 ticket broke even β€” and jackpots that large sell so many tickets that you would probably be splitting it.

* a free Lucky Dip, counted at its Β£2 face value. † jackpots roll over, so Β£5m is a typical draw rather than a fixed figure; every other number here is exact.

Unlikely is not impossible

rare Γ— enormous = routine

One in 45 million is a genuinely tiny chance for you. It is not a tiny chance for the draw, because tens of millions of lines are bought each week, and somebody wins most weeks. Nothing surprising has happened when they do.

Scale does this everywhere. An event with a one-in-a-million chance of happening to a given person on a given day happens to about 8,200 people a day on a planet of 8.2 billion. The miracle is in the headcount, not the event.

Randomness looks clumpy, and people expect it not to

a run of six in 100 flips: 81% likely

Ask someone to write down a fake sequence of 100 coin flips and they will alternate too often and avoid long runs. Real coins do not. In 100 fair flips the chance of a run of six or more identical results in a row is 80.7% β€” a calculation, not an estimate β€” and a run of five is near certain at 97.2%.

Scattered points cluster for the same reason: true randomness has no rule against bunching. Streaks in sport, clusters of illness on a map and a shuffle that plays two songs by the same band are all usually this and nothing more.

What a p-value does and does not mean

read this one twice

A p-value is the probability of seeing a result at least as extreme as yours if the effect you are testing for does not exist. It is a statement about the data given a hypothesis.

It is not the probability that the hypothesis is false, not the probability the result was a fluke, and not a measure of how large or important the effect is. A p of 0.04 on a tiny sample of an implausible claim is much weaker evidence than the number suggests, which is most of why so many published findings fail to replicate.