Chance is not the opposite of order β it is where a very reliable kind of order comes from. Below is a real experiment you can run as many times as you like, from a single trial up to a million. Watch the bars climb towards the theoretical line, then look carefully at the two error numbers, because they move in opposite directions and almost everyone guesses one of them wrong.
Both measure the same disagreement with theory β one in whole trials, one as a share of them. Keep running and they part company.
The gap measured in whole trials. It grows roughly with the square root of the number of trials β the dashed line is that prediction, not a fit.
The same disagreement expressed as a share of all trials. It shrinks, roughly with one over the square root of the number of trials.
After a run of tails, people say heads is βdueβ. It is not. The coin has no memory and no mechanism for repayment.
What actually happens is dilution. The count of surplus heads or tails wanders further from zero as you flip more β the left-hand chart climbs β while that same surplus becomes a smaller and smaller slice of the total, so the proportion settles at one half.
Nothing corrects the imbalance. It is simply outgrown. That is the whole of the law of large numbers, and it is why a losing streak is not evidence that a win is on its way.
Do not take these on trust. Each one below is a simulation you run yourself, with the exact mathematics drawn alongside so you can check that the two agree.
Asked how many people you need in a room before two share a birthday, most people say somewhere near 183 β half of 365. The right answer is 23. The trick is that you are not comparing yourself with everyone else; you are comparing everyone with everyone, and 23 people make 253 pairs.
The simulation gives each person a birthday drawn uniformly from 365 days. Real birthdays are slightly bunched β September is busier than February β and that bunching only makes a match more likely, so 23 is if anything a shade pessimistic.
Three doors, one car. You pick one. The host β who knows where the car is β opens a different door to show a goat, then offers you the swap. Two doors left, so it feels like an even chance either way. Play it and count.
Pick a door to start.
Your first pick is right one time in three. That never changes. So the other two doors hold the car two times in three, and the host has helpfully emptied one of them for you. Switching collects the whole of that two-thirds.
Here are ten thousand people. The condition is rare, the test is good, and most of the positive results are still false. Count them yourself: every dot below is one person, and the dots are grouped so you can see the two kinds of positive side by side.
A Galton board is a plank of pegs. Every ball meets twelve of them and bounces left or right at each one, which is twelve coin flips and nothing more. No ball is told where to land, yet the pile at the bottom always takes the same shape.
The bounce at each peg is a fair coin flip, not a physics collision β the ball's path is drawn, the decision is real. The pale line over the bins is the exact binomial distribution for twelve flips, which is the shape the board has to make. The bins open with 1,500 balls already counted; empty them to watch the pile build from nothing.
Twelve fair flips have a mean of 6 and a standard deviation of β(12 Γ ΒΌ) = 1.732. The board finds those numbers on its own, by dropping beads down a plank.
Add more rows and the steps get finer until the outline is indistinguishable from the smooth curve below. That is the same central limit theorem the ten-dice experiment showed at the top of the page.
Shading shows one, two and three standard deviations either side of the mean. Those bands always hold 68.27%, 95.45% and 99.73% of everything, whatever the mean and spread are β that is the only thing you need to remember about the curve. The vertical scale is fixed, so widening the curve flattens it: the area underneath is always exactly 1.
Beyond three standard deviations lies 27 people in every 10,000. Beyond four, 6 in every 100,000. The tails thin out ferociously fast, which is exactly why the curve is the wrong model for anything that produces genuine extremes.
Adult height is settled by hundreds of genes and a childhood of nutrition, each contributing a little in either direction. Add up enough small independent nudges and you get a bell β British men average about 175 cm with a spread of roughly 7 cm, and nobody is twice the average height.
Weigh the same object a thousand times and the readings scatter symmetrically about the true value, because each reading collects a heap of tiny independent errors. Astronomers noticed this before anyone had a name for it, which is why it was once called the error curve.
Median full-time pay in the UK was Β£37,430 in 2024. The mean is several thousand pounds higher, dragged up by a thin tail of very large salaries. When the mean and the median disagree that much, the distribution is skewed and the mean is describing almost nobody.
London holds about 8.9 million people; Birmingham, the next largest, about 1.1 million. An eightfold gap between first and second place is impossible under a bell curve and completely ordinary under a power law. Word frequencies, earthquake energies and book sales behave the same way β the "average" of any of them is a number to distrust.
0.5 is an even chance, 0.001 is one in a thousand. Percentages are the same number times a hundred. Odds are a different notation for the same thing: "5 to 1 against" means one favourable outcome for every five unfavourable, so a probability of 1/6 β not 1/5. The probabilities of every possible outcome always add to exactly 1, which is a surprisingly good way of catching your own mistakes.
Two events are independent when one happening tells you nothing about the other. Coins, dice and roulette wheels are independent: no result changes the next one.
Cards are not. Drawing two aces from a shuffled deck is 4/52 Γ 3/51 = 1 in 221, because the first ace is gone. If you replaced it and reshuffled, it would be 4/52 Γ 4/52 = 1 in 169. Same deck, same question, different answer β because the second draw remembers the first.
Multiply each prize by its chance and add up the lot. There are C(59,6) = 45,057,474 possible tickets, and the fixed prizes are these:
| Match | Chance | Prize | Worth |
|---|---|---|---|
| 2 | 1 in 10.3 | Β£2* | Β£0.195 |
| 3 | 1 in 96.2 | Β£30 | Β£0.312 |
| 4 | 1 in 2,180 | Β£140 | Β£0.064 |
| 5 | 1 in 144,415 | Β£1,750 | Β£0.012 |
| 5 + bonus | 1 in 7,509,579 | Β£1m | Β£0.133 |
| 6 | 1 in 45,057,474 | Β£5mβ | Β£0.111 |
| Every Β£2 ticket | Β£0.83 |
You get back about 41p in the pound. The jackpot would have to reach Β£58 million before a Β£2 ticket broke even β and jackpots that large sell so many tickets that you would probably be splitting it.
* a free Lucky Dip, counted at its Β£2 face value. β jackpots roll over, so Β£5m is a typical draw rather than a fixed figure; every other number here is exact.
One in 45 million is a genuinely tiny chance for you. It is not a tiny chance for the draw, because tens of millions of lines are bought each week, and somebody wins most weeks. Nothing surprising has happened when they do.
Scale does this everywhere. An event with a one-in-a-million chance of happening to a given person on a given day happens to about 8,200 people a day on a planet of 8.2 billion. The miracle is in the headcount, not the event.
Ask someone to write down a fake sequence of 100 coin flips and they will alternate too often and avoid long runs. Real coins do not. In 100 fair flips the chance of a run of six or more identical results in a row is 80.7% β a calculation, not an estimate β and a run of five is near certain at 97.2%.
Scattered points cluster for the same reason: true randomness has no rule against bunching. Streaks in sport, clusters of illness on a map and a shuffle that plays two songs by the same band are all usually this and nothing more.
A p-value is the probability of seeing a result at least as extreme as yours if the effect you are testing for does not exist. It is a statement about the data given a hypothesis.
It is not the probability that the hypothesis is false, not the probability the result was a fluke, and not a measure of how large or important the effect is. A p of 0.04 on a tiny sample of an implausible claim is much weaker evidence than the number suggests, which is most of why so many published findings fail to replicate.