"There are three kinds of lies: lies, damned lies, and statistics.", Mark Twain
Take a set of numbers. Add them up. Divide by how many there are. That is the mean. It is the balancing point of the data, the value where everything would sit in equilibrium if placed on a seesaw.
A shop with daily revenues of 200, 400, 300, 500, 100 has a mean of 300. That number hides a lot. It says nothing about how irregular the days were. That is what variance is for.
Variance measures how spread out the data is around the mean. For each value, compute how far it sits from the mean, square that distance, and average the result. Squaring does two things: it makes all distances positive, and it punishes large deviations more than small ones.
Standard deviation \(\sigma\) is just the square root of variance. It brings the number back to the original unit, so you can read it as a typical distance from the mean.
| xi | xi − μ | (xi − μ)² |
|---|---|---|
| , |
Two datasets can share the same mean and look nothing alike. Variance tells you whether the data clusters tightly or sprawls. Still, both mean and variance describe a single variable in isolation. When you have two variables, a third question appears.
When one variable changes, does another tend to change with it? A lemonade stand that sells more on warm days has something going on between temperature and sales. A car that loses value as mileage increases has something going on between kilometres and price. The question is whether that something is consistent, and in what direction.
Start with covariance. For each observation, take the deviation of x from its mean and the deviation of y from its mean, multiply them together, and average the result.
When x and y are both above their means, the product is positive. When one is above and the other below, it is negative. Average those products: a positive result means they tend to move together, a negative result means they tend to move in opposite directions.
The problem with covariance is the unit. Multiply euros by kilograms and the number is hard to read. So divide by both standard deviations to normalise everything onto a fixed scale. That gives Pearson's r.
The result always sits between -1 and 1. At 1, a perfect positive linear relationship. At -1, a perfect negative one. At 0, no linear relationship. The word linear is doing real work in that sentence. A curved relationship, a cluster with one extreme outlier, a vertical spike of points: all of these can produce the same r while looking completely different from each other.
In 1973, statistician Francis Anscombe built a trap. He constructed four datasets.
Calculate the following for each dataset below. Mean of x. Mean of y. Variance of x. Variance of y. Correlation r between x and y. Compare your results across all four.
| i | x | y |
|---|---|---|
| 1 | 10 | 8.04 |
| 2 | 8 | 6.95 |
| 3 | 13 | 7.58 |
| 4 | 9 | 8.81 |
| 5 | 11 | 8.33 |
| 6 | 14 | 9.96 |
| 7 | 6 | 7.24 |
| 8 | 4 | 4.26 |
| 9 | 12 | 10.84 |
| 10 | 7 | 4.82 |
| 11 | 5 | 5.68 |
| i | x | y |
|---|---|---|
| 1 | 10 | 9.14 |
| 2 | 8 | 8.14 |
| 3 | 13 | 8.74 |
| 4 | 9 | 8.77 |
| 5 | 11 | 9.26 |
| 6 | 14 | 8.10 |
| 7 | 6 | 6.13 |
| 8 | 4 | 3.10 |
| 9 | 12 | 9.13 |
| 10 | 7 | 7.26 |
| 11 | 5 | 4.74 |
| i | x | y |
|---|---|---|
| 1 | 10 | 7.46 |
| 2 | 8 | 6.77 |
| 3 | 13 | 12.74 |
| 4 | 9 | 7.11 |
| 5 | 11 | 7.81 |
| 6 | 14 | 8.84 |
| 7 | 6 | 6.08 |
| 8 | 4 | 5.39 |
| 9 | 12 | 8.15 |
| 10 | 7 | 6.42 |
| 11 | 5 | 5.73 |
| i | x | y |
|---|---|---|
| 1 | 8 | 6.58 |
| 2 | 8 | 5.76 |
| 3 | 8 | 7.71 |
| 4 | 8 | 8.84 |
| 5 | 8 | 8.47 |
| 6 | 8 | 7.04 |
| 7 | 8 | 5.25 |
| 8 | 19 | 12.50 |
| 9 | 8 | 5.56 |
| 10 | 8 | 7.91 |
| 11 | 8 | 6.89 |
You have done the calculation. The four datasets produce the same numbers. Enter the password to see what the numbers are hiding.
Enter the password to reveal the plots.
His point was not subtle: a number is not a description. Before you report a statistic, you need to see what you are summarising. A correlation of 0.816 can describe a clean linear trend, a perfect curve, a linear relationship with one rogue point, or a vertical stack of values with a single outlier pulling the line. All four look nothing alike. All four produce r = 0.816.
Four datasets. The dashed line is the same regression line on every chart. The points are not.
Think of a report, a presentation, or a dashboard you have seen. A number was shown. A correlation, a mean, a trend line. Was the raw data shown alongside it? What might a plot have revealed that the number did not?