Regularisation: Shift and Spread

Class 5, module one of three. Why a model that is a little bit wrong on purpose can beat one that is right on average.

The model says email loses money

The meeting is at ten. You have two years of weekly numbers for one e-commerce brand. That is a hundred and four rows. Revenue on the left. On the right, spend on Google Search, Google Shopping, Meta, TikTok, display and affiliate, plus how many marketing emails went out that week.

You add adstock lags, because money spent this week keeps selling next week, and everybody who builds these models includes them. Thirteen weekly lags per channel, seven channels. Ninety-one columns. Still a hundred and four rows.

You fit least squares. It takes half a second. The coefficient on email comes back at minus four point two. Two rows further down, TikTok sits at plus eleven point eight.

Minus four point two means that every send loses money. At ten o'clock you have to say that out loud to the woman who runs email. She has four people and a campaign calendar and she built the welcome sequence everybody quotes.

Nobody in the room believes it. Nobody in the room can argue with it either. The arithmetic is correct, and while you are still talking two people check it on their laptops and get the same answer.

Waiting does not help. The data arrives at one row per week and there are fifty-two weeks in a year.


Least squares has nothing left to say

Budgets are not set independently. Somebody decides in September that the fourth quarter is going to be big. Then Search goes up, Shopping goes up, Meta goes up, TikTok goes up and the email calendar fills. All in the same weeks. January is the same thing in reverse.

Now ask least squares what each channel did on its own. The columns never moved separately, so nothing in the file records what happens when they do. Least squares does not stop and tell you that. It minimises error, and it finds that a large positive number on TikTok and a large negative number on email produce the same fitted line as two modest sensible ones.

Those two coefficients are not two findings. They are one finding, cut in half and given opposite signs. They cancel every week.

Here it is on five rows you can check by hand. Search and Shopping are the same column, because in this brand the two budgets are set in the same meeting by the same person.

Week Search spend Shopping spend Revenue
1112
2224
3336
4448
55510

Try two on Search and zero on Shopping. Week one predicts two, and revenue was two. Week five predicts ten, and revenue was ten. Every week exact. Now try zero and two. Exact. Now one and one. Exact. Now a thousand and minus nine hundred and ninety-eight: week three predicts three thousand minus two thousand nine hundred and ninety-four, which is six, and revenue was six.

Search coefficient Shopping coefficient Error on all five weeks
200
020
110
1000−9980

Any two numbers that add up to two fit this data perfectly. There are infinitely many of them and the data prefers none of them.

The data cannot choose. Something else has to. That is what the rest of this module is about.


Shift and spread

Forget coefficients for a minute. You are throwing darts.

Two things can go wrong. The darts land in a tight group in the wrong place. Or they land all over the board with their average in the right place. Those are different problems and they take different repairs.

Call the first one shift: how far the centre of your group sits from the bullseye. Call the second one spread: how far the darts sit from each other.

Every shot below is one refit. Same brand, same model, a slightly different two years of weeks. The shot marks where the model landed. In real life you get one shot. The cloud is what you would have got.

Drag lambda, the amount of restraint. On the left the shots move. On the right the same shots are flattened onto one line, so you can read the two quantities apart: where the centre of the cloud sits, and how wide the cloud is.

The board: one shot per refit

The same shots on one line

λ = 0.00
Shift0.00
Spread1.00
Total error1.00
$$\text{total error} \;=\; \text{shift}^2 \;+\; \text{spread}$$

The bar on the right panel is one standard deviation. Spread in that sum is that bar squared.

Commit before the curve is drawn. Which setting has the lowest total error?

Shift squared goes up the whole way. Spread comes down the whole way. Their sum is the thing you are paid to make small, and a rising curve plus a falling curve does not have to be smallest at the left edge.

Those two words have proper names. Shift is bias. Spread is variance. There is a third piece, noise, which is the part of next week's revenue that nothing in your file predicts. Noise does not move, whatever you do to the model.

You turn the restraint up. What happens to shift and to spread?

The bargain

You accept a little shift to buy a lot less spread.

At the left edge of that slider the model is right on average and wild in practice. You get one draw from it. Being right on average is a statement about infinitely many parallel universes, and that is not a comfort anybody has needed on a Tuesday.

A wrong-but-steady model can beat a right-on-average-but-wild one. That sentence is the whole idea of a penalty.

A penalty is a price on size. Until now the model was asked one question: how close can you get to the data. From here it is asked two. How close can you get, and how big are the numbers you used to get there. The second question has a price attached, and the price is the dial you just dragged.

Hoerl and Kennard proved something stronger in 1970. There always exists a penalty greater than zero whose total error is lower than the total error of least squares. Not usually. Always, for every dataset that has ever been collected and every one that ever will be. They cannot tell you which penalty it is. It is there anyway.


DuPont, and twenty years of the same complaint

Arthur Hoerl was hired in 1950 as DuPont's first statistician. Robert Kennard joined five years later. The job meant plant logs rather than textbook data. Yield, temperature, pressure, feed rate, catalyst concentration, recorded by people who were running a chemical process and not designing an experiment.

Twenty years later the two of them published the complaint. Their words:

"the least squares estimates often do not make sense when put into the context of the physics, chemistry, and engineering of the process"

Regress yield on temperature, pressure, feed rate and catalyst. The engineers standing next to the reactor already know the sign of every one of those. Least squares hands back a number saying that heating the reactor lowers the yield. Not because it does. Because temperature and pressure move together in the plant log, every day, all year, and nothing in the log records a day when one moved without the other.

Nobody at DuPont made a mistake and nobody in your marketing meeting made a mistake. A business that changes one channel at a time while holding the others fixed is a business that is not being run. When a calculation and a person with twenty years on the floor disagree, the useful question is what the calculation was fed.


The picture an engineer looked at

Cross-validation did not exist yet. Neither did grid search, and nobody had a validation fold. The model-selection rule in the 1970 paper was this: plot every coefficient against the penalty, all on one set of axes, and look at it.

That plot is called the ridge trace. Its stated purpose was to show how sensitive the answers are to the data, and to let the analyst pick the point where the system stabilises.

The ridge trace, built from the marketing model at the top of the page. Seven coefficients, one horizontal axis, no cross-validation anywhere.

Read it from the left. The penalty is zero and the coefficients are flailing. Email at minus four point two, TikTok at plus eleven point eight. Absurd on both ends, and absurd together, because they are cancelling. Now move right. The wild ones come screaming in toward zero, because they were only ever large in order to cancel each other, and the moment size costs anything the deal stops being worth it. The ones with something behind them settle and hold.

Somewhere the whole picture stops thrashing. That is where the engineer put the dial. A picture, not a formula, and it is the forgotten half of the paper.


Where the word ridge comes from

Almost every course says ridge refers to a ridge in the loss surface. It does not, and this one is easy to check, because the paper says where the word came from.

Eleven years earlier Hoerl published "Optimum Solution of Many Variables Equations" in Chemical Engineering Progress, in November 1959. Then "Application of Ridge Analysis to Regression Problems" in 1962. That 1962 paper is reference nine of the 1970 paper, cited there as the origin of the idea.

Ridge analysis was a plant optimisation tool. You have a fitted response surface, some yield as a function of your settings, and you want the best settings you can reach. So you find the best point on a sphere of radius one. Then radius two. Then radius three. You trace those constrained optima outward. When the surface is elongated rather than sharply peaked, that path climbs along a raised spine. A ridge, in the walking-boots sense of the word.

A plausible explanation travels faster than a real one, because it is easier to remember. When somebody gives you the origin of a word, notice whether they are reporting it or reconstructing it. Those two sound identical from the outside.

Where does the word ridge come from?

A correction

What everyone says

"Tikhonov invented regularization in 1943"

He did not, and the paper is four pages long and free, so you can settle this yourself in twenty minutes.

The 1943 paper is a stability theorem. It proves that stability of an inverse problem follows from the uniqueness theorem, by a topological mapping argument. There is no penalty in it. There is no stabilizing functional. Nothing anywhere in it is being penalised. Its context is geophysics, which is to say guessing what is underground from what you can measure standing on top of it.

Regularization proper, with a stabilizing functional, is 1963. And David Phillips got there in 1962, in the Journal of the ACM, which is why careful people write Phillips-Tikhonov.

The error is not the interesting part. Hoerl and Kennard's reference list cites neither Tikhonov nor Phillips. The only antecedent they cite is Hoerl's own 1962 paper. Two research communities, Soviet ill-posed inverse problems and American industrial statistics, arrived at the same estimator about seven years apart, for unrelated reasons, without contact.

When two people who never met arrive at the same answer, the answer was sitting inside the problem the whole time, waiting for anybody who turned up with the right complaint.


For the curious: batting averages and the price of tea in Shanghai

In May 1977 Bradley Efron and Carl Morris published a piece in Scientific American about eighteen baseball players. They took the first forty-five at-bats of the 1970 season and asked one question. What will each man hit for the rest of the year?

The obvious answer is his own average. Roberto Clemente was hitting .400, so predict .400.

They did something else. They took the grand mean across all eighteen players.265, and pulled every player toward it by a factor of .212. You can check it in one line:

$$.265 + 0.212 \times (.400 - .265) = .2936$$

Two ninety-four, for a man who was hitting four hundred. Clemente's actual average over the rest of that season was .346. Four hundred was too high. Two ninety-four was closer.

The same procedure beat the raw average for sixteen of the eighteen players, and total squared error on the raw averages was about three and a half times worse. Those last two figures come from a source quoting the article rather than from the article itself, so this page says so out loud instead of printing them clean.

The theory is from 1956. Charles Stein showed that the obvious estimator of the mean of a multivariate Gaussian, which is to use the sample mean separately for each coordinate, is inadmissible in three or more dimensions. There exists another estimator whose total squared error is lower for every possible value of the true means. Not lower on average over some prior. Not lower usually. Lower always.

The theorem does not require the quantities to be related.

Estimate a batting average, the price of tea in Shanghai and the mass of an electron at the same time. Shrink all three toward their common mean. You beat estimating each one separately. The three have nothing to do with each other and it does not matter, because the result is about the geometry of estimating several things at once.

Your best week and your worst review are the same arithmetic. Every extreme reading is extreme partly because something real happened and partly because that was the week the dice landed that way. The truth about you sits closer to the middle than any single reading says, and that applies to the good readings exactly as hard as the bad ones.


The penalty needs a shape

You have the idea now. Of all the answers that fit the data, prefer the small one. What is missing is what small means, and there are two famous answers.

One keeps every feature and makes them all small. Nothing ever reaches exactly zero, so you come out holding ninety-one channels and trusting none of them individually.

The other deletes features outright. Coefficients land on zero and stay there, so you come out with a short list you can read aloud in a meeting.

Those are two different meetings. Keep everything and trust nothing on its own is one conversation. Drop TikTok is a budget line somebody loses. Choosing between them is not only a technical preference.