
Regression to the Mean: 3 Rules to Spot It in Sports, Medicine, and ML

Regression to the mean is the statistical tendency for an extreme measurement to be followed by a less extreme one, simply because that measurement blends a real signal with random noise. It is not a force pulling things back into place. It is what probability predicts once you understand that unusually high or low results are partly luck, and luck does not repeat. The practical lesson: never credit a program, a treatment, or a hot streak with an improvement until you have ruled out this effect through proper controls.
TL;DR:
- Smaller sample sizes and lower measurement reliability amplify the regression to the mean effect, making short-term extreme results less indicative of true change.
- Enrolling participants based on extreme scores guarantees some reversion, which can be mistaken for treatment effects or performance improvements.
- Randomized controlled trials with multiple baseline measurements are the most effective method to differentiate genuine effects from regression to the mean.
- In sports, medicine, and machine learning, validating results across multiple seasons, control groups, or repeated measurements prevents misinterpretation of noise as real progress.
- Recognizing that extreme outcomes are partly noise helps avoid overestimating streaks, short-term wins, or model performance peaks driven by chance.
Table of Contents
- Everyday examples that make regression to the mean click
- The math behind regression to the mean
- How Galton and Pearson uncovered the pattern
- Where people misread the pattern
- How to detect it and rule it out
- Regression to the mean across sports, medicine, AI, and finance
- Why transparent validation matters in sports prediction
- What years of watching models slip past this rule taught me
- Sources
- FAQ
Everyday examples that make regression to the mean click
The clearest way to understand regression to the mean is to watch it happen in familiar settings, where extreme results almost never stay extreme.
Take a classroom test. Students who score unusually high on one exam typically score lower, though still above average, on a retest. It is tempting to think they got complacent. The more accurate explanation is that an unusually high score is a mix of real ability and a favorable roll of the dice: easy questions they happened to know well, a good night’s sleep, lucky guesses. On the retest, the luck component resets, but the ability component stays. The result lands closer to their true average. The same logic runs in reverse for students who scored unusually low.
Sports offers the same pattern under a different name. A basketball player who shoots 55% from three-point range over five games is not suddenly a 55% shooter. League averages hover in the mid-30s, and a five-game sample is small enough that variance dominates. Cooling off toward their real percentage is not a slump. It is the expected outcome once the streak includes fewer lucky bounces. Practical examples from sports and business show this pattern across ticket sales, batting averages, and quarterly business results, wherever a short run of extreme numbers gets mistaken for a new normal.
Clinical research runs into the same trap in a more consequential way. Patients are often enrolled in a study precisely because their blood pressure, cholesterol, or symptom score is unusually high at a screening visit. When they are measured again, some of that severity reflects a bad day rather than their stable baseline. Their numbers often improve at the second visit even with no treatment at all, purely because extreme entry criteria selected for people sitting on a temporary high point.
A few patterns show up across all three settings:
- Extreme scores at time one are partly signal and partly noise, and only the signal persists.
- The smaller the sample or the noisier the measurement, the stronger the pull toward the average at time two.
- Selecting people or teams because they were extreme guarantees some reversion, independent of any treatment or coaching.
Regression toward the mean is a statistical phenomenon in which extreme values are followed by measurements closer to the mean whenever a mix of stable signal and random error is present, a pattern first identified by Francis Galton. That single sentence explains why “sophomore slumps,” “curses,” and “miracle turnarounds” show up so often in headlines.
The single most useful chart for spotting this pattern is a scatterplot with baseline value on the horizontal axis and the change from baseline to follow-up on the vertical axis. When regression to the mean is present, the plot shows a clear downward slope: high baseline scores pair with negative change, low baseline scores pair with positive change, converging toward zero change near the average baseline. Seeing that slope is often the fastest way to confirm the effect before running any formal test.

The math behind regression to the mean
Regression to the mean is not a metaphor. It falls directly out of the algebra of correlated variables, and once you see the formula, the intuitive examples above stop feeling mysterious.
Start with two standardized variables, X (the baseline measurement) and Y (the follow-up measurement), both expressed as z-scores with mean 0 and standard deviation 1. If X and Y follow a bivariate normal distribution with correlation ρ, the expected value of Y given a specific value of X is:
E[Z_Y | Z_X = z] = ρ · z
This is the core expression behind most study-design literature on the topic, including the tutorial by Barnett and colleagues in the International Journal of Epidemiology. Because correlation ρ is always between negative 1 and positive 1, and in most repeated-measurement contexts sits well below 1, the predicted follow-up score is always pulled partway back toward zero, the standardized mean.
The intuition sits entirely in that ρ. A correlation of 1 would mean the two measurements are identical, no noise anywhere, and no regression at all. A correlation of 0 would mean the two measurements are unrelated, so the best prediction for follow-up is simply the average, regardless of how extreme the baseline was. Real-world measurements sit somewhere in between, which is why regression is partial rather than complete.
A numeric example makes the pull concrete. Assume a test-retest correlation of 0.6, a common figure for many psychological and clinical instruments, and a student who scored well above the mean on the first test.
The pattern holds regardless of direction: lower correlation means more regression, and the effect is symmetric for extreme highs and extreme lows.
A frequently cited derivation shows that closed-form sample-size methods can decompose expected pre-post change into an expected regression-to-the-mean component and a residual treatment effect, an approach implemented in tools like the Stata command power one mean_rtm, as detailed in a 2026 methods paper. That decomposition matters because it lets researchers plan a study knowing in advance how much apparent change is mathematically guaranteed before any intervention does a thing.
Two factors change the size of the effect. When the two variables have different variances, the raw-score version of the formula includes a ratio of standard deviations alongside ρ, which can amplify or dampen the standardized pull once you convert back to original units. When the measurement’s test-retest reliability is lower (meaning more of the “signal” is actually noise), ρ drops, and the regression effect grows correspondingly larger. This is why a single blood pressure reading regresses more than the average of three readings taken on separate days: averaging raises the effective reliability of the baseline measure.
- Higher correlation between measurements means less regression toward the mean.
- Lower reliability of either measurement means more apparent regression.
- The effect is proportional to how far the baseline sits from the mean, not a fixed amount.
How Galton and Pearson uncovered the pattern
The phrase traces back to Sir Francis Galton, a Victorian scientist studying inheritance in the late 1800s. Galton measured the heights of parents and their adult children and noticed something odd: unusually tall parents tended to have children who were tall but closer to average height, and unusually short parents had children who were short but, again, closer to average. He observed the same pattern in sweet pea seeds, where the offspring of unusually large seeds tended to be smaller than their parent seeds, though still above the overall average.
Galton originally called this “regression toward mediocrity,” a name that unintentionally suggested some biological force actively pulling extreme traits back down.
Karl Pearson, working with and after Galton, reframed the observation in purely algebraic terms using correlation coefficients, stripping away the implication that anything was being pulled anywhere. Pearson’s version showed the pattern was a mathematical consequence of imperfect correlation between two variables, not a biological tendency toward the ordinary.
A few landmarks in that evolution:
- Galton’s height and sweet-pea studies gave the phenomenon its first documented description and its original, misleading name.
- Pearson’s correlation framework removed the teleology and reduced the effect to algebra.
- The phrase later generalized from heredity to any pair of correlated, imperfectly reliable measurements, which is how it entered psychology, medicine, and sports analytics.
The name stuck even after the biological framing fell away, which is part of why so many people still misread the effect as a causal pull rather than a statistical expectation.
Where people misread the pattern
Confusing regression to the mean with other kinds of statistical reasoning is common, and the mix-ups lead to some genuinely bad decisions.
The gambler’s fallacy assumes that independent random events, like coin flips, owe a correction after a run of one outcome. Regression to the mean involves no such debt. It applies to measurements with a real, stable signal component, where an extreme observation is partly explained by temporary noise that will not repeat. A coin has no signal component to regress toward. Mean reversion in finance is a related but distinct claim, an empirical or model-based assertion that asset prices tend to move back toward a historical average or fair value over time. That is a hypothesis about market behavior, not the mathematical certainty that regression to the mean represents for correlated, noisy measurements.
Plenty of supposed “curses” are regression to the mean wearing a costume. An athlete who wins an award for an outstanding season often performs worse the following year, a pattern popularly blamed on complacency or media pressure. The more mundane explanation is that award-worthy seasons are, by definition, extreme, and extreme seasons regress. The same logic explains why companies featured on “best performer” lists often underperform afterward, and why medical treatments tested only on patients with extreme symptoms look more effective than they are.
Selection on extremes is the mechanism behind nearly every misleading version of this story. Whenever a group is chosen because its members scored unusually high or low, some improvement or decline on remeasurement is baked in mathematically, with no intervention required at all. A weight-loss program that enrolls only the heaviest applicants will show average weight loss at follow-up even if the program does nothing, because the heaviest people at any single point in time include some who were temporarily above their stable weight.
- The gambler’s fallacy involves independent events with no signal to regress toward.
- Mean reversion in finance is an empirical market claim, not a guaranteed statistical property.
- “Curses” following awards or record-breaking performances are usually regression, not decline.
- Enrolling participants because they scored at an extreme guarantees some apparent improvement on its own.
Pro Tip: Before crediting any program, coach, or treatment with an improvement, ask whether participants were selected because they were already at an extreme. If so, assume some of the “effect” is regression to the mean until proven otherwise.
How to detect it and rule it out
Ruling out regression to the mean takes deliberate design choices before data collection and specific analysis techniques afterward. Neither step alone is enough.
Randomized controlled trials remain the most reliable defense, because regression to the mean affects the treatment and control groups equally. If both groups start from the same extreme baseline and are randomly assigned, any regression toward the mean happens in both arms. What is left in the difference between arms is the treatment effect, not a mathematical artifact. This is the core argument in the PMC methods discussion of program evaluation, which notes that ignoring regression to the mean is a common cause of false claims about program effectiveness.
Taking multiple baseline measurements is the second major design fix. Averaging two or three readings before treatment, rather than relying on a single screening value, raises the effective reliability of the baseline and shrinks the regression component correspondingly.
On the analysis side, several techniques help:
- Use ANCOVA (analysis of covariance) with baseline as a covariated rather than a simple paired t-test on the difference scores; the Barnett tutorial notes this generally offers higher statistical power when baseline and follow-up are correlated.
- Apply conditional regression adjustments that explicitly model the expected follow-up value given the baseline, rather than assuming any observed change is real.
- Estimate the expected regression effect directly from a bivariate model using the baseline standard deviation, the correlation between measurements, and how far the sample was selected into the tail.
- Build a diagnostic scatterplot of change against baseline and check for the characteristic downward slope before interpreting any result.
For prospective study design, closed-form sample-size methods that partition expected regression to the mean from the residual treatment effect are increasingly used in cutoff-based, single-arm pre-post studies. These formulas account for the fact that regression reduces the apparent treatment component and changes the conditional variance of the outcome, and they tend to run conservative at very small sample sizes, which is worth knowing before trusting a borderline result.
A short checklist for anyone reviewing a claimed improvement:
- Was the sample selected because of an extreme score, rather than at random?
- Is there a control or comparison group that went through the same selection process?
- Was the baseline measured once, or averaged across multiple readings?
- Does a baseline-versus-change scatterplot show the telltale downward slope?
When any answer raises a flag, treat the reported effect as an upper bound rather than a confirmed result until an appropriate adjustment has been applied.
Regression to the mean across sports, medicine, AI, and finance
The mechanics stay the same everywhere, but the practical response differs by field.
In sports analytics, the danger is chasing short-term noise as though it were a durable trend. A quarterback who throws for 350 yards in three straight games, or a team that covers the spread five times in a row, is producing a sample too small to separate skill from variance. The fix is validation across seasons and holdout samples rather than reacting to the most recent streak, a discipline covered in more depth in a walkthrough of building an NBA prediction model. Sports analytics writeups generally agree that regression to the mean is the main reason hot streaks rarely persist at the same intensity.
In clinical research, enrollment by extreme symptom score is close to unavoidable, since that is often how eligibility criteria are written. The safeguard is a placebo or standard-of-care control arm measured under the same enrollment criteria, so that the regression baked into both arms cancels out when comparing outcomes. Pre/post designs without a control arm are especially prone to overstating benefit for exactly this reason.
In machine learning, regression to the mean shows up as selection bias and what researchers call the winner’s curse: when a model or hyperparameter setting is chosen because it scored best on a validation set, that top score is partly luck specific to that sample, and it typically will not repeat on fresh data. Selection-aware evaluation frameworks described in 2026 research, including post-selection distributional model evaluation (PS-DME) and a method called SIREN, correct for this optimistic shift by estimating the procedure-level performance expected on new data rather than trusting the single observed winner. Advanced practitioners now routinely separate the tuning and search stage from the deployment target, reporting selection-aware metrics instead of the best score seen during search, a shift documented across recent practical analytics writeups on avoiding overfit performance claims.
In finance, it is worth separating two distinct ideas that share a name. Statistical regression to the mean is the mathematical property of any noisy, correlated measurement. Mean reversion in asset pricing is a modeled empirical claim about how prices behave over time, resting on assumptions about market efficiency and historical patterns rather than the guaranteed algebra behind regression to the mean. Treating the two as interchangeable is a common but avoidable error.
- Sports: validate models across multiple seasons rather than trusting a hot or cold stretch.
- Medicine: pair extreme-criteria enrollment with a control arm measured under identical selection rules.
- Machine learning: report selection-aware, fresh-data performance rather than the best score seen during tuning.
- Finance: keep statistical regression to the mean conceptually separate from model-based mean reversion claims.
Why transparent validation matters in sports prediction
Any predictive system that reports an unusually strong short-term result faces the same question this article has been asking throughout: how much of that number is real skill, and how much is regression waiting to happen. A model that goes on a hot run across a small batch of games can look extraordinary for a few weeks purely by chance, the same way a shooter can hit an unsustainable percentage over five games.
The only credible defense is a large, transparent, and permanent sample. A model’s true performance only becomes trustworthy once it has been tracked across enough picks, seasons, and bet types that a short lucky or unlucky run cannot dominate the average. Public grading matters for the same reason peer review matters in research: it prevents cherry-picking the best-looking stretch of results and presenting it as the norm.
The platform keeps a permanent, publicly graded archive of picks made across multiple sports, rather than surfacing only its best-performing stretches. That structure exists specifically to let bettors check performance the way a researcher would check for regression to the mean: over a large sample, not a curated one.
- Look for a track record that includes losses, not only a highlight reel of winning picks.
- Prefer platforms that publish permanent, dated archives rather than editable or removable histories.
- Treat any short hot streak, including your own, as a small sample until it has survived multiple seasons.
- Compare a model’s claimed edge against its full graded history, covered in more detail on the NBA results and track record page and in a guide to clearing the validation gates most bettors skip.
Applying this same skepticism to model performance generally, including the ideas behind selection-aware evaluation in machine learning, is good practice whether you are reading a betting tracker, a fund’s quarterly return, or a study’s headline result.
What years of watching models slip past this rule taught me
Regression to the mean is easy to accept in a textbook and remarkably hard to accept in the moment, when a streak feels real and a headline number feels earned. The pattern I keep seeing, across sports data, clinical papers, and machine learning benchmarks alike, is that the extreme result gets the attention and the underlying sample size gets ignored. A five-game streak, a single clinical trial arm, a top validation score: each one is treated as a discovery when it is often just a snapshot of noise sitting on top of a smaller, less dramatic signal.
Three rules are worth keeping close whenever an extreme number crosses your desk:
- Assume regression to the mean is present whenever a group or model was selected because it scored at an extreme, and adjust your expectations downward before celebrating.
- Trust a control-group comparison over a raw pre/post number every time the two disagree.
- Prefer a result built on repeated measurement over a single, dramatic data point, even when the single data point is more exciting.
Test these rules against the sources below rather than taking this article’s word for it. The math holds up regardless of which field you apply it to, which is exactly why it has survived, largely unchanged, since Galton first stumbled onto it with sweet peas and family height records.
— Manuel
Sources
- Prospective sample-size and RTM decomposition — PMC (2026)
- Regression to the mean: what it is and how to deal with it — IJE (Barnett et al.)
- Regression toward the mean — Wikipedia
- Fs
FAQ
What is the regression tends to the mean?
Regression to the mean describes the tendency for an unusually high or low measurement to be followed by a value closer to average, because extreme results mix a stable signal with temporary noise that does not repeat. It is a mathematical expectation rather than an active force, and it applies whenever two measurements are correlated but imperfectly reliable.
What is regression to the mean in AI?
In machine learning, this effect shows up as selection bias or the winner’s curse: a model chosen for scoring best on a validation set often performs worse on new data, because part of that top score was luck specific to the sample. Selection-aware evaluation methods like PS-DME and SIREN correct for this by estimating expected fresh-data performance instead of trusting the single best result seen during tuning.
How to rule out regression to the mean?
The strongest safeguard is a randomized controlled trial with a control group, since both arms experience the same regression and any real difference between them reflects the treatment. Taking multiple baseline measurements and using ANCOVA instead of a simple pre/post comparison, as recommended in the Barnett tutorial, further reduces the risk of misreading regression as a genuine effect.
What is the regression to the mean strategy?
There is no single “strategy” in the formal sense; the practical approach is to expect extreme results to move back toward average on remeasurement and to avoid crediting any intervention with that movement without a control comparison. Analysts commonly apply this by validating models or programs across large, repeated samples rather than reacting to a single standout result.