MEDIA LITERACY

Confounding Variable: A Treatment That Won Every Subgroup and Lost

A 1986 study found one treatment beat the other on small stones and on large stones, then lost once both groups were combined. That reversal has a name.

LAST UPDATED 2026-09-04

Bar chart of Charig et al.'s 1986 BMJ kidney stone study. Open surgery beats PCNL on small stones (93% vs. 87%) and large stones (73% vs. 69%), but PCNL shows the higher combined rate (83% vs. 78%) once both size groups are merged.

CORE SUMMARY

A confounding variable is a hidden third factor connected to both the thing being tested and the outcome, and it can flip a result once you stop looking at subgroups separately. The clearest documented case is a 1986 British Medical Journal study of two kidney stone treatments: open surgery succeeded 93% of the time on small stones (81 of 87) and 73% of the time on large stones (192 of 263), beating percutaneous nephrolithotomy (PCNL) in both categories, 87% and 69%. Combined across all 350 patients per treatment, though, PCNL's overall rate, 83% (289 of 350), beat open surgery's 78% (273 of 350). PCNL only looked better because doctors had used it on far more of the easier, small-stone cases, 270 of its 350 patients versus open surgery's 87. Stone size was the confound: correlated with which treatment a patient got, and independently deciding how likely that treatment was to work. The same reversal, real inside every subgroup, gone once the subgroups are combined, is called Simpson's paradox, and it has shown up in university admissions data and in decades of research on coffee and cancer.

The short version

A confounding variable is something that affects both sides of a comparison you're making, without you realizing it's in the picture. Ice cream sales and drowning deaths rise and fall together through the year, and neither one causes the other. Both track summer heat. Pull the temperature out of the picture and the connection between ice cream and drowning disappears.

That version is easy to spot once someone points it out. The harder version is when the confound doesn't just weaken a correlation, it reverses one. A group can win a comparison in every single subcategory and still lose when the categories get combined, because the confound decided which cases ended up in which group to begin with. It's the same kind of trap as an argument whose logical form looks airtight while a premise underneath it is quietly false: the surface structure of the comparison is sound, the number feeding into it isn't measuring what it claims to measure.

The surgery that won every subgroup and still lost the study

In 1986, four London urologists published a comparison of two ways to remove kidney stones in the British Medical Journal: open surgery, the older and more invasive option, against percutaneous nephrolithotomy (PCNL), a newer technique that reaches the stone through a small puncture instead of a full incision. They tracked 350 patients who got each treatment and recorded success by stone size.

For small stones, open surgery succeeded in 81 of 87 cases, 93%. PCNL succeeded in 234 of 270 cases, 87%. Open surgery wins. For large stones, open surgery succeeded in 192 of 263 cases, 73%. PCNL succeeded in 55 of 80 cases, 69%. Open surgery wins again. Then look at all 350 patients per treatment combined, no split by stone size: open surgery's overall rate is 78% (273 of 350). PCNL's overall rate is 83% (289 of 350). The treatment that lost in both size categories now has the higher number.

Bar chart of Charig et al.'s 1986 BMJ kidney stone study. Open surgery beats PCNL on small stones (93% vs. 87%) and large stones (73% vs. 69%), but PCNL shows the higher combined rate (83% vs. 78%) once both size groups are merged.

Why the reversal happens: stone size decided both the treatment and the outcome

Nobody in this study faked a number. The reversal happens because stone size wasn't handed out at random between the two treatments. Of open surgery's 350 patients, 263 (75%) had large, harder-to-treat stones. Of PCNL's 350 patients, 270 (77%) had small, easier-to-treat stones. Physicians were more likely to reach for the newer, less invasive PCNL on the cases that were already easier to resolve, and more likely to use the established, invasive procedure on the harder ones.

That single fact, stone size, meets both conditions that make something a confounding variable rather than a coincidence: it's correlated with which treatment a patient received, and it independently affects the odds of success regardless of which treatment gets used. Average the results without accounting for it, and PCNL's total gets padded with easy wins while open surgery's total gets weighed down with the hard cases it was disproportionately assigned. The subgroup numbers were never wrong. The combined number was measuring stone-size mix as much as it was measuring the treatments.

The same shape of error, outside medicine

The same reversal turned up in a 1973 University of California, Berkeley graduate admissions dataset. Of 8,442 men who applied, 44% were admitted. Of 4,321 women who applied, 35% were admitted, a gap wide enough that Berkeley was sued for sex discrimination. Statisticians Peter Bickel, Eugene Hammel, and J. William O'Connell re-analyzed the same applications broken out by department instead of pooled together, and published the result in Science in 1975.

Once department was accounted for, the pattern didn't just shrink. It reversed to "a small but statistically significant bias in favor of women," in the paper's own words. Women had disproportionately applied to the most competitive departments on campus, the ones with low admission rates for men and women alike (English was one cited example), while men had disproportionately applied to departments with much higher admission rates for everyone, engineering among them. Which department someone applied to was the confound: correlated with gender in the applicant pool, and independently a huge driver of whether any given application got accepted, regardless of the applicant's gender.

A confound that took three decades to fully correct

Not every confound gets caught inside one dataset. In 1981, epidemiologist Brian MacMahon and colleagues published a case-control study in the New England Journal of Medicine linking coffee drinking to pancreatic cancer, questioning 369 patients with the disease and 644 hospital controls about their coffee, tea, alcohol, and tobacco habits. The paper reported the coffee association held up even after adjusting for cigarette use, and the finding made international headlines.

Later scrutiny found a different problem with the study: many of its hospital controls had been referred by physicians for gastrointestinal complaints and had been told to cut back on coffee, which meant the comparison group's coffee habits weren't representative of the general population to begin with. Smoking remained the harder confound to fully rule out, because heavy smokers tend to drink more coffee and smoking is independently one of the strongest known risk factors for pancreatic cancer. It took large, later prospective studies, including the UK's Million Women Study, to test the link specifically among people who had never smoked, and a 2012 meta-analysis pooling 54 studies and 10,594 cases (published in Annals of Oncology) found no appreciable overall association between coffee and pancreatic cancer once smoking was properly accounted for. A hypothesis that made front-page news in 1981 took roughly three decades of further research, much of it aimed at isolating smoking as the real driver, to work its way back out of the medical literature. A single well-designed study can carry a number that sounds more settled than the underlying method actually supports, and coffee-and-cancer headlines rode that gap for years before the correction caught up.

How to actually check for one

A number that compares two groups is only as good as whatever decided who ended up in each group. Before treating a comparison as settled, it's worth asking three things. First: was assignment to each group actually random, or did something, doctor judgment, self-selection, geography, income, decide who ended up where? The kidney stone study wasn't randomized, physicians chose the treatment, which is exactly how stone size snuck in as a confound. Second: is there a third factor plausibly connected to both sides of the comparison, not just one? Stone size affected treatment choice and treatment success. Department choice affected both gender distribution and admission odds. A real confound has to touch both variables, not just one of them.

Third, and easiest to check from the outside: does the reported number hold up inside subgroups, or does it only exist in the combined total? Reporters and researchers who stratify by the plausible confound, whether that's department, stone size, or smoking status, and still find the same pattern have a much stronger claim than one built entirely on an aggregate. That's the same instinct behind checking who produced a number and what got left out of it: a hidden third variable is one of the most common things left out.

Surveys have their own close relative of this problem: nonresponse bias, where the hidden factor isn't a third variable acting on a comparison, but whether the decision to answer a specific question is itself correlated with the answer. A 2017 Pew Research benchmarking study found that same signature, a low-response phone poll landing within a point or two of high-response government data on some questions and missing by 38 points on others, depending entirely on which question was being asked.

Frequently asked questions

What is a confounding variable, in plain terms?

A third factor linked to a comparison's two sides at once, correlated with what you're studying (a treatment or a group, say), that also shapes the outcome you're measuring on its own. It can weaken it, inflate it, or (in the Simpson's paradox cases) fully reverse an apparent relationship.

What is Simpson's paradox and how does it relate to confounding?

Simpson's paradox is what happens when a trend appears in every subgroup of a dataset but disappears or reverses once you merge those subgroups back together. It's caused by a confound, in the surgery-versus-PCNL example above that confound was stone size, that's unevenly distributed across the groups being compared.

What's the difference between a confounding variable and a mediating variable?

A confound sits outside the relationship being studied and distorts it from the outside, like stone size distorting the kidney stone comparison. A mediating variable sits inside the causal chain and explains how one thing leads to another, it's part of the actual mechanism, not an outside factor obscuring it.

How did the kidney stone treatment actually reverse when the data were combined?

Per the BMJ's 1986 paper tracking 350 patients on each treatment: small-stone success was 93% for open surgery against 87% for PCNL, and large-stone success was 73% against 69%, open surgery ahead both times. But PCNL had been used on 270 of its 350 patients for the easier small-stone cases, while open surgery took on 263 of its 350 patients with the harder large-stone cases. That case-mix difference gave PCNL the higher combined rate, 83% vs. 78%, even though it lost both individual comparisons.

How do researchers actually control for confounding variables?

The strongest method is randomization, assigning subjects to groups by chance so a confound can't correlate with group assignment in the first place, which is standard in randomized controlled trials. When randomization isn't possible, as in the kidney stone and Berkeley admissions data, researchers instead stratify results by the suspected confound (like stone size or department) or use statistical adjustment to estimate what the comparison would look like with the confound held constant.

Why did it take so long to sort out whether coffee causes pancreatic cancer?

Because the leading suspected confound, smoking, is hard to fully separate from coffee drinking, since heavy smokers are typically heavier coffee drinkers too. It took large studies specifically comparing never-smokers, plus a later Annals of Oncology review that combined 54 separate studies (10,594 cancer cases total), to establish that the original 1981 finding didn't survive once researchers isolated smoking's effect correctly.

READER VERDICT

Did this entry hold up?

Written and edited by the Hollowvane Editorial Team