
What this covers
Doug Rivers, a political scientist at Stanford and founder of YouGov.com, joins Russ Roberts to examine the gap between what polling margins of error claim and what they deliver. The conversation spans the mechanics of modern polling, the steady decline of telephone response rates, and Rivers' argument that internet panels matched to consumer databases can produce more reliable samples than traditional random-digit-dial surveys. Roberts asks the questions; Rivers walks through the statistical and practical reasoning, tracing how non-response skews poll samples in unmeasured ways, how weighting corrects for some of those skews while inflating others, and why the reported margins of error systematically misrepresent accuracy.
The discussion covers several distinct territories. Rivers opens with the fundamental problem: cooperation rates have fallen from roughly 70% to around 20%, turning ostensibly random samples into systematically biased ones, yet standard confidence intervals pretend the procedure remains sound. He examines concrete cases—the 2008 New Hampshire primary where all thirty polls missed Clinton's win, the 2000 Florida election-night retractions, exit-poll biases in presidential elections—to show how these failures point to structural bias rather than chance error. He defends matched internet panels as a practical alternative, argues that averaging polls does not reliably wash out error if pollsters make similar mistakes, and revisits the Bradley effect, which he shows largely disappeared by the late 1990s. Throughout, Rivers argues that professional polling has grown complacent with outdated weighting methods and that transparency about unmeasured bias matters more than the false precision of reported margins.
Rivers argues that modern polling's reported margins of error are misleadingly precise because falling response rates force heavy weighting that introduces unmeasured bias, and that consumer-data-matched internet panels can match or beat random-digit-dial telephone polls.
- Plummeting phone response rates (now ~20%) mean samples are systematically skewed, not random, so standard statistical theory no longer strictly applies
- Weighting both fails to remove unobserved skews and roughly doubles real variability, yet reported margins of error ignore this
- Matching opt-in internet panels to randomly selected voter/consumer profiles can remove biases better than low-response telephone polls
This asset isn't compiled yet
You're seeing its claims, ranked. Compile it to build the argument threads, weight them, and check each claim against your library — the full view.
The clean statistical results (law of large numbers and central limit theorem) only hold for perfectly executed random sampling; in practice sampling plans are rarely executed perfectly because you never get near 100% cooperation, and the non-responders differ systematically from responders in unobservable ways, producing skews rather than mere random noise.
“the problem is that the sampling plans are rarely executed perfectly”
Calling a 46-40 poll lead (±3% margin of error) a 'statistical dead heat' or 'statistical tie' is incorrect: the best estimate is a 6-point lead, not zero, and while the 95% confidence standard means you can't formally declare a winner, it is far more likely the lead is real than that the true gap is zero, since extreme errors are unlikely.
“your best guess if it's 46 to 40 is it's a six-point lead not a zero point lead”
Pollsters typically select 10 to 20 times as many phone numbers as the number of completed interviews they want, meaning the people who actually respond are a small, non-random slice that skews toward more women, higher education, and higher income, requiring weighting and adjustment.
“it's typically 10 to 20 times as many phone numbers are selected as the number of interviews that you wish to do”
Of about 1,300 polls Rivers examined from the 2008 presidential primaries, fewer than 50 reported anything other than the margin of error for a simple random sample with no weighting—meaning the reported margins of error are systematically misleading, which Rivers calls scandalous.
“I looked at 1300 polls in the presidential primaries in 2008 and fewer than 50 of them were reporting anything other than the margin of error for a simple random sample with no weighting”
Internet panel respondents are compensated with a reward they are not told in advance, and are not told what answers qualify them for a poll, in order to avoid biasing results; this contrasts with the phone tradition of not compensating, where rising marketing-call fatigue may eventually require incentives that introduce their own biases.
“we do compensate people that is they get a reward for taking a poll we don't tell them in advance what that reward is we don't tell them what they need to tell us to qualify for the poll”
Weighting creates two distinct problems: it roughly doubles the real variability compared to a simple random sample, and it leaves residual skews (unobserved or uncorrected biases) that do NOT shrink as sample size grows, so conventional sampling-error estimates tell you how the procedure varies sample-to-sample but not how systematically wrong the procedure is.
“the weighting adds a lot of variability probably about twice what the normal margin of error calculations would suggest”
Weighting underrepresented groups by large factors (5-10x for some groups) drastically inflates variability: with a weight of 10 in a sample of 1,000, a single miscoded respondent shifts the whole sample by 1%, so a single error can destroy a nominal 3% margin of error.
“if you have a weight of ten and a sample of a thousand that means one person is representing ten people if that one person's answer is recorded incorrectly it moves the whole sample by one percent”
Random digit dialing became the dominant phone-polling method because roughly 30% of the population had unlisted numbers, making listing-based sampling infeasible; it worked well for 20-30 years but broke down in the last 10-15 years as marketing calls, cell phones, and distrust drove cooperation rates from ~70% down to ~20% or less, hurting accuracy.
“about 30% of the population has unlisted phone numbers so the most popular method was thing called random digit dialing”
Massive consumer and voter databases that did not exist 25 years ago now provide detailed information (income, home value, etc.) on most people, allowing pollsters to draw a randomly selected target sample from a voter list and then match closely-similar respondents from a large opt-in internet panel, creating a sample that mimics a random sample across many dimensions and removes skews that simple demographic weighting cannot.
“there are now massive consumer databases that give you fairly detailed information about most people in the country that enable you to form samples that can be representative on many dimensions”
Party identification shifted dramatically: Democrats had an ~18-point lead from the New Deal through the 1960s, Republicans closed it to a few points by 2004, but 80% of those Republican gains were erased since 2004 (half of them in fall 2007/spring 2008), which in principle should make 2008 a terrible year for Republicans.
“80% of those Republican gains in the last 30 years have been a race since 2004 including half of them in the fall of 2007 in the spring of 2008”
In the 2008 New Hampshire primary, all roughly 30 pre-election polls failed to show Hillary Clinton ahead of Obama, which she won, indicating something systematically wrong affecting all polls rather than random chance.
“there are approximately thirty polls and the week before the primary and not a single one of those polls had Hillary Clinton ahead of Barack Obama”
The 2000 Florida election-night debacle—calling it for Gore, retracting, calling for Bush, retracting—was caused by a combination of erroneous data feeds, a data-entry error behind the Bush call, and an exit poll too close to have been called; since then networks have staffed decision desks with inside and outside experts and made no state-level miscall.
“the call in Florida in 2000 the one for Bush that shouldn't have been made because the election was too close was due to a data entry error”
Network exit polls (high-quality, ~50% response rate) have consistently overstated the Democratic share of the vote in presidential elections (2000, 2004), but because this bias appeared uniformly nationwide—including areas controlled by neither party—it indicates a systematic polling tendency, not the voter fraud that naive observers inferred.
“they frequently overstated the Democratic proportion of vote and presidential elections”
Apparent poll movements late in a campaign are often illusory: a methodology shift (e.g. Gallup switching from registered-voter to likely-voter weighting) creates a bump unrelated to real opinion change, and pollsters who err have a ready ex-post story of a late undetected surge—what Rivers calls the last refuge of scoundrels in polling—when in fact the same poll repeated would have produced the same wrong answer.
“the Gallup Organization at one point in the middle of a campaign shifts from a registered voter sample to a likely voter sample... there's a bump in those polls that's due to a methodology shift”
Averaging polls (as RealClearPolitics does) does not necessarily wash out error: if pollsters make the same systematic errors or herd toward each other, the average is no more reliable than a single poll; averaging only clearly helps when polls use the same methodology, in which case combining them is equivalent to a larger sample, and ideally polls should be weighted by the reciprocal of their variability.
“an average a meaning estimate of a candidate's strengths is not gonna be any necessarily any better than any one poll”
Weighting pollsters by their historical accuracy (as fivethirtyeight.com does) is questionable because past success may be partly random luck and because pollsters who erred likely changed their methodology afterward, so historical outcomes don't reliably predict future accuracy—analogous to picking mutual funds by past manager returns.
“that's like waiting I don't know stock picks based on the past returns of mutual fund managers”
The Bradley/Wilder effect—the claim that voters tell pollsters they will vote for a black candidate but then don't—was a real overstatement of the black-candidate vote in the early 1980s but disappeared by around the late 1990s; a study Rivers cites found no systematic over-reporting across black congressional and Senate candidates and no comparable 'Whitman effect' for women candidates.
“there is in fact an overstatement of the black candidate vote in the early 80s but it disappeared by about ten years ago”
A truly random sample of about 1,000 people gives a margin of error of about plus or minus 3% with 95% confidence regardless of total population size, so a tiny fraction (one in 300,000 Americans) suffices; quadrupling to ~4,000 tightens it to about 1% at 99% confidence.
“a thousand person sample is is good to about plus or minus three percent or a little less”
The convention bounce—a candidate gaining several points (about five per party, net swing ~ten) during/after their party's convention—is partly an artifact of differential response: supporters of the convening party are energized and answer polls while opposition partisans, irritated by the 'phony love fest,' decline to participate, so YouGov controls for prior party identification to neutralize the effect.
“there's some evidence that a lot of the effect comes from the people whose party is having a convention will answer their phone and participate in the poll and the people of the opposite party... go on strike”
Obama's extraordinarily high favorability in fall 2007 faded by spring 2008 as Hillary Clinton ramped up rhetoric, leaving him looking like a run-of-the-mill (not weak) candidate without the magical appeal he had even among Democratic core constituencies; black Democrats, who initially didn't strongly favor him because of the Clintons' ties to blacks, swung to ~90% support once they recognized he was black.
“initially people didn't know Obama was black and he didn't actually do much better among blacks than among whites because of the Clintons long-standing ties to blacks”
In 2006, the matched internet-panel method produced election forecasts whose average error was substantially less than the average reported telephone poll, mainly because it removed biases better (not because sampling variability was smaller), and Rivers argues telephone polling could be improved by similarly substituting a respondent who resembles the missed person rather than drawing another number from the same skewed population.
“in 2006 the average error on those estimates was quite a bit less than the average telephone poll”
Typical media polls correct only for age, race, and gender (sometimes education, rarely income) using weighting methods about 60 years old, well behind modern statistical theory, partly from a long-standing belief that random digit dialing was good enough.
“the typical media poll is corrected for age race and gender sometimes education income rarely income and in fact is uses weighting methods that are about 60 years old”
Likely the New Hampshire 2008 polling miss came from over-representation of college- and graduate-degree voters (who favored Obama, over-represented by 2-3x and under-corrected) and from unreliable self-reports of voting intention, rather than from racism (the Bradley effect).
“Obama does very well among people with college degrees and graduate degrees and they are over represented by a factor of two or three typically and not enough weighting is done to correct for that”
Some organizations (Associated Press, New York Times) refuse to report polls using nonprobability sampling—respondents selected without known probabilities of selection—but Rivers argues they are in denial because their own low-response telephone polls also lack known selection probabilities.
“there's some organizations like The Associated Press New York Times that will not report a poll that uses nonprobability sampling methods”
Self-reported voting intention is unreliable: about 90% of people say they will vote in a primary but actual turnout is much lower (US presidential turnout has hovered in the mid-50% range, congressional under 40%), and likely-voter screens cannot fully fix this because people answer based on what they think they should say.
“if you call people up and ask are you going to vote around 90 percent of the people will tell you they're gonna vote the actual percentage that they're gonna vote is much smaller than that”
Roughly 80-90% of polling errors come from sample skews rather than from sampling (random) error.
“there are these skews that probably account for 80 percent maybe nine percent of the errors and polls”
Final pre-election polls tend to be too similar to each other relative to their known sampling variability, suggesting pollsters herd—choosing among multiple defensible weighting schemes the one that places them in the pack rather than the one that yields a divergent answer—though Rivers attributes this to risk-aversion and the art of weighting rather than dishonesty.
“if you look at final polls they're too similar relative to we what we know their sampling variability is and that suggests that given the choice of two weighting schemes one that puts you in the pack of other polls it's probably hey that looks like a better weighting scheme”
There would seem to be an incentive for a pollster to trust divergent weights and look like a genius (as the lone correct caller of an upset), but in practice small subsamples (e.g. a NH poll with ~100 Republicans) make divergent results look like noise that gets discounted, which is why having 30 polls all wrong in New Hampshire was an extraordinary wake-up call.
“most of the samples people were dealing with in New Hampshire were fairly small I remember talking to one pollster that had something like a little over a hundred Republicans”
Experiments Rivers ran with Shanto Iyengar (lightening/darkening Obama's skin, and swapping black vs white candidate photos in a hypothetical congressional race) found essentially no candidate-race effect on voting, though about 15% of Democratic primary voters—even outside the South—held racial views considered conservative or racist.
“we lightened and darkened Obama skin to see if there was any skin color effect word voting we also took a hypothetical congressional election race and put a black candidate in and the same candidate description but a picture of a white candidate and we didn't really find any effects”
Interactive voice response (robocall) polls make up roughly 80-90% of polling done in 2008 because newspapers cut back on expensive live-interviewer polling, and despite low response rates and not knowing who is answering, their record is not bad—largely because IVR organizations pay closer attention to weighting than traditional phone pollsters.
“that's probably you know 80 or 90 percent of the polling being done this year is newspapers cut back on traditional interview or run polling which is much more expensive”
Rivers agrees with Brian Caplan that the average voter couldn't pass Econ 101—reacting to problems by passing a law to ban them without thinking about consequences and blaming convenient targets—as illustrated by a YouGov poll where many people blamed 'speculators' for gas price increases despite probably not knowing what a speculator is, and a sample of Congress would likely respond similarly.
“your average voter couldn't pass act 10 you know if you put out sort of standard view use of things like trade and price floors and ceilings”
Telephone polling, which arose in the 1970s, replaced in-person area probability sampling because in-person interviewing (door-to-door) was the dominant method from the 1930s until then.
“from about the 1930s when polling started to the 1970s interviewing was done in person”
The fraction of people with internet access at home, school, or work now probably exceeds the number with landline access once you exclude cell-phone-only households and those with no phone, making internet-based polling a viable alternative to telephone polling.
“the fraction of people who have internet access at home school or work is these days probably exceeds the number of people who have landline telephone access when you take out the cell phone only”
Welcome to EconTalk and the closing sign-off acknowledging the sound engineer.
“welcome to econ talk part of the library of economics and liberty I'm your host Russ Roberts”