Doug Rivers
About
Professor of political science at Stanford, senior fellow at Hoover Institution, polling expert, affiliated with YouGov
Cast within
No topic-region cast yet — this appears once Doug Rivers's compiled claims are aligned into a topic region's argument tree.
Claims by Doug Rivers (20 of 33)
Weighting creates two distinct problems: it roughly doubles the real variability compared to a simple random sample, and it leaves residual skews (unobserved or uncorrected biases) that do NOT shrink as sample size grows, so conventional sampling-error estimates tell you how the procedure varies sample-to-sample but not how systematically wrong the procedure is.
Apparent poll movements late in a campaign are often illusory: a methodology shift (e.g. Gallup switching from registered-voter to likely-voter weighting) creates a bump unrelated to real opinion change, and pollsters who err have a ready ex-post story of a late undetected surge—what Rivers calls the last refuge of scoundrels in polling—when in fact the same poll repeated would have produced the same wrong answer.
The convention bounce—a candidate gaining several points (about five per party, net swing ~ten) during/after their party's convention—is partly an artifact of differential response: supporters of the convening party are energized and answer polls while opposition partisans, irritated by the 'phony love fest,' decline to participate, so YouGov controls for prior party identification to neutralize the effect.
Party identification shifted dramatically: Democrats had an ~18-point lead from the New Deal through the 1960s, Republicans closed it to a few points by 2004, but 80% of those Republican gains were erased since 2004 (half of them in fall 2007/spring 2008), which in principle should make 2008 a terrible year for Republicans.
Massive consumer and voter databases that did not exist 25 years ago now provide detailed information (income, home value, etc.) on most people, allowing pollsters to draw a randomly selected target sample from a voter list and then match closely-similar respondents from a large opt-in internet panel, creating a sample that mimics a random sample across many dimensions and removes skews that simple demographic weighting cannot.
Averaging polls (as RealClearPolitics does) does not necessarily wash out error: if pollsters make the same systematic errors or herd toward each other, the average is no more reliable than a single poll; averaging only clearly helps when polls use the same methodology, in which case combining them is equivalent to a larger sample, and ideally polls should be weighted by the reciprocal of their variability.
In 2006, the matched internet-panel method produced election forecasts whose average error was substantially less than the average reported telephone poll, mainly because it removed biases better (not because sampling variability was smaller), and Rivers argues telephone polling could be improved by similarly substituting a respondent who resembles the missed person rather than drawing another number from the same skewed population.
The clean statistical results (law of large numbers and central limit theorem) only hold for perfectly executed random sampling; in practice sampling plans are rarely executed perfectly because you never get near 100% cooperation, and the non-responders differ systematically from responders in unobservable ways, producing skews rather than mere random noise.
Random digit dialing became the dominant phone-polling method because roughly 30% of the population had unlisted numbers, making listing-based sampling infeasible; it worked well for 20-30 years but broke down in the last 10-15 years as marketing calls, cell phones, and distrust drove cooperation rates from ~70% down to ~20% or less, hurting accuracy.
Pollsters typically select 10 to 20 times as many phone numbers as the number of completed interviews they want, meaning the people who actually respond are a small, non-random slice that skews toward more women, higher education, and higher income, requiring weighting and adjustment.
Final pre-election polls tend to be too similar to each other relative to their known sampling variability, suggesting pollsters herd—choosing among multiple defensible weighting schemes the one that places them in the pack rather than the one that yields a divergent answer—though Rivers attributes this to risk-aversion and the art of weighting rather than dishonesty.
Of about 1,300 polls Rivers examined from the 2008 presidential primaries, fewer than 50 reported anything other than the margin of error for a simple random sample with no weighting—meaning the reported margins of error are systematically misleading, which Rivers calls scandalous.
Likely the New Hampshire 2008 polling miss came from over-representation of college- and graduate-degree voters (who favored Obama, over-represented by 2-3x and under-corrected) and from unreliable self-reports of voting intention, rather than from racism (the Bradley effect).
Self-reported voting intention is unreliable: about 90% of people say they will vote in a primary but actual turnout is much lower (US presidential turnout has hovered in the mid-50% range, congressional under 40%), and likely-voter screens cannot fully fix this because people answer based on what they think they should say.
My Notes
Loading notes...