Elie Hassenfeld on GiveWell
What this covers
Elie Hassenfeld, CEO of GiveWell, joins Russ Roberts to discuss how his organization identifies which charities deliver the most measurable good per dollar donated. The conversation centers on GiveWell's method: narrowing thousands of potential recipients to a small recommended list by applying academic evidence, a funding threshold benchmarked against direct cash transfers, and scrutiny of organizational track record. Hassenfeld traces GiveWell's evolution from its 2006 founding, when detailed data on charitable impact barely existed, through a sobering early mistake—over-reliance on a charity's self-reported numbers—that forced a recalibration toward disciplined evaluation. The core argument emerging is not that numbers are wrong, but that they require the ballast of qualitative judgment: transparency matters less than legibility, and cost-effectiveness estimates must account for how studies conducted under controlled conditions differ from real-world deployment.
Roberts presses on tensions Hassenfeld navigates as the organization has grown. He surfaces the psychological pull toward quantification—how certainty soothes while uncertainty provokes anxiety—and how this bias intensifies in a research-driven workplace. The two examine concrete failures, including GiveWell's funding of No Lean Season, a scaling problem that revealed how small, hand-monitored interventions fracture when run through large institutions. They also debate whether donors should segregate giving into two buckets: utility-maximizing contributions to proven global-health charities, and personal or local giving driven by opportunity cost, human connection, and what Hayek might call superior local knowledge. Roberts argues against extreme utilitarianism by insisting that the relational dimensions of flourishing—fatherhood, gift-giving, observing local impact—are themselves non-trivial human goods, not mere luxuries to trade away for marginal effectiveness gains. Hassenfeld acknowledges the tension but resists collapsing it, defending GiveWell's counterfactual approach to measuring whether its own funding genuinely displaces other donors or adds value.
Hassenfeld argues that the most effective philanthropy combines rigorous quantification of impact-per-dollar with disciplined qualitative judgment, and that donors should mentally separate utility-maximizing giving (best directed to a few evidence-backed global-health charities) from personal or local giving that serves human flourishing.
- GiveWell narrows thousands of charities to a few using academic evidence, a cost-effectiveness threshold, and a funding track record.
- Over-reliance on numbers is a danger; qualitative factors like organizational track record and transparency carry the day in close calls.
- Donors can 'bucket' giving so that local/personal giving for flourishing coexists with maximally effective global giving.
GiveWell ranks and funds programs by cost-effectiveness relative to cash transfers, using a threshold that depends on available capital.
- GiveWell estimates it costs approximately $5,000 to avert the death of someone who would otherwise have died from malaria, lack of immunization, or other childhood illness, serving as a benchmark for cost-effectiveness.
“we would estimate about $5,000 to avert the death of someone who would have otherwise died from malaria, lack of immunization, other childhood illness”
- GiveWell sets its funding threshold by ranking all possible programs by cost-effectiveness measured in multiples of direct cash transfers (where giving a poor person $1,000 equals a value of 1) and funding down the ranked list as far as available money allows—currently to a multiple of 10—so the bar that top charities must clear depends on how much money GiveWell has raised.
“We have an estimate for how good it is to just give a very poor person $1,000. That gets you a value of 1. And then, we're going to fund our list all the way down currently to 10.”
- When GiveWell compared poverty-reducing programs like direct cash transfers against childhood-mortality-averting programs, it generally found the mortality-averting programs offered more impact per dollar, though this rests on contestable philosophical judgments about trading off across types of good.
“we've generally felt that the childhood mortality averting programs look better to us--like, we are getting more bang for our buck”
- To become a GiveWell top charity, an organization must meet three criteria: evidence strong enough that it is significantly more likely than not to be having large effect (high confidence, not a 25% chance of large effect); estimated impact per dollar exceeding GiveWell's current funding threshold (currently 10x direct cash transfers); and having received significant prior funding (at least $10 million for at least a year) so GiveWell has experience with it.
“Number One: the evidence is strong enough that we believe it is significantly more likely than not that they're having a lot of effect... Number Two: it's programs whose estimated impact per dollar exceeds our current threshold... the third part, before it gets on the list, we want to have provided significant funding.”
- GiveWell's four recommended top charities account for approximately two-thirds of the funds it directs, with the remainder going to organizations that don't qualify as top charities, such as those funded through the All Grants Fund.
“those four organizations account for approximately two-thirds of the funds we direct”
- The needs and opportunities to help people in low-income countries are dramatically larger in magnitude than the possibilities for helping people in high-income settings like New York City, which drove GiveWell to focus internationally.
“the needs and then the opportunities to help people in low-income countries are just so different than the possibilities for helping people in New York City”
- Over its 15 years, more than 100,000 donors have given more than a billion dollars to GiveWell's recommendations, with the organization currently raising approximately $500 million per year.
“more than a 100,000 donors have given more than a billion dollars to our recommendations”
Quantification and modeling alone are insufficient; judgment about track record, transparency, and organizational capacity must inform borderline decisions.
- GiveWell's early recommendation of PSI rested on taking the organization's self-reported distribution numbers at face value—numbers that were true but not rigorously gathered—reflecting an over-enthusiasm for quantification as the path to the best answer, a mistake that pushed GiveWell toward critical evaluation of data rather than data alone.
“we supported them in large part because, I think, we took those numbers at face value in a way that I think was overly enthusiastic about quantification as a path to getting the best answer”
- Being transparent is not enough if the reasoning isn't legible: about 20% of contest submitters misunderstood what GiveWell was doing because the spreadsheets were convoluted, so GiveWell is now investing in making its research legible enough that people can understand and critique it without spending hundreds of hours.
“about 20% of the people who submitted misunderstood what we were doing... we were doing Y: the spreadsheets were just convoluted.”
- In a highly uncertain investigation like supporting Evidence Action's chlorinated water work with the Indian government, where quantification is impossible to pin down, it is the organization's decade-long track record of successful program delivery and honest transparency that carries the decision; the same numbers from a less-trusted organization would not earn support.
“if it had been a different organization, I think we wouldn't have supported, we wouldn't support the program because we wouldn't have that qualitative overlay”
- When a program's quantified cost-effectiveness number lands near the threshold (e.g., 7 or 13 versus 10), GiveWell relies on qualitative considerations: unmodeled upside (supporting a young person/organization that could grow), confidence the organization will be transparent about what's working and failing, and the track record of the people and organization—pushing borderline decisions up or down.
“Seven and 13 are not that different from 10, given all the uncertainty that's baked in.”
- As GiveWell grew from a handful of researchers to about 70 people, staff showed an increasing tendency to rely more heavily on the numbers because clear rules ('build a model, get a number, the number is the answer') are easier to understand and execute than overlaying disciplined, bias-free judgment, making it organizationally hard to maintain the tension between quantification and judgment.
“I see this tendency now that we're a larger organization for staff to want to rely more on the numbers.”
- Publishing all research and reasoning on GiveWell's website both imposes rigor on the organization's own thinking and enables outsiders to understand, evaluate, and critique its work where they disagree.
“we put all of our research and the reasoning behind the research on our website so that--I think that imposes some rigor on ourselves in thinking about how we do our work, but also enables outsiders to understand what we're doing and why and critique it”
GiveWell measures impact by causal effect on the world, adjusting for fungibility and displacement of other donors' gifts.
- GiveWell cares not about what its literal dollars accomplish but about the causal impact of its work on the world; if its donation merely displaces another donor's gift to the same charity, nothing has changed, so GiveWell applies a 'Fungibility Adjustment' estimating how likely its giving is to displace others' funds and what those others would have done.
“We care about, ideally, the causal impact that our work has on the world.”
- Just as a GiveWell donation may crowd out another donor's gift (the crowding-out phenomenon studied in economics), a GiveWell stamp of approval can crowd in additional donations by reassuring donors, and can help small programs build the track record needed for large governmental funders to scale them.
“if GiveWell--puts a stamp of approval on an organization, that could easily increase donations”
- Both Roberts and Hassenfeld endorse 'bucketing' charitable giving: treating personal and local giving (synagogue, PTA, things that fulfill you as a human being or for which you use the service) as a separate shared-responsibility bucket from the bucket directed at maximizing global impact, since extra money toward flourishing activities yields little additional flourishing while the vast majority can go where it does the most good.
“I don't bucket it into my mental charitable-giving budget... When I give to the synagogue or my school's PTA, I see it as a shared responsibility to support a service that I use.”
Early mistakes in scaling and exiting programs led GiveWell to commit longer-term funding with patience for iterative improvement.
- GiveWell's funding of No Lean Season—a program incentivizing seasonal migration in Bangladesh based on Mobarak's small RCT—failed at scale because the small, hands-on model became harder to deliver through large microfinance institutions, and GiveWell made two mistakes: underestimating the difficulty of scaling and exiting too quickly rather than committing to a longer, iterative timeframe.
“in fact, the program did not work at scale, when we first funded it... we didn't see that the treatment group in this randomized trial, they were not more likely to migrate than the control group.”
- Learning from premature exits, GiveWell now makes initial funding decisions for experimental programs with a longer timeframe in mind and is willing to stick with them even when they don't look good at first, giving them the opportunity to improve.
“now, we tend to commit to longer timeframes with programs that are more experimental”
GiveWell prioritizes interventions where empirical learning and measurable feedback loops are possible over high-risk bets without measurable outcomes.
- GiveWell deliberately directs the vast majority of funds to interventions where it can learn empirically whether it was right or wrong and improve, creating feedback loops; this leads it to avoid high-risk opportunities like funding African think tanks for pro-growth policy, not because such bets are bad but because measurable learning is GiveWell's distinctive competence.
“we tend to, with the vast majority of the funds we direct, put them in places where we will be able to know if we were right or if we were wrong. We will try to learn and then we will try to improve.”
- GiveWell sometimes supports a person or program based on a strong track record—a 'good bet'—even without a quantified estimate of the good that will result, as when it funded Mushfiq Mobarak's Y-RISE center studying the science of scaling programs.
“it's not a case where we had a quantified estimate of the good that would come from supporting this individual and their team at Yale. But it is a case where I think this person's track record makes it a, quote, "good bet"”
Real-world program delivery differs from controlled trial conditions, requiring adjustment of cost-effectiveness estimates and consideration of implementation challenges.
- GiveWell's 'Change Our Mind' contest, offering prizes to people who found errors in its public analysis, drew more than 50 high-quality submissions including from prestigious academics, surfaced small errors across programs, and—most importantly—revealed GiveWell had been insufficiently transparent about the uncertainty inherent in its analysis.
“people called us out for being insufficiently transparent about the uncertainty inherent in our analysis”
- Mosquitoes have built up resistance to the insecticide used in bed nets, so nets used a decade ago are less effective today, requiring GiveWell to account for whether older or newer, more resistance-effective nets are being distributed.
“Mosquitoes are building up resistance to the insecticide used in the nets. The nets that were being used 10 years ago would not be as effective today”
- GiveWell takes seriously 'Room for More Funding'—how much money an organization can use well—and would not put an organization on its list that could only absorb $100,000 well if listing it might channel $10 million, tracking this over time and reducing support when organizations have already received large funding.
“The first thing that we take very seriously is a question that we call... Room for More Funding, which is: 'How much funding do we think they could use well?'”
- When Hassenfeld and Karnofsky sought information on what charities accomplish per dollar given, they found that detailed, useful data on charitable impact simply did not exist, with charities unable to justify claims like '$20 provides a child water for life.'
“we were really surprised that we couldn't find useful information that would help guide our decisions”
- Cost-effectiveness estimates must be qualified because randomized controlled trials are conducted under tightly monitored conditions (e.g., researchers checking daily that nets are used) and at higher baseline mortality and poverty levels than today's national-scale distributions, raising the question of how valid old study results are to current real-world deployment.
“Randomized controlled trials are conducted very differently than national-level malaria net distributions”
Childhood mortality prevention offers greater cost-effectiveness than poverty reduction, though this involves contestable philosophical trade-offs.
- GiveWell would gladly fund programs that address the root cause of poverty if it knew how, but since it doesn't know how to alleviate the cause, it directs funds to alleviating symptoms where impact can be measured—acknowledging it would be better to fix causes than symptoms.
“We would love to know how to give money to end poverty now... I think, unfortunately, we just don't know the answer.”
- Reducing serious childhood illness likely yields better outcomes for the child through healthier early development plus benefits to family and community, but these effects play little role in GiveWell's actual benefit calculations even though they are important considerations.
“Reducing serious illness in childhood probably leads to better outcomes for that child itself by having a healthier young early childhood development.”
- Vitamin A supplementation given to children under five twice a year significantly reduces child mortality—by about 25%—especially in populations that are vitamin-A deficient, a context very different from supplementation in high-income countries.
“this program reduces child mortality by about 25%”
- A huge proportion of deaths in low-income countries occur in the first month, year, and five years of life; once people survive those early years, they tend to live relatively long lives, which justifies focusing on averting childhood mortality rather than treating it as merely delaying death from other causes.
“a huge proportion of deaths that occur in low-income countries occur in those very early years. A shocking amount in the first month, first year, first five years. Once people get past those years, they tend to live long lives”