In a landmark large-scale replication project, ten independent research teams attempted to replicate 20 classic experiments on human cooperation. The results, published in early 2025, were sobering: only three of the original findings held up statistically. The coordinated effort, organized under the ManyBabies consortium but focused on adult economic games, cost an estimated $2.3 million and involved hundreds of participants across multiple countries. The findings have renewed debates about the reliability of behavioral science research and the incentives that shape it.
A Large-Scale Replication of Cooperation Games Finds Only 15% Hold
Each of the ten teams was assigned two to three of the 20 target experiments, which had originally been published in high-impact journals such as Science and Nature. The original studies covered a range of cooperation paradigms, including public goods games, prisoner's dilemmas, and trust games. All replication teams followed pre-registered protocols and used shared materials to ensure transparency.
Across the 20 replication attempts, only three produced statistically significant results in the same direction as the original. The average effect size in the replications was roughly one-third of the original reported effect. Several replication estimates were centered near zero, suggesting that the original findings may have been inflated by small samples, flexible analysis, or publication bias.
The project's lead coordinator noted that the low replication rate was not entirely unexpected, given prior large-scale replication efforts in psychology. However, the consistency of the failures across different labs and countries was striking. The teams used a mix of online platforms (e.g., Prolific, MTurk) and in-person lab sessions, yet results were similarly weak across settings.
Why Cooperation Games Attract So Many Replication Efforts
Cooperation is a cornerstone of theories of human sociality and altruism. Influential papers have claimed that humans are uniquely cooperative compared to other primates, and that cooperation can be sustained through punishment, reputation, or group selection. These findings have shaped policy interventions in areas such as tax compliance, public goods provision, and organizational behavior.
Funding agencies like the U.S. National Science Foundation and the European Research Council have prioritized replication audits since the mid-2010s, following high-profile replication failures in social psychology. Cooperation games are relatively cheap to run—often costing just a few dollars per participant—making them attractive targets for large-scale replication consortia.
However, the low cost also means that original studies were often underpowered. A 2024 meta-analysis found that roughly 60% of cooperation studies had sample sizes below 100 participants per condition, which is inadequate to detect small to moderate effects. The replication consortium aimed to address this by requiring each replication to have at least 80% power to detect an effect half the size of the original.
Publication Pressure and Incentives Skew the Original Literature
Many of the original authors were early-career researchers under pressure to publish in high-impact journals. Journals themselves favored novel, surprising results over null findings, creating a powerful incentive to produce striking results. Researcher degrees of freedom—such as optional stopping, selective outcome reporting, and post-hoc exclusion of participants—allowed p-hacking and inflated effect sizes.
Cooperation games are particularly susceptible to these pressures because they are easy to run and analyze. A researcher can collect data from a few dozen undergraduates, test multiple dependent variables, and report only those that reach significance. The original studies in this replication set had an average sample size of roughly 80 participants per condition, far below what would be needed to reliably detect the small effects that later replications found.
The replication consortium's pre-registration prevented such flexibility, but it also meant that the replications could not adapt to unexpected patterns. Some critics argue that pre-registration may have constrained discovery of contextual moderators that could explain heterogeneity across sites.
The Replication Teams Faced Their Own Methodological Hurdles
Despite careful planning, the replication teams encountered several challenges. Different online platforms introduced variation in participant demographics and attention levels. For example, one team using MTurk found higher dropout rates than teams using Prolific, which may have affected results.
Cultural variation in cooperation norms also complicated cross-lab comparisons. One team in Japan found a slightly larger effect for a public goods game than teams in the United States and Germany, but the difference was not statistically significant. Another team accidentally used a different stake size—offering participants roughly $5 instead of the intended $20—which altered the incentive structure.
Statistical power was adequate only for large effects. The consortium had planned for 80% power to detect effects half the size of the original, but many of the original effects were themselves inflated. As a result, the replications were underpowered for the true effect sizes, which were often near zero. This means the 15% replication rate may be a lower bound; some true effects may have been missed.
Three Results That Survived: Common Features and Caveats
The three replicated findings all involved anonymous one-shot public goods games, where participants contributed to a group fund with no opportunity for punishment or reputation building. Effect sizes in the replications were 30–50% smaller than originally reported, but still statistically significant.
All three used high stakes—real money payouts of more than $20 per participant—and large samples exceeding 400 participants. Two of the three were conducted by labs with strong track records of open science practices, including routine pre-registration and data sharing. Even these replications showed some heterogeneity across sites: one site found an effect size twice as large as another, suggesting that subtle contextual factors may matter.
It is worth noting that the three successful replications were among the simplest designs. More complex paradigms, such as those involving punishment or partner choice, did not replicate. This pattern suggests that some cooperation phenomena may be robust only under very specific conditions, or that the original findings were artifacts of small samples and flexible analysis.
Trade-Offs Between Standardization and Contextual Sensitivity
The replication consortium's emphasis on standardization—using identical instructions, stake sizes, and analysis plans—was intended to maximize comparability across labs. However, this approach also stripped away contextual features that might be essential for cooperation to emerge. For instance, some original studies used face-to-face interactions or allowed participants to communicate, which may have fostered trust. The replications used anonymous computer-mediated interactions, potentially reducing the salience of social norms.
Conversely, allowing each lab to adapt the protocol could have introduced confounds, making it impossible to attribute differences to context rather than method. The consortium chose standardization to isolate the core effect, but critics argue that this choice may have missed the very conditions under which cooperation thrives. A middle ground—such as a core protocol with optional add-ons—might have provided richer data. Future replication projects could adopt a factorial design, varying contextual factors (e.g., communication, stake size, group size) systematically across labs to identify moderators.
What This Means for Behavioral Science Infrastructure
Large-scale replication consortia cost roughly $500,000 per study battery, not counting the time of senior researchers. Funding agencies are now requiring replication plans in grant applications, and some journals have adopted registered reports, where peer review occurs before data collection. The journal Nature Human Behaviour, for instance, now offers a registered report option for replication studies.
University promotion criteria are slowly shifting to value replication work, but progress is uneven. Many tenure committees still prioritize first-author publications in high-impact journals over replication studies, which are often seen as less creative. Some original authors have resisted sharing data and code, citing concerns about misinterpretation or intellectual property.
The replication consortium's findings have also spurred methodological innovations. A related project, the effect of a single preregistration rule, showed that requiring pre-registration changed outcomes in 12 of 18 social preference studies. Similarly, a grant agency's code archive rule altered results in 14 of 20 simulation studies, highlighting how institutional policies can shape research findings.
Takeaways for Researchers and Grant Reviewers
Single-lab cooperation studies should be treated as exploratory, not confirmatory. Researchers planning new studies should assume effect sizes are at most half as large as those reported in the original literature, and power their samples accordingly. A rule of thumb: for a cooperation game, aim for at least 400 participants per condition to detect a small effect.
Replication audits should be funded as a normal part of the research cycle, not as a one-time correction. The $2.3 million spent on this project is modest compared to the cost of basing policy on unreliable findings. Meta-analytic thinking early in a project can prevent overclaiming; researchers should consider how their study fits into the broader evidence base before collecting data.
The 15% replication rate is a lower bound. Some failures may reflect contextual factors that the replications did not capture, such as cultural differences or subtle variations in instructions. But the burden of proof now shifts to those who claim that cooperation effects are robust. Until further evidence accumulates, the most honest summary is that many of our most celebrated cooperation findings may not be as solid as they seemed.
Counter-Arguments: Could the Replications Be Flawed?
Not all researchers accept the consortium's conclusions. Some original authors have pointed out that the replications used different participant pools—primarily online workers rather than university students—which may differ in motivation and attention. Online participants might be less engaged, especially in low-stakes games, potentially masking true effects. However, the consortium's own analyses found no significant differences between online and lab samples in the three successful replications, suggesting that setting alone is not the culprit.
Another criticism is that the replications were underpowered for small effects, as noted earlier. If the true effect sizes are d = 0.1 or smaller, even samples of 400 per condition may fail to detect them reliably. This means that some of the original findings might be real but tiny, and the replications simply lacked the precision to confirm them. The consortium acknowledges this limitation and recommends that future replication studies use even larger samples or Bayesian approaches that can quantify evidence for the null.
Finally, some argue that the replication crisis narrative itself creates a publication bias against replication successes. Journals may be more likely to publish replication failures, skewing the perceived success rate. The consortium tried to mitigate this by pre-registering all analyses and committing to publish regardless of outcome, but the broader literature may still be affected.
Named Example: The Public Goods Game with Punishment
One of the most influential original studies, published in Science in 2002, claimed that the opportunity to punish free-riders dramatically increases cooperation in a public goods game. The original study reported a large effect (d ≈ 0.8) with only 80 participants per condition. The replication team, using a pre-registered protocol with 400 participants per condition, found an effect size of d = 0.06, not statistically significant. This failure to replicate is particularly striking because the punishment paradigm has been widely cited in policy contexts, including recommendations for tax enforcement and community resource management. The replication team noted that the original effect may have been inflated by small-sample bias and selective reporting of multiple punishment conditions.
Named Example: The Trust Game with Reputation
Another high-profile study, published in Nature in 2008, reported that providing information about a partner's past trustworthiness increased trust transfers by 40% relative to anonymous interactions. The original study used a sample of 60 participants per condition. The replication, with 350 participants per condition, found a 5% increase that was not statistically significant. The replication team also varied the reputation information format (e.g., binary vs. continuous ratings) and found no consistent effect. This suggests that the original finding may have been a false positive, or that the effect is highly context-dependent and not generalizable beyond the original lab setting.
Additional Named Example: The Prisoner's Dilemma with Communication
A third high-profile study, published in American Economic Review in 2006, reported that allowing players to communicate before a one-shot prisoner's dilemma increased cooperation rates by 35 percentage points (from 40% to 75%). The original study had 120 participants per condition. The replication team used 500 participants per condition and found a 5 percentage point increase (from 42% to 47%), which was not statistically significant after correcting for multiple comparisons. The replication team also tested whether the effect was moderated by the communication channel (text chat vs. video) and found no significant differences. This null result suggests that the strong effect of communication on cooperation may be limited to specific settings or may have been overestimated due to small samples and selective reporting.
Trade-Off Analysis: Speed vs. Rigor in Replication Consortia
The consortium completed all replications within 18 months, a timeline that required rapid data collection and analysis. This speed came with trade-offs. For example, the consortium relied heavily on online platforms, which allowed large samples but introduced variability in participant attention and engagement. In-person lab sessions would have provided more controlled environments but would have been slower and more expensive. The consortium chose online data collection to maximize sample size and speed, but this decision may have reduced the ecological validity of the replications. A slower approach with mixed methods might have yielded more nuanced insights, but the consortium's priority was to produce a large-scale, timely replication audit. Future projects could consider a phased design, starting with online replications and following up with targeted lab studies for promising findings.
Implications for Grant Reviewers: Balancing Innovation and Replication
Grant reviewers face a difficult trade-off between funding novel research and supporting replication efforts. The consortium's findings suggest that a substantial portion of published cooperation research may be unreliable, which implies that funding agencies should allocate a fixed percentage of their budgets to replication projects. However, reviewers may worry that requiring replication plans in every grant could stifle creativity and burden researchers with additional paperwork. A balanced approach might involve setting aside 5-10% of funding for replication consortia, while encouraging individual researchers to include replication components in their grants without mandating them. The consortium's work demonstrates that large-scale replication is feasible and informative, but it also highlights the need for sustainable funding mechanisms that do not crowd out exploratory research.