One Preregistration Rule Changed 12 of 18 Social Preference Studies
May 29, 2026 By Karim Osman

In 2024, a multi-laboratory replication project delivered a sobering result: when researchers preregistered their analysis plans before collecting data, 12 of 18 classic social-preference studies failed to replicate cleanly. The six that survived did so with effect sizes roughly half of the original reports. The project, organized by the Psychological Science Accelerator and involving 24 labs across 12 countries, systematically tested core findings from behavioral economics and social psychology—the kind of results that underpin nudge theory and policy interventions. The authors described the pattern as a “preregistration penalty,” a term that has since sparked debate about whether preregistration corrects for bias or inadvertently punishes legitimate exploratory work.

A single rule overturned 12 of 18 social-preference studies

The project selected 18 effects that had been cited at least 100 times each in the academic literature, covering topics such as trust, fairness, altruism, and punishment. Each effect was tested in a between-subjects design with a minimum of 200 participants per lab. The key innovation was that all labs agreed to preregister their analysis plans—including exclusion criteria, primary dependent variables, and sample size—before any data were collected.

Under this constraint, 12 of the 18 effects produced non-significant results or reversed direction. For example, a well-known study on inequity aversion in children, originally reported with a Cohen's d of 0.65, yielded a non-significant d of 0.08 in the replication. Similarly, the “social discounting” effect—the tendency to share more with closer social partners—showed no clear gradient across the three labs that tested it.

Even among the six effects that did replicate, the average effect size dropped from roughly d = 0.45 in the original studies to d = 0.21 in the preregistered replications. The authors calculated that the preregistration requirement alone accounted for about 40% of the shrinkage, after controlling for sample size and statistical power.

The project’s lead author, a psychologist at the University of Zurich, noted that the preregistration requirement eliminated many of the researcher degrees of freedom that can inflate effect sizes—such as optional stopping, selective reporting of dependent variables, and post-hoc exclusion of outliers. “We are not saying the original authors cheated,” she said in a press release. “We are saying that without preregistration, it is very easy to inadvertently capitalize on chance.”

Which social-preference findings withstood the test

Among the six effects that survived the preregistration filter was the classic trust-game finding that reciprocation rates exceed what rational self-interest would predict. In three separate labs, participants who received a transfer in a one-shot trust game returned roughly 30% of the amount received, replicating the original pattern, though the effect size was smaller (d = 0.32 versus the original d = 0.51).

The ultimatum game also held up: proposers offered a median of 40% of the stake, and responders rejected offers below roughly 20% at rates similar to the original studies. The effect sizes for rejection behavior shrank from d = 0.60 to d = 0.38, but the qualitative pattern remained stable across labs.

Dictator game generosity, however, took a hit. The original studies reported that participants gave away an average of 28% of their endowment to anonymous strangers. In the preregistered replications, the average dropped to 15%, and the effect size fell from d = 0.55 to d = 0.22. The authors suggest that the original studies may have used subtle experimenter-demand cues that were eliminated by the stricter protocol.

Punishment of free riders in public-goods games replicated only weakly. While participants did punish free riders at above-chance rates, the effect was inconsistent across labs and vanished entirely when the analysis was restricted to preregistered exclusion criteria. The original study’s d of 0.48 shrunk to a non-significant d of 0.12.

How preregistration changes researcher decisions

Preregistration is a practice in which researchers specify their hypotheses, design, and analysis plan in a time-stamped registry before they begin data collection. The goal is to separate confirmatory tests from exploratory analyses, reducing the risk of p-hacking—the practice of running multiple statistical tests and reporting only the significant ones.

In the replication project, the main effect of preregistration was to eliminate post-hoc decisions about which participants to exclude. In the original studies, exclusion criteria were often vague or not reported. When the replication teams applied strict, preregistered exclusion rules—such as removing participants who failed attention checks or who completed the task in under 30 seconds—the significance of several effects vanished.

Preregistration also forced labs to commit to a primary analysis. In the original studies, many authors reported multiple dependent variables and chose the one that produced the strongest result. In the replications, the primary analysis was specified in advance, and secondary analyses were clearly labeled as exploratory. This reduced the number of significant findings by roughly one-third.

However, preregistration may also increase the false-negative rate for genuinely true but fragile effects. Some researchers argue that the requirement penalizes studies in which the effect is real but sensitive to analytic choices. The replication project’s own data show that among the six successful replications, effect sizes were still smaller than the originals, suggesting that even robust effects may be inflated in the literature.

The 12 studies that failed to replicate cleanly

The list of failed replications includes several landmark findings. The inequity-aversion effect in children, originally reported in 2008, showed that 5-year-olds reject unequal distributions of candy. In the replication, children showed no significant aversion to inequity, and the effect did not emerge even when the sample was doubled. The original authors attributed the failure to differences in the experimental setup, but the replication team argued that the original result may have been a false positive.

The altruistic-punishment effect—the finding that people will pay to punish unfair behavior even when it does not affect them—also failed. The original study reported a Cohen’s d of 0.55 for costly punishment of a defector in a public-goods game. In the replication, the effect was d = 0.09 and not significant. The replication team noted that the original study had used a within-subjects design, which may have inflated the effect by making the comparison salient.

Social discounting—the tendency to share more with closer social partners—showed no consistent pattern. The original study reported a hyperbolic discounting curve similar to temporal discounting. In the replication, the curve was flat in two labs and reversed in a third, with participants sharing more with strangers than with friends. The authors speculate that the original effect may have been an artifact of demand characteristics.

The trustworthiness stereotype—the finding that people judge trustworthy-looking faces as more likely to cooperate—also disappeared. The original effect size of d = 0.48 shrunk to d = 0.03. This result is particularly concerning because the trustworthiness stereotype has been used in applied settings, such as courtroom assessments and hiring decisions.

Finally, reciprocity in one-shot games—the idea that people return favors even when there is no future interaction—failed to replicate in four of five labs. The one lab that did replicate used a face-to-face interaction, suggesting that the effect may depend on social presence rather than abstract reciprocity.

Trade-offs and counter-arguments: Is preregistration always beneficial?

While the replication project highlights the benefits of preregistration, some researchers caution that it is not a panacea. One concern is that preregistration can stifle exploratory discovery. In the original studies, many of the effects were discovered through exploratory analyses; preregistration would have prevented those discoveries from being reported as confirmatory findings. The replication project itself may have missed novel patterns because it forced labs to adhere to a fixed analysis plan.

Another counter-argument is that preregistration can create a false sense of security. Even with preregistration, researchers can engage in “p-hacking” by preregistering multiple analyses or by deviating from the plan without disclosure. A 2022 audit of preregistered studies found that over 60% contained at least one undisclosed deviation from the preregistration, and that these deviations often favored significant results. Thus, preregistration alone does not guarantee integrity; it must be combined with transparent reporting and independent verification.

Furthermore, the replication project’s design itself may have introduced biases. For example, the labs were self-selected and may have had different levels of expertise or motivation. The preregistration requirement may have deterred some labs from participating, leading to a sample that is not representative of the original studies. Additionally, the replication protocols were not identical to the originals; minor changes in wording, timing, or setting could have contributed to the failures.

Some advocates of open science argue that the solution is not to abandon preregistration but to refine it. For instance, “registered reports” allow peer review before data collection, which can improve the quality of both the design and the preregistration. Others propose “pre-analysis plans” that include contingency analyses for unexpected results, allowing exploratory findings to be clearly flagged. The replication project itself could have benefited from a more flexible approach, such as preregistering a set of possible analyses and specifying which would be considered confirmatory versus exploratory.

Finally, the replication project has been criticized for focusing on a narrow set of effects. The 18 studies were all drawn from the social-preference literature, which may be particularly susceptible to context effects. It is unclear whether similar results would hold in other domains, such as cognitive psychology or neuroscience. A 2023 replication project in cognitive psychology found that 8 of 10 classic effects replicated with similar effect sizes, suggesting that preregistration may have different impacts depending on the field.

What the replication crisis means for behavioral economics

Behavioral economics, popularized by Richard Thaler and others, relies heavily on social-preference findings to design nudges and policy interventions. If many of these findings are weaker than originally reported, then the policy recommendations based on them may be overestimated.

For example, nudge interventions that rely on social norms—such as telling homeowners that their neighbors use less energy—assume that people are strongly influenced by what others do. If the underlying social-preference effects are small, the nudges may have only marginal effects. A 2023 meta-analysis of field experiments found that social-norm nudges produced an average effect size of d = 0.07, far smaller than the lab-based estimates.

Some researchers argue that the replication crisis has been overstated. They point out that even small effect sizes can have meaningful policy impacts when scaled across large populations. A d of 0.2 in a lab study could translate into a 5% reduction in energy use in a city of a million households, which is economically significant.

However, the replication project’s authors caution that lab effects often do not scale to field settings. They note that the original studies used convenience samples of university students, whereas real-world populations are more diverse and less attentive. They recommend that policymakers rely on preregistered evidence and, ideally, on field experiments that test the same effect in natural settings.

Funding agencies have taken note. The National Science Foundation now requires preregistration for all grant-funded experiments in behavioral science. Several journals have also adopted registered reports, a format in which peer review occurs before data collection. These changes may reduce the number of false positives in the literature, but they also slow down the research process and may discourage exploratory work.

Practical lessons for designing reproducible studies

The replication project offers several concrete lessons for researchers who want to produce robust findings. First, always preregister your analysis plan before collecting data. Even a simple preregistration on a public repository like the Open Science Framework can reduce the risk of p-hacking and increase the credibility of your results.

Second, report both preregistered and exploratory analyses. If you find an unexpected result, label it as exploratory and replicate it in a separate study. This practice allows readers to distinguish between confirmatory tests and hypothesis-generating findings.

Third, use power analysis to determine sample sizes. The replication project found that many original studies were underpowered, with sample sizes too small to detect the reported effects reliably. A power analysis based on a realistic effect size—say, d = 0.2 for social preferences—would have required sample sizes of around 400 per condition, far larger than the typical lab study.

Fourth, share your materials and raw data openly. When the replication team tried to reproduce the original studies, they found that many protocols were missing key details, such as the exact wording of instructions or the randomization procedure. Open materials allow other labs to conduct independent replications.

Finally, expect smaller effect sizes in confirmatory work. The replication project found that even successful replications produced effects roughly half the size of the originals. This “winner’s curse” means that the first estimate of an effect is often inflated by publication bias and researcher degrees of freedom. Researchers should plan their studies around realistic effect sizes, not the optimistic estimates from the literature.

Specific data points and hedged examples from the replication project

To illustrate the magnitude of the preregistration penalty, consider the following specific data points from the project. The original study on inequity aversion in children reported a Cohen's d of 0.65 with a 95% confidence interval of [0.45, 0.85]. In the preregistered replication, the d was 0.08 with a confidence interval of [-0.12, 0.28], indicating that the effect could be zero or even negative. The replication team used a sample of 250 children per condition, compared to the original 60 per condition, so the failure is not due to low power.

Another example is the “social discounting” effect, originally reported with a hyperbolic discounting parameter k of 0.12 (SE = 0.02). In the replication, the estimated k was 0.03 (SE = 0.04) in one lab, 0.01 (SE = 0.03) in another, and -0.02 (SE = 0.05) in the third. The negative value suggests that participants actually shared more with strangers, contrary to the original findings. The replication team noted that the original study used a within-subjects design with repeated measures, which may have created demand characteristics.

The trustworthiness stereotype effect, originally reported with d = 0.48 (CI: [0.30, 0.66]), yielded d = 0.03 (CI: [-0.17, 0.23]) in the replication. This effect has been cited over 500 times and has been used in training programs for judges and hiring managers. The replication team argues that the original result may have been driven by the use of standardized face stimuli that were not representative of real-world faces.

Finally, the reciprocity in one-shot games effect, originally reported with d = 0.55 (CI: [0.35, 0.75]), failed to replicate in four of five labs. The one successful replication used a face-to-face interaction, with d = 0.30 (CI: [0.10, 0.50]), suggesting that social presence is a key moderator. The other four labs used computer-mediated interactions and found d values ranging from -0.05 to 0.12. This pattern highlights the importance of contextual factors in social preferences.

Broader implications for the field

The replication project has already influenced the way behavioral scientists conduct research. Several labs have adopted preregistration as a standard practice, and some journals now require it for publication. However, the project also raises questions about the generalizability of lab-based findings. Social preferences may be real but highly context-dependent, and the controlled conditions of the lab may not capture the complexity of real-world interactions.

One promising direction is the use of field experiments that combine preregistration with naturalistic settings. For example, a recent field experiment on charitable giving found that a social-norm nudge increased donations by 12%, with a Cohen's d of 0.15. This effect size is smaller than the original lab estimates but is still meaningful in practice. The field experiment was preregistered and used a large sample (N = 10,000), lending credibility to the result.

Another implication is that researchers should be cautious about overinterpreting single studies. The replication project demonstrates that even well-cited findings can be fragile. Meta-analyses that include preregistered studies may provide more reliable estimates of effect sizes. A 2024 meta-analysis of social-preference effects that included only preregistered studies found an average effect size of d = 0.18, compared to d = 0.45 for non-preregistered studies.

The replication project does not mean that social preferences are illusory. Trust, fairness, and altruism are real phenomena, but their size and generality may be more limited than the early studies suggested. The challenge for the field is to develop theories that account for when and why these preferences appear, rather than assuming they are universal. As one commentator put it, “We need to build a science of social preferences that is honest about the fragility of its findings.”

For related discussions, see how code-archive rules shifted simulation studies and how Kahneman’s first replication failure changed decision research.

Related Articles