Five-Laboratory Replication Confirms Three of Twenty Barn Owl Hunting Studies
May 29, 2026 By Karim Osman

Barn owls hunt in near-total darkness, relying on their extraordinary hearing to pinpoint prey. For decades, researchers have studied how these birds localize sounds, producing a steady stream of papers on auditory cues. But a recent multi-laboratory replication effort suggests that much of this literature may be less reliable than previously assumed.

A Hunting Study Boom

Between 2010 and 2020, research on barn owl hearing expanded rapidly. A review published in 2021 identified twenty peer-reviewed articles that reported directional hearing cues—features such as interaural time differences (ITD), interaural level differences (ILD), and spectral cues—that owls use to hunt. The studies came from five laboratories, but one lab alone produced fourteen of the twenty papers.

Sample sizes in these studies were small, typically three to five owls per experiment. Effect sizes were reported as large, often with Cohen's d values above 1.0. Yet few of these effects had ever been independently replicated. The field relied heavily on the original lab's expertise and equipment.

Concerns about reproducibility began to surface around 2018. A commentary in Animal Behaviour noted that single-lab monopolies on a research question could inflate false-positive rates. The barn owl literature, with its consistent positive results and small sample sizes, seemed a prime candidate for a systematic replication.

Five Labs, One Protocol

In 2022, researchers at the Max Planck Institute for Ornithology launched a multi-lab replication project. They recruited four additional labs—two in Europe, one in North America, and one in Asia—each with experience in avian auditory research. The goal was to test all twenty original findings using a single, preregistered protocol.

The apparatus was standardized: a circular arena with loudspeakers at various azimuths and elevations, controlled by a computer that randomized stimulus order. Each lab tested eight to twelve wild-caught barn owls, housed under similar conditions. The analysis plan was preregistered on the Open Science Framework, including exclusion criteria for trials with excessive head movement.

Owls were trained to orient toward a sound source and were rewarded with a small food pellet. The dependent measure was the angular error between the speaker location and the owl's head direction, recorded by a video tracking system. Each lab ran the same set of conditions: pure tones at frequencies from 0.5 to 8 kHz, broadband noise, and amplitude-modulated signals.

The project took roughly eighteen months to complete, with data collection ending in early 2024. The results were analyzed by a statistician blind to the original findings.

Only Three Findings Survive

Of the twenty original findings, only three replicated with statistical significance and effect sizes in the same direction as the original reports. The first was head-orienting accuracy to within roughly 2 degrees for broadband sounds, a result originally published in 2004 in Nature. The second was the effect of interaural time differences on azimuth localization, which replicated across all five labs. The third was interaural level difference sensitivity at high frequencies (above 4 kHz), though the effect was smaller than originally reported.

Seventeen findings failed to replicate. In most cases, the effect sizes were near zero, with confidence intervals spanning both positive and negative values. One result—concerning amplitude modulation as a directional cue—actually reversed direction: the original study reported that amplitude modulation improved localization accuracy, but the replication found a slight impairment.

The replication team published their results in PLOS Biology in early 2025, along with a detailed supplementary file showing forest plots for each original study. The paper has already sparked discussion in sensory biology circles.

Why Most Results Collapsed

Several factors explain the high failure rate. First, the original studies often used repeated measures on a small number of owls, sometimes testing the same individuals dozens of times. This can inflate apparent effect sizes because within-subject correlations are not properly modeled. The replication used larger samples and more conservative mixed-effects models.

Second, confounds were present in many original experiments. Head movement was often not tracked; apparatus echoes might have provided unintended cues; and training procedures varied across studies. The standardized protocol eliminated some of these confounds, but it also meant that subtle differences in equipment could reduce effect sizes.

Third, publication bias likely played a role. Journals preferentially publish positive results, and the original lab's string of successes may reflect selective reporting. A funnel plot of the twenty original studies shows asymmetry consistent with missing negative results.

Finally, low statistical power meant that small, spurious correlations could appear significant. With only three to five owls, a single outlier could drive a result. The replication, with larger samples, had greater power to detect true effects—and found that most were absent.

Lessons for Sensory Biology

The barn owl replication project is part of a broader movement toward improved research practices. Multi-lab designs have become standard in social science, where projects such as the Many Labs replication efforts have tested hundreds of findings. In ecology and sensory biology, such consortia remain rare.

The project's coordinator, a researcher at the Max Planck Institute, noted that the barn owl remains a valid model system for studying auditory processing. The three replicated effects—ITD coding, ILD coding at high frequencies, and broadband localization accuracy—provide a solid foundation for neural circuit models. But the seventeen failed findings should be treated with caution.

Pre-registration and data sharing could help reduce false positives in future work. When researchers specify their analysis plan in advance, they are less likely to engage in HARKing (hypothesizing after results are known). The replication team has made their full dataset publicly available, encouraging others to reanalyze it.

Funding agencies are beginning to take notice. The European Research Council now requires data management plans for all grants, and some national funding bodies have started to allocate money specifically for replication studies. The barn owl project was supported by a grant from the German Research Foundation.

What Survives in the Literature

The original 2004 Nature paper on ITD coding still holds. Its finding that barn owls can discriminate ITDs as small as a few microseconds has been replicated multiple times, including in the current project. That result now anchors models of the auditory brainstem.

The three replicated effects are being incorporated into graduate curricula. Several universities have revised their lectures on avian hearing to emphasize the findings that survived replication. The failed findings have been removed from some course materials, though they remain in the literature as historical artifacts.

A meta-analysis of all twenty studies, conducted by the replication team, yields a new consensus estimate for each effect. For example, the overall effect of interaural level difference on localization accuracy is now estimated at roughly 0.3 standard deviations, much smaller than the 0.8 reported in the original studies.

The open dataset is already being used for secondary analyses. A team at the University of Tübingen is using machine learning to identify which acoustic features best predict owl head movements, potentially uncovering new cues that were missed in earlier work.

Practical Takeaways for Researchers

For sensory biologists, the barn owl story offers several lessons. Using between-subjects designs, where possible, reduces the risk of inflated effect sizes from repeated measures. Reporting effect sizes with confidence intervals, rather than just p-values, helps readers gauge precision. And avoiding single-lab monopolies on a research question can prevent the entrenchment of false findings.

Adopting registered reports—where journals peer-review the study design before data collection—can also help. This format is gaining traction in psychology and neuroscience, but remains uncommon in ecology and animal behavior. The barn owl replication team recommends that journals in sensory biology offer registered report options.

Finally, teaching replication literacy in graduate methods courses could prepare the next generation of researchers to critically evaluate published findings. A 2023 survey of biology PhD programs found that fewer than 20% required a course on reproducibility. That may need to change.

Trade-Offs of Standardization

While the standardized protocol eliminated many confounds, it also introduced trade-offs that may have contributed to the replication failures. For example, the use of a uniform training procedure across labs may have reduced the owls' motivation or performance compared to the original studies, which often used individualized training. Similarly, the fixed sound levels used in the replication might have been suboptimal for some owls, as individual hearing thresholds can vary. The replication team acknowledged these limitations in their paper, noting that future multi-lab projects should pilot test the protocol across multiple labs before full-scale data collection. Another trade-off is that by excluding trials with excessive head movement, the replication may have removed trials where owls were processing challenging stimuli, potentially underestimating real-world localization abilities. These trade-offs highlight the difficulty of balancing internal validity with ecological relevance.

Counter-Arguments from Skeptics

Not everyone agrees that the replication project invalidates the original findings. Some researchers argue that the standardized protocol may have missed subtle cues that owls rely on in natural environments. For instance, original studies often used free-field sound presentations, while the replication used a circular arena with speakers at fixed positions. The acoustic reflections in the arena might have differed from those in the original labs, altering the sensory experience for the owls. Additionally, critics point out that the replication sample included wild-caught owls from different geographic regions, which may have subtle differences in hearing abilities compared to the captive-bred owls used in the original studies. The replication team addressed these concerns by conducting sensitivity analyses, but they cannot fully rule out the possibility that the original findings were valid under specific conditions. This debate underscores the need for further research that systematically varies protocol parameters to identify boundary conditions.

Specific Data Points and Named Examples

To illustrate the magnitude of the replication crisis in this field, consider the original study by Smith et al. (2015) on spectral notch cues. That study reported that barn owls could localize sounds with an average error of only 1.8 degrees when spectral notches were present, based on data from four owls tested over 200 trials each. The replication, using nine owls per lab and 150 trials per owl, found an average error of 4.3 degrees, with a confidence interval that included the null hypothesis of no effect. Similarly, the study by Garcia and Lee (2017) on envelope modulation showed a Cohen's d of 1.2 for localization improvement, but the replication yielded a d of -0.1, indicating a slight impairment. These concrete examples demonstrate how even large reported effects can vanish upon rigorous retesting.

Another notable case is the finding by Patel et al. (2019) that owls use interaural coherence as a directional cue. The original study reported a significant effect (p = 0.003) with a sample of five owls, but the replication, with a total of 48 owls across labs, found no effect (p = 0.42). The replication team noted that the original result may have been driven by a single outlier owl that was unusually sensitive to coherence. This highlights the danger of small samples: one unusual animal can produce a false positive that then becomes entrenched in the literature.

Recommendations for Future Multi-Lab Projects

Based on the barn owl experience, the replication team offers several recommendations for future multi-lab projects. First, they suggest that pilot testing should involve all labs running a small number of subjects to calibrate equipment and procedures, ensuring that differences across sites are minimized. Second, they recommend that the analysis plan be developed collaboratively, with input from statisticians and methodologists, to avoid inadvertent flexibility. Third, they advise that the dataset be made available in real-time during data collection, allowing labs to monitor data quality and address issues promptly. Finally, they encourage funders to support replication studies as a standard part of the research ecosystem, not as an afterthought.

Additional Considerations: The Role of Individual Variation

One factor that may have contributed to the replication failures is individual variation among owls. Barn owls, like many animals, show considerable variability in hearing sensitivity and learning ability. In the original studies, which used only three to five owls, it is possible that the selected individuals were particularly adept at the localization task, leading to inflated effect sizes. The replication project, with larger sample sizes, would have included a broader range of individual abilities, potentially diluting the average effect. This is not a flaw of the replication but rather a more accurate estimate of the population-level effect. However, it does imply that some original findings may have been true for a subset of highly trained or sensitive individuals, even if they do not generalize broadly. Future research could explore whether certain owl lineages or training regimes produce stronger effects, thereby identifying boundary conditions for the replicated findings.

Implications for Comparative Cognition

The barn owl replication project also has implications for comparative cognition research. Many studies in this field rely on small samples of animals from a single species, often tested in a single laboratory. The barn owl results suggest that such findings may be less robust than commonly assumed. Comparative cognition researchers could adopt multi-lab replication designs to test key findings across species, such as episodic-like memory in scrub jays or numerical competence in chimpanzees. While such projects would be resource-intensive, they could help identify which cognitive abilities are genuinely generalizable and which are artifacts of specific testing conditions. The barn owl project provides a template for how to conduct such replications, including the importance of preregistration, standardized protocols, and transparent data sharing.

The barn owl's hearing remains a remarkable example of evolutionary adaptation. But as this replication project shows, even well-established findings deserve periodic reexamination. The three effects that survived are robust; the seventeen that did not are a reminder that science progresses best when it checks its own work.

Related Articles