Before You Cite a Beloved Study, Run This Five-Point Autopsy
A useful finding can outlive its evidence; here is a practical autopsy for studies you have cited, taught, or built decisions on.
Some findings survive because they are true. Others survive because they are convenient. The difference matters when a result has already entered your syllabus, your team's operating assumptions, or your citation network. A beloved study can feel like infrastructure: once it is taught, cited, and reused, questioning it costs time, credibility, and comfort. That is exactly when a postmortem is needed.
A Psychological Science paper reported a failure to replicate Study 2 of Ariely and Wertenbroch's influential procrastination/deadlines article. The original article had been widely used and cited, with more than 2,100 Google Scholar citations and use as assigned reading in economics and psychology courses
The finding was intuitive: deadlines can change behavior. It was also widely useful: managers, students, and teachers could use it to structure work. Intuition is not proof. Usefulness is not proof. A result can be practically valuable and still rest on data that would not survive a careful inspection.
What the autopsy found
Data Colada described the reported deadline effect as implausibly large, with a Cohen's d of 2.5 for proofreading performance
The data also contained suspicious patterns. Data Colada found duplicate observations in the Last Day Deadline condition, where 18 of 20 participants had another participant with identical correction counts across all three proofreading tasks. It is not automatically fraud, but it is the kind of pattern that demands explanation, and it becomes much harder to explain when the data are not transparent.
Data Colada concluded that the data in Studies 1 and 2 were tampered with, and Ariely and Wertenbroch requested retraction of the article from Psychological Science while the process was ongoing. This is not a story about a bad idea. It is a story about a useful idea that was carried by data that could not stand up to scrutiny.
Five-point study autopsy
Before you cite a beloved study again, run these five checks. They are not a replacement for peer review, but they are a replacement for nostalgia.
- Is the effect size plausible for the domain? Ask whether the magnitude is consistent with what you know about the outcome.
- Are raw data, code, and analysis decisions available? A result is easier to trust when the path from raw observations to final claim is visible. If only the paper is available, note that limitation. If data, code, or analysis notes are missing, treat the result as provisional, not canonical.
- Do duplicates, outliers, or suspicious patterns survive inspection? Look for repeated values, impossible distributions, condition-level anomalies, and participants who are too similar to be believable. You do not need to be a data detective to ask whether the pattern has a plausible explanation.
- Has an independent group replicated it? Replication is not a vote. It is a test of whether the result appears outside the original laboratory, sample, and analytical choices. A single failure to replicate is not a death sentence, but it is a strong signal to slow down.
- What incentives made the result hard to challenge? Ask who benefits from the finding remaining unquestioned: departments, curricula, management frameworks, popular books, or the researchers' own reputations. Incentives do not prove fraud, but they explain why a weak result can become a load-bearing wall.
These checks are especially important when the study is already embedded in teaching or practice. A canonical paper can become a kind of social fact. Once it is assigned reading, cited in reviews, or used to justify a policy, the cost of questioning it rises. That is not a reason to stop questioning. It is a reason to question more carefully.
What to do differently
If you have cited, taught, or acted on a beloved study, do not pretend the autopsy was not needed. The people who used it were not foolish for trusting a result that looked plausible and was widely endorsed. For researchers, the practical step is to build a small evidence file for any study you plan to rely on: effect size, data availability, replication status, and known criticisms. For editors, ask authors to distinguish established findings from popular but contested ones. For managers, treat a single study as one input, not a policy. If a finding is central to a decision, require at least two independent lines of evidence or a clear statement of uncertainty.
If you have cited, taught, or built a deadline policy on this article, mark it as contested, add a caveat, and require a second independent line of evidence before relying on it again.
