All podcasts / People I (Mostly) Admire / Summary

The Data Sleuth Taking on Shoddy Science

2025-08-02 - source - Read full transcript
Steve Levitt (host)Uri SimonsohnMorgan Levey

Key insights

Small, undisclosed researcher choices (measuring multiple outcomes, stopping data collection early, adding covariates) can turn a hypothesis that is one hundred percent false into a 'statistically significant' finding over 60 percent of the time.
Simonsohn's 2011 'False-Positive Psychology' paper demonstrated this with a fake experiment claiming the Beatles song 'When I'm Sixty-Four' makes listeners younger. By testing multiple outcome measures, multiple comparison groups, and optionally controlling for covariates like the participant's father's age, the team generated a significant result from a hypothesis they knew was untrue, showing how common 'best practices' at the time functioned as undisclosed cheating.
statistical-malpractice
Stopping data collection as soon as a result becomes significant is a subtly biased practice, even though it feels intuitively less dishonest than cherry-picking outcome variables.
Simonsohn explains this with a tennis analogy: if a match ends the instant one player is ahead rather than after a fixed number of sets, that player's win probability rises far above 50 percent even between equally skilled players. Researchers who add more subjects only when a result is 'almost significant' and stop once it clears 0.05 are doing the statistical equivalent, and Simonsohn says many experimentalists still argue this is legitimate practice.
statistical-malpractice
P-curve analysis exploits the fact that researchers chasing the 0.05 significance threshold tend to stop exactly when they clear it, producing an unnatural cluster of P-values just under 0.05 rather than the very low P-values a true effect would generate.
Simonsohn compares this to running toward a fixed-time goal: a marathoner aiming for four hours finishes at 3:58 or 3:59, not 3:30. Applying P-curve to the power-posing literature (roughly 30 published papers) found no genuine evidence for the effect, despite its being the subject of a hugely popular TED talk.
statistical-malpractice
Fraud investigation is emotionally and professionally costly enough that Simonsohn had sworn it off before being pulled back in only for cases involving very famous researchers or very influential papers.
He describes the work as fun for the first two days of discovery and then 'dreadful' for the year of drudgery required to reach near-certainty, because an accusation demands not just personal confidence but evidence strong enough that any outside party would be convinced; being wrong, or merely a whistleblower, creates lasting professional and social costs.
research-fraud-detection
The Francesca Gino fraud case was proven by cross-referencing a numeric rating column suspected of being altered against a separate free-text column describing the same events, which could not be faked without also rewriting the numbers consistently.
Participants who supposedly gave a rating of 1 (out of 7) had written text like 'best time of my life,' while participants who supposedly gave a 7 had written 'I felt disgusted.' This mismatch let Simonsohn's team tell Harvard exactly which rows on Gino's original Qualtrics server data to check, and the university's own access confirmed the numbers had been swapped.
research-fraud-detection
Uncovering Francesca Gino's alleged fraud led Simonsohn's team, while examining a co-authored paper, to separately notice apparent fraud by Dan Ariely in an unrelated insurance-honesty study, which they could only fully confirm years later via investigative journalism.
The insurance dataset showed a perfectly uniform distribution of self-reported miles driven from zero to 50,000 with zero people above 50,000, a statistically impossible pattern. Ariely immediately took sole responsibility when contacted, but Simonsohn's team believed only an investigative reporter with access to the insurer's original data could prove alteration; a New Yorker reporter later obtained that data and found what Simonsohn calls irrefutable proof it had been changed after being sent to Ariely.
research-fraud-detection
Institutions facing fraud accusations against their own faculty have strong incentives not to find them guilty, and their investigation processes and disclosure practices diverge accordingly.
Harvard ran a public investigation and stripped Gino's tenure; Duke ran an undisclosed investigation into Ariely, never published its findings, and let Ariely characterize the outcome himself (losing only his named professorship). Simonsohn notes Duke likely never had access to insurance-company data as damning as what Harvard obtained on Gino, so the difference in punishment reflects available evidence as much as institutional posture.
institutional-accountability
Because the only career consequence most confirmed fraudsters face is losing a job they would have lost anyway for lack of results, there is effectively no rational deterrent against committing fraud.
Simonsohn argues that since being fired for poor (non-fraudulent) performance was already a live outcome, being fired for fraud instead is a 'win-win' from the perpetrator's expected-value standpoint; he concludes the burden falls on the field to make committing and hiding fraud structurally harder rather than relying on punishment to deter it.
institutional-accountability
Simonsohn does not blame academic incentives (rewarding publication and interesting results) for fraud; he blames the near-total absence of any penalty for low-quality or fabricated inputs.
He says he is comfortable with a system that rewards successful, prolific researchers over unsuccessful ones. The actual defect is that peer reviewers cannot easily detect P-hacking or fraud, so transparency tools (raw data sharing, pre-registration, and his new 'data receipt' platform AsCollected) exist to make bad inputs cheap to catch rather than to change what gets rewarded.
academic-incentives
Requiring raw data to be publicly posted, a practice that barely existed in academia until roughly the last decade, has been transformative for catching errors, fraud, and non-robust results.
Simonsohn credits the shift partly to his own and colleagues' advocacy but mostly to the internet simply making data sharing logistically easy; before this norm, academics were not expected to let anyone see the raw data behind a published finding.
scientific-transparency
AsCollected, Simonsohn's new funded platform, asks researchers structured questions about data provenance (public or private source, who received and cleaned the data, whether a nondisclosure contract exists) up front, modeled on requiring a receipt before a reimbursement is approved.
The output is a public URL with two tables (a 'when, what, how' table and a 'who' table); Simonsohn says all roughly 15 fraud cases his team has handled would have been substantially harder to pull off if the perpetrator had first been forced to answer these questions, and he is pitching journals, deans, and granting agencies as the target customers who would require the URL at submission.
scientific-transparency
A live, low-stakes replication of the same 'researcher degrees of freedom' problem happened on the show itself: Levitt tried to find a statistically significant pattern in a listener poll about climate optimism by cutting the data by age, gender, and country, and found only two significant results out of many attempts, both plausibly spurious.
Levitt explicitly frames this as demonstrating Simonsohn's point in real time: given enough demographic cuts (age, gender inferred from name, country, response timing), you can almost always find something that looks 'significant' by chance alone, even in genuinely uninteresting data; the only two significant findings (Canadians and Australians responding disproportionately often) were about survey engagement, not the actual optimism/pessimism question.
statistical-malpractice

Media referenced

Companies

Techniques and frameworks

Summary

Steve Levitt talks with Uri Simonsohn, the Data Colada blogger and behavioral science professor who has spent over a decade catching bad statistics and outright fraud in academic psychology. The conversation opens with Simonsohn's origin story: reviewing a paper claiming people with matching name initials are more likely to marry each other, he traced the "effect" to a mundane artifact, couples who divorce and remarry each other after one spouse had taken the other's surname, which is exactly the kind of hidden confound his later work made a career of hunting down. From there, Levitt walks through Simonsohn's landmark 2011 paper "False-Positive Psychology," co-authored with Joe Simmons and Leif Nelson, which used a deliberately absurd fake study (that a Beatles song makes listeners younger) to show that ordinary researcher flexibility, testing multiple outcomes, adjusting for covariates, and collecting more data only when results are "almost significant," can turn a completely false hypothesis into a "statistically significant" one more than 60 percent of the time.

Much of the middle of the episode is a plain-language tour of why these practices are so damaging and so hard to see as cheating. Simonsohn uses vivid analogies throughout: throwing multiple dice to explain outcome-shopping, a tennis match with no fixed endpoint to explain biased early stopping, and a marathoner hitting their goal time almost exactly to explain why P-hacked literatures cluster suspiciously near the 0.05 significance threshold rather than showing the very low P-values a genuine effect would produce. That last insight became P-curve analysis, a tool Simonsohn's team used to debunk the entire "power posing" literature, a widely cited, TED-talk-famous body of research that turned out to have no real supporting evidence once examined this way.

The episode's center of gravity shifts to Simonsohn's highest-profile fraud investigations: Francesca Gino, his own former Wharton colleague, and Dan Ariely, whom his team stumbled onto almost by accident while investigating Gino. The Gino case turned on a striking piece of detective work, cross-checking a numeric satisfaction rating that appeared altered against a free-text description of the same event; participants supposedly rating an event a 1 had written "best time of my life," and vice versa, a mismatch that let the team tell Harvard exactly which server rows to check. Harvard's investigation confirmed the alteration and stripped Gino's tenure. The Ariely case involved an impossible, perfectly uniform distribution of self-reported driving miles in an insurance-honesty study; Ariely immediately claimed sole responsibility, but full proof only came years later when a New Yorker reporter obtained the insurer's original data. Simonsohn is candid that Gino's lawsuit against him, Data Colada, and Harvard for $25 million was frightening and expensive, made survivable only by his Spanish university's willingness to fund a motion to dismiss and by a spontaneous academic-community GoFundMe that moved him to tears.

Levitt and Simonsohn close the substantive discussion on institutional incentives: Harvard's public handling of Gino contrasts sharply with Duke's secretive investigation of Ariely, which Simonsohn attributes partly to a real asymmetry in evidence quality rather than gender or prestige bias alone. Both agree that because the worst realistic outcome for a caught fraudster is losing a job that low performance might have cost them anyway, there is essentially no rational deterrent, so the responsible move is structural: make fraud harder to commit and easier to detect. Simonsohn describes his new platform, AsCollected, a spinoff of his pre-registration tool AsPredicted, which forces researchers to document data provenance before publication (echoing an expense-reimbursement receipt) precisely so that the kind of undocumented, unverifiable data-handling that enabled the Gino and Ariely cases becomes structurally harder to hide.

The episode closes with the recurring listener-mailbag segment with producer Morgan Levey, where Levitt turns his own climate-optimism listener poll into a live, low-stakes demonstration of Simonsohn's core lesson: cutting the same small dataset by age, gender, and country repeatedly, he finds almost nothing statistically significant except two results (Canadians and Australians responding to the poll far more often than their download share would predict) that are about engagement, not climate views, and are exactly the kind of noise-masquerading-as-signal the whole episode has been warning about.

Notable Quotes

"We talk about red flags versus smoking guns. So red flags, that gives you, like, probable cause, so to speak. But it's not enough to raise an accusation of fraud. That's a smoking gun." - Uri Simonsohn

"It's not enough to be right in your head with a hundred percent certainty. You need to be certain that others will be certain." - Uri Simonsohn

"If the worst thing that can happen to you is that you're fired, but without fraud, you would be fired, it's still a win-win. For somebody like that to commit fraud, there's no real disincentive." - Uri Simonsohn

"You should read a paper that surprises you, and you should update... And we were not experiencing that." - Uri Simonsohn

"So if you have enough dice, even if it's a 20-sided dye, like only one in 20 chance, if you keep rolling that dye, eventually it's going to work out." - Uri Simonsohn