How We Rate the Science
Reading the Badge
Every article covering a peer-reviewed paper carries a two-part rating. Here is what those grades mean, how I assign them and how to argue with one.
The Problem
A Signal, Not a Verdict
Science journalism has a credibility problem. Readers see a headline saying "study shows X" and reasonably ask how much weight to give it. A single mouse study and a decade of replicated human trials get reported in almost the same voice, and there's rarely anything to help you tell them apart.
This rating is my attempt at that signal. It doesn't tell you what to think about a finding; it tells you how confident to be in it right now, and why. Every rating I assign is published against the criteria on this page, so you can hold me to them and push back when you think I've called one wrong.
Six cells, one per criterion, each lit or dark. This badge has the sample and power cell dark, which is why the letter reads B rather than A.
The Two Questions
How well was the work done?
A letter, A to D. It asks whether the study was designed, conducted and reported to the standard the field expects: peer review, sample size, preregistration, data availability, honest declaration of conflicts and limitations.
A well-conducted study can still turn out to be wrong. A badly conducted one can accidentally be right. This grade tells you which type you're looking at.
How much should we believe it yet?
A word: Emerging, Supported, Robust or Settled. A completely different question. How many independent teams have looked at this, how much converging or contradictory work surrounds it, and how well established the underlying mechanism is.
A single excellent study on a novel question sits at Emerging however good it is, until other people have had a chance to check.
The Six Cells
Each side of the hexagon is one criterion, and each is either lit or dark. There is no half. A middle state becomes a dumping ground for every awkward call, and within a year the badge stops telling you anything.
Design fit
Can this design answer the question being asked? A study can be run perfectly and still be pointed at the wrong thing.
Sample and power
Is there enough evidence for this particular claim? For a model that means asking whether the uncertainty it reports actually covers its dominant assumption.
Bias controls
Are the main routes to a false result closed off? Randomisation and blinding for an experiment, confounder adjustment for an observational study, sensitivity analysis for a model.
Transparency
Can somebody else check it? Data, analysis code, and methods reported in enough detail to repeat the work.
Independence
Does anyone with a commercial stake in the answer control the study? A research council or a charity does not darken this cell. A company selling the thing under test does.
Honest reporting
Do the paper and its press release claim only what the design can show, with the limitations said out loud?
How the Letter Follows
The letter does not come from adding anything up. Summing separate kinds of flaw into one score lets a paper offset broken randomisation by publishing its data, which is a well-documented error rather than a matter of taste. What matters is which cells are dark, and in particular how many of the three critical ones.
A if all six are lit. B if one or two non-critical cells are dark. C if all three are.
B at best, however clean the rest of the badge looks.
C at best.
D.
Before Any of That
Five checks run before a single cell is scored, because a paper can be written to look immaculate and still be fabricated. A gate can only lower a rating, never raise one.
No rating is issued and the article does not run.
A credible unresolved case on PubPeer caps the letter at D.
Tortured phrases and similar fabrication markers cap at D.
Pay-to-publish with no real review caps at D.
Caps at B. Unreviewed is not the same as untrustworthy.
The Evidence Tier
A single study, however good, or a first-of-its-kind result. Nobody else has independently replicated or extended the work yet. Most single papers land here at first coverage. Read it as a serious opening move rather than a settled fact.
Several independent studies converging on the same finding, with sound methodology across most of them. The result has been checked and it largely holds. Open questions remain; the shape of the answer is clear.
Many well-conducted studies and a mechanism that is understood. Only fine-grained questions are still contested. Most well-tested drugs, most established physics and most core climate findings sit here.
Textbook material, replicated across decades by independent groups using different methods. Only cranks dispute it. Evolution, plate tectonics, the germ theory of disease, vaccine efficacy for well-studied vaccines.
Limits and Revisions
What the rating doesn't cover
Three things I deliberately leave out, because they're yours to judge rather than mine.
How a rating can change
A rating is a snapshot of the evidence when the article was published, and science moves. An Emerging finding that gets independently replicated becomes Supported. A Supported finding that fails to replicate can drop back, or the paper can be retracted outright.
This is one reason the rating lives in the article rather than in the headline. Headlines don't get retroactively updated; ratings do.
Disagree With Me
How to argue with a rating
Every rating here is an editorial judgement made against the criteria above, and any editorial judgement can be wrong. If you think one of mine is off, I want to hear about it: new replication evidence, a methodological point I missed, a conflict of interest I didn't spot, or a reasoned disagreement about how I applied the rubric.
Every rated article carries a "Revisiting This Rating" card in its sidebar. Use it. If you've made a good case, the rating changes.
Get in touch →