// Methodology

How We Rate the Science

Reading the Badge

2 Questions
6 Cells
4 Tiers

Every article covering a peer-reviewed paper carries a two-part rating. Here is what those grades mean, how I assign them and how to argue with one.

The Problem

// Why bother

A Signal, Not a Verdict

Science journalism has a credibility problem. Readers see a headline saying "study shows X" and reasonably ask how much weight to give it. A single mouse study and a decade of replicated human trials get reported in almost the same voice, and there's rarely anything to help you tell them apart.

This rating is my attempt at that signal. It doesn't tell you what to think about a finding; it tells you how confident to be in it right now, and why. Every rating I assign is published against the criteria on this page, so you can hold me to them and push back when you think I've called one wrong.

// The badge

Six cells, one per criterion, each lit or dark. This badge has the sample and power cell dark, which is why the letter reads B rather than A.

The Two Questions

01 // Methodology

How well was the work done?

A letter, A to D. It asks whether the study was designed, conducted and reported to the standard the field expects: peer review, sample size, preregistration, data availability, honest declaration of conflicts and limitations.

A well-conducted study can still turn out to be wrong. A badly conducted one can accidentally be right. This grade tells you which type you're looking at.

Judges the work, not the conclusion.
02 // Evidence

How much should we believe it yet?

A word: Emerging, Supported, Robust or Settled. A completely different question. How many independent teams have looked at this, how much converging or contradictory work surrounds it, and how well established the underlying mechanism is.

A single excellent study on a novel question sits at Emerging however good it is, until other people have had a chance to check.

One brilliant paper is still one paper.

The Six Cells

Each side of the hexagon is one criterion, and each is either lit or dark. There is no half. A middle state becomes a dumping ground for every awkward call, and within a year the badge stops telling you anything.

1

Design fit

Critical

Can this design answer the question being asked? A study can be run perfectly and still be pointed at the wrong thing.

2

Sample and power

Critical

Is there enough evidence for this particular claim? For a model that means asking whether the uncertainty it reports actually covers its dominant assumption.

3

Bias controls

Critical

Are the main routes to a false result closed off? Randomisation and blinding for an experiment, confounder adjustment for an observational study, sensitivity analysis for a model.

4

Transparency

Can somebody else check it? Data, analysis code, and methods reported in enough detail to repeat the work.

5

Independence

Does anyone with a commercial stake in the answer control the study? A research council or a charity does not darken this cell. A company selling the thing under test does.

6

Honest reporting

Do the paper and its press release claim only what the design can show, with the limitations said out loud?

How the Letter Follows

// No arithmetic

The letter does not come from adding anything up. Summing separate kinds of flaw into one score lets a paper offset broken randomisation by publishing its data, which is a well-documented error rather than a matter of taste. What matters is which cells are dark, and in particular how many of the three critical ones.

0 dark

A if all six are lit. B if one or two non-critical cells are dark. C if all three are.

1 dark

B at best, however clean the rest of the badge looks.

2 dark

C at best.

3 dark

D.

An A means all six lit. There is no A with a gap in it.

Before Any of That

// Gates

Five checks run before a single cell is scored, because a paper can be written to look immaculate and still be fabricated. A gate can only lower a rating, never raise one.

Retracted

No rating is issued and the article does not run.

Manipulated data or images

A credible unresolved case on PubPeer caps the letter at D.

Paper mill signature

Tortured phrases and similar fabrication markers cap at D.

Predatory venue

Pay-to-publish with no real review caps at D.

Preprint

Caps at B. Unreviewed is not the same as untrustworthy.

Where a paper was published earns it nothing. It can only cost.

The Evidence Tier

// How far it has been checked
Emerging

A single study, however good, or a first-of-its-kind result. Nobody else has independently replicated or extended the work yet. Most single papers land here at first coverage. Read it as a serious opening move rather than a settled fact.

Supported

Several independent studies converging on the same finding, with sound methodology across most of them. The result has been checked and it largely holds. Open questions remain; the shape of the answer is clear.

Robust

Many well-conducted studies and a mechanism that is understood. Only fine-grained questions are still contested. Most well-tested drugs, most established physics and most core climate findings sit here.

Settled

Textbook material, replicated across decades by independent groups using different methods. Only cranks dispute it. Evolution, plate tectonics, the germ theory of disease, vaccine efficacy for well-studied vaccines.

Limits and Revisions

03 // Out of scope

What the rating doesn't cover

Three things I deliberately leave out, because they're yours to judge rather than mine.

How important the finding is. A settled fact about a minor topic and an emerging finding about a major one might each matter more, depending on why you're reading.
How much it applies to you. A robust finding in mice is still a finding in mice. A settled finding about population averages tells you very little about your own situation.
How interesting it is. Some robust findings are dull and some emerging ones are electrifying. Interest is a matter of taste, not of evidence.
04 // Revisions

How a rating can change

A rating is a snapshot of the evidence when the article was published, and science moves. An Emerging finding that gets independently replicated becomes Supported. A Supported finding that fails to replicate can drop back, or the paper can be retracted outright.

This is one reason the rating lives in the article rather than in the headline. Headlines don't get retroactively updated; ratings do.

Every change is logged on the article.

Disagree With Me

05 // Right of reply

How to argue with a rating

Every rating here is an editorial judgement made against the criteria above, and any editorial judgement can be wrong. If you think one of mine is off, I want to hear about it: new replication evidence, a methodological point I missed, a conflict of interest I didn't spot, or a reasoned disagreement about how I applied the rubric.

Every rated article carries a "Revisiting This Rating" card in its sidebar. Use it. If you've made a good case, the rating changes.

Get in touch →