Skip to content

How we score a paper.

Each paper receives one adjusted score between 0 and 100. That score combines two independent judgements: the quality of the work, and the evidential weight that the study design carries for the paper's claims. This page explains how we build each judgement, how they combine, and why the combination has this structure.

Quality, 0 to 100

A review-based assessment of the paper's rigour, clarity, transparency, and the validity of its conclusions. Built from issues (scored by criticality) and merits (scored by calibre) across eight categories.

Design weight, 0 to 100

The evidential weight that the study design carries for the paper's claims: how fully the design shields them from alternative explanations. What counts as strong evidence depends on the field.

Adjusted score, 0 to 100

The two ratings combined through an interaction-aware adjustment. Quality stays the dominant signal. The design weight increases or decreases it by a maximum of a few points, more when the two agree.

§ 1

Tiers

The tiers

The adjusted score also names a tier. The tiers are a reading of the same 0–100 scale, so you can see the position of a score immediately.

0FLAWED10TIN30COBALT50BRONZE60SILVER70GOLD85PLATINUM1000/100
  • 85–100
    Platinum

    Exceptional work whose claims withstand scrutiny.

  • 70–84
    Gold

    Strong work whose claims hold up, bar a few minor points to address.

  • 60–69
    Silver

    Broadly sound work, though with substantive points still to resolve.

  • 50–59
    Bronze

    Largely sound work, but with issues that temper its claims.

  • 30–49
    Cobalt

    Notable weaknesses that weigh on the strength of its claims.

  • 10–29
    Tin

    Deeper problems that leave several of the central claims undersupported.

  • 0–9
    Flawed

    Fundamental issues that undercut the paper's central claims.

§ 2

Quality

Quality: the Q score

Q rates the execution of the paper, independently of the type of study. Two parallel streams supply it: the issues a reviewer raises, each scored by criticality, and the merits a reviewer identifies, each scored by calibre.

The numbers are a navigation aid. One high-criticality issue can matter more than a dozen minor ones. One exceptional-calibre merit can justify a paper that has noticeable issues elsewhere. The written reasoning that accompanies each score holds the real evaluative work.

Criticality

Issues · 0-10

How seriously an issue undermines the work. Not the difficulty of a repair, and not its prominence in the paper. The measure is its effect on the validity and integrity of the contribution.

  • 0Negligible.A trivial concern. We record it for completeness, but it has almost no effect on the soundness of the work. Typographical issues, minor stylistic preferences, small omissions that a careful reader can barely notice.
  • 5Moderate.A genuine problem that a reasonable reviewer will expect the authors to address. The issue does not invalidate the work, but it weakens a claim, a conclusion, or part of the methodology. A thoughtful response or a revision is appropriate.
  • 10Fatal.A defect so severe that the reader cannot trust the work in its current form. Causes include a fundamental design flaw, a breakdown in the chain of evidence, an analytical error that invalidates the central conclusions, or an ethical breach. Polish cannot rescue the paper. It needs substantive rework.

Calibre

Merits · 0-10

How impressive or valuable a strength is, compared with the norms of the field. Calibre is the positive mirror of criticality: it records the strengths of the paper, and how exceptional each one is.

  • 0Negligible.A strength so minor that it almost does not register. The work does the task competently, but at a level that does not distinguish it from other adequately executed papers in the field.
  • 5Moderate.A solid, real merit. The work does something clearly above baseline competence. A reviewer will want to identify it, and a reader will benefit from it.
  • 10Exceptional.A standout strength. A new methodological standard, evidence of rare quality, an insight that changes the approach to a question, or execution at the top of the field.

The review files each issue and each merit into one of eight categories. This makes the shape of a paper's strengths and weaknesses easy to read.

  1. 01

    Research Design

    The foundational architectural decisions: choice of methodology, study structure, selection of participants or materials, and the fit of the overall approach to the question.

  2. 02

    Data and Evidence

    The raw material the work rests on: the soundness of data collection, the assembly of evidence, the measurement methods, and the quality of underlying sources.

  3. 03

    Analytical Approach

    What the author does with the collected data: statistical models, qualitative coding, computational pipelines, formal reasoning, or argumentative structure, both choice and execution.

  4. 04

    Scholarly Grounding

    How fully the work situates itself in its field: engagement with relevant literature, strength of theoretical underpinnings, proper credit to prior work, and a clear sense of where the contribution sits.

  5. 05

    Reporting Quality

    Transparency and completeness: is there enough detail (methods, data, code, justification) for a reader to understand, evaluate, and (where applicable) reproduce the work?

  6. 06

    Interpretive Rigour

    The leap from results to conclusions. Do the findings support the claims? Does the author acknowledge limitations honestly, hedge claims appropriately, and examine plausible alternative explanations seriously?

  7. 07

    Ethical Conduct

    Responsible-research dimensions: ethical treatment of subjects, informed consent, declared conflicts of interest, data integrity, and adherence to the broader norms of responsible conduct.

  8. 08

    Contribution

    What the work adds. A paper can be technically impeccable and contribute little. Another can be rough in execution but advance understanding in important ways.

§ 3

Design weight

Design weight: the D score

D is the evidential weight a study design carries for the paper's claims. What counts as strong evidence depends on the field: resistance to confounding in medicine, logical validity in mathematics, binding authority in law, community legitimacy in Indigenous studies. The sections that follow show the hierarchy applied to several research domains, and why that hierarchy is the right one for that type of knowledge.

What the hierarchy measures: how fully a study design shields its claim from error. The hierarchy ranks one thing: epistemic warrant for the claim type of that field, given the study design.

What it does not measure: the novelty, importance, or contribution of the work. A landmark qualitative study can matter more than a mediocre RCT. D rates how fully the design shields a claim, never how much the claim is worth.

D is not a value judgement. A lower position on the hierarchy does not mean worse research. It means a different type of warrant, often the strongest available for that question. A score of 3 in ecology is not “mediocre”. It means the paper uses a natural experiment or a quasi-experimental design. That design is appropriate, and often the only ethical option, for many ecological questions.

The core question is different in each field

Different fields ask fundamentally different questions, and the question determines which hierarchy is applicable.

Field

The question it asks

Hierarchy applied

Clinical / Health

“Does X cause Y in humans?”

Hierarchy of causal strength

Mathematical Sciences

“Is X necessarily true?”

Hierarchy of proof rigour

Physical / Chemical Sciences

“Can X be measured reliably?”

Hierarchy of experimental reproducibility

Social / Economic

“Does X generalise?”

Hierarchy of internal + external validity

Humanities / Interpretive

“Is the reading of X warranted?”

Hierarchy of argumentative quality

Law

“Is X binding?”

Hierarchy of legal authority

Engineering

“Does X work in practice?”

Hierarchy of validation rigour

Indigenous Studies

“Is X legitimate to this community?”

Hierarchy of relational / cultural validity

The schema labels

In the list of research domains below, each discipline carries exactly one of three labels. The label gives the source of its hierarchy.

Formal
An existing, named, published schema, applied exactly as its source defines it. We cite the schema name and the primary source verbatim.
Adapted
An adaptation of an existing named schema for this field. We show each mapping from an original level name to its adapted name.
Informal
No published schema exists for this field. We assembled a hierarchy from discipline norms to show the epistemic standards of the field.

Expand any domain to see its disciplines and the applicable hierarchy. When disciplines in a domain use different schemas, the page shows them as separate blocks.

23 domains

§ 4

Adjusted score

The adjusted score: Q and D combined

An interaction-aware adjustment combines the quality score Q and the design-weight score D into one adjusted score S, also on the 0-100 scale. The construction treats Q and D as two independent signals that interact, not as two numbers to sum. The adjustment rewards well-executed work in a demanding design, and penalises poorly-executed work in a forgiving design. Off-diagonal cases (high Q with low D, or low Q with high D) receive only modest adjustments.

Definitions

Δ=D−50\Delta = D - 50
m(Q,D)={Q100if D≥50100−Q100if D<50m(Q, D) = \begin{cases} \dfrac{Q}{100} & \text{if } D \geq 50 \\[6pt] \dfrac{100 - Q}{100} & \text{if } D < 50 \end{cases}

Adjusted score

S=clip ⁣( Q+Δ⋅[ s⋅m(Q,D)+b⋅(1−m(Q,D)) ],  0,  100 )S = \mathrm{clip}\!\left(\, Q + \Delta \cdot \bigl[\, s \cdot m(Q, D) + b \cdot (1 - m(Q, D)) \,\bigr],\; 0,\; 100 \,\right)

Q

review-based quality score, 0 to 100

D

study-design weight, 0 to 100

Delta

centred design weight, -50 to +50

m

alignment magnitude, 0 to 1

s = 0.15

interaction strength

b = 0.075

baseline strength

What the adjustment does

The adjustment to Q depends on how much Q and D align: they align when both are high or both are low, and diverge when one is high and the other is low. In the formula, the interaction term (s · m) grows as the two align, while the baseline term (b · (1 - m)) carries the off-diagonal cases. The four cases below show which term dominates in each condition.

  • High quality with a high-weight design (e.g. a well-executed meta-analysis of RCTs): alignment is high, the interaction term dominates, and the paper receives the largest positive adjustment. The system rewards good execution of a demanding study design.

  • Low quality with a low-weight design (e.g. a poorly-executed expert opinion): alignment is again high, but in the opposite direction. The interaction term dominates, and the paper receives the largest negative adjustment. Poor execution of easy work compounds against the score, because work of this type is straightforward to do correctly.

  • Off-diagonal cases (high Q with low D, or low Q with high D): alignment is low, the baseline term dominates, and the adjustment is small. A well-executed expert opinion keeps most of its quality credit. The system treats a meta-analysis that struggles more leniently, because the format is genuinely harder.

  • A neutral design weight. If D = 50, then the centred weight is 0 and the quality score does not change: S = Q.

Worked example: medicine

The table below takes the medical hierarchy as the reference. It shows the adjusted score S for a strong paper (Q = 80), an average paper (Q = 50), and a weak paper (Q = 20). The size of each adjustment is in parentheses.

Study design

D

Strong (Q = 80)

Average (Q = 50)

Weak (Q = 20)

Systematic review / meta-analysis of RCTs

D

100

Strong (Q = 80)

86.75 (+6.75)

Average (Q = 50)

55.62 (+5.62)

Weak (Q = 20)

24.50 (+4.50)

Individual RCT or observational with large effect

D

75

Strong (Q = 80)

83.38 (+3.38)

Average (Q = 50)

52.81 (+2.81)

Weak (Q = 20)

22.25 (+2.25)

Non-randomised controlled cohort

D

50

Strong (Q = 80)

80.00 (0.00)

Average (Q = 50)

50.00 (0.00)

Weak (Q = 20)

20.00 (0.00)

Case-series / case-control

D

25

Strong (Q = 80)

77.75 (-2.25)

Average (Q = 50)

47.19 (-2.81)

Weak (Q = 20)

16.62 (-3.38)

Mechanism-based reasoning / expert opinion

D

0

Strong (Q = 80)

75.50 (-4.50)

Average (Q = 50)

44.38 (-5.62)

Weak (Q = 20)

13.25 (-6.75)

Bounds. The maximum possible adjustment is +/-7.5 points. It occurs at the perfectly aligned corners (Q = 100 with D = 100, or Q = 0 with D = 0). At the perfectly misaligned corners (Q = 100 with D = 0, or Q = 0 with D = 100), the adjustment is exactly half: +/-3.75 points.

Classification: ANZSRC 2020 Fields of Research, Australian Bureau of Statistics & Stats NZ, released 30 June 2020

OCEBM schema: Oxford Centre for Evidence-Based Medicine, Levels of Evidence Working Group (2011). "The Oxford Levels of Evidence 2." cebm.ox.ac.uk

ESSA / WWC schema: US Dept of Education, What Works Clearinghouse Procedures and Standards Handbook v5.0 (2022)

Melnyk & Fineout-Overholt schema: Melnyk, B.M. & Fineout-Overholt, E. (2023). EBP in Nursing & Healthcare, 5th ed. Wolters Kluwer

IPCC schema: IPCC Guidance Note for Lead Authors on the Use of Expert Judgment and Treatment of Uncertainty (2010)

APA Div 12: Chambless, D.L. et al. (1998). Update on empirically validated therapies, II. The Clinical Psychologist, 51(1), 3-16

EBSE: Kitchenham, B. et al. (2004). Evidence-based software engineering. Proc. ICSE 2004

AIATSIS: AIATSIS Code of Ethics for Aboriginal and Torres Strait Islander Research (2020)

UNDRIP: United Nations General Assembly (2007). United Nations Declaration on the Rights of Indigenous Peoples, Resolution 61/295, adopted 13 September 2007. un.org/development/desa/indigenouspeoples

CARE Principles: Research Data Alliance International Indigenous Data Sovereignty Interest Group (September 2019). CARE Principles for Indigenous Data Governance. Global Indigenous Data Alliance. gida-global.org

GIDA: Global Indigenous Data Alliance, the network hosting the CARE Principles; see Carroll, S.R. et al. (2020), "The CARE Principles for Indigenous Data Governance," Data Science Journal 19(43). gida-global.org

Mathematical typesetting: KaTeX

The weight of a study comes from its position in the applicable hierarchy for its field, not the novelty, importance, or contribution of the work.

See it on your own paper. Free.