How we score a paper.
Each paper receives one adjusted score between 0 and 100. That score combines two independent judgements: the quality of the work, and the evidential weight that the study design carries for the paper's claims. This page explains how we build each judgement, how they combine, and why the combination has this structure.
Quality, 0 to 100
A review-based assessment of the paper's rigour, clarity, transparency, and the validity of its conclusions. Built from issues (scored by criticality) and merits (scored by calibre) across eight categories.
Design weight, 0 to 100
The evidential weight that the study design carries for the paper's claims: how fully the design shields them from alternative explanations. What counts as strong evidence depends on the field.
Adjusted score, 0 to 100
The two ratings combined through an interaction-aware adjustment. Quality stays the dominant signal. The design weight increases or decreases it by a maximum of a few points, more when the two agree.
Tiers
The tiers
The adjusted score also names a tier. The tiers are a reading of the same 0–100 scale, so you can see the position of a score immediately.
- 85–100Platinum
Exceptional work whose claims withstand scrutiny.
- 70–84Gold
Strong work whose claims hold up, bar a few minor points to address.
- 60–69Silver
Broadly sound work, though with substantive points still to resolve.
- 50–59Bronze
Largely sound work, but with issues that temper its claims.
- 30–49Cobalt
Notable weaknesses that weigh on the strength of its claims.
- 10–29Tin
Deeper problems that leave several of the central claims undersupported.
- 0–9Flawed
Fundamental issues that undercut the paper's central claims.
Quality
Quality: the Q score
Q rates the execution of the paper, independently of the type of study. Two parallel streams supply it: the issues a reviewer raises, each scored by criticality, and the merits a reviewer identifies, each scored by calibre.
The numbers are a navigation aid. One high-criticality issue can matter more than a dozen minor ones. One exceptional-calibre merit can justify a paper that has noticeable issues elsewhere. The written reasoning that accompanies each score holds the real evaluative work.
Criticality
Issues · 0-10
How seriously an issue undermines the work. Not the difficulty of a repair, and not its prominence in the paper. The measure is its effect on the validity and integrity of the contribution.
- 0Negligible.A trivial concern. We record it for completeness, but it has almost no effect on the soundness of the work. Typographical issues, minor stylistic preferences, small omissions that a careful reader can barely notice.
- 5Moderate.A genuine problem that a reasonable reviewer will expect the authors to address. The issue does not invalidate the work, but it weakens a claim, a conclusion, or part of the methodology. A thoughtful response or a revision is appropriate.
- 10Fatal.A defect so severe that the reader cannot trust the work in its current form. Causes include a fundamental design flaw, a breakdown in the chain of evidence, an analytical error that invalidates the central conclusions, or an ethical breach. Polish cannot rescue the paper. It needs substantive rework.
Calibre
Merits · 0-10
How impressive or valuable a strength is, compared with the norms of the field. Calibre is the positive mirror of criticality: it records the strengths of the paper, and how exceptional each one is.
- 0Negligible.A strength so minor that it almost does not register. The work does the task competently, but at a level that does not distinguish it from other adequately executed papers in the field.
- 5Moderate.A solid, real merit. The work does something clearly above baseline competence. A reviewer will want to identify it, and a reader will benefit from it.
- 10Exceptional.A standout strength. A new methodological standard, evidence of rare quality, an insight that changes the approach to a question, or execution at the top of the field.
The review files each issue and each merit into one of eight categories. This makes the shape of a paper's strengths and weaknesses easy to read.
01
Research Design
The foundational architectural decisions: choice of methodology, study structure, selection of participants or materials, and the fit of the overall approach to the question.
02
Data and Evidence
The raw material the work rests on: the soundness of data collection, the assembly of evidence, the measurement methods, and the quality of underlying sources.
03
Analytical Approach
What the author does with the collected data: statistical models, qualitative coding, computational pipelines, formal reasoning, or argumentative structure, both choice and execution.
04
Scholarly Grounding
How fully the work situates itself in its field: engagement with relevant literature, strength of theoretical underpinnings, proper credit to prior work, and a clear sense of where the contribution sits.
05
Reporting Quality
Transparency and completeness: is there enough detail (methods, data, code, justification) for a reader to understand, evaluate, and (where applicable) reproduce the work?
06
Interpretive Rigour
The leap from results to conclusions. Do the findings support the claims? Does the author acknowledge limitations honestly, hedge claims appropriately, and examine plausible alternative explanations seriously?
07
Ethical Conduct
Responsible-research dimensions: ethical treatment of subjects, informed consent, declared conflicts of interest, data integrity, and adherence to the broader norms of responsible conduct.
08
Contribution
What the work adds. A paper can be technically impeccable and contribute little. Another can be rough in execution but advance understanding in important ways.
Design weight
Design weight: the D score
D is the evidential weight a study design carries for the paper's claims. What counts as strong evidence depends on the field: resistance to confounding in medicine, logical validity in mathematics, binding authority in law, community legitimacy in Indigenous studies. The sections that follow show the hierarchy applied to several research domains, and why that hierarchy is the right one for that type of knowledge.
What the hierarchy measures: how fully a study design shields its claim from error. The hierarchy ranks one thing: epistemic warrant for the claim type of that field, given the study design.
What it does not measure: the novelty, importance, or contribution of the work. A landmark qualitative study can matter more than a mediocre RCT. D rates how fully the design shields a claim, never how much the claim is worth.
D is not a value judgement. A lower position on the hierarchy does not mean worse research. It means a different type of warrant, often the strongest available for that question. A score of 3 in ecology is not “mediocre”. It means the paper uses a natural experiment or a quasi-experimental design. That design is appropriate, and often the only ethical option, for many ecological questions.
The core question is different in each field
Different fields ask fundamentally different questions, and the question determines which hierarchy is applicable.
Field
The question it asks
Hierarchy applied
Clinical / Health
“Does X cause Y in humans?”
Hierarchy of causal strength
Mathematical Sciences
“Is X necessarily true?”
Hierarchy of proof rigour
Physical / Chemical Sciences
“Can X be measured reliably?”
Hierarchy of experimental reproducibility
Social / Economic
“Does X generalise?”
Hierarchy of internal + external validity
Humanities / Interpretive
“Is the reading of X warranted?”
Hierarchy of argumentative quality
Law
“Is X binding?”
Hierarchy of legal authority
Engineering
“Does X work in practice?”
Hierarchy of validation rigour
Indigenous Studies
“Is X legitimate to this community?”
Hierarchy of relational / cultural validity
The schema labels
In the list of research domains below, each discipline carries exactly one of three labels. The label gives the source of its hierarchy.
- Formal
- An existing, named, published schema, applied exactly as its source defines it. We cite the schema name and the primary source verbatim.
- Adapted
- An adaptation of an existing named schema for this field. We show each mapping from an original level name to its adapted name.
- Informal
- No published schema exists for this field. We assembled a hierarchy from discipline norms to show the epistemic standards of the field.
Expand any domain to see its disciplines and the applicable hierarchy. When disciplines in a domain use different schemas, the page shows them as separate blocks.
23 domains
Adjusted score
The adjusted score: Q and D combined
An interaction-aware adjustment combines the quality score Q and the design-weight score D into one adjusted score S, also on the 0-100 scale. The construction treats Q and D as two independent signals that interact, not as two numbers to sum. The adjustment rewards well-executed work in a demanding design, and penalises poorly-executed work in a forgiving design. Off-diagonal cases (high Q with low D, or low Q with high D) receive only modest adjustments.
Definitions
Adjusted score
Q
review-based quality score, 0 to 100
D
study-design weight, 0 to 100
Delta
centred design weight, -50 to +50
m
alignment magnitude, 0 to 1
s = 0.15
interaction strength
b = 0.075
baseline strength
What the adjustment does
The adjustment to Q depends on how much Q and D align: they align when both are high or both are low, and diverge when one is high and the other is low. In the formula, the interaction term (s · m) grows as the two align, while the baseline term (b · (1 - m)) carries the off-diagonal cases. The four cases below show which term dominates in each condition.
High quality with a high-weight design (e.g. a well-executed meta-analysis of RCTs): alignment is high, the interaction term dominates, and the paper receives the largest positive adjustment. The system rewards good execution of a demanding study design.
Low quality with a low-weight design (e.g. a poorly-executed expert opinion): alignment is again high, but in the opposite direction. The interaction term dominates, and the paper receives the largest negative adjustment. Poor execution of easy work compounds against the score, because work of this type is straightforward to do correctly.
Off-diagonal cases (high Q with low D, or low Q with high D): alignment is low, the baseline term dominates, and the adjustment is small. A well-executed expert opinion keeps most of its quality credit. The system treats a meta-analysis that struggles more leniently, because the format is genuinely harder.
A neutral design weight. If D = 50, then the centred weight is 0 and the quality score does not change: S = Q.
Worked example: medicine
The table below takes the medical hierarchy as the reference. It shows the adjusted score S for a strong paper (Q = 80), an average paper (Q = 50), and a weak paper (Q = 20). The size of each adjustment is in parentheses.
Study design
D
Strong (Q = 80)
Average (Q = 50)
Weak (Q = 20)
Systematic review / meta-analysis of RCTs
D
100
Strong (Q = 80)
86.75 (+6.75)
Average (Q = 50)
55.62 (+5.62)
Weak (Q = 20)
24.50 (+4.50)
Individual RCT or observational with large effect
D
75
Strong (Q = 80)
83.38 (+3.38)
Average (Q = 50)
52.81 (+2.81)
Weak (Q = 20)
22.25 (+2.25)
Non-randomised controlled cohort
D
50
Strong (Q = 80)
80.00 (0.00)
Average (Q = 50)
50.00 (0.00)
Weak (Q = 20)
20.00 (0.00)
Case-series / case-control
D
25
Strong (Q = 80)
77.75 (-2.25)
Average (Q = 50)
47.19 (-2.81)
Weak (Q = 20)
16.62 (-3.38)
Mechanism-based reasoning / expert opinion
D
0
Strong (Q = 80)
75.50 (-4.50)
Average (Q = 50)
44.38 (-5.62)
Weak (Q = 20)
13.25 (-6.75)
Bounds. The maximum possible adjustment is +/-7.5 points. It occurs at the perfectly aligned corners (Q = 100 with D = 100, or Q = 0 with D = 0). At the perfectly misaligned corners (Q = 100 with D = 0, or Q = 0 with D = 100), the adjustment is exactly half: +/-3.75 points.
Classification: ANZSRC 2020 Fields of Research, Australian Bureau of Statistics & Stats NZ, released 30 June 2020
OCEBM schema: Oxford Centre for Evidence-Based Medicine, Levels of Evidence Working Group (2011). "The Oxford Levels of Evidence 2." cebm.ox.ac.uk
ESSA / WWC schema: US Dept of Education, What Works Clearinghouse Procedures and Standards Handbook v5.0 (2022)
Melnyk & Fineout-Overholt schema: Melnyk, B.M. & Fineout-Overholt, E. (2023). EBP in Nursing & Healthcare, 5th ed. Wolters Kluwer
IPCC schema: IPCC Guidance Note for Lead Authors on the Use of Expert Judgment and Treatment of Uncertainty (2010)
APA Div 12: Chambless, D.L. et al. (1998). Update on empirically validated therapies, II. The Clinical Psychologist, 51(1), 3-16
EBSE: Kitchenham, B. et al. (2004). Evidence-based software engineering. Proc. ICSE 2004
AIATSIS: AIATSIS Code of Ethics for Aboriginal and Torres Strait Islander Research (2020)
UNDRIP: United Nations General Assembly (2007). United Nations Declaration on the Rights of Indigenous Peoples, Resolution 61/295, adopted 13 September 2007. un.org/development/desa/indigenouspeoples
CARE Principles: Research Data Alliance International Indigenous Data Sovereignty Interest Group (September 2019). CARE Principles for Indigenous Data Governance. Global Indigenous Data Alliance. gida-global.org
GIDA: Global Indigenous Data Alliance, the network hosting the CARE Principles; see Carroll, S.R. et al. (2020), "The CARE Principles for Indigenous Data Governance," Data Science Journal 19(43). gida-global.org
Mathematical typesetting: KaTeX
The weight of a study comes from its position in the applicable hierarchy for its field, not the novelty, importance, or contribution of the work.
See it on your own paper. Free.
Get started free