How the score is calculated
Every paper carries three separate judgements. Rigor asks can I believe this?, clinical importance asks how much could this matter to patients?, and editorial priority asks why should I read it this week?
merit = [clinical importance + 0.5(editorial priority − 45)] × credibility(rigor) × synthesis certainty
heat = merit − excess missingness uncertainty − recency
credibility(r) = 0.35 + 0.65 (r/100)0.7
Clinical importance anchors the rank. Editorial priority contributes only its deviation from a neutral 45, so novelty can change reading order without inflating the estimated clinical effect. The floor of 0.35 means weak methods discount a finding rather than annihilating it, and credibility reaches exactly 1.00 at rigor 100 so nothing is ever inflated. The recency term takes off 1.5 points a week after the first week, capped at 12; it orders the live feed only, and the band a paper displays is taken before it, so nothing quietly drops a band three weeks after publication and an archived paper reads the same as it did on the day. A retraction zeroes all four numbers, and nothing else can.
Three rules shape everything below. Design sets a starting point and a ceiling, not a verdict. Gates fire only on observed defects, never on absence. Signals add up inside a domain, and domains cap rather than sum. The last is the shape of RoB 2, ROBINS-I and AMSTAR-2, and it is the answer to the 1999 demonstration that additive quality scales give contradictory answers about the same trials.
Drive it yourself
This runs the actual scoring module, not a simplified copy of it. Set a design and an appraisal and watch the numbers move. Three things are worth trying: switch a criterion between no and not reported and watch the score refuse to drop; drag the confidence interval from a decisive null to a demonstrated benefit; and turn on a spin flag to watch a cap clamp the score rather than nudge it.
PThe paper
RHow it was done
IWhat it found
Heat · clinical importance, editorial priority, and credibility
- Rigor
- 95
- Clinical
- 72
- Editorial
- 54
merit starts with clinical 72+ 0.5 × (editorial 54− 45), then applies credibility(95) 0.977
coverage 0.95 · uncertainty half-width ±7 · rank debit 1
A precise null: the 95% CI -1.33 to 2.84 rules out a difference as large as the MCID the paper cites of 5 in either direction.
Zone Z5 · magnitude score 63.9
What each criterion is worth right now
Each row re-runs the scorer with that one field changed. Note that not reported never scores below no.
| Criterion | Yes | Silent | No |
|---|---|---|---|
| Allocation concealment | 95 | 90 | 81 |
| Blinded outcome assessment | 95 | 88 | 79 |
| Random sequence generation | 95 | 93 | 86 |
| Intention-to-treat analysis | 95 | 92 | 86 |
| Prospective registration | 95 | 90 | 90 |
| Aligned time zero | 95 | 95 | 95 |
| Competing risks handled | 99 | 95 | 95 |
| Harms reported | 95 | 95 | 95 |
Allocation concealment was actually performed in 100% of surveyed trials but reported in 41%, and power calculations performed in 76% but reported in 16%. Reporting is a poor proxy for conduct, so silence contributes exactly zero. Only an affirmed defect subtracts.
The widget exposes the criteria that move a score most, not all seventy-odd fields of the appraisal, and the confidence-interval slider is read directly in MCID units. Everything else is the shipped code: the same computeHeat the weekly pipeline calls.
Rigor, in full
Rigor starts from a design prior drawn from a 47-family table and is held under a ceiling from the same row. The starting points are compressed and the ceilings wide, so execution decides where in the corridor a paper lands. A sham-controlled trial starts at 68.6 with no ceiling below 100; an adjusted registry cohort starts at 37.4 with a ceiling of 80 and therefore has 42.6 points of headroom; a retrospective case series starts at 8.8 under a ceiling of 42.
The ceiling is structural rather than a judgement about quality. A cross-sectional study cannot establish temporality however well it is conducted, and a cadaveric construct cannot speak to patient outcomes. Preclinical families are exempt from the spread that separates the clinical hierarchy, because their limitation belongs to impact, not to rigor: good bench science should read as good science answering a different question.
Each credited criterion then adds or subtracts exactly what the tables below say, scaled once by a constant 0.85. There is no variable discounting: an earlier build shrank the sum toward the prior as the appraisal thinned, and that multiplier could make an observed defect raise the score, so it was removed. A thin appraisal now sits near its prior simply because few criteria moved it. Sparseness widens the uncertainty band and subtracts the excess uncertainty from the weekly rank; a paper appraised on under 60% of its applicable criteria cannot lead the week. Journal tier is the only proxy left in the model, worth at most 6 points, and it decays to exactly zero once 80% of the applicable criteria have been read.
A property test now holds the whole axis to that contract: observing a strength can only raise a score, observing a defect can only lower it, and silence moves nothing.
| Concealed allocation | +6 |
| Allocation not concealedon a subjective outcome; −5 on an objective one; −8 when the outcome's objectivity was not reported, the midpoint rather than the worst case | −11 |
| Random sequence generation | +3 |
| Quasi-random allocation | −8 |
| Sham-controlleda bonus, and does not enlarge the coverage denominator | +8 |
| Adjustment-method labelmatching, regression, propensity scores, IPTW, g-methods and IV are neutral by themselves; sophistication is not adequacy | 0 |
| Unadjusted causal comparisona significant effectiveness claim also caps rigor at 50 | −8 to −10 |
| Post-adjustment balance demonstratedfor matching, propensity scores or weighting; important residual imbalance −6 | +5 |
| New-user designfor an intervention comparison; including prevalent users −4 | +4 |
| Exposure defined after cohort entryimmortal time bias, and it caps rigor at 25 | −14 |
| Comparatorplacebo or sham +6, current standard +4, outdated −10 | −10 to +6 |
| Blinded outcome assessmenton a subjective outcome; +4 on an objective one; +6 when objectivity was not reported | +8 |
| Unblinded assessment of a subjective outcomethe largest single criterion penalty in the model; −4 objective, −7.5 unreported | −11 |
| Registry coding validated | +3 |
| Independently adjudicated outcomes | +4 |
| Registry capturea continuous ramp from nothing at 85% capture to the full credit at 98% | 0 to +5 |
Blinding carries the heaviest penalty because the evidence says it should. Unblinded assessors exaggerate odds ratios by about 36% and standardised mean differences by 68%, and in orthopedics specifically the gap is a standardised mean difference of 0.76 unblinded against 0.25 blinded.
| Intention-to-treat analysis | +4 |
| Per-protocol analysis only | −7 |
| Follow-up short of the window the question needsa continuous ramp, largest when nothing of the window was observed. Below half the window a cap also engages, relaxing linearly as the shortfall closes. Arthroplasty survivorship needs 60 months; fracture union 9; periprosthetic infection 24 | 0 to −12 |
| Follow-up well past the thresholdramping up toward the generous benchmark for the question | 0 to +4 |
| Loss to follow-upa continuous ramp from +2 at no loss to −10 at 35%. The old bins charged 7 points for one patient crossing 20% | +2 to −10 |
| Crossover between armsramping down to −14 at 30% crossover; at 40% or more, rigor is capped at 65, because the arms as analysed are no longer the arms as randomised. Operative-versus-conservative crossover routinely exceeds 20% in orthopedic trials and dilutes the estimate toward the null | +2 to −14 |
| Adequately powered | +7 |
| Underpowered | −10 |
| No power calculation reportednot scored at all: it holds on 78% of the corpus, so it is an offset rather than a discriminator | 0 |
| Prospectively registereda bonus and never a penalty, because only about a quarter of orthopedic surgical trials report it | +6 |
| Headline is not the primary outcome | −8 |
| No correction for multiple comparisons | −4 |
| A positive result from under 50 patientsunder 100, −4 | −8 |
| A positive result from a single centre | −5 |
| An unpowered null from under 50 patientsunder 150, −4. A different defect from exaggeration: this debits an uninformative null for being uninformative, and never touches a null that was adequately powered | −7 |
| A single centre's own patients, in observational workselection, not exaggeration, so it applies to cohorts and series whichever way the result went | −4 |
| Cases selected rather than consecutiveconsecutive enrolment earns +3 | −6 |
| Fragility index at or below the subspecialty medianabove it +1, at three times it +3. Medians: arthroplasty 1, sports and spine 2, trauma 3, foot and ankle 6 | −5 |
| More patients lost than the result can absorb | −8 |
| Reverse fragility index of 5 or morefires only on a non-significant result, where a high value means the null is robust | +4 |
The size and single-centre penalties on positive results fire only when the primary outcome was significant, because small-study and single-centre inflation are directional: they exaggerate positive findings and cannot manufacture a null.
| Pools randomised trialsmixed designs −4, observational −8, uncontrolled case series −14 | +6 |
| I² and reported GRADE certaintyinconsistency and certainty affect ranking credibility once; neither substitutes for review conduct | 0 conduct points |
| Comprehensive literature search | +4 |
| Risk of bias assessed in the included studiesand carried into the synthesis rather than parked in a table, +3 more | +4 |
| Excluded studies listed and justified | +2 |
| Publication bias assessedaffirmatively NOT assessed across ten or more pooled studies, −8 | +3 |
| Two or more AMSTAR-2 critical failuresone caps at 64. Checked domains are protocol registration, search, excluded-study justification, risk-of-bias assessment and interpretation, plus publication-bias assessment when enough studies were pooled | cap 42 |
A synthesis is no longer charged criteria it cannot satisfy. An earlier build expected assessor blinding, a follow-up window and a power calculation of every meta-analysis, none of which a meta-analysis has, and the missing denominators quietly barred the one design that can close a question from ever leading the week.
| Immortal time bias | 25 |
| A bench study drawing a clinical conclusion | 30 |
| Data leakage between training and test sets | 30 |
| Two-gate sampling, conclusions the results do not support, or a prediction model with no external validation | 40 |
| A pilot making an efficacy claim | 45 |
| An effectiveness claim with no adjustment for confounding | 50 |
| A conclusion asserting a benefit the primary outcome did not shownever applied to a non-inferiority or equivalence trial, where concluding the treatments are comparable is correct reporting rather than spin | 55 |
| The outcome window was not reachedrelaxing linearly as the shortfall closes, so a month of follow-up cannot move the score 33 points | 60 up |
| Crossover of 40% or more between arms | 65 |
Clinical importance, in full
Clinical importance starts from a flat 45 for every paper, with no design prior at all. Design enters only at the bottom, as an absolute ceiling. The dominant term is not the author's conclusion but the confidence interval, normalised into units of the minimal clinically important difference and signed so that positive means benefit.
That choice is deliberate. Conclusion strength is exactly what spin inflates: 44.8% of orthopedic trials with a non-significant primary outcome contain spin, and spun papers are cited more, so scoring the rhetoric would rank the most spun papers highest by construction. The authors calling their own work practice-changing is worth zero points; what the data show drives the axis.
| The whole interval clears the MCIDa demonstrated clinically important effect, rising with how far the near bound clears it | 82 to 100 |
| A precise null · both bounds inside the MCIDthe question is answered: an important difference is ruled out in either direction, and the tighter the interval the higher the score | 56 to 69 |
| Significant but everything inside the MCIDreal, but noise wearing a p-value | 30 to 48 |
| Straddles the null and still allows more than the MCIDinconclusive, not negative, falling as the interval widens | 20 to 56 |
| Between those polesgraded by how decisively the data separate from the null, so a marginally significant result is treated as marginally significant instead of thrown to one side of a cliff | a blend |
| A demonstrated harma harm of a given size scores exactly what a benefit of that size scores | the benefit surface, mirrored |
| Not scoreable, or declined for a stated reasona refusal names its reason on the card: no concurrent comparator, a risk factor rather than a treatment, a ratio with no baseline risk, a hazard ratio without time-specific baseline survival, a risk difference with no unit, or an incoherent estimate | 50 |
The surface is continuous everywhere, and a property test holds it to a maximum step of a few points against any small change in a bound. An earlier build classified the interval into discrete zones by the signof its lower bound, which put a 36-point cliff exactly where a transcription error in the last digit of a confidence limit lands, and an interval whose bound was printed as exactly 1.00 fell through every zone and scored as if no effect had been reported. Non-inferiority and equivalence trials never reach this surface: non-inferiority is one-sided, while equivalence requires the whole interval inside both margins. A clinically acceptable demonstrated margin scores 72, an overly generous one scores 60, and a demonstrated result with no verifiable clinical threshold is capped at 58. Failure to meet the study's own margin scores 38; a missing margin declines to score rather than falling through to superiority logic.
Where a null sits, and why. A trial that demonstrated a benefit patients can feel ranks slightly above an otherwise identical trial that returned a decisive null. Holding a multicentre trial's design, blinding, registration and guideline collision fixed and changing only the effect estimate, the demonstrated benefit scores clinical importance 78 and the decisive null 72. The margin is small on purpose: both answer their question, and one of them also changes what we do.
What that re-ranking deliberately did not touch is more important than the re-ranking itself. A decisive null starts at 56 while a clearly significant but clinically trivial result tops out at 48, so the null wins by construction, and no amount of p-value polishing lifts a trivial finding past it. A decisive null also still outranks every observational design. A null is not weak evidence. It is just not, on its own, a reason to change an operation.
| The MCID the paper itself citesit is the bar the authors were judged against. But it is checked before it is trusted: a stated threshold outside 2% to 50% of the instrument's range, or wildly outside the published values we hold, is rejected and the library used instead. Untested, a cited MCID of 0.667 on QuickDASH bought the top of the scale while citing the honest 15.9 scored below saying nothing | wins when plausible |
| A published library valueKOOS JR spans 0.5 to 36.6 across sixteen derivation methods. Where published thresholds genuinely disagree, the verdict is scored on the best-matched value, damped toward neutral in proportion to how much it would move across the range, and the card says it is threshold-sensitive. Refusing outright was worse: it made effect size invisible on KOOS, ODI, ASES and QuickDASH, the instruments orthopedics uses most | population-matched |
| A risk ratio or odds ratiousing the measure-correct formula and control-arm event rate, then judged against the same absolute anchors as a risk difference: 1 point for mortality, 2 for revision, 5 for minor complications. Equivalent OR and risk-difference representations are regression-tested to receive the same magnitude | converted to absolute risk |
| A hazard ratioa hazard ratio is a relative event-rate comparison over time, not a risk ratio, so a simple 2x2 event table cannot identify its absolute effect | declines without time-specific baseline survival |
| A standardised mean differencerather than Cohen's arbitrary 0.5 | the median anchor-based important change |
| Mortality or major morbidity | 1.00 |
| Revision or reoperation | 0.95 |
| Patient-reported function or pain | 0.85 |
| Clinician-measured | 0.65 |
| Radiographic or physiologic surrogatean affirmative claim on a surrogate primary also caps the magnitude at 55, because such findings are inflated roughly fourfold | 0.45 |
| Laboratory or biomechanical | 0.30 |
The spread at the top is wider than it used to be, on the argument that revision is the outcome that actually drives implant and technique choice: a patient-reported difference is not a substitute for it, and the two were previously two hundredths apart. The factor is applied after the magnitude is centred on its neutral 50, which is what makes a lower tier scale the finding rather than penalise the reporting of it. Applied before centring, as an earlier build did, a modest surrogate result scored worse than reporting no result at all.
| Contradicts itStrong +20, Moderate +18, Limited +14, Consensus +12 | +12 to +20 |
| Fills a gap it statespays most where the guideline concedes the evidence is limited | +8 to +18 |
| Confirms itconfirming a weak recommendation is worth more than confirming a strong one, because a Strong recommendation needs two or more concordant high-quality studies. That row is what stops the model punishing replication | +2 to +10 |
| A relation we cannot tie to a real recommendationit is then a claim about the literature, not a collision with guidance | × 0.4 |
| A powered, registered null that settles the question | +5 |
| First evidence on the question (editorial)settles conflicting trials +5, first randomised evidence +4, adds to already consistent evidence −2 | +6 |
| Collision and novelty together (editorial)contradicting a recommendation and being first are partly the same editorial signal | capped at +24 |
| Comparatorsham +9, current standard +6, accepted alternative +3, outdated −3, none −5 | −5 to +9 |
| How many patients this affectsfloored at +3 for a limb- or life-threatening problem, so prevalence weighting cannot bury sarcoma or a mangled extremity | −2 to +7 |
| Needs no new equipmentrequires equipment most centres lack, −4 | +4 |
| Not reproducible from the description | −12 |
| Harms reportednot reported, −2 | +2 |
| Five or more centres | +4 |
| Under 10% of screened patients were eligibleenrolled most of what it screened, +3. This is GRADE's indirectness domain on the population: a trial that screened 4,000 patients to randomise 60 answered its question for a sliver of the people a reader treats | −5 |
| IPD meta-analysis of randomised trialsmeta-analysis of trials 98, network meta-analysis 94, systematic review of trials 92. Only synthesis of randomised evidence can reach the top of the scale | 100 |
| Sham-controlled or multicentre trialregistry-based trial 90, cluster 88, single-centre 85, crossover 80 | 92 |
| Target trial emulation | 78 |
| Prospective or adjusted registry cohortthe anchor: a 17-centre prospective cohort halving reoperation is clearly important and clearly not a landmark | 72 |
| Meta-analysis of observational studies | 70 |
| Retrospective comparative cohort | 65 |
| Cross-sectional | 55 |
| Cadaveric constructanimal and finite element 35 | 40 |
| Case report, editorial, narrative review | 25 |
These ceilings are not a statement that a given paper is poor. An clinical importance of 99 means a finding that changes orthopedic surgery: one that overturns a guideline or becomes one. A single trial, however flawless, is one concordant high-quality study, and a Strong recommendation needs two or more, so no single primary study reaches the top of the scale; the headroom above 92 belongs to synthesis, the only design that can close a question on its own. The same logic runs down the hierarchy: GRADE starts observational evidence at low certainty, which is why a cohort stops at 72. Reaching above clinical importance 85 also requires that at least 60% of the applicable criteria could actually be assessed. A sparse paper can outrank a fully appraised bad one, which is correct because we do not know it is bad, but it cannot win on ignorance.
Why you see a band, not a decimal
Scores are computed continuously and displayed as one of five bands: exceptional from 85, strong from 68, sound from 48, limited from 30, weak below that. Composite critical-appraisal agreement runs around an intraclass correlation of 0.65 to 0.70, which supports four to six distinguishable strata and not a decimal place. Every card also carries an explicit uncertainty half-width of 6 + 14 × (1 − coverage): about ±6 when the whole paper could be appraised, and ±20 when there was no full text to read. The portion above ±6 is also subtracted from the live rank, so missingness is not merely cosmetic. A bare 78 out of 100 asserts a precision the underlying judgement cannot support.
Two things are deliberately absent. Citation counts, because at one or two weeks old they measure indexing speed rather than importance. And reader-relevance signals such as Canadian authorship or subspecialty, because they say nothing about how good a paper is; they inform editorial triage instead. Heat is a screening tool, not a verdict. Cold papers can still be great papers.