ortHOTalk home

How the score is calculated

Every paper carries three separate judgements. Rigor asks can I believe this?, clinical importance asks how much could this matter to patients?, and editorial priority asks why should I read it this week?

merit = [clinical importance + 0.5(editorial priority − 45)] × credibility(rigor) × synthesis certainty
heat = merit − excess missingness uncertainty − recency
credibility(r) = 0.35 + 0.65 (r/100)0.7

Clinical importance anchors the rank. Editorial priority contributes only its deviation from a neutral 45, so novelty can change reading order without inflating the estimated clinical effect. The floor of 0.35 means weak methods discount a finding rather than annihilating it, and credibility reaches exactly 1.00 at rigor 100 so nothing is ever inflated. The recency term takes off 1.5 points a week after the first week, capped at 12; it orders the live feed only, and the band a paper displays is taken before it, so nothing quietly drops a band three weeks after publication and an archived paper reads the same as it did on the day. A retraction zeroes all four numbers, and nothing else can.

Three rules shape everything below. Design sets a starting point and a ceiling, not a verdict. Gates fire only on observed defects, never on absence. Signals add up inside a domain, and domains cap rather than sum. The last is the shape of RoB 2, ROBINS-I and AMSTAR-2, and it is the answer to the 1999 demonstration that additive quality scales give contradictory answers about the same trials.

Drive it yourself

This runs the actual scoring module, not a simplified copy of it. Set a design and an appraisal and watch the numbers move. Three things are worth trying: switch a criterion between no and not reported and watch the score refuse to drop; drag the confidence interval from a decisive null to a demonstrated benefit; and turn on a spin flag to watch a cap clamp the score rather than nudge it.

PThe paper

Design family
Patients randomised
Centres
Journal tier

RHow it was done

Primary outcome
Outcome type
Allocation concealmentneutral when silent
Blinded outcome assessmentneutral when silent
Intention-to-treat analysis
Prospective registrationneutral when silent
Power calculation
Adjustment for confounding
Spin: conclusions the primary outcome does not support

IWhat it found

Confidence interval, in MCID units
Outcome tier
Relation to existing guidance
Topic, matched against the recommendation table
How many patients this affects

Heat · clinical importance, editorial priority, and credibility

Strong73.8roughly 6781
Rigor
95
Clinical
72
Editorial
54

merit starts with clinical 72+ 0.5 × (editorial 54− 45), then applies credibility(95) 0.977

coverage 0.95 · uncertainty half-width ±7 · rank debit 1

A precise null: the 95% CI -1.33 to 2.84 rules out a difference as large as the MCID the paper cites of 5 in either direction.

Zone Z5 · magnitude score 63.9

What each criterion is worth right now

Each row re-runs the scorer with that one field changed. Note that not reported never scores below no.

CriterionYesSilentNo
Allocation concealment959081
Blinded outcome assessment958879
Random sequence generation959386
Intention-to-treat analysis959286
Prospective registration959090
Aligned time zero959595
Competing risks handled999595
Harms reported959595

Allocation concealment was actually performed in 100% of surveyed trials but reported in 41%, and power calculations performed in 76% but reported in 16%. Reporting is a poor proxy for conduct, so silence contributes exactly zero. Only an affirmed defect subtracts.

The widget exposes the criteria that move a score most, not all seventy-odd fields of the appraisal, and the confidence-interval slider is read directly in MCID units. Everything else is the shipped code: the same computeHeat the weekly pipeline calls.

Rigor, in full

Rigor starts from a design prior drawn from a 47-family table and is held under a ceiling from the same row. The starting points are compressed and the ceilings wide, so execution decides where in the corridor a paper lands. A sham-controlled trial starts at 68.6 with no ceiling below 100; an adjusted registry cohort starts at 37.4 with a ceiling of 80 and therefore has 42.6 points of headroom; a retrospective case series starts at 8.8 under a ceiling of 42.

The ceiling is structural rather than a judgement about quality. A cross-sectional study cannot establish temporality however well it is conducted, and a cadaveric construct cannot speak to patient outcomes. Preclinical families are exempt from the spread that separates the clinical hierarchy, because their limitation belongs to impact, not to rigor: good bench science should read as good science answering a different question.

Each credited criterion then adds or subtracts exactly what the tables below say, scaled once by a constant 0.85. There is no variable discounting: an earlier build shrank the sum toward the prior as the appraisal thinned, and that multiplier could make an observed defect raise the score, so it was removed. A thin appraisal now sits near its prior simply because few criteria moved it. Sparseness widens the uncertainty band and subtracts the excess uncertainty from the weekly rank; a paper appraised on under 60% of its applicable criteria cannot lead the week. Journal tier is the only proxy left in the model, worth at most 6 points, and it decays to exactly zero once 80% of the applicable criteria have been read.

A property test now holds the whole axis to that contract: observing a strength can only raise a score, observing a defect can only lower it, and silence moves nothing.

D1 · Allocation and comparison validity
Concealed allocation+6
Allocation not concealedon a subjective outcome; −5 on an objective one; −8 when the outcome's objectivity was not reported, the midpoint rather than the worst case−11
Random sequence generation+3
Quasi-random allocation−8
Sham-controlleda bonus, and does not enlarge the coverage denominator+8
Adjustment-method labelmatching, regression, propensity scores, IPTW, g-methods and IV are neutral by themselves; sophistication is not adequacy0
Unadjusted causal comparisona significant effectiveness claim also caps rigor at 50−8 to −10
Post-adjustment balance demonstratedfor matching, propensity scores or weighting; important residual imbalance −6+5
New-user designfor an intervention comparison; including prevalent users −4+4
Exposure defined after cohort entryimmortal time bias, and it caps rigor at 25−14
Comparatorplacebo or sham +6, current standard +4, outdated −10−10 to +6
D2 · Measurement integrity
Blinded outcome assessmenton a subjective outcome; +4 on an objective one; +6 when objectivity was not reported+8
Unblinded assessment of a subjective outcomethe largest single criterion penalty in the model; −4 objective, −7.5 unreported−11
Registry coding validated+3
Independently adjudicated outcomes+4
Registry capturea continuous ramp from nothing at 85% capture to the full credit at 98%0 to +5

Blinding carries the heaviest penalty because the evidence says it should. Unblinded assessors exaggerate odds ratios by about 36% and standardised mean differences by 68%, and in orthopedics specifically the gap is a standardised mean difference of 0.76 unblinded against 0.25 blinded.

D3 · Completeness
Intention-to-treat analysis+4
Per-protocol analysis only−7
Follow-up short of the window the question needsa continuous ramp, largest when nothing of the window was observed. Below half the window a cap also engages, relaxing linearly as the shortfall closes. Arthroplasty survivorship needs 60 months; fracture union 9; periprosthetic infection 240 to −12
Follow-up well past the thresholdramping up toward the generous benchmark for the question0 to +4
Loss to follow-upa continuous ramp from +2 at no loss to −10 at 35%. The old bins charged 7 points for one patient crossing 20%+2 to −10
Crossover between armsramping down to −14 at 30% crossover; at 40% or more, rigor is capped at 65, because the arms as analysed are no longer the arms as randomised. Operative-versus-conservative crossover routinely exceeds 20% in orthopedic trials and dilutes the estimate toward the null+2 to −14
D4 · Reporting integrity
Adequately powered+7
Underpowered−10
No power calculation reportednot scored at all: it holds on 78% of the corpus, so it is an offset rather than a discriminator0
Prospectively registereda bonus and never a penalty, because only about a quarter of orthopedic surgical trials report it+6
Headline is not the primary outcome−8
No correction for multiple comparisons−4
D5 · Precision and selection
A positive result from under 50 patientsunder 100, −4−8
A positive result from a single centre−5
An unpowered null from under 50 patientsunder 150, −4. A different defect from exaggeration: this debits an uninformative null for being uninformative, and never touches a null that was adequately powered−7
A single centre's own patients, in observational workselection, not exaggeration, so it applies to cohorts and series whichever way the result went−4
Cases selected rather than consecutiveconsecutive enrolment earns +3−6
Fragility index at or below the subspecialty medianabove it +1, at three times it +3. Medians: arthroplasty 1, sports and spine 2, trauma 3, foot and ankle 6−5
More patients lost than the result can absorb−8
Reverse fragility index of 5 or morefires only on a non-significant result, where a high value means the null is robust+4

The size and single-centre penalties on positive results fire only when the primary outcome was significant, because small-study and single-centre inflation are directional: they exaggerate positive findings and cannot manufacture a null.

Synthesis · judged on its own criteria, not a trial's
Pools randomised trialsmixed designs −4, observational −8, uncontrolled case series −14+6
I² and reported GRADE certaintyinconsistency and certainty affect ranking credibility once; neither substitutes for review conduct0 conduct points
Comprehensive literature search+4
Risk of bias assessed in the included studiesand carried into the synthesis rather than parked in a table, +3 more+4
Excluded studies listed and justified+2
Publication bias assessedaffirmatively NOT assessed across ten or more pooled studies, −8+3
Two or more AMSTAR-2 critical failuresone caps at 64. Checked domains are protocol registration, search, excluded-study justification, risk-of-bias assessment and interpretation, plus publication-bias assessment when enough studies were pooledcap 42

A synthesis is no longer charged criteria it cannot satisfy. An earlier build expected assessor blinding, a follow-up window and a power calculation of every meta-analysis, none of which a meta-analysis has, and the missing denominators quietly barred the one design that can close a question from ever leading the week.

Hard caps · every one fires on an observed defect, never on silence
Immortal time bias25
A bench study drawing a clinical conclusion30
Data leakage between training and test sets30
Two-gate sampling, conclusions the results do not support, or a prediction model with no external validation40
A pilot making an efficacy claim45
An effectiveness claim with no adjustment for confounding50
A conclusion asserting a benefit the primary outcome did not shownever applied to a non-inferiority or equivalence trial, where concluding the treatments are comparable is correct reporting rather than spin55
The outcome window was not reachedrelaxing linearly as the shortfall closes, so a month of follow-up cannot move the score 33 points60 up
Crossover of 40% or more between arms65

Clinical importance, in full

Clinical importance starts from a flat 45 for every paper, with no design prior at all. Design enters only at the bottom, as an absolute ceiling. The dominant term is not the author's conclusion but the confidence interval, normalised into units of the minimal clinically important difference and signed so that positive means benefit.

That choice is deliberate. Conclusion strength is exactly what spin inflates: 44.8% of orthopedic trials with a non-significant primary outcome contain spin, and spun papers are cited more, so scoring the rhetoric would rank the most spun papers highest by construction. The authors calling their own work practice-changing is worth zero points; what the data show drives the axis.

The magnitude surface · one continuous function of the interval, in MCID units
The whole interval clears the MCIDa demonstrated clinically important effect, rising with how far the near bound clears it82 to 100
A precise null · both bounds inside the MCIDthe question is answered: an important difference is ruled out in either direction, and the tighter the interval the higher the score56 to 69
Significant but everything inside the MCIDreal, but noise wearing a p-value30 to 48
Straddles the null and still allows more than the MCIDinconclusive, not negative, falling as the interval widens20 to 56
Between those polesgraded by how decisively the data separate from the null, so a marginally significant result is treated as marginally significant instead of thrown to one side of a cliffa blend
A demonstrated harma harm of a given size scores exactly what a benefit of that size scoresthe benefit surface, mirrored
Not scoreable, or declined for a stated reasona refusal names its reason on the card: no concurrent comparator, a risk factor rather than a treatment, a ratio with no baseline risk, a hazard ratio without time-specific baseline survival, a risk difference with no unit, or an incoherent estimate50

The surface is continuous everywhere, and a property test holds it to a maximum step of a few points against any small change in a bound. An earlier build classified the interval into discrete zones by the signof its lower bound, which put a 36-point cliff exactly where a transcription error in the last digit of a confidence limit lands, and an interval whose bound was printed as exactly 1.00 fell through every zone and scored as if no effect had been reported. Non-inferiority and equivalence trials never reach this surface: non-inferiority is one-sided, while equivalence requires the whole interval inside both margins. A clinically acceptable demonstrated margin scores 72, an overly generous one scores 60, and a demonstrated result with no verifiable clinical threshold is capped at 58. Failure to meet the study's own margin scores 38; a missing margin declines to score rather than falling through to superiority logic.

Where a null sits, and why. A trial that demonstrated a benefit patients can feel ranks slightly above an otherwise identical trial that returned a decisive null. Holding a multicentre trial's design, blinding, registration and guideline collision fixed and changing only the effect estimate, the demonstrated benefit scores clinical importance 78 and the decisive null 72. The margin is small on purpose: both answer their question, and one of them also changes what we do.

What that re-ranking deliberately did not touch is more important than the re-ranking itself. A decisive null starts at 56 while a clearly significant but clinically trivial result tops out at 48, so the null wins by construction, and no amount of p-value polishing lifts a trivial finding past it. A decisive null also still outranks every observational design. A null is not weak evidence. It is just not, on its own, a reason to change an operation.

The MCID divisor
The MCID the paper itself citesit is the bar the authors were judged against. But it is checked before it is trusted: a stated threshold outside 2% to 50% of the instrument's range, or wildly outside the published values we hold, is rejected and the library used instead. Untested, a cited MCID of 0.667 on QuickDASH bought the top of the scale while citing the honest 15.9 scored below saying nothingwins when plausible
A published library valueKOOS JR spans 0.5 to 36.6 across sixteen derivation methods. Where published thresholds genuinely disagree, the verdict is scored on the best-matched value, damped toward neutral in proportion to how much it would move across the range, and the card says it is threshold-sensitive. Refusing outright was worse: it made effect size invisible on KOOS, ODI, ASES and QuickDASH, the instruments orthopedics uses mostpopulation-matched
A risk ratio or odds ratiousing the measure-correct formula and control-arm event rate, then judged against the same absolute anchors as a risk difference: 1 point for mortality, 2 for revision, 5 for minor complications. Equivalent OR and risk-difference representations are regression-tested to receive the same magnitudeconverted to absolute risk
A hazard ratioa hazard ratio is a relative event-rate comparison over time, not a risk ratio, so a simple 2x2 event table cannot identify its absolute effectdeclines without time-specific baseline survival
A standardised mean differencerather than Cohen's arbitrary 0.5the median anchor-based important change
Outcome tier · scales the demonstrated effect after centring, so a neutral verdict is worth zero at every tier and a large effect on a surrogate cannot out-score a modest one on revision
Mortality or major morbidity1.00
Revision or reoperation0.95
Patient-reported function or pain0.85
Clinician-measured0.65
Radiographic or physiologic surrogatean affirmative claim on a surrogate primary also caps the magnitude at 55, because such findings are inflated roughly fourfold0.45
Laboratory or biomechanical0.30

The spread at the top is wider than it used to be, on the argument that revision is the outcome that actually drives implant and technique choice: a patient-reported difference is not a substitute for it, and the two were previously two hundredths apart. The factor is applied after the magnitude is centred on its neutral 50, which is what makes a lower tier scale the finding rather than penalise the reporting of it. Applied before centring, as an earlier build did, a modest surrogate result scored worse than reporting no result at all.

Editorial priority · guideline collision by the recommendation actually matched
Contradicts itStrong +20, Moderate +18, Limited +14, Consensus +12+12 to +20
Fills a gap it statespays most where the guideline concedes the evidence is limited+8 to +18
Confirms itconfirming a weak recommendation is worth more than confirming a strong one, because a Strong recommendation needs two or more concordant high-quality studies. That row is what stops the model punishing replication+2 to +10
A relation we cannot tie to a real recommendationit is then a claim about the literature, not a collision with guidance× 0.4
Clinical-importance inputs, except where labelled editorial
A powered, registered null that settles the question+5
First evidence on the question (editorial)settles conflicting trials +5, first randomised evidence +4, adds to already consistent evidence −2+6
Collision and novelty together (editorial)contradicting a recommendation and being first are partly the same editorial signalcapped at +24
Comparatorsham +9, current standard +6, accepted alternative +3, outdated −3, none −5−5 to +9
How many patients this affectsfloored at +3 for a limb- or life-threatening problem, so prevalence weighting cannot bury sarcoma or a mangled extremity−2 to +7
Needs no new equipmentrequires equipment most centres lack, −4+4
Not reproducible from the description−12
Harms reportednot reported, −2+2
Five or more centres+4
Under 10% of screened patients were eligibleenrolled most of what it screened, +3. This is GRADE's indirectness domain on the population: a trial that screened 4,000 patients to randomise 60 answered its question for a sliver of the people a reader treats−5
Absolute clinical-importance ceilings · every family carries one, whatever the execution
IPD meta-analysis of randomised trialsmeta-analysis of trials 98, network meta-analysis 94, systematic review of trials 92. Only synthesis of randomised evidence can reach the top of the scale100
Sham-controlled or multicentre trialregistry-based trial 90, cluster 88, single-centre 85, crossover 8092
Target trial emulation78
Prospective or adjusted registry cohortthe anchor: a 17-centre prospective cohort halving reoperation is clearly important and clearly not a landmark72
Meta-analysis of observational studies70
Retrospective comparative cohort65
Cross-sectional55
Cadaveric constructanimal and finite element 3540
Case report, editorial, narrative review25

These ceilings are not a statement that a given paper is poor. An clinical importance of 99 means a finding that changes orthopedic surgery: one that overturns a guideline or becomes one. A single trial, however flawless, is one concordant high-quality study, and a Strong recommendation needs two or more, so no single primary study reaches the top of the scale; the headroom above 92 belongs to synthesis, the only design that can close a question on its own. The same logic runs down the hierarchy: GRADE starts observational evidence at low certainty, which is why a cohort stops at 72. Reaching above clinical importance 85 also requires that at least 60% of the applicable criteria could actually be assessed. A sparse paper can outrank a fully appraised bad one, which is correct because we do not know it is bad, but it cannot win on ignorance.

Why you see a band, not a decimal

Scores are computed continuously and displayed as one of five bands: exceptional from 85, strong from 68, sound from 48, limited from 30, weak below that. Composite critical-appraisal agreement runs around an intraclass correlation of 0.65 to 0.70, which supports four to six distinguishable strata and not a decimal place. Every card also carries an explicit uncertainty half-width of 6 + 14 × (1 − coverage): about ±6 when the whole paper could be appraised, and ±20 when there was no full text to read. The portion above ±6 is also subtracted from the live rank, so missingness is not merely cosmetic. A bare 78 out of 100 asserts a precision the underlying judgement cannot support.

Two things are deliberately absent. Citation counts, because at one or two weeks old they measure indexing speed rather than importance. And reader-relevance signals such as Canadian authorship or subspecialty, because they say nothing about how good a paper is; they inform editorial triage instead. Heat is a screening tool, not a verdict. Cold papers can still be great papers.