Guideline Development | GRADE | Implementation | Critical Appraisal
- Clinical Practice Guideline (CPG): Systematically developed statements to assist practitioner and patient decisions about appropriate care for specific clinical circumstances.
- GRADE System: Separates evidence quality (High/Moderate/Low/Very Low) from recommendation strength (Strong/Weak).
- Strong Recommendation: Clinicians should follow in most patients. Requires large benefit, high-quality evidence, or ethical imperative.
- Weak Recommendation: Different choices for different patients. Requires shared decision-making based on patient values.
- Implementation Gap: Guidelines often not followed in practice due to barriers (awareness, agreement, adoption, adherence).
- “Strong recommendation does NOT always require high-quality evidence (can have low quality if large effect)
- “Guideline development panels should be multidisciplinary and free from conflicts of interest
- “AGREE II tool assesses guideline quality across 6 domains (scope, stakeholder involvement, rigor, clarity, applicability, independence)
- “Guidelines are population-level tools - they inform, but never override, individualised clinical judgement and shared decision-making
- “International guidelines require local adaptation using regional registry data (AOANJRR, NJR, AJRR) and resource setting
Overview
What a guideline is. Clinical practice guidelines (CPGs) are systematically developed statements that link the best available evidence to explicit, actionable recommendations for specific clinical circumstances. They sit at the junction of evidence-based medicine, implementation science and medicolegal practice, and they recur in basic-science and clinical vivas because they test whether a candidate can appraise evidence rather than merely recall it.
Why appraisal is not optional. This page teaches AGREE II and GRADE, and its own citations say why you need them; the figure is worse than most readers expect. When 279 guidelines published between 1985 and 1997 were scored against a 25-item methodological standard, mean adherence was 43.1%, about 10.8 of the 25 items (Shaneyfelt, JAMA 1999, PMID 10349893). The worst-performing domain was identification and synthesis of the evidence, at 33.6%, the very step a guideline exists to perform, and adherence improved only modestly across the period studied, from 36.9% to 50.4%.
The examinable position follows from that. A guideline is an artefact to be appraised, not an authority to be obeyed, and "which guideline says so" is a weaker answer than "how was that guideline made". That is the justification for the critical appraisal skills tested alongside this topic, and for knowing where a recommendation sits in the levels of evidence and which study design produced it.
The three pillars. The exam-relevant understanding rests on how a guideline grades its evidence and recommendations (GRADE), how its methodological quality is appraised (AGREE II), and why good guidelines still fail at the bedside (implementation science). A consultant-level answer treats a guideline as a population-level tool that informs but never dictates the care of the individual in front of you.
Anatomy of a Guideline
The components. A guideline has five essential components, and each can be inspected on its own:
- Clinical question (PICO): population, intervention, comparison, outcome
- Evidence summary: a systematic review of the available evidence
- Evidence quality assessment: GRADE or a similar system
- Recommendation statements: clear, actionable guidance
- Rationale: the explanation linking the evidence to the recommendation
The question. PICO frames the question the evidence summary answers.
- Definition
- Target patient group
- Example (VTE Prophylaxis)
- Adults undergoing major orthopaedic surgery
- Definition
- Treatment or exposure
- Example (VTE Prophylaxis)
- Pharmacological thromboprophylaxis
- Definition
- Alternative (if applicable)
- Example (VTE Prophylaxis)
- No prophylaxis or mechanical only
- Definition
- Measurable health outcomes
- Example (VTE Prophylaxis)
- Symptomatic DVT/PE, major bleeding
Types of guideline. By scope, a guideline is condition-specific (a single disease or injury, such as ACL rupture), procedure-specific (a single intervention, such as TKA) or cross-cutting (applying across conditions, such as VTE prophylaxis). By developer, it comes from a professional society (AAOS, AOA), a government agency (NHMRC, NICE), a health system (hospital-specific) or a Cochrane group.
Appraising Guideline Quality
AGREE II. The instrument assesses the methodological rigour and transparency of a guideline across six domains and 23 items, and it is the gold standard for guideline appraisal. It answers how the guideline was made, and each domain asks a short set of questions of the document.
Domain 1: Scope and Purpose (23%). Clarity on who the guideline applies to prevents misapplication, and a well-defined scope prevents guideline creep.
- Are the objectives clearly described?
- Are the health questions covered by the guideline specified?
- Is the target population clearly described?
Scoring. Each item is rated 1-7, and domain scores are calculated as a percentage of the maximum possible. The appraisal closes with an overall assessment: would you recommend this guideline for use? Yes, yes with modifications, or no.
The rigour threshold. A score above 60% in Rigour of Development indicates a methodologically sound guideline. Many orthopaedic guidelines score poorly on this domain, so always check the evidence grading and the systematic review methodology before applying the recommendations.
The quick assessment. Before a formal AGREE II appraisal, four questions screen a guideline:
- Who developed it? (reputation, expertise)
- When was it updated? (currency; within 5 years is ideal)
- What evidence grading was used? (GRADE preferred)
- Are conflicts declared? (funding source, disclosures)
Guidelines more than 5 years old may be outdated. Always check the publication date and whether there have been subsequent updates or superseding guidelines. Key orthopaedic guidelines (AAOS, NICE) are typically reviewed every 3-5 years.
A living guideline is continuously surveilled and updated as new evidence emerges, rather than on a fixed 3-5 year cycle.
Finding guidelines. Start with the specialty society guidelines, check the government or national guideline databases, then search PubMed for "clinical practice guideline" plus the topic. The TRIP Database searches multiple guideline databases at once; the other main sources each have a characteristic strength and limitation.
- Coverage
- National guideline collections
- Strengths
- Local relevance and applicability
- Limitations
- Limited orthopaedic-specific content
- Coverage
- UK NHS guidelines
- Strengths
- Rigorous methodology, well-maintained
- Limitations
- UK-specific recommendations
- Coverage
- US orthopaedic practice
- Strengths
- Specialty-specific, GRADE methodology
- Limitations
- May not apply to all populations
- Coverage
- Systematic reviews
- Strengths
- Gold standard methodology
- Limitations
- Reviews, not recommendations
- Coverage
- Global guidelines
- Strengths
- Comprehensive, searchable
- Limitations
- Variable quality
Implementation and Barriers
Why guidelines fail. A high-quality guideline changes nothing unless it changes practice, and the reasons it does not are diagnosable. There are four barriers, and each has its own remedy.
AAAAImplementation Barriers - The 4 As
Hook:Diagnose which barrier dominates locally, then match the intervention - tailoring to identified barriers outperforms generic rollout.
Strategies. Passive dissemination, publication and mailing, has low effectiveness. The active strategies are:
- Education and dissemination through multiple channels
- Clinical decision support: electronic alerts, order sets
- Audit and feedback: compare performance to the guideline and feed the result back
- Academic detailing: one-on-one education with opinion leaders
- Local champions and opinion leaders
- Multifaceted interventions: combining strategies, the most effective
What the evidence says about uptake. The reviews of what actually changes practice (Grimshaw 2004; Cabana 1999) point to a small number of consistent factors:
- Effect
- Passive dissemination has minimal effect; active strategies modest-moderate
- Practical Action
- Do not rely on publication or email alone
- Effect
- Small to moderate improvement, larger when baseline performance is low
- Practical Action
- Feed back individual or unit data against a clear target
- Effect
- Among the more reliable single interventions
- Practical Action
- Build prompts into order sets and the EHR
- Effect
- Improve adoption by addressing the 'agreement' barrier
- Practical Action
- Recruit respected colleagues as champions
- Effect
- Tailoring to identified barriers outperforms generic rollout
- Practical Action
- Diagnose barriers first, then match the intervention
Read "multifaceted is most effective" with care: Grimshaw found no reliable relationship between the number of components in a multifaceted intervention and its effect: more is not automatically better, and fit to the local barrier matters more. The same review of 235 studies and 309 comparisons found only about 29% reported any cost information (PMID 14960256). Implementation is not free and is poorly costed, and a recommendation without an implementation plan is a wish.
A guideline mandated is not a guideline implemented. The WHO surgical safety checklist is the best worked example on this page. It has three parts: sign in (identity, consent, site marking), time out (team briefing, antibiotic timing) and sign out (counts, recovery plan). Introducing the 19-item checklist across eight hospitals in eight cities was associated with inpatient death falling from 1.5% to 0.8% (p=0.003) and complications from 11.0% to 7.0% (p less than 0.001) (Haynes 2009, PMID 19144931).
Quote it with its companion. Those eight hospitals volunteered for a WHO improvement programme and there was no concurrent control; when Ontario mandated checklists across 101 hospitals, the effect vanished, with mortality 0.71% to 0.65% (OR 0.91, 95% CI 0.80-1.03, p=0.13) and complications 3.86% to 3.82% (OR 0.97, p=0.29) across 215,711 procedures (Urbach 2014, PMID 24620866). The difference between the two results is the whole of implementation science.
- The clinical audit cycle (a closed loop): (1) set a standard/criterion - usually taken directly from a guideline (e.g. "antibiotics within 60 minutes of incision in 100 percent of cases"); (2) measure current practice against it; (3) compare performance to the standard and identify the gap; (4) implement change (the barrier-matched intervention); and crucially (5) RE-AUDIT to close the loop. An audit that never re-measures is not a true audit - closing the loop is the examiner's favourite point.
- PDSA (Plan-Do-Study-Act) is the rapid, small-scale iterative cousin: plan a change, do it on a small scale, study the effect, act (adopt/adapt/abandon), then repeat - multiple fast cycles rather than one big annual audit.
- Audit vs research: audit measures practice against an existing standard (no new knowledge, usually no ethics approval); research generates new knowledge and needs ethical review. Confusing the two is a common error.
Exam point: implement a guideline through the closed-loop clinical audit cycle (set standard → measure → compare → change → re-audit) or iterative PDSA cycles - the defining feature is re-auditing to close the loop, and audit (against a standard) is distinct from research (new knowledge).
Measuring whether it worked. Implementation is judged on process measures (guideline awareness, compliance rates) and outcome measures (complication rates, patient-reported outcomes, cost-effectiveness). The benefits it is meant to deliver are reduced variation in care, improved adherence to best practice and measurable outcome improvements; VTE guidelines, for instance, reduced symptomatic PE rates. In postoperative care the usual quality indicators are VTE prophylaxis timing, antibiotic compliance, SSI rates and readmission rates.
Most guideline-impact studies report process measures (e.g. proportion receiving prophylaxis on time), not patient outcomes. A high compliance rate is necessary but not sufficient - it does not prove patient benefit unless linked to a hard outcome. When appraising an implementation study, ask whether it measured what patients actually care about.
- Structure - the fixed attributes of the setting: staffing, equipment, theatre/ICU availability, whether a guideline/protocol exists at all. Easy to measure but the weakest link to patient benefit.
- Process - what is actually done to the patient: proportion receiving timely antibiotics, VTE prophylaxis prescribed, guideline-concordant care delivered. The most sensitive and actionable measure of guideline uptake, and what most implementation studies report.
- Outcome - what happens to the patient: mortality, surgical-site infection, symptomatic VTE, patient-reported outcome measures (PROMs). What patients ultimately care about, but confounded by case-mix and often needs risk-adjustment and large numbers.
The key insight (linking to the "process vs outcome" pearl above): good structure and process are necessary but not sufficient - high compliance (process) only proves benefit when it is shown to move a hard outcome. Use a balanced set across all three, and risk-adjust outcomes before comparing units.
Exam point: frame quality measurement with the Donabedian triad - structure, process and outcome - process measures best capture guideline uptake, outcome measures (risk-adjusted) capture patient benefit, and you need both because high process compliance does not by itself prove improved outcomes.
Guidelines in Orthopaedic Practice
What guidelines address. Indications for surgery, VTE and antibiotic prophylaxis, perioperative care protocols and enhanced recovery pathways. Specific surgical techniques, implant selection and approach comparisons are often not addressed.
Most surgical technique recommendations are consensus-based. Surgical technique relies on training and observational evidence.
AAOS examples.
- Recommendation
- Early surgery within 24-48h
- Strength
- Strong
- Recommendation
- Against arthroscopic debridement
- Strength
- Strong
- Recommendation
- Pharmacological or mechanical
- Strength
- Moderate
- Recommendation
- Exercise before surgery
- Strength
- Moderate
Postoperative care. Guideline-directed postoperative care covers VTE prophylaxis (duration and agent), antibiotic prophylaxis, pain management protocols and ERAS pathways.
Enhanced Recovery After Surgery bundles combine multiple guideline recommendations. Shown to reduce complications, length of stay, and costs in arthroplasty.
VTE prophylaxis by guideline. The agents and durations differ by body:
- Agents
- LMWH, fondaparinux, warfarin, aspirin, DOAC
- Duration
- 10-35 days
- Agents
- Pharmacological or mechanical
- Duration
- Variable
- Agents
- LMWH, rivaroxaban, aspirin
- Duration
- 14-35 days
- Agents
- LMWH, rivaroxaban, aspirin
- Duration
- 28-35 days
Limitations, Deviation and the Medicolegal Position
Not cookbook medicine. Guidelines inform decisions; they do not dictate them. Individual patient factors may override a recommendation, and rare complications or comorbidities may not be addressed at all. Wise clinicians use a guideline as the starting point and then individualise on patient-specific factors.
When a guideline does not apply. Deviation is expected when:
- The patient is atypical (age, comorbidities) or has a contraindication
- The patient's preferences differ from the guideline's assumptions
- Local resources are unavailable
- New evidence has been published since the guideline was developed
Strong recommendations don't mean every patient must receive the intervention. Individualise based on patient factors and preferences, documenting rationale when deviating.
How guidelines go wrong. Misapplication means applying a guideline to the wrong population, using an outdated one, or ignoring individual factors. Overreliance is cookbook medicine, defensive practice and ignoring clinical judgement. The guidelines themselves carry inherent limits (evidence gaps, lag time in development, a population rather than individual focus) and quality concerns (conflicts of interest, industry influence, variable methodology).
Always check conflict of interest disclosures. Industry-funded guidelines may overestimate treatment benefits.
The medicolegal position. Guidelines are evidence of, but do not define, the standard of care.
- Courts (Bolam/Bolitho in the UK, comparable standards elsewhere) ask whether a responsible body of practitioners would have acted similarly; a guideline is strong evidence of that body's practice but is not automatically binding.
- Documented, reasoned deviation, based on patient-specific factors and shared decision-making, is defensible. Unaware deviation is far harder to defend.
- Guidelines can cut both ways: following an outdated or poor-quality guideline is not a defence if a reasonable clinician would have recognised it as superseded.
- The safest position is awareness of the relevant guideline, an explicit rationale when departing from it, and a contemporaneous note of the shared decision.
Exam Focus
Differential: Guideline vs Related Evidence Documents
A frequent viva trap is conflating a clinical practice guideline with other documents that look similar but carry different authority. Know the distinctions.
- Basis
- Systematic review of evidence + explicit benefit-harm appraisal
- How Recommendations Are Made
- GRADE or similar; strength separated from evidence quality
- Authority
- Highest - evidence-based, transparent, updatable
- Basis
- Expert opinion, often without systematic review
- How Recommendations Are Made
- Voting / Delphi process; evidence link often implicit
- Authority
- Lower - reflects opinion, useful where evidence is sparse
- Basis
- Systematic synthesis of primary studies
- How Recommendations Are Made
- Summarises evidence but makes NO clinical recommendation
- Authority
- Evidence source - feeds guidelines, not a recommendation itself
- Basis
- Local operationalisation of guidance
- How Recommendations Are Made
- Translates recommendations into step-by-step local actions
- Authority
- Local - implementation tool, not de novo evidence
- Basis
- Auditable minimum requirements derived from guidelines
- How Recommendations Are Made
- Short, mandatory-style statements for audit
- Authority
- Sets a measurable floor for quality
Controversies and Areas of Uncertainty
Critics argue rigid guideline adherence promotes "cookbook medicine" and erodes clinical judgement; proponents counter that they reduce unwarranted variation. The resolution is that strong recommendations apply to populations, while individuals are still managed by shared decision-making.
Many guideline panels historically included members with financial conflicts. Whether full recusal, a conflict-free chair, or transparent declaration is sufficient remains debated - AGREE II's editorial-independence domain captures this concern.
Traditional 3-5 year update cycles leave guidelines outdated between revisions. Living guidelines with continuous surveillance address this but are resource-intensive and not yet standard in orthopaedics.
Guidelines built largely on trial populations may not apply to the comorbid, elderly, or under-represented patients seen in practice - a key reason a guideline informs but never replaces individualised judgement.
Guidelines, Registries & Global Practice
Major Guideline Developers Worldwide
American Academy of Orthopaedic Surgeons. GRADE-based orthopaedic CPGs (hip fracture, knee/hip OA, VTE, rotator cuff, distal radius). Strength downgraded transparently when evidence is limited.
National Institute for Health and Care Excellence plus British Orthopaedic Association standards. Highly rigorous methodology, economic modelling (cost-per-QALY), and short, auditable BOAST standards for trauma.
AO Foundation fracture-management principles and EFORT European consensus statements. Strong on operative technique consensus where RCT evidence is sparse.
NHMRC (Australia), AOA (Australia), SIGN (Scotland), WHO (global surgical safety), and society guidelines elsewhere. Therapeutic Guidelines-type formularies localise antibiotic and VTE prophylaxis.
Side-by-Side: Where Guidelines Genuinely Differ
- Position
- AAOS and many US/Australian protocols accept aspirin; NICE/ACCP historically favoured LMWH or DOACs
- Reason for Divergence
- Weighting of bleeding vs thrombosis, cost, and registry/RCT evidence (e.g. CRISTAL)
- Position
- AAOS strong recommendation AGAINST; broadly concordant across NICE/EFORT
- Reason for Divergence
- Consistent high-quality RCT evidence (Moseley, Kirkley) - rare global agreement
- Position
- NICE favours cemented for fragility hip fracture; registries inform nuance
- Reason for Divergence
- Registry signals (cement implantation syndrome vs periprosthetic fracture) weighted differently
Registry Evidence Feeding Guidelines
Arthroplasty and fixation guidance increasingly draws on national joint registries rather than RCTs alone:
- NJR (England, Wales, NI, IoM, Guernsey), AJRR (US), AOANJRR (Australia), SHAR (Sweden), Norwegian and NZJR registries supply implant survival and revision-rate data on millions of procedures.
- Registries provide external validity and rare-event detection (e.g. metal-on-metal failure) that trials cannot, but are observational - they inform but cannot prove causation.
- Guideline panels combine registry survival data with RCT functional outcomes and GRADE the resulting body of evidence.
High- vs Limited-Resource Practice Variation
International guidelines require local adaptation; the ADAPTE and GRADE-ADOLOPMENT frameworks formalise this.
Key drivers of variation:
- Resource setting: availability of DOACs, intra-operative imaging, implant range, and theatre capacity changes what is feasible (a strong recommendation is meaningless if the drug or implant is unavailable).
- Local epidemiology: injury patterns, comorbidity burden, and antimicrobial resistance differ by region and alter the benefit-harm balance.
- Local registry data: revision rates for a given implant in the local population may justify departing from an international default.
Example: an international guideline may default to a DOAC for thromboprophylaxis, but a limited-resource setting may justifiably adopt aspirin plus mechanical prophylaxis based on cost, monitoring capacity, and acceptable local outcome data - a defensible local adaptation, not a deviation in error.
MCQ Practice Points
Q: Can a guideline make a strong recommendation based on low-quality evidence? A: Yes - GRADE separates evidence quality (confidence in effect) from recommendation strength (should we do it). Strong recommendation possible with low-quality evidence if there is a large magnitude of effect, ethical imperative, or clear benefit-harm balance favoring intervention. Example: Strong recommendation for surgery in displaced fractures despite lack of RCTs.
Q: What is the most important AGREE II domain for assessing guideline quality? A: Rigor of Development - assesses whether systematic methods were used to search for evidence, appraise quality, link evidence to recommendations, and formulate recommendations using explicit criteria. This distinguishes evidence-based guidelines from expert consensus documents.
Q: What are the main barriers to guideline implementation? A: The 4 As: Awareness (clinicians do not know guideline exists), Agreement (disagree with recommendations), Adoption (too difficult to implement due to resources or system barriers), Adherence (forget to apply or revert to old habits). Multifaceted active implementation strategies needed.
Q: How often should clinical practice guidelines be updated? A: Guidelines should be reassessed every 2-3 years and formally updated every 3-5 years. Living guidelines use continuous surveillance to update recommendations as new evidence emerges. A guideline is considered outdated if it has not been updated within 5 years or if substantial new evidence contradicts current recommendations.
Q: How should conflicts of interest be managed in guideline development? A: Panel members should declare all financial and intellectual COI at the outset. Those with significant COI should recuse from voting on related recommendations. The chair of the guideline panel should ideally be free from relevant COI. All declarations should be publicly available in guideline documentation. COI management is a key domain assessed by AGREE II.
Exam Viva Scenarios
Practise clinical reasoning and management decisions out loud
“How do you critically appraise a clinical practice guideline?”
“A patient asks about a new treatment they read about online. How do you use clinical practice guidelines to inform your discussion?”
“How do you approach shared decision-making when guidelines make a weak recommendation?”
“You deviate from a guideline and the patient has a complication. How do you defend your decision?”
“How would you implement a new guideline in your department?”
Guideline Definition and Purpose
- CPG = Systematically developed statements to guide clinical decisions
- Based on systematic review of evidence and explicit consideration of benefits/harms
- Purpose: Reduce unwarranted variation, improve quality, inform policy
- Should be updated every 3-5 years as new evidence emerges
- Distinguish from consensus statements (opinion-based, not systematic)
GRADE System
- Evidence Quality: High/Moderate/Low/Very Low (confidence in effect estimate)
- Recommendation Strength: Strong/Weak (should we do it?)
- RCT starts at High, Observational starts at Low, then apply modifiers
- Downgrade for: RIIIP (Risk of bias, Inconsistency, Indirectness, Imprecision, Publication bias)
- Strong recommendation possible with low evidence if large effect or ethical imperative
Strong vs Weak Recommendations
- Strong: We recommend / Most patients should receive / Can use as quality measure
- Weak: We suggest / Different choices for different patients / Shared decision-making
- Strong requires: Large benefit, minimal harm, aligned values, feasible, OR ethical imperative
- Weak: Close benefit-harm balance, varied patient values, high cost, or uncertain evidence
- Wording signals strength - clinicians must recognize difference
AGREE II Quality Appraisal
- 6 domains: Scope, Stakeholder involvement, Rigor (most important), Clarity, Applicability, Independence
- Rigor domain: Systematic search, explicit methods, evidence-to-recommendation link, external review, update plan
- Editorial independence: Funding declared, conflicts managed, majority non-conflicted
- Patient involvement essential for patient-centered guidelines
- Overall assessment: Recommend for use / With modifications / Do not recommend
Implementation Barriers and Solutions
- 4 As: Awareness, Agreement, Adoption, Adherence
- Passive dissemination (publication, mailing) = ineffective
- Active strategies: Clinical decision support, audit-feedback, academic detailing, reminders
- Multifaceted interventions (combine strategies) most effective
- Local adaptation needed to address barriers and context
Guidelines, Registries & Global Practice
- Major developers: AAOS (US), NICE/BOA-BOAST (UK), AO Foundation/EFORT (Europe), NHMRC/AOA, SIGN, WHO
- Registries (NJR, AJRR, AOANJRR, SHAR, NZJR) supply implant survival and revision data feeding guidelines
- Genuine divergence: aspirin vs LMWH/DOAC for VTE; consensus AGAINST knee arthroscopic debridement
- Adaptation frameworks: ADAPTE and GRADE-ADOLOPMENT formalise local tailoring
- Adapt to local epidemiology, registry data, and resource setting - a strong rec is meaningless if the drug/implant is unavailable
GRADE: Quality of Evidence
Two separate judgements. GRADE rates the quality of the evidence (High, Moderate, Low, Very Low) separately from the strength of the recommendation (Strong or Weak, for or against). Quality is confidence in the effect estimate; strength answers whether we should do it. The recommendation side is taken in the next section.
The four levels. Each level carries a statement of how much further research would change the picture:
- Typical evidence
- RCTs without serious limitations
- What it means
- Further research is unlikely to change confidence in the effect
- Typical evidence
- RCTs with limitations, or strong observational studies
- What it means
- Further research is likely to change confidence
- Typical evidence
- Observational studies
- What it means
- Further research is very likely to change the effect estimate
- Typical evidence
- Case series, expert opinion
- What it means
- The estimate of effect is very uncertain
Where a body of evidence starts. A randomised trial starts at High; an observational study starts at Low. That baseline is established before any judgement is applied, and the explicit modifiers then move it down or up.
Downgrading. Five factors lower the rating.
RIIIPGRADE Downgrade Factors - RIIIP
Hook:Risk of bias, inconsistency, indirectness and imprecision can each drop quality by 1 or 2 levels; publication bias drops it by 1.
Upgrading. The mirror-image factors raise the rating, and they apply typically to observational studies with exceptional features.
- Criterion
- RR greater than 2 or less than 0.5 with no plausible confounders
- Upgrade By
- Plus 1 level
- Criterion
- RR greater than 5 or less than 0.2
- Upgrade By
- Plus 2 levels
- Criterion
- Clear dose-response gradient
- Upgrade By
- Plus 1 level
- Criterion
- All plausible confounders would REDUCE effect (bias toward null)
- Upgrade By
- Plus 1 level
Oxford levels of evidence. A separate scheme ranks study designs by level:
- Study Design
- Systematic review of RCTs with homogeneity
- Example in Orthopaedics
- Cochrane review of TXA in arthroplasty
- Study Design
- Individual RCT with narrow CI
- Example in Orthopaedics
- SPRINT trial (reaming in tibial nailing)
- Study Design
- Systematic review of cohort studies
- Example in Orthopaedics
- Meta-analysis of arthroplasty registry data
- Study Design
- Individual cohort study or low-quality RCT
- Example in Orthopaedics
- Registry study of bearing surface outcomes
- Study Design
- Systematic review of case-control studies
- Example in Orthopaedics
- Meta-analysis of risk factors for PJI
- Study Design
- Individual case-control study
- Example in Orthopaedics
- Case-control study of implant loosening
- Study Design
- Case series, poor cohort/case-control
- Example in Orthopaedics
- Case series of new surgical technique
- Study Design
- Expert opinion without critical appraisal
- Example in Orthopaedics
- Consensus statement

GRADE: From Evidence to Recommendation
Two strengths. A recommendation is strong or weak, for or against. Strong means the benefits clearly outweigh the harms; weak (also called conditional) means the balance of benefits and harms is close. The wording is the signal, "we recommend" against "we suggest", and surgeons must recognise the difference.
- Strong Recommendation
- We recommend... / Clinicians should...
- Weak Recommendation
- We suggest... / Clinicians might consider...
- Strong Recommendation
- Most patients would want this intervention
- Weak Recommendation
- Different choices for different patients based on values
- Strong Recommendation
- Most patients should receive this intervention
- Weak Recommendation
- Engage in shared decision-making, individualise
- Strong Recommendation
- Can be used as performance measure or quality indicator
- Weak Recommendation
- Should NOT be used as performance measure (patient choice matters)
Do not state that a strong recommendation requires high-quality evidence. GRADE explicitly decouples the two - a strong recommendation can rest on low-quality evidence when the effect is large, harm is minimal, or there is an ethical imperative. Reciting this decoupling correctly is a high-yield mark.
When low-quality evidence still earns a strong recommendation. The situations are (1) benefits that clearly outweigh harms even with the uncertainty, (2) ethical considerations that mandate action, and (3) resource implications that favour the intervention. The standing example is prophylactic antibiotics for open fractures: a strong recommendation despite limited RCT evidence.
Evidence to decision. Beyond evidence quality, the GRADE evidence-to-decision framework weighs six things:
- Benefits and harms: the balance of desirable and undesirable effects
- Values and preferences: how important the outcomes are to patients
- Resources: cost and feasibility
- Equity: impact on health disparities
- Acceptability: to stakeholders
- Feasibility: implementation barriers
Two guidelines can therefore make different recommendations from the same evidence, according to local values, resources or priorities.
Applying the strength in practice. A strong recommendation is applied to most patients unless contraindicated: "We recommend VTE prophylaxis for major orthopaedic surgery" means giving prophylaxis to essentially all patients. A weak recommendation requires shared decision-making: a recommendation worded "We suggest arthroscopic debridement may be considered for mechanical symptoms in early OA" means discussing the alternatives, because patient values matter. Where there is no recommendation the evidence is insufficient to guide practice; use clinical judgement and tell the patient about the uncertainty.
Shared decision-making. For a weak recommendation the management algorithm is explicitly shared decision-making:
- Present the options, including the option of no intervention.
- Communicate the evidence: the magnitude of benefit and harm in absolute terms (natural frequencies, not relative risk), plus the GRADE confidence in those estimates.
- Elicit values and preferences: what outcomes matter most to this patient (function versus avoiding surgery versus avoiding bleeding).
- Reach and document a shared decision that integrates the evidence with the patient's goals.
Decision aids improve knowledge and reduce decisional conflict; they are the practical tool that operationalises a weak recommendation at the bedside.
Evidence Base
GRADE: an emerging consensus on rating quality of evidence and strength of recommendations
- Landmark statement introducing the GRADE system to a broad clinical audience
- Separates evidence quality (confidence in the effect estimate) from recommendation strength (should we do it)
- Evidence rated High/Moderate/Low/Very Low; recommendations Strong or Weak
- RCTs start High and observational studies start Low, then are up- or downgraded by explicit factors
- Now adopted by WHO, Cochrane and over 100 organisations worldwide
GRADE guidelines: 15. Going from evidence to recommendation - determinants of a recommendation's direction and strength
- Defines how panels move from a body of evidence to a graded recommendation
- Four determinants: balance of desirable vs undesirable effects, confidence in estimates, values and preferences, and resource use
- Strength reflects confidence that net desirable effects outweigh undesirable effects
- Explicitly allows a strong recommendation despite low-confidence evidence in defined situations
- Requires explicit panel judgement integrating all four domains