Research Methodologies | Study Hierarchy | Evidence Quality
- RCT: Random allocation eliminates selection bias and balances known/unknown confounders
- Cohort Study: Follows exposed and unexposed groups forward in time to measure outcomes
- Case-Control Study: Starts with disease (cases) and no disease (controls), looks backward for exposures
- Cross-Sectional Study: Snapshot in time - measures exposure and outcome simultaneously
- Case Series: Descriptive study of patients with similar condition - no comparison group
- “RCT is gold standard for therapeutic interventions but not always ethical or feasible
- “Cohort studies are best for rare exposures; Case-control studies are best for rare outcomes
- “Observational studies are prone to confounding and bias - must use statistical adjustment
- “Registry studies provide real-world effectiveness data but lack randomisation
Overview and Classification
Experimental or observational. This is the fundamental division. In an experimental study, the RCT, the investigator assigns the intervention. In an observational study the investigator only observes what happens naturally, and observational designs are prone to confounding and bias that call for statistical adjustment.
Direction in time. A prospective study collects data going forward from the day it starts and follows participants forward in time. A retrospective study looks back at existing data from past records. A cross-sectional study takes a single point in time.
Internal and external validity. Internal validity asks whether the results are valid within the study. External validity asks whether they can be generalised to other populations.
- Type
- Randomised Controlled Trial
- Investigator Role
- Assigns intervention
- Examples
- Drug trial, surgical technique comparison
- Type
- Cohort Study
- Investigator Role
- Observes only
- Examples
- Smoking and nonunion, registry studies
- Type
- Case-Control Study
- Investigator Role
- Observes only
- Examples
- Rare disease risk factors
- Type
- Cross-Sectional
- Investigator Role
- Observes only
- Examples
- Prevalence surveys
- Type
- Case Series
- Investigator Role
- Observes only
- Examples
- Novel technique reports
Randomised Controlled Trials
The design. Participants are randomly allocated to an intervention or a control group, then followed prospectively to measure outcomes. The control group provides the comparison against which the treatment effect is measured, and blinding can be single, double or triple.
Why randomisation matters. Random allocation eliminates selection bias and balances known and unknown confounders, creating groups that are comparable at baseline. It is what gives the RCT the highest level of evidence for therapeutic questions: bias and confounding are minimised and causality can be established.
Its limits. RCTs are expensive and time-consuming, and their narrow inclusion criteria mean they may not reflect real-world practice. They are not ethical for harmful exposures and not feasible for rare outcomes.
- Description
- Two separate groups compared
- Advantage
- Simple analysis, most common
- Disadvantage
- Requires large sample size
- Description
- Each participant receives both treatments
- Advantage
- Smaller sample needed, controls for individual variation
- Disadvantage
- Requires washout period, carryover effects
- Description
- Tests 2 or more interventions simultaneously
- Advantage
- Efficient, can assess interactions
- Disadvantage
- Complex analysis, increased sample size
- Description
- Groups (hospitals, clinics) randomised, not individuals
- Advantage
- Prevents contamination, practical
- Disadvantage
- Larger sample needed, complex statistics
The question a trial is framed to answer. Beyond these structural variations, RCTs differ in their hypothesis:
- Superiority - is the new treatment better than the comparator? This is the default framework, and failing to show superiority does not prove equivalence.
- Non-inferiority - is the new treatment not worse than the standard by more than a pre-specified margin (Δ)? It is used when the new option has other advantages (cheaper, safer, less invasive, oral vs IV), and the margin must be defined and clinically justified in advance.
- Equivalence - are the two treatments within a margin of each other in both directions (e.g. a biosimilar)?
In a superiority trial, intention-to-treat (ITT) is the conservative analysis. In a non-inferiority trial the opposite is true: ITT (which dilutes differences by including crossovers/non-adherers) biases toward falsely concluding non-inferiority, so non-inferiority trials must report both ITT and per-protocol analyses and require both to agree. Also remember: "no statistically significant difference" in a superiority trial does not establish non-inferiority — absence of evidence is not evidence of absence.
Observational Analytical Study Designs
Cohort studies
The design. A cohort study follows groups with and without an exposure forward in time and compares the incidence of outcomes. Establishing temporality cleanly is the cohort's advantage, since exposure is recorded before any outcome occurs.
Prospective cohort. Identify exposed and unexposed groups at baseline, follow both forward in time, measure the incidence of outcomes and calculate the relative risk (RR). Following surgeons who operate (exposed) and those who do not (unexposed) to measure radiation exposure and cancer risk is an example, and the design gives Level II evidence.
What a prospective cohort buys and costs. It yields incidence and relative risk, can study multiple outcomes, makes the temporal relationship clear and is less prone to recall bias. It is time-consuming and expensive, loses patients to follow-up and is not efficient for rare outcomes, and confounding is possible.
Retrospective cohort. Existing records identify past exposures, the cohort is followed forward through those records, and outcomes that have already occurred are measured to calculate a relative risk. Reviewing registry data for patients who received cemented or uncemented THA in the 1990s and measuring revision rates to the present day is an example.
Efficient, but only as good as the records. A retrospective cohort is faster and cheaper than a prospective one, can study rare exposures and draws large samples from databases. It depends on the quality of existing records, missing data are common, the investigator cannot control data collection, and confounding by indication is a limitation.
Case-control studies
The design. Start with cases, who have the disease or outcome, and controls, matched or unmatched, who do not. Measure past exposure in both groups and calculate an odds ratio (OR). Comparing patients with AVN (cases) to those without AVN (controls) to see whether steroid use was more common in the cases is an example.
Where it earns its place. A case-control study is efficient for rare diseases, faster and cheaper than a cohort study, can study multiple exposures and needs a small sample. It gives Level III evidence: useful for rare outcomes, but inferior to a cohort for establishing causality.
Its weaknesses. It cannot calculate incidence or relative risk, only an odds ratio. It is prone to recall bias and selection bias, and confounding is common.
Temporality is a weakness here. Because you start from people who already have the outcome and look backwards for exposure, it is often unclear whether the exposure truly preceded the disease or was an early consequence of it. Case-control sits above cross-sectional on this point only because it at least attempts a time direction.
Observational Descriptive Study Designs
Cross-sectional studies. Exposure and outcome are measured at a single point in time, a snapshot. The design serves prevalence surveys, screening studies and hypothesis generation; surveying orthopaedic surgeons for the prevalence of burnout and correlating it with work hours is an example. It is quick, inexpensive and good for prevalence data.
What a snapshot cannot do. A cross-sectional study cannot measure incidence. With exposure and outcome measured together the temporal relationship is unclear (which came first?), so it cannot establish causality, and it is subject to survival bias.
Case series and case reports. A case series is a descriptive study of patients with a similar condition and no comparison group. It is simple to conduct and is used to describe new diseases or rare conditions, report novel surgical techniques and generate hypotheses. It has no control group, cannot establish causality and carries selection bias; a case series is Level IV evidence, and a case report sits at Level V with expert opinion.
Systematic Reviews and Meta-Analysis
Systematic review. A comprehensive, reproducible synthesis of all available evidence on a specific question. It uses explicit, pre-specified methods, a comprehensive literature search and critical appraisal of the included studies, and its synthesis may be qualitative or quantitative. It is reported to the PRISMA standard (Preferred Reporting Items for Systematic Reviews and Meta-Analyses), set out with the other checklists under Limitations and Pitfalls.
Meta-analysis. The statistical combination of results from multiple studies, providing a pooled effect estimate with a confidence interval. It is appropriate when the studies are clinically and methodologically similar and heterogeneity is acceptable (I² less than 75%).
Reading a forest plot. Each study is represented by its point estimate and confidence interval, and its weight is proportional to its sample size or precision. The diamond at the bottom is the pooled estimate. The line of no effect is RR = 1 or OR = 1, and a confidence interval that crosses it is not significant.
Heterogeneity. The I² statistic is the percentage of variation due to heterogeneity: less than 25% is low, 25-75% moderate, and greater than 75% high, at which point pooling should be reconsidered. The Q statistic is a chi-square test for heterogeneity, and a random-effects model is used when heterogeneity is present.
Meta-analysis of low-quality studies produces low-quality evidence. "Garbage in, garbage out" - pooling biased studies does not eliminate bias. Quality assessment (risk of bias) is essential.
Registry Studies in Orthopaedics
What a registry offers. Registry studies are large-scale observational studies using data from national or regional registries. They bring sample sizes in the hundreds of thousands, real-world effectiveness data and long follow-up, and they detect rare outcomes and complications and track implant performance. They are observational only, with no randomisation, and their limits are confounding by indication, variable data quality and limited clinical detail.
Major orthopaedic registries. These are the national registries to know:
- AOANJRR (Australia) - over 500,000 THAs and TKAs
- Swedish Hip Arthroplasty Register - established 1979, with the longest follow-up
- National Joint Registry (UK) - over 3 million procedures recorded
- American Joint Replacement Registry (AJRR) - a growing database
Reading registry survival. Implant survival is shown on Kaplan-Meier curves with revision for any reason as the endpoint. Groups are compared with hazard ratios, and competing-risk analysis separates death from revision.
Propensity score matching. A statistical technique to reduce confounding, it matches treated and untreated patients on their probability of receiving the treatment, creating pseudo-randomised groups. It does not control for unmeasured confounders.
Registries provide EFFECTIVENESS data (real-world outcomes), while RCTs provide EFFICACY data (outcomes under ideal conditions). Both are valuable but answer different questions.
Choosing a Design
Match the design to the question. Each kind of clinical question has a best design, an alternative, and the measure that design reports.
- Best Design
- RCT (if ethical and feasible)
- Alternative
- Prospective Cohort
- Measure
- RR, NNT, ARR
- Best Design
- Prospective Cohort Study
- Alternative
- Retrospective cohort from registry, Case Series
- Measure
- Survival rates, hazard ratio
- Best Design
- Cohort or Case-Control
- Alternative
- RCT (if ethical)
- Measure
- RR, OR
- Best Design
- Case-Control Study
- Alternative
- Large registry cohort
- Measure
- OR
- Best Design
- Cross-sectional survey
- Alternative
- Registry analysis
- Measure
- Prevalence
- Best Design
- Cross-Sectional
- Alternative
- Cohort
- Measure
- Sensitivity, Specificity, LR
- Best Design
- Cost-effectiveness study
- Alternative
- Decision analysis
- Measure
- ICER, QALY
The FINER criteria help choose the research question and the design to answer it.
FINERChoosing the Right Study Design
Hook:FINER criteria help you choose the right research question and design!
Study Design Components
Population and sampling. The target population is the group about whom conclusions will be drawn, and the study sample is the subset of it actually studied. Participants are selected by random, consecutive or convenience sampling.
Exposure and outcome. The exposure or intervention is what is being studied, a treatment or a risk factor. The outcome is what is being measured: disease, recovery or complication.
Control groups. The comparison can take several forms:
- Placebo control - an inactive treatment, with ethical concerns in surgery
- Active control - comparison with standard treatment
- Historical control - comparison with past data, a weak design
- Within-subject control - crossover designs
Bias and Confounding
Selection bias. Systematic error in how participants are selected, for example including only patients who survived long enough to be studied. Random sampling and consecutive enrolment prevent it.
Information bias. Also called measurement bias, this is systematic error in how data are collected. In recall bias cases remember exposures better than controls; in observer bias the assessor is influenced by knowledge of group allocation. Blinding and standardised measurement prevent it.
Confounding. A confounder is a third variable associated with both the exposure and the outcome. It can create a spurious association or mask a true one, and it is controlled by randomisation, matching, stratification or multivariable analysis (regression).
Screening biases. These are high-yield and often missed:
- Lead-time bias - screening detects disease earlier, so survival measured from diagnosis appears longer even if the time of death is unchanged; the apparent survival gain is just an earlier start point, not a real benefit
- Length-time bias - screening preferentially detects slow-growing, indolent disease, which spends longer in a detectable asymptomatic phase, so screen-detected cases look as if they do better. Overdiagnosis, detecting disease that would never have caused harm, is its extreme
Judge a screening programme by disease-specific mortality in an RCT, not by survival from diagnosis or 5-year survival comparisons.
- Main Bias Risks
- Performance bias, detection bias
- Prevention Strategies
- Blinding of participants, assessors, analysts
- Main Bias Risks
- Confounding, loss to follow-up
- Prevention Strategies
- Matching, multivariable adjustment, sensitivity analysis
- Main Bias Risks
- Recall bias, selection bias
- Prevention Strategies
- Blinded interviewing, multiple control groups
- Main Bias Risks
- Survivor bias, temporal ambiguity
- Prevention Strategies
- Cannot fully address - inherent limitation
Statistical Measures by Design
Relative risk. The incidence in the exposed divided by the incidence in the unexposed; an RR greater than 1 means increased risk with exposure. It is used in cohort studies and RCTs, and cannot be calculated from a case-control study.
Odds ratio. The odds of exposure in cases divided by the odds of exposure in controls. It can be calculated in cohorts and RCTs as well, but it is the only measure available from a case-control design, and it approximates the relative risk when the outcome is rare (less than 10%).
Hazard ratio. Used in survival (time-to-event) analysis, the HR is the instantaneous risk of the event at any time point, and it accounts for censoring and time-varying exposure.
Absolute measures. The absolute risk reduction (ARR) is the control event rate minus the treatment event rate (CER - EER), and it is more clinically meaningful than the RR. The number needed to treat (NNT = 1/ARR) is the number of patients treated to prevent one event, so a lower NNT means a more effective treatment. The number needed to harm (NNH = 1/absolute risk increase) is the number treated before one is harmed, and a higher NNH means a safer treatment.
- Cohort
- Yes
- Case-Control
- No
- RCT
- Yes
- Cohort
- Yes
- Case-Control
- Yes
- RCT
- Yes
- Cohort
- Yes
- Case-Control
- No
- RCT
- Yes
- Cohort
- Yes (computable, but confounded)
- Case-Control
- No
- RCT
- Yes
Why a cohort is "yes" for NNT. NNT is 1/ARR, and ARR is the difference between the two event rates. A cohort measures incidence in both its exposed and unexposed groups, so the arithmetic is available, and denying it NNT would contradict the incidence and relative-risk rows of the table. What a cohort cannot supply is the causal warrant: without randomisation the difference between the groups carries their confounding with it, so a cohort NNT describes an association, not the effect of treating.
Why case-control is a genuine "no". Sampling is done on the outcome, so the investigator, not nature, fixes the ratio of events to non-events. Incidence is therefore not recoverable, and with no absolute risks there is no ARR and no NNT.
One structural fact, three entries. Sampling on outcome instead of exposure is also the reason case-control yields only an odds ratio. It explains three entries in this table at once, and it is the most examinable idea on the page.
Outcomes and Endpoints
Primary outcome. The main outcome the study is powered to detect and the one used to calculate the sample size; it should be clinically meaningful. There is only one primary outcome, because multiple primary outcomes inflate the type I error.
Secondary outcomes. Additional outcomes of interest. They are exploratory, not powered to detect, and generate hypotheses for future studies.
Surrogate or patient-centred. A surrogate outcome is a laboratory value or radiograph, such as radiographic union. A patient-centred outcome is function, pain or quality of life, such as PROMIS scores. Surrogate outcomes may not correlate with patient-centred ones.
Patient-reported outcome measures. PROMs are generic (SF-36, EQ-5D, VAS pain) or joint-specific (WOMAC, Oxford Hip/Knee Score, DASH), and are validated, reliable and responsive to change.
Composite outcomes. These combine multiple outcomes into a single endpoint, as MACE (major adverse cardiac events) does. Combining them increases the event rate and reduces the sample size, and the components should be of similar importance.
Minimal clinically important difference. The MCID is the smallest change that matters to patients, used to interpret whether a statistical difference is clinically meaningful. For VAS pain it is about 1-2 points.
A result can be statistically significant (p less than 0.05) but not clinically significant if the difference is smaller than the MCID. Always consider whether the effect size is meaningful to patients.
Limitations and Pitfalls
Pitfalls by design. Recall bias, selection bias and the case-control limits on incidence and relative risk are covered above; each design also has its own traps:
- RCT - underpowered studies (type II error), poor allocation concealment, unblinded outcome assessors, per-protocol analysis instead of ITT, and narrow inclusion criteria limiting generalisability
- Cohort - loss to follow-up (over 20% is concerning), confounding by indication, immortal time bias, and the selection of exposed and unexposed groups
- Case-control - inappropriate control selection
Reporting and appraisal checklists. Match the checklist to the design:
- CONSORT (RCTs) - 25-item checklist for trial reporting; flow diagram required; endorsed by major journals
- STROBE (observational studies) - 22-item checklist, with separate versions for cohort, case-control and cross-sectional designs; focuses on transparent reporting
- PRISMA (systematic reviews) - 27-item checklist with a flow diagram of the study selection process; required for publication
- MINORS (non-randomised studies) - 12-item methodological index, a scoring system for quality assessment
Guidelines, Registries & Global Practice
Study-design methodology is governed by international reporting standards and evidence-grading frameworks rather than disease-specific clinical guidelines. The dominant frameworks are convergent worldwide: CONSORT for trials, STROBE for observational studies, PRISMA for systematic reviews, and GRADE for rating certainty of evidence. National bodies layer their own evidence hierarchies on top of these, and large national arthroplasty registries supply the real-world observational evidence that complements the trial literature.
Reporting Standards and Evidence Frameworks (Side-by-Side)
- Region
- International (EQUATOR)
- Purpose
- Reporting of RCTs
- Output
- 25-item checklist + flow diagram
- Region
- International (EQUATOR)
- Purpose
- Reporting of observational studies
- Output
- 22-item checklist (cohort/case-control/cross-sectional)
- Region
- International (EQUATOR)
- Purpose
- Reporting of systematic reviews
- Output
- 27-item checklist + flow diagram
- Region
- International (WHO, Cochrane)
- Purpose
- Rating certainty of evidence + recommendation strength
- Output
- High / Moderate / Low / Very low
- Region
- UK / international
- Purpose
- Level of evidence by question type
- Output
- Levels 1-5, question-specific
- Region
- Australia
- Purpose
- Evidence hierarchy + recommendation grades
- Output
- Levels I-IV, Grades A-D
- Region
- UK
- Purpose
- Guideline development using GRADE
- Output
- GRADE-based evidence profiles
Position of the Major Guideline Bodies
The American Academy of Orthopaedic Surgeons Clinical Practice Guidelines grade each recommendation (Strong, Moderate, Limited, Consensus) according to the level of evidence underpinning it, using a system derived from the Oxford/CEBM hierarchy and explicit risk-of-bias appraisal.
NICE develops guidance using the GRADE approach, separating certainty of evidence from strength of recommendation. The British Orthopaedic Association Standards (BOASTs) translate this evidence into auditable practice standards.
The AO Foundation and EFORT promote structured evidence appraisal and education across Europe, applying CONSORT/STROBE/PRISMA to trauma and arthroplasty literature and supporting multinational registry collaboration.
The National Health and Medical Research Council evidence hierarchy spans Levels I-IV with recommendation Grades A-D under the FORM framework. It mirrors international standards but explicitly incorporates Australian registry evidence.
National Arthroplasty Registries (Global Practice Variation)
- Country
- Sweden
- Established
- 1975 (knee) / 1979 (hip)
- Scale / Notable Feature
- Longest continuous follow-up; pioneered registry methodology
- Country
- Australia
- Established
- 1999
- Scale / Notable Feature
- Near-complete national capture; mandatory reporting; early outlier-implant detection
- Country
- UK (Eng/Wales/NI/IoM)
- Established
- 2003
- Scale / Notable Feature
- Over 3 million procedures; surgeon- and unit-level outcomes
- Country
- USA
- Established
- 2009
- Scale / Notable Feature
- Largest by annual volume; voluntary participation, growing coverage
Registries demonstrate practice variation in real time: the AOANJRR famously identified poorly performing metal-on-metal hip resurfacing and large-head designs years before they were withdrawn, illustrating how high-completeness observational data can detect rare device failures that no individual RCT is powered to find. Registry effectiveness data (real-world, all-comers) complements RCT efficacy data (selected populations, ideal conditions).
For the exam you must be able to:
- Critically appraise a published study against the appropriate reporting standard (CONSORT/STROBE/PRISMA)
- Match the appropriate design to a clinical question (therapy, prognosis, harm, diagnosis)
- Explain why GRADE can downgrade an RCT or upgrade observational data
- Interpret national registry survival data (Kaplan-Meier, hazard ratios, revision endpoints) including AOANJRR and NJR
- Distinguish statistical significance from clinical significance (MCID)
Distinguishing Look-Alike Designs
A frequent exam trap is mislabelling a study design. Use the table below to separate designs that are commonly confused, based on the direction of enquiry and the measures they permit.
- Prospective Cohort
- Exposure status
- Retrospective Cohort
- Exposure status (past records)
- Case-Control
- Outcome (disease) status
- Cross-Sectional
- Neither - sampled population
- Prospective Cohort
- Exposure → outcome (forward)
- Retrospective Cohort
- Exposure → outcome (forward, in records)
- Case-Control
- Outcome → exposure (backward)
- Cross-Sectional
- Simultaneous snapshot
- Prospective Cohort
- Yes
- Retrospective Cohort
- Yes
- Case-Control
- Often unclear
- Cross-Sectional
- No
- Prospective Cohort
- Relative risk, incidence
- Retrospective Cohort
- Relative risk, incidence
- Case-Control
- Odds ratio
- Cross-Sectional
- Prevalence, prevalence OR
- Prospective Cohort
- Rare exposures, prognosis
- Retrospective Cohort
- Rare exposures with existing data
- Case-Control
- Rare outcomes
- Cross-Sectional
- Prevalence / hypothesis generation
- Prospective Cohort
- Loss to follow-up, confounding
- Retrospective Cohort
- Data quality, missing data
- Case-Control
- Recall and selection bias
- Cross-Sectional
- Survivor bias, temporal ambiguity
MCQ Practice Points
Q: A researcher wants to study the association between high BMI and knee osteoarthritis. She measures BMI and presence of knee OA in 500 patients at a single clinic visit. What type of study is this? A: Cross-sectional study. Exposure (BMI) and outcome (OA) are measured at the same point in time. This design can measure prevalence but cannot establish causality or temporal relationship.
Q: What is the main advantage of randomization in an RCT? A: Balances both known and unknown confounders between groups. Randomization creates groups that are comparable at baseline, eliminating selection bias and confounding, allowing isolation of treatment effect.
Q: When is a case-control study the preferred design? A: For rare diseases or outcomes. Case-control studies are efficient because you start with cases (already have the rare disease) and look backward for exposures. Much faster than waiting for rare outcome to occur in a cohort.
Exam Viva Scenarios
Practise clinical reasoning and management decisions out loud
“You want to study whether smoking increases the risk of nonunion after tibial fracture. What study design would you choose and why?”
“You are reviewing an RCT comparing operative vs non-operative treatment for displaced ankle fractures. What key features would you look for to assess the quality of this trial?”
Study Design Hierarchy
- Level I = RCT, Systematic Review of RCTs
- Level II = Prospective Cohort, Lesser RCTs
- Level III = Case-Control, Retrospective Cohort
- Level IV = Case Series, no control group
- Level V = Expert Opinion, lowest evidence
Key Design Features
- RCT = Randomization + Prospective + Control group
- Cohort = Exposure → Outcome (forward in time)
- Case-Control = Outcome → Exposure (backward in time)
- Cross-sectional = Snapshot (exposure and outcome at same time)
- Case Series = Descriptive only, no comparison
Design Selection
- Therapeutic question + Ethical + Feasible = RCT
- Rare exposure = Cohort study
- Rare outcome = Case-control study
- Prevalence question = Cross-sectional survey
- Harmful exposure = Observational (cohort), NOT RCT
RCT Critical Features
- Randomization eliminates selection bias
- Allocation concealment prevents manipulation
- Blinding prevents performance and detection bias
- Intention-to-treat preserves randomization
- CONSORT = reporting guidelines for RCTs
Common Pitfalls
- Cross-sectional cannot establish causality (temporal relationship unclear)
- Case-control cannot calculate relative risk (only OR)
- Cohort studies prone to loss to follow-up
- Case series have selection bias and no comparison
- Confounding common in all observational designs
The Evidence Hierarchy
For therapeutic questions the designs form a hierarchy of five levels. Randomised controlled trials sit at the apex because randomisation eliminates selection bias and balances both known and unknown confounders.
- Level I - systematic reviews and meta-analyses of RCTs, or individual high-quality RCTs. The strongest evidence for causation, and the gold standard for therapeutic questions.
- Level II - prospective cohort studies and lesser-quality RCTs. A cohort cannot prove causation, only association, and is prone to confounding and selection bias; it is the appropriate design when an RCT is not ethical or feasible.
- Level III - case-control studies and retrospective cohort studies. High risk of recall bias and selection bias; best for rare diseases or outcomes.
- Level IV - case series and cross-sectional studies. The case series has no comparison group, neither can establish a temporal relationship, and both are useful for describing disease characteristics.
- Level V - expert opinion and case reports. The lowest level, subject to individual bias and experience, and may generate hypotheses for future research.
What the hierarchy does not tell you. Learn the hierarchy, then learn its most important qualification, because examiners reward the second far more than the first. The premise that observational studies systematically exaggerate treatment effects has been tested, and it did not hold. Concato and colleagues took five clinical topics where meta-analyses of randomised trials and of observational studies had addressed the same question, compared 99 reports, and found the average observational estimates remarkably similar to the randomised ones.
For BCG and active tuberculosis, 13 RCTs gave a relative risk of 0.49 (95% CI 0.34-0.70) while 10 case-control studies gave an odds ratio of 0.50 (95% CI 0.39-0.65), agreement to the second decimal place.
The finding that most unsettles the pyramid. The point estimates were more widely scattered across the randomised trials (0.20 to 1.56) than across the observational studies (0.17 to 0.84). Randomisation removes confounding; it does not remove the imprecision of a small trial, and a hierarchy built on design alone cannot see that difference.
How to use this correctly. It is no argument that design does not matter, and no licence to prefer a database study to a trial. It is the reason modern frameworks such as GRADE rate the quality of a body of evidence rather than the label on a study: a large, consistent, precise set of observational studies can be upgraded, and a small, inconsistent, imprecise set of randomised trials can be downgraded. The level of evidence tells you where to start your appraisal, never where to finish it.
Evidence Base
CONSORT 2010 Statement for Reporting Randomised Trials
- CONSORT 2010 provides a 25-item checklist for transparent reporting of parallel-group RCTs
- Mandates a flow diagram documenting participant flow through enrolment, allocation, follow-up and analysis
- Updated from the 2001 version to incorporate new methodological evidence on bias
- Published simultaneously across BMJ, Lancet, Annals of Internal Medicine and other major journals to maximise dissemination
STROBE Statement for Reporting Observational Studies
- STROBE provides a 22-item checklist covering cohort, case-control and cross-sectional designs
- Eighteen items are common to all three designs; four are design-specific
- Developed at a 2004 methodologists' workshop with iterative consensus revision
- Accompanied by a separate Explanation and Elaboration document with worked examples
RCTs, Observational Studies and the Hierarchy of Research Designs
- Compared meta-analyses of RCTs against observational studies addressing the same five clinical topics (99 reports)
- Average effect estimates from well-designed observational studies were remarkably similar to those of RCTs
- Example: BCG vaccine RR 0.49 (95% CI 0.34-0.70) from 13 RCTs versus OR 0.50 (95% CI 0.39-0.65) from 10 case-control studies
- The spread of point estimates was actually wider across RCTs (0.20 to 1.56) than across the observational studies (0.17 to 0.84) - the trials disagreed with each other MORE than the cohort and case-control studies did
- Five clinical topics were compared, drawn from meta-analyses published in five major journals between 1991 and 1995
PRISMA 2020 Statement for Reporting Systematic Reviews
- PRISMA 2020 replaces the 2009 statement with a 27-item checklist plus an abstract checklist
- Revised flow diagrams document study identification, screening, eligibility and inclusion
- Updated to reflect advances in search, selection, appraisal and synthesis methods
- Includes expanded item-level reporting guidance to aid implementation
GRADE: Rating Quality of Evidence and Strength of Recommendations
- GRADE rates evidence as high, moderate, low or very low quality, separately from strength of recommendation
- RCTs start as high-quality but can be downgraded for risk of bias, inconsistency, indirectness, imprecision or publication bias
- Observational studies start as low-quality but can be upgraded for large effect, dose-response or plausible residual confounding
- Adopted by WHO, NICE, Cochrane and numerous guideline developers worldwide
User's Guide to the Orthopaedic Literature: Article About a Surgical Therapy
- Frames critical appraisal of a surgical therapy study around validity, results and applicability
- Validity hinges on randomisation, allocation concealment, blinding and intention-to-treat analysis
- Stresses complete follow-up and analysis of patients in their assigned groups
- Translates generic evidence-based-medicine appraisal into surgical decision-making