Statistical Significance | Effect Estimation | Clinical Interpretation
- P-Value: Probability of observing data as extreme as yours IF null hypothesis is true. NOT probability that null is true.
- Confidence Interval (95% CI): Range of plausible values for true effect. If repeated many times, 95% of CIs would contain true value.
- Statistical Significance (p less than 0.05): Does NOT equal clinical importance. Must compare effect to MCID.
- CI Interpretation: If 95% CI excludes null (0 for difference, 1 for ratio), result is statistically significant at p less than 0.05.
- CI Width: Narrow CI = precise estimate. Wide CI = imprecise, underpowered study.
- “p = 0.05 is arbitrary threshold - not magic cutoff between real and unreal
- “p-value depends on sample size - large studies find significance in trivial differences
- “CI provides effect size AND significance - more informative than p-value alone
- “CI that crosses MCID suggests effect may not be clinically meaningful
Overview/Introduction
P-values and confidence intervals (CIs) are the two primary tools for statistical inference in orthopaedic research, and they answer different but complementary questions. The p-value tests whether the observed data are compatible with the null hypothesis of no effect or difference. The confidence interval estimates the range of plausible values for the true effect.

Why it matters. Misinterpretation of p-values is pervasive in the medical literature. Understanding these tools prevents overconfident claims from underpowered studies, and separates statistically significant but clinically trivial findings from truly meaningful ones. In orthopaedics they matter for:
- Distinguishing statistical significance from clinical importance (MCID)
- Interpreting RCT results for treatment decisions
- Evaluating diagnostic test accuracy studies
- Assessing prognostic factor analyses
- Critical appraisal, in the exam viva and in clinical practice
History. Ronald Fisher introduced p-values in the 1920s as a continuous measure of evidence against the null hypothesis. The 0.05 cut-off attributed to him became a historical convention, not a scientific law or constant. Jerzy Neyman and Egon Pearson later developed confidence intervals in the 1930s as a complementary approach to estimation.
Current emphasis. Modern statistical guidelines (ASA 2016, CONSORT, STROBE) emphasise reporting effect sizes and confidence intervals over dichotomous p-value thresholds. Journals increasingly require CIs alongside or instead of p-values.
Principles of Statistical Inference
The null hypothesis (H₀) is a statement of no effect, no difference or no association. Its value depends on the measure being tested:
- Mean difference = 0
- Risk ratio = 1
- Correlation coefficient = 0
The alternative hypothesis (H₁) states that there is an effect, difference or association. It can be two-sided (any difference) or one-sided (a specific direction).
P-values and CIs are mathematically related. When a confidence interval excludes the null value (0 for differences, 1 for ratios), the result is statistically significant at the level that matches the interval:
- 95% CI excludes the null: p less than 0.05
- 99% CI excludes the null: p less than 0.01
- 90% CI excludes the null: p less than 0.10
So statistical significance can be read directly from the confidence interval without needing the p-value. This is why modern guidelines emphasise CIs over p-values.
Understanding P-Values
Definition. The p-value is the probability of observing data as extreme as, or more extreme than, what was observed, assuming the null hypothesis is true. Formally, p = P(Data | H₀ is true). It ranges from 0 to 1, and is often expressed as 0 to 100%.
Saying it correctly. "Assuming there is no true difference between groups, the probability of observing a difference as large as or larger than what we observed, purely by chance, is [p-value]." For a comparison of two surgical techniques with p = 0.03, that becomes: if the two techniques are truly equivalent, there is a 3% probability of observing a difference this large or larger by random chance alone.
- Interpretation
- Very strong evidence against null
- Conclusion
- Highly statistically significant
- Action
- Check effect size and clinical relevance
- Interpretation
- Moderate evidence against null
- Conclusion
- Statistically significant
- Action
- Check confidence interval and MCID
- Interpretation
- Weak evidence, borderline
- Conclusion
- Not statistically significant
- Action
- Consider if underpowered, examine trend
- Interpretation
- Little evidence against null
- Conclusion
- Not statistically significant
- Action
- Check power, may be true null or Type II error
The threshold is arbitrary. p = 0.051 is not fundamentally different from p = 0.049, and the 0.05 line is not a natural boundary. Treating the two as categorically different is statistically indefensible, yet remains common in practice and in journal decisions.
What a p-value is not. Each of these readings is wrong:
- The probability that the null is true. The p-value assumes the null is true and then calculates the probability of the data. It is P(Data | H₀), not P(H₀ | Data), and confusing the two is the most common error.
- The probability of a Type I error. The Type I error rate is alpha, set before the study (usually 0.05). The p-value is calculated from the observed data; alpha is the pre-set threshold.
- The size or clinical importance of the effect. The p-value reflects both effect size and sample size, so a large sample can yield p less than 0.05 for a trivial effect.
- Proof of the null when p is greater than 0.05. Failure to reject the null does not prove it is true. The study may be underpowered (Type II error), or the null may be true.
- A measure of power. The p-value does not indicate whether the study was adequately powered.
For the p = 0.03 example, then, "there is a 3% chance the null hypothesis is true", "there is a 3% chance this result is a false positive" and "the techniques differ by 3%" are all incorrect statements.
Understanding Confidence Intervals
Definition. A 95% confidence interval is a range of values that, if the study were repeated many times, would contain the true population parameter in 95% of those studies. It is built as the point estimate (the observed effect) ± a margin of error. The frequentist interpretation forbids the tempting reading: there is not a 95% probability that the true value lies in this particular interval.
What it tells you. The point estimate (mean or median) is the best guess of the true effect, and the interval is the range of plausible values within which the true effect likely lies. A CI therefore shows effect size, direction, precision and statistical significance at once, whereas a p-value shows only significance, not magnitude.
- Meaning
- Best guess of true effect
- Example (Mean Difference)
- Mean difference = 8 points
- Interpretation
- Observed effect in this sample
- Meaning
- Minimum plausible effect
- Example (Mean Difference)
- 95% CI: 2 to 14 points
- Interpretation
- True effect unlikely below 2
- Meaning
- Maximum plausible effect
- Example (Mean Difference)
- 95% CI: 2 to 14 points
- Interpretation
- True effect unlikely above 14
- Meaning
- Precision of estimate
- Example (Mean Difference)
- Width = 12 points (14 minus 2)
- Interpretation
- Wider = less precise, needs larger sample
Commonly Confused Concepts (Differential)
Examiners frequently probe whether candidates can distinguish closely related statistical terms. The table contrasts the concepts most often confused in vivas and the literature.
- What It Is
- Probability of data this extreme IF null is true
- What It Is Often Confused With
- Probability the null is true
- Key Discriminator
- Conditional direction: P(data | H0), not P(H0 | data)
- What It Is
- Pre-set acceptable Type I error rate
- What It Is Often Confused With
- The observed p-value
- Key Discriminator
- Alpha is fixed before the study; p is computed from the data
- What It Is
- Plausible range for the true effect (frequentist)
- What It Is Often Confused With
- Credible interval (Bayesian probability range)
- Key Discriminator
- Only a credible interval gives a direct probability for the parameter
- What It Is
- Effect unlikely under the null (p less than alpha)
- What It Is Often Confused With
- Clinical significance
- Key Discriminator
- Clinical significance needs the effect to exceed the MCID
- What It Is
- False positive: rejecting a true null
- What It Is Often Confused With
- Type II error (false negative)
- Key Discriminator
- Type I is alpha; Type II is beta (1 minus power)
- What It Is
- Spread of individual observations
- What It Is Often Confused With
- Standard error (spread of the mean estimate)
- Key Discriminator
- SE = SD / sqrt(n); SE shrinks with larger samples, SD does not
Bayesian Inference: Credible Intervals and Bayes Factors
The p-value and confidence interval are frequentist tools: they describe long-run behaviour and cannot give the direct probability statement clinicians actually want ("how likely is it that this treatment works?"). The Bayesian framework answers that question directly, at the cost of having to specify a prior.
- Frequentist
- How compatible are the data with the null?
- Bayesian
- How probable is the hypothesis given the data and prior belief?
- Frequentist
- Confidence interval (a long-run coverage statement)
- Bayesian
- Credible interval - genuinely a 95% probability the parameter lies within it
- Frequentist
- P-value (cannot support the null)
- Bayesian
- Bayes factor - ratio of the data's likelihood under each hypothesis, can favour the null
- Frequentist
- Counter-intuitive; gives no direct probability for the hypothesis
- Bayesian
- Requires choosing a prior, which can be subjective
Bayes' theorem. The posterior is proportional to the likelihood times the prior, so the observed data update a prior belief into a posterior. The credible interval is the Bayesian analogue of the confidence interval, and the Bayes factor quantifies how much the data shift the odds between two hypotheses.
In practice. Frequentist CIs and p-values dominate the orthopaedic literature. Bayesian methods are increasingly used in adaptive trials and evidence synthesis, and adoption is growing but uneven. The dependence on the prior remains the central criticism.
Clinical Application
Two kinds of significance. Statistical significance (p less than 0.05, CI excludes the null) and clinical significance (the effect exceeds the MCID) are separate judgements, and a result can have one without the other. Always check whether the point estimate and the CI bounds exceed the MCID, not just whether p is less than 0.05. The four possible scenarios:
- Statistical Significance
- p less than 0.05, CI excludes 0
- Clinical Significance
- Effect exceeds MCID
- Interpretation
- Significant AND clinically meaningful - implement
- Statistical Significance
- p less than 0.05, CI excludes 0
- Clinical Significance
- Effect below MCID
- Interpretation
- Significant but trivial - do NOT implement
- Statistical Significance
- p greater than 0.05, CI includes 0
- Clinical Significance
- Point estimate exceeds MCID
- Interpretation
- Not significant but trend - need larger study
- Statistical Significance
- p greater than 0.05, CI includes 0
- Clinical Significance
- Effect well below MCID
- Interpretation
- No effect - do not implement
Worked Example: THA Study
The study. Cemented versus uncemented THA, compared on WOMAC score at 1 year. The results:
- Mean difference = 8 points (cemented better)
- 95% CI: 1 to 15 points
- p = 0.02
- MCID for WOMAC = 10 points
Reading it. Take each number in turn against the MCID:
- Statistically significant. p = 0.02 is less than 0.05 and the CI excludes 0.
- Point estimate. 8 points is less than the MCID of 10, so not clinically meaningful.
- CI upper bound. 15 points is greater than the MCID, so the effect could be meaningful.
- CI lower bound. 1 point is much less than the MCID, so the effect could be trivial.
The verdict. The result is statistically significant but clinically uncertain. The CI is wide and crosses the MCID threshold, and the point estimate suggests the effect may not be clinically important. A larger study is needed to narrow the CI and determine whether the true effect exceeds 10 points.
Absolute and Relative Effect Measures (NNT)
For binary (yes/no) outcomes the same result can be expressed in absolute or relative terms, and the choice strongly affects how impressive it looks. Confidence intervals and significance apply to all of these measures, but clinical decisions depend on the absolute ones.
- Definition
- Control event rate minus experimental event rate
- Note
- The true clinical benefit; null value is 0
- Definition
- Experimental event rate divided by control event rate
- Note
- Null value is 1; common in cohort studies and RCTs
- Definition
- ARR divided by the control event rate (equals 1 minus RR)
- Note
- Looks large even when the ARR is small
- Definition
- 1 divided by the ARR (rounded up)
- Note
- Patients treated to prevent one extra bad outcome; lower is better
- Definition
- Odds in the exposed divided by odds in the unexposed
- Note
- Case-control and logistic regression; approximates RR only when the outcome is rare
Relative versus absolute. A treatment that cuts risk from 2% to 1% has an impressive relative risk reduction of 50%, but an ARR of only 1% and an NNT of 100. The RRR is identical whether the baseline risk is 2% or 50%, so a relative measure detached from baseline risk can badly mislead: the same RRR gives a very different NNT depending on baseline risk.

Harm and precision. The number needed to harm (NNH) is the analogous 1 divided by the absolute risk increase for adverse effects. The NNT should always be quoted with its confidence interval.
A 50% relative risk reduction means little if the baseline risk is tiny. Quote ARR and NNT (with confidence intervals), not RRR alone, when counselling patients.
Guidelines, Registries & Global Practice
Reporting Standards Across the Global Literature
Statistical reporting expectations are remarkably consistent worldwide because they are driven by international journal and methodology guidelines rather than national bodies. The candidate should know which guideline governs which study type.
- Scope
- Global statistical practice
- Position on P-Values and CIs
- Do not base conclusions on a single p threshold; report effect sizes and CIs
- Scope
- All biomedical manuscripts
- Position on P-Values and CIs
- Quantify findings with appropriate indicators of measurement error such as confidence intervals; avoid relying solely on hypothesis testing
- Scope
- RCT reporting
- Position on P-Values and CIs
- Report effect size and its precision (e.g. 95% CI) for primary and secondary outcomes
- Scope
- Cohort, case-control, cross-sectional
- Position on P-Values and CIs
- Give estimates with confidence intervals; report so as not to over-interpret p-values
- Scope
- Guidelines and meta-analyses
- Position on P-Values and CIs
- Judges imprecision largely by CI width relative to clinical decision thresholds (MCID)
Registries and Effect Estimation
Large national arthroplasty registries (NJR for England, Wales, Northern Ireland and the Isle of Man; AOANJRR Australia; SHAR Sweden; the Norwegian and New Zealand registries; AJRR USA) illustrate the practical primacy of confidence intervals over p-values. Because registry sample sizes run into hundreds of thousands, almost any difference reaches statistical significance. Registries therefore report hazard ratios and revision rates with narrow confidence intervals and interpret them against clinically meaningful thresholds, not against p = 0.05. This is the real-world embodiment of the large-sample problem: with very large n, statistical significance is near-guaranteed and effect size plus CI width become the only meaningful discriminators.
High- vs Limited-Resource Practice Variation
- Well-resourced settings: Routine access to statistical software and biostatistician support; journals enforce CI reporting and pre-registration, reducing selective p-value reporting.
- Limited-resource settings: Smaller single-centre studies predominate, raising the risk of underpowered analyses and Type II errors; wide confidence intervals are common and should be interpreted as inconclusive rather than negative.
- Universal principle: The interpretation of p-values and CIs does not change by country; what varies is study size, access to methodological support, and exposure to publication and reporting bias.
Whichever fellowship you sit, the expected answer is identical: interpret the confidence interval against the MCID, never equate statistical with clinical significance, and treat the p-value as one continuous piece of evidence, not a verdict.
Controversies and Areas of Uncertainty
Statistical inference is an area of genuine, ongoing debate. Candidates who can articulate the controversy, as well as recite the rules, demonstrate consultant-level understanding.
Redefine or abandon significance. Some authors propose lowering the threshold for new discoveries to p less than 0.005 to reduce false positives. Others argue for abandoning fixed thresholds entirely and reporting p-values as continuous evidence. No global consensus exists.
Which MCID. There is no universally agreed MCID for many orthopaedic PROMs. Anchor-based and distribution-based methods can give different values for the same instrument, which complicates the judgement of clinical importance.
Concern over irreproducible findings across biomedical science has been driven substantially by the misuse of p-values: p-hacking (analysing data many ways until p less than 0.05), selective reporting, and underpowered studies. The exam-relevant lesson is that a single significant p-value, especially from a small or post-hoc analysis, is weak evidence until replicated and judged against effect size and prior plausibility.
MCQ Practice Points
Q: What does a p-value of 0.04 mean? A: Assuming null hypothesis is true, there is 4% probability of observing data this extreme or more extreme by chance alone. It does NOT mean 4% probability null is true, nor 4% probability of Type I error, nor 4% effect size.
Q: A 95% CI for mean difference is -2 to 8 points. Is this statistically significant at alpha = 0.05? A: No - the CI includes 0 (no difference), meaning the result is NOT statistically significant. If CI excluded 0, p would be less than 0.05.
Q: Can a result be statistically significant but not clinically significant? A: Yes - large studies can detect tiny differences with p less than 0.05 that are below the MCID threshold. Statistical significance depends on sample size; clinical significance depends on whether effect exceeds MCID.
Q: What does a wide confidence interval indicate? A: Imprecise estimate due to small sample size or high variability. A wide CI crossing both clinically important and trivial effects means the study is inconclusive - you cannot determine if the true effect is meaningful or not. This indicates the study is underpowered and needs a larger sample.
Q: If p = 0.03, what is the probability this result is a false positive (Type I error)? A: Unknown - cannot be determined from p-value alone. Alpha (0.05) is the Type I error rate set BEFORE the study. The p-value (0.03) is calculated FROM the data. Many students confuse these - p-value is NOT the probability of Type I error for THIS specific result.
Q: Study shows no significant difference (p = 0.15) between two treatments. Can you conclude the treatments are equally effective? A: No - failure to reject null does NOT prove null is true. This could be: (1) True null (treatments truly equivalent), OR (2) Type II error (underpowered study missing a real difference). Check the power calculation - if power is below 80%, cannot trust negative result. To prove equivalence, need a specifically designed equivalence or non-inferiority trial.
Exam Viva Scenarios
Practise clinical reasoning and management decisions out loud
“A colleague shows you an RCT comparing two rehab protocols. The study found no significant difference (p = 0.08). She concludes the protocols are equivalent. How do you respond?”
“An RCT of 1000 patients found statistically significant improvement in WOMAC score with new treatment: mean difference = 3 points, 95% CI 1 to 5 points, p = 0.003. The MCID for WOMAC is 10 points. How do you interpret this?”
“A trial of a new fixation device reported no overall difference, but the authors highlight that in the subgroup of smokers aged over 65 the device was significantly better (p = 0.04). The company asks you to adopt the device for this subgroup. How do you respond?”
P-Value Interpretation
- p-value = P(Data | Null is true), NOT P(Null is true | Data)
- p less than 0.05 = statistically significant (arbitrary convention)
- p-value does NOT indicate effect size or clinical importance
- Large sample can yield p less than 0.05 for trivial effects
- p greater than 0.05 does NOT prove null hypothesis (may be underpowered)
Confidence Interval Interpretation
- 95% CI = range of plausible values for true effect
- If 95% CI excludes null (0 or 1), p less than 0.05
- Narrow CI = precise estimate; Wide CI = imprecise, underpowered
- CI provides effect size, precision, AND significance
- Check if entire CI exceeds MCID for clinical relevance
Statistical vs Clinical Significance
- Statistical significance = p less than 0.05, CI excludes null
- Clinical significance = effect exceeds MCID
- Can have statistical significance without clinical importance (large sample, trivial effect)
- Can have clinical importance without statistical significance (small sample, large effect)
- Always compare point estimate AND CI to MCID
Common Misconceptions
- p-value is NOT probability null is true
- p-value is NOT Type I error for this study (that is alpha)
- p greater than 0.05 does NOT prove equivalence (may be Type II error)
- 0.05 threshold is arbitrary, not magic cutoff
- CI contains more information than p-value alone
Clinical Application
- Report effect sizes and CIs, not just p-values
- Check if CI crosses MCID threshold for clinical uncertainty
- Wide CI suggests need for larger study
- Borderline p (0.05-0.10) may indicate trend, check power
- Non-inferiority trials prove equivalence; superiority trials do not
Evidence Base
ASA Statement on Statistical Significance and P-Values
- P-values can indicate how incompatible data are with a specified statistical model, but do NOT measure the probability that the studied hypothesis is true
- Scientific conclusions should NOT be based only on whether a p-value passes a specific threshold such as 0.05
- A p-value does NOT measure the size or importance of an effect, nor provide a good measure of evidence on its own
- Proper inference requires full reporting and transparency, plus effect sizes and confidence intervals
Confidence Intervals Rather Than P Values: Estimation Rather Than Hypothesis Testing
- Overemphasis on hypothesis testing and dichotomising results as significant or non-significant detracts from more useful estimation approaches
- Investigators are usually interested in the SIZE of the difference between groups, not merely whether it is statistically significant
- Confidence intervals present a plausible range for the population value and convey magnitude, direction, and precision
- CIs should be reported for major findings in both the main text and the abstract
A Dirty Dozen: Twelve P-Value Misconceptions
- Reviews twelve common p-value misconceptions and explains why each is wrong, e.g. treating p as the probability the null is true
- The p-value is a measure of evidence that is not part of any formal system of statistical inference, making its meaning easily misconstrued
- Contrasts the p-value with the Bayes factor, which has interpretability properties the p-value lacks
- The most serious error is believing the probability of a wrong conclusion can be calculated from a single experiment without external evidence
Statistical Tests, P Values, Confidence Intervals, and Power: A Guide to Misinterpretations
- Provides an explanatory list of 25 distinct misinterpretations of p-values, confidence intervals, and statistical power
- Selective analysis (choosing what to present based on the p-value obtained) can produce small p-values even when the test hypothesis is correct
- There is no interpretation of these concepts that is simultaneously simple, intuitive, correct, and foolproof
- Concludes with practical guidelines for improving statistical interpretation and reporting