Test Selection | Parametric vs Non-Parametric | Interpretation
- T-test: Compares means between 2 groups. Assumes normality, equal variance, independence.
- ANOVA: Compares means across 3 or more groups. Post-hoc tests needed to identify which groups differ.
- Chi-square: Tests association between categorical variables. Expected count should be over 5 in each cell.
- Regression: Models relationship between outcome and predictor(s). Linear for continuous outcomes, logistic for binary.
- Parametric vs Non-Parametric: Parametric assumes normal distribution (t-test, ANOVA). Non-parametric does not (Mann-Whitney, Kruskal-Wallis).
- “Use paired t-test for before-after comparisons, independent t-test for separate groups
- “ANOVA tells you IF groups differ, not WHICH groups - need post-hoc tests (Tukey, Bonferroni)
- “Fisher exact test preferred over chi-square when expected counts under 5
- “Correlation does NOT imply causation - confounders may explain association
Overview
Why it matters. Statistical tests are the foundation of evidence-based orthopaedics. Knowing when each test applies and what its output means is what lets you appraise the literature critically and conduct research of your own. This topic covers the tests that appear most often in orthopaedic research.
Where they are used.
- Evaluating treatment outcomes, such as surgical against conservative management
- Assessing prognostic factors: identifying the predictors of complications
- Quality improvement, analysing registry data for benchmarking
- Research design, choosing the tests written into a study protocol
Foundations
The hypothesis. Every test begins with a null hypothesis (H0) that there is no difference or no effect between groups, and an alternative hypothesis (H1) that there is one. The test then asks how the data sit against the null, and two mistakes are possible:
- Type I error (α): rejecting H0 when it is true, a false positive; conventionally set at 0.05
- Type II error (β): failing to reject H0 when it is false, a false negative
- Power (1 − β): the probability of detecting a true effect; aim for over 80%
The test statistic. Each test condenses the data into a single calculated value: a t-statistic for t-tests, an F-statistic for ANOVA, a chi-square statistic (χ²), or a Z-score with large samples. The larger its absolute value, the more extreme the result; it is compared against a critical value or used to calculate the p-value.
Degrees of freedom. The number of independent values behind the statistic sets the critical value threshold, and more degrees of freedom mean narrower confidence intervals.
- Independent t-test: df = n₁ + n₂ − 2
- Paired t-test: df = n − 1
- Chi-square: df = (rows − 1) × (columns − 1)
- ANOVA: df between groups and df within groups
The p-value. The probability of obtaining the observed result if the null hypothesis is true. By convention p under 0.05 is "significant", under 0.01 "highly significant" and under 0.001 "very highly significant". It is not the probability that the null is true, and it is not the probability that the result is due to chance.
The confidence interval. A 95% CI is a range within which the true population value likely lies; the correct reading is that 95% of similarly constructed intervals would contain the true value, not that there is a 95% chance the true value sits in this particular one. If the 95% CI excludes the null value the result is significant at p under 0.05.
Why the interval beats the p-value. Its width shows the precision of the estimate, so a CI shows magnitude and precision together and aids clinical interpretation. Width is determined by sample size (larger is narrower), the variability in the data and the confidence level chosen.
The central limit theorem. As sample size increases, the sampling distribution of the mean approaches a normal distribution regardless of the population distribution. This is why parametric tests work with large samples even when the data are skewed.
Parametric or non-parametric. Parametric tests assume an underlying distribution, usually normal, together with equal variance across groups and independent observations; when those assumptions are met they have the higher power. Non-parametric tests make no distribution assumptions: they are less powerful but more robust, and they are the choice for ordinal data, small sample sizes, skewed distributions and data with outliers.
- Parametric: independent samples t-test, paired t-test, one-way ANOVA, two-way ANOVA, Pearson correlation, linear regression. Use for continuous data with a normal distribution (or a large n) and equal variance across groups
- Non-parametric: Mann-Whitney U (rank-sum), Wilcoxon signed-rank, Kruskal-Wallis (H test), Friedman test, Spearman correlation
Checking the assumptions. Normality is judged first by eye, a bell-shaped histogram and Q-Q plot points that follow the diagonal, and then by test: Shapiro-Wilk for n under 50, where it has the best power, and Kolmogorov-Smirnov for n over 50. With a large n minor deviations can be "significant", so visual inspection is often more informative than the tests. Equal variance (homoscedasticity) is checked with the Levene test; independence comes from the study design, observations that are not clustered or repeated; and sample size adequacy is checked at the same time.
Handling violations. If normality fails, use the non-parametric alternative, transform the data (log, square root) or bootstrap the confidence intervals. If equal variance fails, use the Welch t-test, which does not assume it, report robust standard errors, or fall back to a non-parametric test.
Diagnostic Test Statistics
Sensitivity and specificity. Sensitivity is the true positive rate, the proportion of diseased patients the test correctly identifies, TP / (TP + FN); a highly sensitive test has few false negatives, and at a sensitivity of 95% 5% of cases will be missed. Specificity is the true negative rate, the proportion of non-diseased correctly identified, TN / (TN + FP); a highly specific test has few false positives, and at 90% specificity 10% of results will be false alarms.
Predictive values. These answer the patient's question. The positive predictive value, TP / (TP + FP), is the probability of disease if the test is positive: "my patient tested positive, how likely are they to actually have it?" The negative predictive value, TN / (TN + FN), is the probability of no disease if the test is negative: "how confident am I they are disease-free?" Both depend on prevalence, PPV rising with higher prevalence and NPV rising with lower prevalence.
- Disease Present
- True Positive (TP)
- Disease Absent
- False Positive (FP)
- Column 4
- PPV = TP/(TP+FP)
- Disease Present
- False Negative (FN)
- Disease Absent
- True Negative (TN)
- Column 4
- NPV = TN/(TN+FN)
- Disease Present
- Sens = TP/(TP+FN)
- Disease Absent
- Spec = TN/(TN+FP)
- Column 4
Likelihood ratios. The positive likelihood ratio is sensitivity / (1 − specificity) and the negative likelihood ratio is (1 − sensitivity) / specificity.
- LR+ greater than 10: strong evidence for disease; 5-10 moderate; 2-5 weak
- LR− less than 0.1: strong evidence against disease; 0.1-0.2 moderate
ROC curves and AUC. A receiver operating characteristic curve plots sensitivity against (1 − specificity) at every threshold, and the area under it runs from 0.5 to 1.0. It is used to compare diagnostic tests and to choose the optimal cut-off.
- AUC 0.9-1.0: excellent discrimination
- AUC 0.8-0.9: good
- AUC 0.7-0.8: fair
- AUC less than 0.7: poor
Pre-test and post-test probability. The Bayesian approach to diagnosis starts from a pre-test probability, the prior likelihood from clinical suspicion, and updates it with the test result: post-test odds = pre-test odds × likelihood ratio. Clinical judgement plus the test result makes the better decision. With a pre-test probability of 50% and an LR+ of 10, pre-test odds of 1:1 become post-test odds of 10:1, a post-test probability of 91%.
- SnNOut: Highly Sensitive test, Negative result rules OUT disease (low FN rate means negatives are reliable)
- SpPIn: Highly Specific test, Positive result rules IN disease (low FP rate means positives are reliable)
Use a sensitive test for screening (don't miss cases) and a specific test for confirmation (don't label incorrectly).
Performing the Analysis
The workflow. Following a systematic sequence is what makes an analysis rigorous and reproducible.
- Formulate null and alternative hypotheses
- Identify outcome variable(s) and predictor/exposure variable(s)
- Determine the comparison type (difference, association, prediction)
- Check data type (continuous, categorical, ordinal)
- Assess distribution (histogram, Q-Q plot)
- Identify outliers, missing data and data entry errors
- Use the DINGO questions to match the test to data type and design
- Choose parametric or non-parametric
- Consider confounders (multivariable analysis)
- Normality (Shapiro-Wilk, Q-Q plot), equal variance (Levene test)
- Independence of observations
- Sample size adequacy
- Run the analysis in software (SPSS, R, Stata)
- Report test statistic, df, p-value, effect size and confidence interval
- Present results clearly (tables, figures)
- Cost
- Expensive
- Learning Curve
- Easy
- Best For
- Beginners, basic analyses
- Cost
- Free
- Learning Curve
- Steep
- Best For
- Advanced users, custom analyses
- Cost
- Moderate
- Learning Curve
- Moderate
- Best For
- Epidemiology, panel data
- Cost
- Common
- Learning Curve
- Easy
- Best For
- Simple calculations only
- Cost
- Expensive
- Learning Curve
- Steep
- Best For
- Clinical trials, pharma
Power and sample size. A power calculation has four components, the expected effect size, the alpha level (usually 0.05), the power (usually 0.80, and higher for important outcomes) and the sample size, and knowing any three gives the fourth. Done a priori it gives the sample size needed; done post hoc it gives the power achieved. An underpowered study misses true effects (a Type II error) and, if it does reach significance, overestimates the effect size, which is why power belongs in the planning. Rule-of-thumb sizes for a medium effect:
- t-test (d = 0.5): n = 64 per group
- Chi-square 2×2 (w = 0.3): n = 88 total
- Correlation (r = 0.3): n = 84 total
Worked Orthopaedic Examples
Two treatment groups. Cemented against uncemented THA with the Harris Hip Score (continuous, 0-100) as the outcome: an independent samples t-test, or Mann-Whitney U if the scores are not normal. Reported as "Mean HHS was 85.2 (SD 12.1) in cemented vs 87.4 (SD 11.8) in uncemented group (t = −1.42, df = 98, p = 0.16)".
Before and after. Knee range of motion before and after TKA in the same patients, ROM in degrees: a paired t-test, or Wilcoxon signed-rank if not normal. Reported as "ROM improved from 92° (SD 18) to 115° (SD 12), mean difference 23° (95% CI: 18-28, p less than 0.001)".
Several groups. VAS pain across four fracture classifications: one-way ANOVA with Tukey or Bonferroni post hoc. Reported as "Significant difference in VAS between groups (F = 5.23, df = 3,96, p = 0.002). Post-hoc: Type D higher than Types A, B (p less than 0.05)".
Two categorical variables. Smoking status (smoker/non-smoker) and nonunion (yes/no): a chi-square test, or Fisher exact if any expected count is under 5. Reported as "Nonunion rate was 15% in smokers vs 5% in non-smokers (χ² = 6.8, df = 1, p = 0.009)".
Linear regression. Predictors of WOMAC score after TKA from age, BMI, preoperative pain and comorbidities. Each year of age cost 0.3 WOMAC points (95% CI: −0.5 to −0.1) and each unit of BMI 1.2 points (95% CI: −1.8 to −0.6), with R² = 0.35, the model explaining 35% of the variance; BMI was the strongest predictor of a poorer outcome.
Logistic regression. Risk factors for surgical site infection (yes/no) from diabetes, smoking, operative time and ASA grade. Diabetes OR 2.4 (95% CI: 1.3-4.4, p = 0.005), smoking OR 1.8 (95% CI: 1.0-3.2, p = 0.048) and each hour of operative time OR 1.3 (95% CI: 1.1-1.5, p = 0.001); diabetes carried the highest independent risk.
Kaplan-Meier and log-rank. Implant survivorship, implant A against implant B, with time to revision as a censored outcome. Ten-year survivorship was 92% for A and 88% for B, log-rank χ² = 4.2, p = 0.041: a significant difference in the survival curves.
Cox regression. Predictors of THA revision from age, sex, fixation type and diagnosis. Age under 55 HR 1.8 (95% CI: 1.3-2.5), inflammatory arthritis HR 2.1 (95% CI: 1.4-3.2), uncemented fixation HR 0.8 (95% CI: 0.6-1.1, not significant); young age and an inflammatory diagnosis increased the risk of revision.
Registry data. An AOANJRR analysis follows the same pattern: survival analysis with Kaplan-Meier curves, comparison with the log-rank test, Cox regression for adjusted hazard ratios, an account of competing risks (death before revision) and cumulative percent revision reported with its 95% CI. The registry's sample sizes are so large that tiny differences can be "significant", so the reading has to focus on clinical importance.
Errors, Pitfalls and Misinterpretation
Producing a false positive. Type I errors are manufactured by multiple comparisons without correction, by p-hacking (testing outcomes until one gives p under 0.05) and by selective outcome reporting. The defences are to pre-specify the primary outcome, correct multiple tests with Bonferroni or FDR, and register the study protocol before data collection.
Producing a false negative. Type II errors come from an underpowered study, high variability in the data or a small true effect. The defences are an a priori power calculation, power greater than 80% and sensitive outcome measures.
The multiple comparisons problem. Test 20 outcomes at α = 0.05 and you should expect one false positive, and "p-hacking" inflates the false positive rate. The corrections:
- Bonferroni: α' = α/n, conservative
- Holm: a step-down procedure, less conservative
- FDR: controls the false discovery rate, for many tests
The best approach is to pre-specify a single primary outcome.
Confounding and bias. A confounder is a third variable that explains an observed association; selection bias is non-random sampling that shapes the result. The solutions are randomisation (an RCT), multivariable adjustment by regression, propensity score matching and stratification. Association is not causation.
Ignoring the assumptions. Running a parametric test on data that violate its assumptions gives biased p-values, invalid confidence intervals and unreliable conclusions. Two further habits do quiet damage: ignoring clustering (several joints per patient) and dichotomising a continuous variable, which loses information.
- What People Think
- 4% chance null is true
- What It Actually Means
- 4% chance of result if null true
- What People Think
- 95% chance true value in range
- What It Actually Means
- 95% of such intervals contain true value
- What People Think
- Effect is large/important
- What It Actually Means
- Effect unlikely due to chance alone
- What People Think
- No effect exists
- What It Actually Means
- Insufficient evidence to reject null
- What People Think
- First effect is larger
- What It Actually Means
- Only tells us about evidence, not magnitude
When reading an orthopaedic paper, check:
- Was sample size justified with power calculation?
- Was the primary outcome pre-specified, and does it match the study aims?
- Were appropriate tests used for data type?
- Were assumptions checked and reported?
- Are effect sizes and CIs reported (not just p-values)?
- Is clinical significance discussed?
- Are limitations acknowledged?
Red flags: multiple outcomes with only one "significant"; post-hoc subgroup analyses driving the conclusions; incomplete reporting of non-significant results.
Reporting and Publishing Results
What every paper reports. Under the CONSORT and STROBE standards, a result is an exact p-value (not "p less than 0.05"), a 95% confidence interval for the estimate, an effect size (Cohen's d, OR, HR), the statistical software and version, a description of how missing data were handled, and a primary outcome and analysis plan that were pre-specified.
CONSORT, for randomised trials.
- Participant flow diagram, with losses to follow-up
- Sample size calculation stated
- All outcomes reported, primary and secondary
- Confidence intervals for the main results
- Intention-to-treat analysis
- Baseline characteristics table
STROBE, for observational studies.
- Clear statement of the design: case-control, cohort or cross-sectional
- Setting, dates and eligibility described
- Numbers at each stage
- Outcome data with denominators
- Confounding addressed, bias assessed
- Sensitivity analyses
- Reporting Format
- Mean (SD) or median (IQR)
- Example
- Pain score: 3.2 (SD 1.4)
- Reporting Format
- n (%) with denominator
- Example
- 23/50 (46%) achieved union
- Reporting Format
- RR or OR with 95% CI
- Example
- RR 0.65 (95% CI 0.48-0.88)
- Reporting Format
- HR with 95% CI, survival curve
- Example
- HR 0.72 (95% CI 0.55-0.94)
- Reporting Format
- Exact value to 2-3 decimal places
- Example
- p = 0.034 (not p less than 0.05)
Viva Point: "What must be reported alongside any p-value?" Answer: The effect size (difference between groups) and 95% confidence interval - p-values alone do not indicate clinical importance or precision of the estimate.
What the journals ask for. JBJS requires sample size justification, a pre-specified primary outcome, multiplicity adjustments for multiple comparisons, an explanation of missing data handling and effect sizes with confidence intervals; its level of evidence is set initially by study design and can be modified by quality assessment. JAAOS and BJJ are similar, asking for a CONSORT or STROBE checklist on submission, recommending a statistical analysis plan, requiring all pre-specified outcomes to be reported and publishing negative results appropriately. The common reasons for rejection are inadequate sample size justification, missing confidence intervals and multiple testing without correction.
- Problem
- Multiplicity inflation, false positives
- Correct Approach
- Pre-specify, limit number, report interaction tests
- Problem
- Selection bias if only per-protocol
- Correct Approach
- Report both; ITT is primary analysis
- Problem
- Can bias results either direction
- Correct Approach
- Report missing data rate, use multiple imputation
- Problem
- Increased Type I error risk
- Correct Approach
- Alpha spending functions, clear stopping rules
- Problem
- Demonstrate robustness
- Correct Approach
- Report how results change with different assumptions
Data sharing. Many journals now require a data availability statement, individual patient data are increasingly requested, and sharing code and analysis scripts is encouraged. The practical obstacles are de-identification, institutional review board considerations and choosing a repository such as Dryad or Figshare.
Pre-registration. Trials are registered on ANZCTR for Australian trials or ClinicalTrials.gov internationally, with the primary outcome and analysis plan pre-specified. Registration prevents selective outcome reporting, addresses publication bias and is required by ICMJE journals.
Tests for Continuous Outcomes
Independent samples t-test. The most common test in orthopaedic research. It compares the means of a continuous outcome between two independent groups, for example WOMAC scores in cemented against uncemented THA, under the null hypothesis that the mean is equal in both; p under 0.05 indicates a significant difference in means. It assumes a normal distribution in each group, independence of observations and equal variance. When the distribution is not normal use the Mann-Whitney U test; when the variances are unequal use the Welch t-test.
t = (x̄₁ - x̄₂) / SE(difference) where SE = √[(s₁²/n₁) + (s₂²/n₂)]
Paired samples t-test. The same subjects measured at two time points, a before-after or crossover design, such as pain scores before and after surgery in the same patients. The null hypothesis is that the mean difference is zero. Because it accounts for within-subject correlation it controls for individual variation and is more powerful than the independent test. The assumptions are a normal distribution of the differences (not of the raw values), independence of the pairs and no order effect in a crossover; if the differences are not normal, use the Wilcoxon signed-rank test.
t = d̄ / (SD_d / √n) where d̄ = mean of the differences
One-way ANOVA. Compares the means of three or more independent groups, for example functional scores across three surgical approaches, under the null hypothesis that all group means are equal. The assumptions are those of the t-test: normal distribution in each group, independence and equal variance; the non-parametric alternative is Kruskal-Wallis. Never run multiple independent t-tests instead: they inflate the family-wise Type I error.
Post-hoc tests. ANOVA is an omnibus test: it tells you IF any groups differ, not WHICH. If it is significant, a post-hoc test with a correction finds the pairs:
- Tukey HSD: all pairwise comparisons, controls the family-wise error
- Bonferroni: conservative, divides alpha by the number of comparisons
- Dunnett: compares every group to the control group only
Repeated measures ANOVA. The same subjects measured at three or more time points, essential for longitudinal orthopaedic outcome studies: functional scores at baseline, 3, 6 and 12 months after surgery. Accounting for within-subject correlation makes it more powerful. Its extra assumption is sphericity, that the variance of the differences between every pair of time points is equal, tested with Mauchly's test; if it is violated apply the Greenhouse-Geisser or Huynh-Feldt correction.
Mixed-effects models. The more flexible modern alternative for longitudinal data. A mixed model handles missing data better, has no sphericity assumption, can include time-varying covariates and accounts for clustering, such as by surgeon or hospital, and is preferred for complex longitudinal designs.
Two-way (factorial) ANOVA. Compares a continuous outcome across groups defined by two categorical factors at once, and tests each factor and whether they interact: functional score by fixation type (cemented against uncemented) and age band (under 65 against 65 and over), asking whether the effect of fixation depends on age. It is three tests in one model:
- Main effect of factor 1 (fixation)
- Main effect of factor 2 (age band)
- Interaction effect (factor 1 by factor 2): does the effect of one factor change across the levels of the other?
Reading the interaction. The interaction term is the whole reason to run a two-way ANOVA rather than two separate one-way ANOVAs. A significant interaction means the main effects cannot be interpreted in isolation, since the effect of fixation differs by age group; plot the cell means to see it, parallel lines meaning no interaction and diverging or crossing lines meaning there is one. The assumptions are those of one-way ANOVA (normality of residuals, equal variance, independence) applied within each cell, and reasonably balanced cell sizes help. There is no simple non-parametric two-way ANOVA; options include the aligned rank transform (rank-based) or a generalised or mixed model.
Tests for Categorical Outcomes
Chi-square test. Compares proportions between two or more independent groups on a categorical outcome, for example complication rates (yes/no) across three surgical techniques, or infection rates in smokers against non-smokers. The null hypothesis is that there is no association between the variables, so the proportions are equal across groups. It is fast and widely available, but the approximation needs an expected count greater than 5 in every cell of the contingency table and is inaccurate with a small sample or low expected counts.
χ² = Σ [(O - E)² / E] O = observed frequency, E = expected frequency
E = (row total × column total) / grand total
Fisher exact test. The alternative when the expected-count rule is broken. It gives an exact p-value with no assumption about expected counts and works at any sample size, at the cost of being computationally intensive for large tables. Choosing correctly between the two is what keeps the p-value honest.
McNemar test. Chi-square and Fisher are for independent groups, so paired categorical data need a different test. McNemar is for a binary outcome measured on the same subjects twice, before and after, or in matched pairs, where the 2x2 table cross-tabulates the paired responses rather than two independent groups: whether a clinical sign is present before and after an intervention in the same patients, or whether a finding is positive under two paired test conditions.
How it works. McNemar tests only the discordant pairs, the subjects who changed (yes-to-no and no-to-yes), and asks whether the two directions of change are equally likely; concordant pairs carry no information about change and are ignored. A standard chi-square on paired data ignores the within-subject correlation and gives an invalid p-value. McNemar is to chi-square what the paired t-test is to the independent t-test. For paired categorical data with more than two categories the extension is the Stuart-Maxwell test, and for quantifying paired agreement rather than testing change use Cohen's kappa, covered under agreement and reliability below.
Associations and Regression
Correlation. Measures the strength and direction of the relationship between two continuous variables on a scale from −1 to +1: +1 is a perfect positive correlation, 0 none and −1 a perfect negative one. Descriptively, r from 0.0 to 0.3 is weak, 0.3 to 0.7 moderate and 0.7 to 1.0 strong. Correlation does not imply causation; confounders may explain the association.
Pearson (r). The most common measure for a linear relationship, such as age against functional score after THA. It needs both variables continuous, a linear relationship and a bivariate normal distribution.
Spearman (rho). The non-parametric version, with no normality assumption and the same interpretation. Use it for ordinal (ranked) data, a non-normal distribution or a relationship that is monotonic but not linear, for example pain scores on an ordinal 1-10 scale against function scores, and whenever the Pearson assumptions are violated.
Linear regression. Models the relationship between a continuous outcome and one or more predictors and predicts the outcome value from them. Simple linear regression has one predictor, Y = a + b×X, where the slope b is the change in Y for a one-unit increase in X. Multiple linear regression has two or more, Y = a + b₁×X₁ + b₂×X₂ + ..., and each coefficient is adjusted for the other variables, which is how observational studies adjust for confounders: predicting functional score from age, BMI and comorbidities, for instance. The assumptions are a linear relationship, normally distributed residuals, homoscedasticity (constant variance of the residuals) and independence; the outputs are the β coefficients and R².
Logistic regression. The model for a binary outcome (yes/no, success/failure), such as the probability of a complication from age, smoking and diabetes; it is essential for modelling complication risk in orthopaedics. Linear regression is the tempting wrong choice here because it can predict probabilities outside 0-1. It handles multiple predictors, adjusts for confounders and models non-linear relationships with binary outcomes.
Reading the odds ratio. The output is the odds ratio: OR greater than 1 means increased odds of the outcome as the predictor increases, less than 1 decreased odds, and 1 no association. So OR = 2.5 means 2.5 times higher odds for each unit increase in the predictor.
Beyond continuous and binary. The outcome type chooses the multivariable method. Time-to-event data with censoring need survival methods, Kaplan-Meier curves compared with the log-rank test and Cox regression for adjusted hazard ratios; a chi-square on revised against not revised ignores censoring and follow-up time and biases the estimate.
- Method
- Linear regression
- Use Case
- Multiple predictors of continuous outcome
- Key Output
- β coefficients, R²
- Method
- Logistic regression
- Use Case
- Predictors of yes/no outcome
- Key Output
- Odds ratios, AUC
- Method
- Poisson regression
- Use Case
- Predictors of count data
- Key Output
- Rate ratios
- Method
- Cox regression
- Use Case
- Survival analysis with predictors
- Key Output
- Hazard ratios
- Method
- Ordinal logistic
- Use Case
- Ordered categorical outcome
- Key Output
- Cumulative OR
Agreement and Reliability
Choose by the data, not by the comparison. Whether two observers agree on a fracture grade, or one observer agrees with themselves on a second occasion, the test is chosen by the type of data being compared, categorical, ordinal or continuous, and not by whether the comparison is inter- or intra-observer.
- Test
- Cohen's kappa - chance-corrected
- Scale Type
- Categorical
- Interpretation
- Landis & Koch: 0.21-0.40 fair, 0.41-0.60 moderate, 0.61-0.80 substantial, 0.81-1.00 almost perfect
- Test
- WEIGHTED kappa - credits near-misses
- Scale Type
- Ordinal
- Interpretation
- Same bands; unweighted kappa penalises an adjacent-grade error as heavily as a distant one
- Test
- ICC (intraclass correlation)
- Scale Type
- Continuous
- Interpretation
- 0.75-0.90 good, over 0.90 excellent; state which ICC form and model was used
- Test
- Bland-Altman - NOT a correlation coefficient
- Scale Type
- Continuous
- Interpretation
- Bias (mean difference) and 95% limits of agreement; judge clinically whether that spread matters
- Test
- Fleiss' kappa
- Scale Type
- Categorical
- Interpretation
- Interpreted against the same bands
- Test
- Cronbach's alpha
- Scale Type
- Scale items
- Interpretation
- over 0.7 acceptable; a very high value may just mean redundant items
First: kappa versus ICC is decided by the DATA, not by the comparison. Cohen's kappa is used for categorical agreement whether you are comparing two observers (inter-observer) or one observer with themselves on two occasions (intra-observer); ICC is used for continuous measurements in exactly the same two situations. Saying "ICC for intra-observer reliability" of a fracture classification is wrong - a classification is categorical, so it needs kappa however many times you look at it.
Second: correlation is not agreement. Two methods can correlate at r = 0.95 while one reads ten degrees higher than the other every single time. Correlation asks whether they move together; agreement asks whether they give the same answer. Use Bland-Altman.
Fuller treatment, including the kappa paradox, in diagnostic test statistics.
Choosing the Test
The decision. Ask three questions in order: what type the outcome is, how many groups there are, and whether the observations are independent or paired. A fourth, whether the data are normally distributed, picks the parametric or non-parametric row of the table.
- 2 Groups (Independent)
- Independent t-test
- 2 Groups (Paired)
- Paired t-test
- 3+ Groups
- One-way ANOVA
- 2 Groups (Independent)
- Mann-Whitney U
- 2 Groups (Paired)
- Wilcoxon signed-rank
- 3+ Groups
- Kruskal-Wallis
- 2 Groups (Independent)
- Chi-square or Fisher
- 2 Groups (Paired)
- McNemar test
- 3+ Groups
- Chi-square
- 2 Groups (Independent)
- Mann-Whitney U
- 2 Groups (Paired)
- Wilcoxon signed-rank
- 3+ Groups
- Kruskal-Wallis
- 2 Groups (Independent)
- Log-rank test
- 2 Groups (Paired)
- N/A
- 3+ Groups
- Log-rank test
DINGOChoosing the Right Test
Hook:Follow the DINGO trail to find the right statistical test for your data!
Two further questions. Once the table has given a test, ask whether you need to adjust for confounders, which moves the analysis to a multivariable model, and whether there is clustering in the data, several joints per patient or clustering by surgeon or hospital, which calls for a mixed model. Consult a statistician if unsure.
The tempting wrong choice. The tests that are commonly confused, and the discriminator that tells them apart:
- Tempting (Wrong) Choice
- Several independent t-tests
- Correct Choice
- One-way ANOVA then post-hoc
- Discriminator
- Multiple t-tests inflate family-wise Type I error
- Tempting (Wrong) Choice
- Independent t-test
- Correct Choice
- Paired t-test (or Wilcoxon signed-rank)
- Discriminator
- Data are paired, so within-subject correlation must be used
- Tempting (Wrong) Choice
- Pearson chi-square
- Correct Choice
- Fisher exact test
- Discriminator
- Chi-square approximation fails with low expected counts
- Tempting (Wrong) Choice
- Chi-square
- Correct Choice
- McNemar test
- Discriminator
- Chi-square assumes independent observations
- Tempting (Wrong) Choice
- Linear regression
- Correct Choice
- Logistic regression (odds ratios)
- Discriminator
- Linear regression can predict probabilities outside 0-1
- Tempting (Wrong) Choice
- Chi-square on revised vs not
- Correct Choice
- Kaplan-Meier + log-rank / Cox
- Discriminator
- Ignoring censoring and follow-up time biases estimates
- Tempting (Wrong) Choice
- Pearson correlation
- Correct Choice
- Spearman correlation (rho)
- Discriminator
- Pearson assumes linear, bivariate-normal data
Interpreting Research Outcomes
Statistical significance. A p-value below the chosen alpha, usually 0.05, tells you the difference is unlikely to be due to chance alone and nothing about its magnitude or clinical importance; a large sample can make a trivial difference "significant". "Statistically significant" is not "important", and p = 0.04 is not much different from p = 0.06. Report the effect size alongside the p-value.
Clinical significance. A difference large enough to change practice. The threshold is the minimal clinically important difference (MCID), a patient-centred value that varies by outcome measure and is more relevant than the p-value: a WOMAC improvement of 2 points may be significant at p under 0.05 and still not clinically meaningful. Examples in orthopaedics:
- VAS pain: 2 points (or a 30% change)
- WOMAC: 8-12 points; a figure of 15 points is also quoted
- SF-36 Physical: 5 points
- Use
- Continuous outcomes (t-test)
- Interpretation
- 0.2 small, 0.5 medium, 0.8 large
- Formula
- (Mean₁ - Mean₂) / SD pooled
- Use
- ANOVA effect size
- Interpretation
- 0.01 small, 0.06 medium, 0.14 large
- Formula
- SS_between / SS_total
- Use
- Linear relationship
- Interpretation
- 0.1 small, 0.3 medium, 0.5 large
- Formula
- Covariance / (SD_x × SD_y)
Two sets of bands for r. The effect-size bands for r (0.1 small, 0.3 medium, 0.5 large) are not the weak, moderate and strong bands (0.3 and 0.7) used to describe a correlation above.
- Definition
- Risk in exposed / Risk in unexposed: (a/(a+b)) / (c/(c+d)) in a 2×2 table; the effect measure for binary outcomes in a cohort design
- Interpretation
- 1 = no effect, over 1 = increased risk; RR = 2.0 means double the risk
- Definition
- Odds in cases / Odds in controls: (a/b) / (c/d) in a 2×2 table
- Interpretation
- Approximates RR when outcome rare (less than 10%); 1 = no effect
- Definition
- Control rate - Treatment rate
- Interpretation
- Actual percentage point reduction
- Definition
- 1 / ARR
- Interpretation
- Patients to treat to prevent 1 event
- Definition
- Instantaneous risk ratio over time
- Interpretation
- HR = 0.7 means 30% reduction in hazard
Viva Question: "A drug reduces DVT risk from 4% to 2%. What is the NNT?" Answer: ARR = 4% - 2% = 2% = 0.02. NNT = 1/0.02 = 50. You need to treat 50 patients to prevent one DVT.
Using the confidence interval to decide. If the CI excludes the clinically meaningful threshold the result is confident; if it is wide, more data are needed; if it spans both benefit and harm, the study is inconclusive. A mean difference of 5 points with a 95% CI of 2 to 8 against an MCID of 5 includes values below the MCID, so the effect may not be clinically important.
Reading a negative study. A true negative needs an adequate sample, powered for the MCID, and a confidence interval that excludes a clinically important difference; then the treatments can be called equivalent, as with a 95% CI of −3 to +2 points against an MCID of 5. A wide interval spanning the MCID from an underpowered study, say −8 to +12 points against the same MCID, cannot rule out an important difference: absence of evidence is not evidence of absence.
Beyond the mean difference. Forest plots show the effect within subgroups, the route to individual patient prediction. A responder analysis reports the proportion of patients exceeding the MCID, which is more meaningful than a mean difference. The fragility index is the number of events that would have to change to reverse a significant result; the orthopaedic figures are in the closing warning below.
Meta-analysis. The pooled effect is a weighted average of the included studies, and its heterogeneity is measured by I²: over 50% is substantial, over 75% considerable. When studies are heterogeneous the pooled estimate may be misleading, so look for subgroup differences and consider a random effects model.
GRADE. The certainty of the evidence is graded high (further research is unlikely to change confidence in the estimate), moderate (may change the estimate), low (likely to change it) or very low (very uncertain). Evidence is downgraded for risk of bias, inconsistency, indirectness, imprecision and publication bias.
This page teaches which test to run. Its citations answer the harder question: whether the answer means anything.
- The typical significant orthopaedic RCT is fragile. Across 48 randomised trials in sports medicine and arthroscopic surgery, median sample size was 64, median total outcome events 19, and the median Fragility Index was 2 - changing the outcome of just two patients in one arm reversed statistical significance in the typical trial (PMID 27895038). That is the correct first reaction to a small significant p-value in this literature.
- Ten events per variable is the usable rule for a regression. In simulation, at EPV of 10 or more regression coefficients showed no major bias or coverage problems; below 10 the coefficients were biased in both directions, variance estimates became unreliable, and paradoxical wrong-direction associations appeared (PMID 8970487). So count the events, divide by the number of predictors, and be suspicious under 10.
- Adjustment is usually reported too poorly to trust. Of 174 observational intervention studies, over 98% acknowledged confounding, but only 51% reported which observed confounders were included, only 10% how they were selected, and just 9% commented on likely residual confounding (PMID 18693038). Regression adjustment is not a substitute for randomisation, and the reporting rarely lets you judge how close it came.
And a caution about your own reading, from an orthopaedic readership: when the same 40 studies were rated with and without their p-values shown, surgeons rated them as more important when a significant p-value was visible - 10 of 12 reviewers shifted, a mean difference of 0.6 on a 1-3 scale (PMID 16156453). The bias being measured there is the reader's, not the author's.
Where this sits. The tests here are chosen against a study design and read against the levels of evidence; the effect measures themselves are on measures of effect, the sample-size logic on statistical power, the diagnostic-accuracy family on diagnostic test statistics, and the appraisal habit on critical appraisal.
Guidelines, Registries & Global Practice
Global Reporting Standards (Side by Side)
Statistical reporting is governed by internationally harmonised, study-design-specific guidelines endorsed by the EQUATOR Network. These apply regardless of jurisdiction and are referenced by AAOS, NICE, BOA/BOOS, EFORT and most journals.
- Study Design
- Randomised controlled trials
- Key Statistical Mandates
- Pre-specified primary outcome, sample-size calculation, effect size with 95% CI, ITT analysis, flow diagram
- Endorsement
- ICMJE, AAOS/JBJS, BJJ, NICE
- Study Design
- Cohort, case-control, cross-sectional
- Key Statistical Mandates
- Confounder handling, effect estimates with CI, numbers at each stage, sensitivity analyses
- Endorsement
- ICMJE, EFORT, BJJ
- Study Design
- Diagnostic accuracy studies
- Key Statistical Mandates
- Sensitivity/specificity with CI, 2x2 data, ROC/AUC, reference standard
- Endorsement
- Radiology and diagnostic journals
- Study Design
- Systematic reviews and meta-analyses
- Key Statistical Mandates
- Pooled effect with CI, heterogeneity (I-squared), risk-of-bias, GRADE certainty
- Endorsement
- Cochrane, ICMJE
- Study Design
- Prediction/risk models
- Key Statistical Mandates
- Events per variable, calibration, discrimination (C-statistic), internal/external validation
- Endorsement
- Methodology and registry journals
Major Joint Replacement Registries (Global)
- Country/Region
- Australia
- Primary Statistical Output
- Cumulative percent revision (Kaplan-Meier), HR via Cox
- Note
- Near-complete capture, validated revision linkage
- Country/Region
- England, Wales, NI, IoM
- Primary Statistical Output
- Revision rates, PROMs linkage, funnel plots
- Note
- One of the largest registries worldwide
- Country/Region
- Sweden
- Primary Statistical Output
- Implant survivorship, competing-risk analysis
- Note
- Oldest hip registry (from 1979)
- Country/Region
- Norway
- Primary Statistical Output
- Cox regression with adjustment, revision endpoints
- Note
- Strong methodological output
- Country/Region
- United States
- Primary Statistical Output
- Revision burden, cumulative incidence
- Note
- Rapidly growing voluntary registry
Registry follow-up is censored and varies between patients, so a crude revision percentage is misleading. Registries report cumulative percent revision via Kaplan-Meier (or cumulative incidence with competing-risk methods, since death precludes revision) and compare implants with Cox proportional-hazards models adjusted for age, sex and diagnosis. Observational/registry estimates can closely match RCT estimates when confounding is well controlled (Concato et al, NEJM 2000, DOI).
Practice Variation in Statistical Standards
- Significance threshold: p less than 0.05 remains the global convention, but many statisticians (and the American Statistical Association) caution against dichotomising results; some fields advocate reporting exact p-values, confidence intervals and effect sizes instead of a binary "significant/not significant".
- Trial registration: ANZCTR (Australia/NZ), ClinicalTrials.gov (US/international) and ISRCTN/EU-CTR (Europe) all satisfy the ICMJE prospective-registration requirement.
- Fragility of the evidence base: The median Fragility Index of orthopaedic sports-surgery RCTs is only 2 (Khan et al, AJSM 2017, DOI), so reliance on single small trials varies and meta-analysis is often required.
Australian Research Framework
- National Statement on Ethical Conduct in Human Research
- Mandatory ethics approval for human research
- Australian Code for Responsible Conduct of Research
- Human Research Ethics Committee (HREC) approval
- Informed consent documentation
- Data management plans
- Reporting adverse events
- www.anzctr.org.au
- Mandatory for clinical trials
- Required before participant enrollment
- ICMJE requirement for publication
- Primary and secondary outcomes
- Sample size calculation
- Statistical analysis plan
Australian Orthopaedic Data Sources
- Type
- Registry
- Application
- Joint replacement outcomes nationally
- Type
- Quality standards
- Application
- Clinical care standards, indicators
- Type
- Health statistics
- Application
- National injury and disease data
- Type
- Registry
- Application
- Victoria, NSW trauma outcomes
Viva Point: "What is the level of evidence of AOANJRR data?" Answer: Level III (retrospective cohort) - but with very high validity due to near-complete capture (greater than 98%) and validated data linkage for revision endpoints.
Australian trainees should be familiar with national research infrastructure and ethics requirements.
MCQ Practice Points
Q: Why should you NOT run multiple independent t-tests when comparing 3 or more groups? A: Inflates Type I error rate. Each t-test has 5% false positive risk. Three t-tests (Group 1 vs 2, 1 vs 3, 2 vs 3) inflate family-wise error to approximately 14%. ANOVA controls overall Type I error at 5%, then post-hoc tests with correction identify specific differences.
Q: When would you use a paired t-test instead of an independent t-test? A: When comparing same subjects at two time points (e.g., before vs after surgery). Paired t-test accounts for within-subject correlation and is more powerful. Independent t-test is for comparing two separate groups of different subjects.
Q: When should you use Fisher exact test instead of chi-square? A: When expected count is under 5 in any cell of the contingency table. Chi-square approximation is inaccurate with small expected counts. Fisher exact provides exact p-value for any sample size.
Q: How do you interpret the Pearson correlation coefficient r = 0.7? A: Strong positive linear relationship. r = 0.7 means 49% of variance in one variable is explained by the other (r² = 0.49). Clinical significance depends on context. Interpretation: r under 0.3 = weak, 0.3-0.7 = moderate, over 0.7 = strong. Note: correlation does not imply causation.
Q: How do you decide between parametric and non-parametric tests? A: Use non-parametric tests when: (1) data violate normality assumption (check with Shapiro-Wilk test), (2) ordinal data (e.g., Likert scales), (3) small sample size where normality cannot be verified, (4) extreme outliers present. Non-parametric tests are more robust but less powerful.
Exam Viva Scenarios
Practise clinical reasoning and management decisions out loud
“You are comparing functional scores between 3 different surgical approaches for rotator cuff repair. What statistical test would you use and why?”
“A study used logistic regression to identify predictors of nonunion after tibial fracture. Age had an odds ratio of 1.5 (95% CI 1.2 to 1.9, p = 0.001). What does this mean?”
Tests for Continuous Outcomes
- 2 groups, independent, normal = Independent t-test
- 2 groups, paired (before-after), normal = Paired t-test
- 2 groups, independent, non-normal = Mann-Whitney U
- 2 groups, paired, non-normal = Wilcoxon signed-rank
- 3+ groups, independent, normal = One-way ANOVA + post-hoc
- 3+ groups, repeated measures, normal = Repeated measures ANOVA
- 3+ groups, independent, non-normal = Kruskal-Wallis
Tests for Categorical Outcomes
- Comparing proportions, expected count over 5 = Chi-square
- Comparing proportions, expected count under 5 = Fisher exact
- Binary outcome with predictors = Logistic regression
- Multiple categorical outcomes = Chi-square or multinomial regression
Tests for Associations
- Correlation between 2 continuous, normal = Pearson correlation (r)
- Correlation between 2 variables, non-normal or ordinal = Spearman correlation (rho)
- Predicting continuous outcome from predictors = Linear regression
- Predicting binary outcome from predictors = Logistic regression (OR)
- Correlation does NOT imply causation - confounders may explain
Critical Test Selection Rules
- Match test to outcome type (continuous vs categorical)
- Check normality before parametric tests (histogram, Q-Q plot, Shapiro-Wilk)
- Use paired tests for before-after, independent for separate groups
- ANOVA first for 3+ groups, then post-hoc (never multiple t-tests)
- Fisher exact when expected count under 5 (not chi-square)
Interpretation Principles
- ANOVA tells IF groups differ, post-hoc tells WHICH groups
- Pearson r: 0-0.3 weak, 0.3-0.7 moderate, 0.7-1.0 strong correlation
- Logistic regression OR greater than 1 = increased odds, OR less than 1 = decreased odds
- Regression coefficients are adjusted for other variables in model
- Non-parametric tests less powerful but more robust to violations
Common Mistakes
- Multiple t-tests instead of ANOVA (inflates Type I error)
- Independent t-test for paired data (loses power)
- Chi-square with expected count under 5 (inaccurate p-value)
- Not checking parametric assumptions before using t-test or ANOVA
- Confusing correlation with causation (observational data)
Evidence Base
P-values Unduly Influence Perceived Importance of Study Results
- Orthopaedic residents, fellows and surgeons rated importance of 40 published studies with and without p-values
- Of 40 comparative studies, 30 reported p less than 0.05 for the primary comparison
- Mean importance scores were higher when a significant p-value was shown (difference 0.6 on a 1-3 scale, 95% CI 0.1 to 1.1)
- 10 of 12 reviewers perceived results as more important when a significant p-value was presented
- Statistically significant p-values exert undue influence on perceived clinical importance
Number of Events Per Variable in Logistic Regression (the 10-EPV rule)
- Monte Carlo simulation of a 673-patient cardiac cohort with 252 deaths and 7 predictors
- Tested events per variable (EPV) of 2, 5, 10, 15, 20 and 25
- For EPV of 10 or greater, regression coefficients showed no major bias or coverage problems
- For EPV less than 10, coefficients were biased in both directions, variance estimates were unreliable, and paradoxical (wrong-direction) associations increased
- Established the widely used minimum of approximately 10 events per candidate variable
Fragility of Statistically Significant Findings in Sports Surgery RCTs
- Systematic survey of 48 RCTs in sports medicine and arthroscopic surgery (2005-2015)
- Median sample size 64 (IQR 48.5-89.5); median total outcome events 19 (IQR 10-27)
- Median Fragility Index was 2 (IQR 1-2.8)
- Changing the outcome status of just 2 patients in one arm reversed significance in the typical trial
- Most statistically significant orthopaedic RCTs are statistically fragile
RCTs, Observational Studies and the Hierarchy of Research Designs
- Compared meta-analyses of RCTs with meta-analyses of cohort/case-control studies on the same 5 clinical topics (99 reports)
- Average summary estimates from well-designed observational studies were remarkably similar to those of RCTs
- Example: BCG vaccine RR 0.49 (95% CI 0.34-0.70) from 13 RCTs vs OR 0.50 (95% CI 0.39-0.65) from 10 case-control studies
- Point estimates were actually more variable across RCTs than across observational studies
- Well-designed observational studies did not systematically overestimate treatment effects
STROBE Statement: Reporting of Observational Studies (Explanation and Elaboration)
- Consensus reporting guideline for cohort, case-control and cross-sectional studies
- 22-item checklist covering title, abstract, introduction, methods, results and discussion
- 18 items common to all designs; 4 items specific to each design
- Explicitly addresses confounding, bias, effect-size reporting and study-flow numbers
- Now required or recommended by most major orthopaedic and medical journals
Quality of Reporting of Confounding in Observational Intervention Studies
- Systematic review of 174 observational intervention studies in 10 general medical and epidemiological journals (2004-2007)
- Over 98% acknowledged the potential for confounding bias
- Only 51% reported details on inclusion of observed confounders and only 10% on their selection
- Just 9% commented on the likely effect of residual (unobserved) confounding
- Median overall reporting quality score was only 4 of 8, with no improvement across years