Kappa, ICC and Bland-Altman
- RELIABILITY is reproducibility or PRECISION - how consistently a measurement gives the same result; VALIDITY is ACCURACY - whether it measures what it is supposed to. They are independent: a measurement can be RELIABLE WITHOUT being VALID (consistently wrong, a systematic bias), so reliability is necessary but NOT sufficient for validity (the dartboard analogy - tight grouping vs hitting the bullseye).
- Reliability has forms: INTRA-OBSERVER (the same rater repeating a measurement), INTER-OBSERVER (different raters), and TEST-RETEST (the same instrument over time); orthopaedic classification systems and outcome measures must be tested for these.
- For CATEGORICAL data (e.g. a fracture classification), agreement is measured by COHEN'S KAPPA between two raters - which corrects for the agreement expected by CHANCE - with FLEISS' kappa for more than two raters and WEIGHTED kappa for ORDINAL categories; the conventional LANDIS-KOCH interpretation is: below 0.20 slight, 0.21-0.40 fair, 0.41-0.60 moderate, 0.61-0.80 substantial, and 0.81-1.00 almost perfect agreement.
- For CONTINUOUS data (e.g. an angle or a length measurement), reliability is measured by the INTRACLASS CORRELATION COEFFICIENT (ICC), which runs from 0 to 1 (broadly above 0.75 is good and above 0.9 excellent), and method comparison is best shown with a BLAND-ALTMAN PLOT, which plots the DIFFERENCE between two methods against their MEAN to reveal systematic BIAS and the LIMITS OF AGREEMENT (mean difference +/- 1.96 standard deviations).
- A key trap is that CORRELATION (e.g. Pearson's r) is NOT agreement: two methods can correlate almost perfectly yet disagree systematically (one always reads higher), so r should NOT be used to claim two methods/raters agree - use the ICC or a Bland-Altman plot instead.
- Applied to orthopaedics: classification systems are valuable as a SHARED LANGUAGE but have LIMITED stand-alone reliability - increasing the number of categories/subcategories consistently REDUCES kappa/ICC, while a brief rater CALIBRATION session improves agreement; report agreement statistics with confidence intervals and choose the statistic that matches the data type.
- “Reliability = precision (reproducible); Validity = accuracy (true). Reliable can be invalid (systematic bias) - dartboard analogy.
- “Categorical agreement = kappa (Cohen's 2 raters, Fleiss' for more than 2, weighted for ordinal); Landis-Koch: 0.21-0.40 fair, 0.41-0.60 moderate, 0.61-0.80 substantial, 0.81-1.0 almost perfect.
- “Continuous: ICC for reliability; Bland-Altman (difference vs mean; bias + limits of agreement) for method comparison. Correlation is NOT agreement. More categories -> lower kappa.
Consistent, reproducible results. A reliable but invalid measure is consistently wrong (tight cluster, off the bullseye - systematic bias).
Measures the true value. Reliability is necessary but not sufficient for validity - you need both to hit the bullseye.
Reliability vs Validity
Reliability (reproducibility, precision) asks whether repeated measurements agree with each other; validity (accuracy) asks whether the measurement reflects the true value. The classic dartboard analogy makes the relationship clear: tight grouping = reliable, hitting the centre = valid. A measurement can be reliable but not valid - tightly clustered but systematically off-target (a bias) - so reliability is necessary but not sufficient for validity. Reliability comes in forms - intra-observer (same rater repeated), inter-observer (different raters) and test-retest (over time) - all of which matter when validating a classification system or an outcome measure.

Agreement Statistics
- Categorical data (e.g. a fracture classification): Cohen's KAPPA for two raters (corrects for chance agreement), Fleiss' kappa for more than two raters, and weighted kappa for ordinal categories. Interpret with Landis-Koch: below 0.20 slight, 0.21-0.40 fair, 0.41-0.60 moderate, 0.61-0.80 substantial, 0.81-1.00 almost perfect.
- Continuous data (e.g. angle/length measurements): the INTRACLASS CORRELATION COEFFICIENT (ICC) (0-1; broadly above 0.75 good, above 0.9 excellent) for reliability.
- Method comparison: the BLAND-ALTMAN plot - plot the difference between two methods against their mean, revealing systematic bias (the mean difference) and the limits of agreement (mean difference plus or minus 1.96 standard deviations).
- Do NOT use correlation (Pearson's r) to claim agreement - two methods can correlate strongly yet disagree systematically; use ICC or Bland-Altman.

- Statistic
- Cohen's kappa
- Notes
- Corrects for chance; Landis-Koch grades
- Statistic
- Fleiss' kappa
- Notes
- Extension of kappa
- Statistic
- Weighted kappa
- Notes
- Penalises bigger disagreements more
- Statistic
- Intraclass correlation (ICC)
- Notes
- 0-1; above 0.75 good, above 0.9 excellent
- Statistic
- Bland-Altman plot
- Notes
- Bias + limits of agreement (NOT correlation)
Validity Types & Classification Reliability
Validity has several forms: face (does it look reasonable), content (does it cover the construct), construct (does it behave as theory predicts), and criterion validity - concurrent (agrees with a gold standard now) and predictive (predicts a future outcome); sensitivity/specificity against a gold standard are criterion validity (see our Diagnostic Test Statistics topic). In orthopaedics, classification systems (e.g. fracture classifications) are tested for inter- and intra-observer reliability using kappa/ ICC; the evidence shows their stand-alone reliability is often only fair-to-moderate, that increasing granularity (more categories) lowers kappa/ICC, and that rater calibration improves agreement - which is why classifications are best used as a shared language and research scaffold alongside specific radiographic thresholds and patient factors, rather than as the sole basis for decisions.
The Kappa Paradox: Why a High Agreement Can Give a Low Kappa
- The prevalence paradox: when one category is very common (or very rare), the chance-expected agreement is high, so even with high observed agreement the kappa comes out paradoxically low (kappa is "punished" for an unbalanced category mix).
- The bias paradox: when the two raters use the categories at different rates (asymmetric marginals, i.e. rater bias), kappa can be paradoxically raised.
- What to do: always report the raw (observed) percent agreement alongside kappa, and consider the prevalence-and-bias-adjusted kappa (PABAK = 2 × observed agreement − 1) when the categories are unbalanced.
- The cut-offs are arbitrary: the Landis-Koch bands are a convention, not an absolute - interpret kappa with its confidence interval and in the context of prevalence/bias, not as a hard grade.
Kappa is distorted by prevalence (an unbalanced/rare category gives a paradoxically low kappa despite high observed agreement) and by rater bias (asymmetric marginals can raise it). So quote the raw percent agreement and consider PABAK with unbalanced data, and treat the Landis-Koch grades as arbitrary - report kappa with its confidence interval.
ICC in Practice: Which Model, and Turning It Into SEM and MDC
- There is no single ICC - state the model: you must specify one-way vs two-way, random vs fixed effects, single-measure vs average-measure, and consistency vs absolute agreement, because these give different values for the same data - so a reported ICC is uninterpretable without its form.
- Convert reliability into measurement error: the clinically useful quantities derived from the ICC are the Standard Error of Measurement (SEM = SD × √(1 − ICC)) - the typical error in the units of the measure - and the Minimal Detectable Change (MDC; e.g. MDC95 = 1.96 × √2 × SEM) - the smallest change that exceeds measurement error, i.e. a change you can be confident is real, not noise.
- Real vs meaningful: pair the MDC with the Minimal Clinically Important Difference (MCID) - the smallest change a patient perceives as worthwhile - so an observed change should ideally exceed both the MDC (real) and the MCID (meaningful). The MCID/responsiveness of outcome scores is developed in our Outcome Measures (PROMs) topic.
An ICC is meaningless without its model (one-way/two-way, random/fixed, single/average, consistency/absolute agreement). Convert it to SEM = SD × √(1 − ICC) (error in real units) and MDC ≈ 1.96 × √2 × SEM (smallest real change); then compare with the MCID (smallest meaningful change) - a true, useful change exceeds both.
Mnemonics & Memory Aids
PRECISE vs TRUE
Hook:Reliability = PRECISE; Validity = TRUE; you need both.
KIB
Hook:KIB: Kappa (categorical), ICC (continuous), Bland-Altman (method comparison).
Clinical Decision Scenarios
Practise clinical reasoning and management decisions out loud
“What is the difference between reliability and validity, and which statistics would you use to assess the reliability of a fracture classification and of an angle measurement?”
“Why do more detailed (granular) classification systems tend to be less reliable, and how can agreement be improved?”
Concepts
- Reliability = precision/reproducibility; validity = accuracy
- Reliable can be invalid (systematic bias) - dartboard analogy
- Reliability forms: intra-observer, inter-observer, test-retest
Categorical agreement
- Cohen's kappa (2 raters), Fleiss' kappa (more than 2), weighted kappa (ordinal)
- Corrects for chance agreement
- Landis-Koch: 0.21-0.40 fair, 0.41-0.60 moderate, 0.61-0.80 substantial, 0.81-1.0 almost perfect
Continuous agreement
- ICC for reliability (0-1; above 0.75 good, above 0.9 excellent)
- Bland-Altman for method comparison (difference vs mean; bias + limits of agreement)
- Correlation (Pearson r) is NOT agreement
Validity & classifications
- Validity: face, content, construct, criterion (concurrent/predictive)
- More categories -> lower kappa; rater calibration improves agreement
- Classifications = shared language; report kappa/ICC with CIs
Evidence & Key Studies
Distal radius fracture classifications in real life: reliability and how they change treatment
- Interobserver agreement for distal radius fracture classifications was typically fair-to-moderate on radiographs, with only modest improvement on CT.
- Increasing granularity (more categories/subcategories) consistently REDUCED kappa/ICC, whereas a brief rater calibration session improved agreement.
- Classifications remain valuable as a shared language but have limited stand-alone reliability and prognostic power, best combined with instability thresholds and patient factors.
Zero echo time MRI vs CT in intra-articular distal radius fractures: inter/intraobserver agreement
- Inter- and intraobserver agreement were quantified with Cohen's and Fleiss' kappa and intraclass correlation coefficients.
- Classification agreement was 'good' (kappa about 0.68-0.78), with surgeons agreeing more than radiologists; continuous measures showed good ICC for angulation (about 0.76-0.86) but lower for inclination.
- Illustrates the practical use and interpretation of kappa (categorical) and ICC (continuous) for reliability.
The fair-to-moderate reliability of fracture classifications, the reduction of kappa/ICC with greater granularity and the benefit of rater calibration come from the cited Nguyen review, and the worked use of Cohen's/Fleiss' kappa and ICC (with interpretive values) from the cited Kaymakoglu study. The reliability-versus-validity distinction, the Landis-Koch kappa grades, the ICC and the Bland-Altman method- comparison approach are standard, well-established statistical teaching. (See also our Diagnostic Test Statistics, Measures of Effect and Study Design topics.)