Patient-Reported Outcomes | Measurement Properties | Clinical Application
- PROM: Patient-Reported Outcome Measure - patient completes without clinician interpretation. Captures patient perspective.
- MCID: Smallest change in score that patients perceive as meaningful benefit. Essential for clinical interpretation.
- Validity: Does the measure assess what it claims to assess? (content, construct, criterion validity)
- Reliability: Does the measure give consistent results? (test-retest, inter-rater, internal consistency)
- Responsiveness: Can the measure detect clinically meaningful change over time? (ceiling/floor effects)
- “SF-36 has 2 components: Physical (PCS) and Mental (MCS) - scored 0-100, higher is better
- “WOMAC assesses 3 domains: Pain, Stiffness, Function - scored 0-96, lower is better (or normalised 0-100)
- “DASH measures upper extremity disability - 0-100 scale, 0 = no disability
- “Floor/ceiling effects over 15% indicate measure may not detect worsening or improvement
Overview and Introduction
What a PROM is. A patient-reported outcome measure (PROM) is a standardised, validated questionnaire that the patient completes without clinician interpretation. It captures the patient's own perspective on health status, symptoms, function and quality of life.
Why PROMs matter. Surgeon assessment may not match the patient's experience, and pain, function and satisfaction cannot be objectively measured. A PROM quantifies these subjective outcomes and captures what matters to patients: pain, daily activities and quality of life. PROMs are essential in clinical trials to demonstrate treatment efficacy, and they are also used for registry benchmarking and value-based care (see Clinical Application and Relevance).
PROMs and clinician measures. Clinician measures such as range of motion and strength are important, but they may not correlate with patient satisfaction. Best practice is to use both PROMs and objective measures.
Types of Outcome Measures
Generic against specific. Generic PROMs assess overall health status in any condition and allow comparison between diseases, populations and population norms, but they are less sensitive to specific joint pathology. Joint-specific PROMs are highly sensitive to pathology in a single joint but cannot be compared across different joints or with the general population. Use both when possible, to capture the joint and overall health together.

Generic PROMs
The Short Form-36 Health Survey is a 36-item generic health status measure, and the most widely used generic PROM in orthopaedic research. It has eight subscales:
- Physical functioning
- Role physical (work/activities due to physical health)
- Bodily pain
- General health
- Vitality (energy/fatigue)
- Social functioning
- Role emotional (work/activities due to emotional problems)
- Mental health
Scoring. Each subscale runs 0-100, higher being better health. The physical domains aggregate into the Physical Component Summary (PCS) and the mental domains into the Mental Component Summary (MCS), and the MCID is approximately 5 points for each.
Region-, Joint- and Disease-Specific PROMs
The Western Ontario and McMaster Universities Arthritis Index is the most widely used PROM for hip and knee osteoarthritis and the gold standard for hip and knee arthroplasty outcome assessment. Its 24 items fall into three domains:
- Pain - 5 items, pain with various activities
- Stiffness - 2 items, morning and later-day stiffness
- Physical function - 17 items, difficulty with daily activities
Scoring. The Likert version scores each item 0-4, for a total of 0-96, where lower is better; the VAS version scores each item 0-100mm. Scores are often normalised to 0-100, and whether higher or lower is then better depends on the version. The MCID is approximately 10-15 points on the 100-point scale.
Its validity and reliability for hip and knee osteoarthritis are excellent, and it is widely used in arthroplasty research. It was designed for arthritis, so it is less applicable to ligament injuries and fractures.
Choosing Between PROMs
The most common exam error is treating all PROMs as interchangeable. The table contrasts the major instrument types so that a choice can be justified under viva pressure.
- Examples
- SF-36, SF-12
- Key Strength
- Cross-disease comparison, population norms, captures whole-person health
- Key Limitation
- Lower responsiveness to focal joint pathology
- Best Use
- Secondary outcome; comparing burden across conditions
- Examples
- EQ-5D, SF-6D
- Key Strength
- Generates QALY utility (0 to 1) for cost-utility analysis
- Key Limitation
- Coarse (few levels); ceiling effects in healthy people
- Best Use
- Health-economic evaluation, payer/HTA submissions
- Examples
- DASH/QuickDASH, LEFS
- Key Strength
- One score across a whole limb when pathology spans joints
- Key Limitation
- Less sensitive than single-joint scores
- Best Use
- Multi-level or undefined upper/lower limb pathology
- Examples
- WOMAC, OHS/OKS, ASES, ODI
- Key Strength
- Highest responsiveness to the target joint or disease
- Key Limitation
- Cannot compare across joints or to general population
- Best Use
- Primary outcome in arthroplasty/disease-specific trials and registries
Measurement Properties
Three properties decide whether a PROM can be trusted: validity, reliability and responsiveness. A measure with poor properties produces unreliable conclusions.
Validity asks whether the measure assesses what it claims to assess.
- Definition
- Covers all relevant aspects of construct
- How to Assess
- Expert panel review, patient input
- Example
- WOMAC includes pain, stiffness, function for OA
- Definition
- Correlates with related measures, discriminates from unrelated
- How to Assess
- Correlation with similar PROMs (convergent), lack of correlation with dissimilar (discriminant)
- Example
- WOMAC correlates with knee ROM (convergent) but not with mental health scores (discriminant)
- Definition
- Correlates with gold standard
- How to Assess
- Compare to established measure
- Example
- New knee score correlates with WOMAC
Reliability asks whether the measure gives consistent results when the condition is stable.
- Definition
- Same result when repeated in stable patients
- How to Assess
- Intraclass Correlation Coefficient (ICC)
- Target
- ICC greater than 0.70
- Definition
- Different raters get same result
- How to Assess
- ICC for clinician-administered measures
- Target
- ICC greater than 0.70
- Definition
- Items within scale measure same construct
- How to Assess
- Cronbach alpha
- Target
- Alpha 0.70 to 0.95 (too high suggests redundancy)
Responsiveness asks whether the measure can detect clinically meaningful change over time. A floor effect is present when a high proportion of patients, over 15%, score at the minimum, the worst possible score, and the measure cannot detect worsening in them. A ceiling effect is the same at the maximum, the best possible score, and the measure cannot detect improvement.
Quantifying responsiveness. Responsiveness is expressed as the standardised response mean (SRM) or effect size:
- SRM greater than 0.8 - large responsiveness (good)
- SRM 0.5 to 0.8 - moderate responsiveness
- SRM less than 0.5 - small responsiveness (may miss change)
Minimal Clinically Important Difference (MCID)
Definition. The MCID is the smallest change in PROM score that patients perceive as beneficial and that would mandate a change in management. Its purpose is to distinguish a statistically significant change from a clinically meaningful one: a change with p less than 0.05 may not reach the MCID, and then it may not matter to patients.

Deriving it. There are two methods:
- Anchor-based - compare the PROM change with an external anchor, a patient global assessment such as "Compared to before surgery, how would you rate your improvement: much better, better, same, worse?" The MCID is the mean change in the "better" group.
- Distribution-based - use a statistical threshold, either 0.5 × the standard deviation or the standard error of measurement. This is less clinically intuitive than the anchor-based method.
Applying it. If the mean improvement is 8 points and the MCID is 10, the improvement is statistically significant but not clinically meaningful. If the 95% CI runs from 12 to 18 points against an MCID of 10, the entire interval exceeds the MCID and the improvement is clinically meaningful. An interval that crosses the MCID leaves the clinical significance uncertain, as Scenario 2 below works through. Always compare treatment effects with the MCID, not just with p-values.
Real before important. A change score must clear two separate hurdles. The minimal detectable change (MDC) is the smallest change that exceeds measurement error, so that it is real rather than noise; the MCID is the smallest change patients value. They answer different questions and must not be conflated.
- Minimal Detectable Change (MDC)
- Is the change beyond measurement error (real)?
- Minimal Clinically Important Difference (MCID)
- Is the change big enough to matter to the patient?
- Minimal Detectable Change (MDC)
- Reliability/measurement error (distribution-based)
- Minimal Clinically Important Difference (MCID)
- Patient perception (usually anchor-based)
- Minimal Detectable Change (MDC)
- MDC95 = 1.96 x sqrt(2) x SEM, where SEM = SD x sqrt(1 - reliability)
- Minimal Clinically Important Difference (MCID)
- Mean change in the 'a little better' anchor group, or ~0.5 SD
- Minimal Detectable Change (MDC)
- A score must exceed the MDC to be trusted as a true change
- Minimal Clinically Important Difference (MCID)
- A score should reach the MCID to be considered worthwhile
Ideally MCID is greater than or equal to MDC - then a clinically important change is also reliably detectable above measurement error. If the MCID is smaller than the MDC, the instrument cannot reliably distinguish a "clinically important" change from random measurement noise, which undermines its use as an outcome measure. Always check that an instrument's MCID is at least its MDC before relying on it.
Modern PROMs: PROMIS, Item-Response Theory and CAT
Classical and modern instruments. Legacy PROMs use a fixed list of items, each scored equally (classical test theory). The modern direction is item-response theory (IRT) instruments, of which PROMIS (Patient-Reported Outcomes Measurement Information System) is the leading example, increasingly used by US registries.
What PROMIS is. A set of NIH-funded item banks, such as Physical Function and Pain Interference, calibrated by item-response theory so that each question has a known difficulty. Scores are reported on a T-score metric, mean 50 and standard deviation 10, anchored to the US general population, so a score of 40 is one SD worse than average. This common metric lets different item sets from the same bank be compared directly.
Computer-adaptive testing (CAT). Because every item is calibrated, a computer-adaptive test can choose the next question from the previous answer, homing in on the patient's level in typically 4-8 items instead of 30. CAT reduces respondent burden and floor and ceiling effects while maintaining precision. Where adaptive delivery is not feasible, PROMIS can also be given as fixed short forms.
The familiar benchmarks (a specific MCID in points, Cronbach alpha, ICC targets) were derived for fixed classical scores and do not transfer directly to IRT/CAT instruments; PROMIS uses its own MCID estimates on the T-score metric, and cross-walk tables that convert legacy scores (e.g. a knee or hip score) to PROMIS are approximate. This is why a unit cannot simply swap a legacy score for PROMIS and assume the old interpretation thresholds still apply.
Principles of Outcome Measurement
The hierarchy. Outcome measurement rests on what you measure, how you measure it, and how you interpret it. The WHO ICF framework (body structure and function, activity, participation) is a useful map: PROMs predominantly capture the activity and participation levels, while clinician measures (range of motion, strength, radiographs) capture body structure and function.
Types of outcome. Outcome measures fall into four types:
- Patient-reported outcome measures (PROMs) - the patient's own rating of symptoms, function and quality of life, with no clinician interpretation
- Clinician-reported outcomes (ClinROs) - examiner-derived, such as range of motion, Constant strength or neurological grade
- Performance outcomes (PerfOs) - observed task performance, such as the Timed Up-and-Go or the six-minute walk
- Composite scores - blend domains, as the Constant-Murley combines patient pain with examiner-measured strength and range; this improves breadth but can obscure which domain drives a change
Anchoring interpretation. Interpretation rests on four anchoring concepts. The MCID, the smallest change a patient perceives as worthwhile, has its own section below, and floor and ceiling effects, which distort responsiveness when too many patients cluster at the extremes, are covered under Measurement Properties. Two further anchors are defined here:
- PASS (Patient Acceptable Symptom State) - the post-treatment score above which a patient considers their state satisfactory. It is increasingly preferred to MCID because it reports an attainable end state rather than a change.
- SCB (Substantial Clinical Benefit) - a higher threshold than MCID, denoting a large, clearly meaningful improvement.
What a good study does. It pre-specifies a single primary PROM, justifies it on measurement properties, and reports both mean change against the MCID and the proportion of patients reaching MCID or PASS.
Clinical Application and Relevance
Registries. The major arthroplasty registries (NJR, AJRR, AOANJRR, SHAR) collect PROMs at a pre-operative baseline and at post-operative follow-up, at 1 year and 5 years. That allows benchmarking of performance and quality improvement.
Value-based care. Payers increasingly link reimbursement to PROMs. Demonstrating patient-reported improvement justifies procedures, and PROMs are essential for value-based contracts. Linking them to payment has risks of its own, set out under Controversies below.
Guidelines, Registries & Global Practice
PROM collection has shifted from research tool to routine quality infrastructure worldwide, but the chosen instruments, mandate and uptake vary by region.
Global epidemiology and uptake. Joint arthroplasty is among the most-studied PROM settings: large registry programmes consistently show that the majority of hip and knee replacement patients achieve improvements exceeding the MCID at one year, while a meaningful minority (commonly cited around one in five for knees) report being unsatisfied - a gap PROMs make visible that complication rates alone miss. Uptake is high in publicly funded, registry-linked systems and far patchier where collection is voluntary or unfunded.
Guidance and Registry Programmes Side by Side
- Stance on PROMs
- National PROMs programme historically mandated pre/post hip and knee replacement; NJR links implant survival to outcomes
- Typical Instruments
- Oxford Hip/Knee Score, EQ-5D
- Stance on PROMs
- Registry-integrated PROM collection at standardised intervals for benchmarking
- Typical Instruments
- Oxford Hip/Knee Score, EQ-5D, VAS
- Stance on PROMs
- AAOS promotes PROMs and CMS value-based programmes increasingly require them; AJRR collects PROMs
- Typical Instruments
- HOOS/KOOS JR, PROMIS, VR-12
- Stance on PROMs
- Long-standing registry PROM collection underpinning revision and bearing comparisons
- Typical Instruments
- EQ-5D, joint-specific scores, satisfaction VAS
- Stance on PROMs
- Defines standard outcome sets to harmonise PROMs across countries for a given condition
- Typical Instruments
- Condition-specific standard sets (e.g. hip/knee OA)
Methodological standards. The COSMIN initiative (Mokkink and colleagues) provides the most widely cited international consensus framework for selecting and appraising PROMs, complementing earlier quality criteria. ICHOM standard sets push toward globally comparable outcome reporting.
High- versus limited-resource practice variation. In well-resourced, registry-linked systems, electronic PROM (ePROM) capture is increasingly embedded in routine care, enabling case-mix-adjusted benchmarking. In limited-resource settings, barriers include literacy and language (validated translations are not universal), lack of electronic infrastructure, staffing for follow-up, and the cost of licensed instruments - so brief, free, culturally validated tools (and pragmatic VAS/EQ-5D use) are favoured. The principle that statistical significance does not equal clinical significance, and that MCID/PASS should anchor interpretation, applies in every setting regardless of resource level.
Controversies and Areas of Uncertainty
PROM science is evolving, and several issues remain genuinely unsettled. They make useful "areas of debate" answers in a viva.
MCID is not a single number. The same PROM yields different MCIDs depending on whether an anchor-based or distribution-based method is used, the anchor question, baseline severity and follow-up timing. Quoting "the MCID" as if it were fixed is a recognised pitfall: always state the method and population.
MCID versus PASS. A patient can exceed the MCID yet remain symptomatic and dissatisfied, and many groups now favour PASS, or the proportion reaching a "good outcome", as more patient-relevant than mean change.
Ceiling effects in legacy scores. Widely used scores (Constant, Harris Hip, some Oxford items) show marked ceiling effects in well-functioning patients. That masks further improvement and biases comparisons of already good results.
Response shift and missing data. Patients recalibrate their internal standard for "good" over time (response shift), which complicates before-and-after comparisons. Differential loss to follow-up, with sicker patients dropping out, inflates apparent improvement, and complete-case analysis is a common source of bias.
Linking PROMs to payment. Using PROMs for reimbursement or surgeon-level ranking risks gaming, risk-aversion (avoiding complex patients) and inadequate case-mix adjustment. These are the reasons several systems publish PROMs for benchmarking rather than for direct pay-for-performance.
MCQ Practice Points
Q: What is the difference between a generic PROM (SF-36) and a joint-specific PROM (WOMAC)? A: Generic PROMs assess overall health status across any condition, allow comparison between diseases and to population norms, but are less sensitive to specific joint pathology. Joint-specific PROMs are highly sensitive to pathology in a single joint but cannot compare across different joints or to general population.
Q: Why is MCID important when interpreting PROM changes? A: MCID defines clinically meaningful change - the smallest improvement that patients perceive as beneficial. Statistically significant changes (p less than 0.05) may not exceed MCID and thus not be clinically important. Always compare observed change to MCID, not just p-value.
Q: What is a ceiling effect and why does it matter? A: Ceiling effect occurs when high proportion (over 15%) of patients score at maximum (best possible score). This prevents the measure from detecting improvement in these patients and reduces responsiveness. Choose a different measure or add a more challenging domain if ceiling effects are problematic.
Q: What is the difference between validity and reliability? A: Validity = Does the measure assess what it claims to assess? (accuracy). Reliability = Does the measure give consistent results when repeated in stable patients? (precision). A measure can be reliable but not valid (consistently wrong), but cannot be valid without being reliable.
Q: What ICC value indicates good test-retest reliability? A: ICC greater than 0.70 indicates acceptable reliability. ICC (Intraclass Correlation Coefficient) ranges 0-1. ICC greater than 0.90 is excellent, 0.70-0.90 is good, less than 0.70 is poor. This measures consistency when same patient completes PROM twice with stable condition.
Q: How is responsiveness quantified? A: Standardized Response Mean (SRM) or Effect Size. SRM = mean change / SD of change. SRM greater than 0.8 = large responsiveness (good), 0.5-0.8 = moderate, less than 0.5 = small (may miss clinically important change). Responsiveness is essential for detecting treatment effects.
Exam Viva Scenarios
Practise clinical reasoning and management decisions out loud
“You are planning an RCT comparing cemented vs uncemented THA. What outcome measures would you use and why?”
“An RCT of 200 patients found that new rehab protocol improved WOMAC score by mean 8 points (95% CI 5 to 11 points, p = 0.001) compared to standard protocol. The MCID for WOMAC is 10 points. How do you interpret this result?”
“A colleague proposes adopting a brand-new shoulder PROM for your unit. How would you appraise whether it is fit for purpose, and what numbers would you want to see?”
Common Orthopaedic PROMs
- Generic: SF-36 (PCS/MCS, 0-100, higher better), EQ-5D (utility 0-1)
- Hip/Knee: WOMAC (pain/stiffness/function, 0-96 or 0-100, lower or higher better depending on version)
- Upper Extremity: DASH (0-100, 0 = no disability), QuickDASH (11 items)
- Shoulder: ASES (0-100, higher better), Constant score
- Spine: ODI (Oswestry 0-100%, lower better), NDI (Neck Disability)
MCID Values
- SF-36 PCS/MCS: MCID approximately 5 points
- WOMAC: MCID 10-15 points (on 100-point scale)
- DASH: MCID 10-15 points
- VAS Pain: MCID 15-20mm (on 100mm scale)
- Always compare treatment effect to MCID for clinical significance
Measurement Properties
- Validity = Does it measure what it claims? (content, construct, criterion)
- Reliability = Consistent results? (test-retest ICC greater than 0.70, Cronbach alpha 0.70-0.95)
- Responsiveness = Detects change? (SRM greater than 0.8 = large, less than 15% floor/ceiling effects)
- Floor effect = Too many at minimum (cannot detect worsening)
- Ceiling effect = Too many at maximum (cannot detect improvement)
PROM Selection
- Joint-specific for sensitivity (WOMAC for THA trial)
- Generic for cross-disease comparison and population norms (SF-36)
- Utility measure for cost-effectiveness (EQ-5D)
- Use combination: Joint-specific (primary) + Generic (secondary)
- Check floor/ceiling effects (over 15% problematic)
Interpreting PROM Data
- Compare mean change to MCID, not just p-value
- Check if 95% CI excludes MCID threshold
- Report proportion of patients achieving MCID
- Wide CI crossing MCID = uncertain clinical significance
- Large sample with trivial effect (below MCID) = not clinically important
Clinical Application
- Major registries (NJR, AJRR, AOANJRR) collect PROMs (baseline and follow-up)
- Value-based care links reimbursement to PROM improvement
- Statistical significance ≠ Clinical significance
- Generic vs Specific trade-off: Comparison vs Sensitivity
- Pre-specify primary PROM and timing in study protocol
Evidence Base
WOMAC: Original Validation (Landmark)
- WOMAC developed and validated within a double-blind RCT of two NSAIDs in hip and knee osteoarthritis
- Self-administered, disease-specific instrument with pain, stiffness and physical-function subscales
- Subscales fulfilled conventional criteria for face, content and construct validity
- Demonstrated reliability, responsiveness and relative efficiency
- Described by the authors as a 'high-performance' instrument for evaluative OA research
SF-36: Conceptual Framework (Landmark Generic PROM)
- Introduced the 36-item Short Form (SF-36) from the Medical Outcomes Study
- Surveys eight health concepts spanning physical and mental health domains
- Designed for self-administration in people aged 14 years and over, or by trained interviewer
- Built for clinical practice, research, health-policy evaluation and population surveys
- Established the template for generic, profile-based health-status measurement
DASH: Development of an Upper-Extremity PROM
- Joint AAOS, COMSS and Institute for Work and Health initiative to create a region-wide upper-limb measure
- Item generation produced 821 candidate items reduced to a focused symptom and function set
- Single questionnaire spans the whole upper limb rather than an isolated joint
- Field tested across centres in the United States, Canada and Australia
- Provided the basis for the validated DASH and the shortened QuickDASH
Oxford Hip Score: A Joint-Specific Registry PROM
- Developed a 12-item patient-completed questionnaire for total hip replacement (n=220, prospective)
- High internal consistency and satisfactory test-retest reproducibility
- Validity confirmed by correlation with Charnley score, SF-36 and AIMS
- Standardised effect size (responsiveness) compared favourably with SF-36 and AIMS
- Short, practical and sensitive to clinically important change after THR