Definition
A reliability statistic that quantifies the proportion of total observed variance in continuous measurements attributable to between‑subject (or between‑target) variation within a specified measurement model; its numeric value depends on the chosen ICC form (e.g., one‑way random, two‑way mixed) and whether 'consistency' or 'absolute agreement' is defined.
Principle
Principle
ICC expresses how much measured variability reflects true differences between targets rather than measurement error or within‑target variation; higher ICC indicates greater relative between‑subject variance under the model's assumptions.
Demonstration
Demonstration
Illustrative scenario: Ten subjects each measured twice with the same device. A two‑way random‑effects ICC for absolute agreement is computed as 0.88, indicating that 88% of total variance arises from differences between subjects and 12% from within‑subject or measurement error under the model used.
Misapplication
Misapplication
Reporting an ICC without specifying the model, type, or whether it refers to consistency versus absolute agreement. The semantic error is treating ICC values as comparable across studies or designs when they reflect different variance decompositions and assumptions.
Consequence
Consequence
Using ICC to assess reliability informs sample‑size calculations, instrument selection, and interpretation of repeated measures; mis-specified or misinterpreted ICCs can produce invalid conclusions about reproducibility and lead to under‑ or over‑estimation of measurement error.
Reversal
Reversal
When the data violate model assumptions (non‑normality, heteroscedasticity, small number of targets), or when systematic biases exist between raters, the chosen ICC form may be unstable or misleading (e.g., high consistency ICC but low absolute agreement), requiring alternative analyses like Bland–Altman or variance‑component modelling.
Boundary
Boundary
Clearly within: continuous measurements on the same scale, repeated on the same targets under a well‑specified model. Boundary case: ordinal data with many categories—ICC may be used but alternatives (weighted kappa, ordinal models) can be preferable. Clearly outside: nominal categorical outcomes, where ICC is not appropriate.
Semantic Tension
Semantic Tension
Reliability (repeatability expressed by ICC) versus agreement/validity (whether measurements equal true values); a high ICC can coexist with systematic bias, revealing tension between consistency and absolute correctness.
Synthesis
Synthesis
ICC is a model‑dependent summary of how much variability is attributable to between‑target differences; valid use requires explicit specification of the ICC form and complementary analyses of agreement and measurement error to fully characterise reliability.