Definition
A chance‑corrected measure of inter‑rater agreement for two raters classifying items into categorical outcomes; it compares observed agreement to expected agreement under statistical independence of raters and typically ranges from −1 to 1, where 0 indicates agreement no better than chance.

Principle

Principle
Kappa isolates agreement beyond that expected by the marginal distributions (chance agreement); its value depends on both observed concordance and the raters' category prevalences (marginals).

Demonstration

Demonstration
Illustrative scenario: Two pathologists independently classify 100 tissue samples as benign or malignant. Observed agreement is 90%, but expected agreement by chance (given marginals) is 70%, yielding Cohen's kappa ≈ 0.67, interpreted as substantial agreement beyond chance.

Misapplication

Misapplication
Interpreting kappa values without regard to prevalence or imbalance in marginal totals. The common error is treating a low kappa as indicating poor agreement when high percent agreement exists (the 'kappa paradox'), or using kappa for more than two raters without adaptation.

Consequence

Consequence
Kappa is widely used to report inter‑rater reliability for categorical classifications; uncritical use can understate agreement in imbalanced categories or mislead instrument selection and training decisions if marginal effects are ignored.

Reversal

Reversal
When category prevalences are extreme or raters have markedly different marginal distributions, kappa can be low despite high observed agreement; in such cases adjusted statistics (e.g., prevalence‑adjusted measures) or alternative coefficients (Gwet's AC1, percent agreement) may better reflect practical concordance.

Boundary

Boundary
Clearly within: two raters, nominal categorical outcomes. Boundary case: ordinal categories where weighted kappa is more appropriate. Clearly outside: multi‑rater nominal settings without adaptation (use Fleiss' kappa or other multi‑rater measures).

Semantic Tension

Semantic Tension
Raw percent agreement (simple intuitive measure) versus chance‑corrected agreement (kappa); tension arises because percent agreement ignores chance agreement while kappa may be sensitive to prevalence and bias.

Synthesis

Synthesis
Cohen's kappa quantifies agreement beyond chance for two categorical raters but must be interpreted alongside percent agreement, marginal distributions, and, when appropriate, alternative statistics to avoid misleading conclusions in imbalanced or multi‑rater settings.