Description
The current module provides Cohen’s kappa for two raters (and average pairwise kappa for multiple raters) and Fleiss’ kappa for multiple raters. However, both Cohen’s and Fleiss’ kappa can be generalized to two or more raters, differing mainly in their definition of expected agreement. I suggest adding generalized Cohen’s kappa, Fleiss’ kappa, and Brennan–Prediger kappa, regardless of the number of raters. Weighted versions and confidence intervals could also be provided.
Purpose
No response
Use-case
No response
Is your feature request related to a problem?
Users are currently limited to Fleiss’ kappa for more than two raters, which may not always be appropriate, as its definition of chance agreement relies on different assumptions from those underlying Cohen’s kappa.
Is your feature request related to a JASP module?
Reliability
Describe the solution you would like
Use the package irrCAC to mimic what is offered for quantitative outcomes in the module intraclass correlation
Describe alternatives that you have considered
No response
Additional context
The module (like other software) currently propose Cohen’s kappa for two raters (and take the average for more than two raters) and Fleiss kappa for more than two raters. However, both Cohen’s and Fleiss versions can be proposed for two or more than two raters. The difference between these two coefficients lies in different definitions of chance agreement or expected agreement.
(1) Agreement obtained under the statistical independence assumption by assuming that all raters classify the subjects completely randomly (i.e., uniform marginal distribution). This leads to the so-called Brennan-Prediger kappa
(2) Agreement obtained under the statistical independence assumption by assuming that the marginal distribution of the raters is equal to their observed marginal distribution (i.e., using the marginal distribution of each rater). This leads to Cohen’s kappa and its generalization to more than two raters.
(3) Agreement obtained under the statistical independence assumption by assuming that the marginal distribution of the raters is the same for all raters (i.e., using the average marginal distribution). This leads to Scott’s pi and its generalization to more than two raters, known as Fleiss kappa.
In other terms, Fleiss kappa is equivalent (for binary scales) to the intraclass correlation under a one-way ANOVA model while Cohen’s kappa is equivalent to (for binary scales) to the intraclass correlation coefficient for agreement under a two-way ANOVA model.
Weighted coefficients (e.g., linear, quadratic) can also be obtained under all 3 chance definitions. The standard error and confidence intervals of all these coefficients can be obtained using the delta method or a slight modification (R package irrCAC).
The fact that Fleiss kappa is usually used for more than two raters, and Cohen’s kappa and weighted kappas for two raters originates, I believe, in the fact these coefficients were introduced in these particular contexts in the literature and were later implemented in “popular” software only in these particular cases. I think that JASP can therefore gain in audience and participate to better statistical practices by modifying the module.
References
S. Vanbelle, C. H. Engelhart, and E. Blix, “ Measuring Agreement in Diagnostics: A Practical Guide for Researchers,” Statistics in Medicine 44, no. 23-24 (2025): e70299, https://doi.org/10.1002/sim.70299.
Vanbelle, S., Engelhart, C. & Blix, E. A comprehensive guide to study the agreement and reliability of multi-observer ordinal data. BMC Med Res Methodol 24, 310 (2024). https://doi.org/10.1186/s12874-024-02431-y
Description
The current module provides Cohen’s kappa for two raters (and average pairwise kappa for multiple raters) and Fleiss’ kappa for multiple raters. However, both Cohen’s and Fleiss’ kappa can be generalized to two or more raters, differing mainly in their definition of expected agreement. I suggest adding generalized Cohen’s kappa, Fleiss’ kappa, and Brennan–Prediger kappa, regardless of the number of raters. Weighted versions and confidence intervals could also be provided.
Purpose
No response
Use-case
No response
Is your feature request related to a problem?
Users are currently limited to Fleiss’ kappa for more than two raters, which may not always be appropriate, as its definition of chance agreement relies on different assumptions from those underlying Cohen’s kappa.
Is your feature request related to a JASP module?
Reliability
Describe the solution you would like
Use the package irrCAC to mimic what is offered for quantitative outcomes in the module intraclass correlation
Describe alternatives that you have considered
No response
Additional context
The module (like other software) currently propose Cohen’s kappa for two raters (and take the average for more than two raters) and Fleiss kappa for more than two raters. However, both Cohen’s and Fleiss versions can be proposed for two or more than two raters. The difference between these two coefficients lies in different definitions of chance agreement or expected agreement.
(1) Agreement obtained under the statistical independence assumption by assuming that all raters classify the subjects completely randomly (i.e., uniform marginal distribution). This leads to the so-called Brennan-Prediger kappa
(2) Agreement obtained under the statistical independence assumption by assuming that the marginal distribution of the raters is equal to their observed marginal distribution (i.e., using the marginal distribution of each rater). This leads to Cohen’s kappa and its generalization to more than two raters.
(3) Agreement obtained under the statistical independence assumption by assuming that the marginal distribution of the raters is the same for all raters (i.e., using the average marginal distribution). This leads to Scott’s pi and its generalization to more than two raters, known as Fleiss kappa.
In other terms, Fleiss kappa is equivalent (for binary scales) to the intraclass correlation under a one-way ANOVA model while Cohen’s kappa is equivalent to (for binary scales) to the intraclass correlation coefficient for agreement under a two-way ANOVA model.
Weighted coefficients (e.g., linear, quadratic) can also be obtained under all 3 chance definitions. The standard error and confidence intervals of all these coefficients can be obtained using the delta method or a slight modification (R package irrCAC).
The fact that Fleiss kappa is usually used for more than two raters, and Cohen’s kappa and weighted kappas for two raters originates, I believe, in the fact these coefficients were introduced in these particular contexts in the literature and were later implemented in “popular” software only in these particular cases. I think that JASP can therefore gain in audience and participate to better statistical practices by modifying the module.
References
S. Vanbelle, C. H. Engelhart, and E. Blix, “ Measuring Agreement in Diagnostics: A Practical Guide for Researchers,” Statistics in Medicine 44, no. 23-24 (2025): e70299, https://doi.org/10.1002/sim.70299.
Vanbelle, S., Engelhart, C. & Blix, E. A comprehensive guide to study the agreement and reliability of multi-observer ordinal data. BMC Med Res Methodol 24, 310 (2024). https://doi.org/10.1186/s12874-024-02431-y