grantscience.com · reliability, coded
It depends who you count, and how
The same review process yields a median ICC of 0.28 for one reviewer but 0.55 once a panel is averaged, because averaging raters raises reliability (the Spearman-Brown relationship). Chance-corrected kappa sits near 0.19. Reporting a single “reliability of peer review” without saying single reviewer or panel is close to meaningless, which is why every coefficient below is coded to its exact estimand.
Read with care: every measurement was extracted from the paper’s full text by two independent AI coders, and their disagreements were adjudicated by a third model against the source, with each value carrying a verbatim quote and locator. Human verification of the rows has not yet been done: values are provisional and to be verified before citing. A sample size is reported for 80% of coefficients and a confidence interval for 16%; the exact ICC form is often unstated in the papers themselves. Headline medians above use only independent inter-rater agreement on submission scores, outside intervention arms. Self-authored studies are included and flagged. Method after Vacha-Haase (1998) and Bornmann, Mutz and Daniel (2010); n = number of objects rated, k = reviewers per object.