grantscience.com

Codings, summaries and estimates on this site are AI-assisted and provisional, pending verification against primary sources. About the data

grantscience.com · reliability, coded

It depends who you count, and how

Coding each study’s exact coefficient and sample makes the reliability estimates comparable. The headline is not one number but a fork: a single reviewer’s score is barely reliable, the panel average looks respectable, and most of the variation sits between studies.
175
studies coded in full text
0.28
median single-reviewer ICC (285 coefficients)
0.55
median panel-average ICC (116 coefficients)

The same review process yields a median ICC of 0.28 for one reviewer but 0.55 once a panel is averaged, because averaging raters raises reliability (the Spearman-Brown relationship). Chance-corrected kappa sits near 0.19. Reporting a single “reliability of peer review” without saying single reviewer or panel is close to meaningless, which is why every coefficient below is coded to its exact estimand.

Read with care: every measurement was extracted from the paper’s full text by two independent AI coders, and their disagreements were adjudicated by a third model against the source, with each value carrying a verbatim quote and locator. Human verification of the rows has not yet been done: values are provisional and to be verified before citing. A sample size is reported for 80% of coefficients and a confidence interval for 16%; the exact ICC form is often unstated in the papers themselves. Headline medians above use only independent inter-rater agreement on submission scores, outside intervention arms. Self-authored studies are included and flagged. Method after Vacha-Haase (1998) and Bornmann, Mutz and Daniel (2010); n = number of objects rated, k = reviewers per object.