grantscience.com

Codings, summaries and estimates on this site are AI-assisted and provisional, pending verification against primary sources. About the data

grantscience.com · about

About this site

grantscience.com is the public presentation of a doctoral synthesis of the research literature on the reliability of peer review, grant review first, with journal review and neighbouring judgment contexts as the comparison. It accompanies the PhD thesis When Experts Disagree: The Reliability of Grant Peer Review (University of Oslo).

The corpus

The evidence base is a systematic scoping corpus assembled under PRISMA-ScR: 9,783 records identified, then 9,242 screened (dual AI screen, kappa 0.866; recovered arm 0.890), then 1,627 included, then 1,123 in the corpus after de-duplication and retrieval. Each paper carries its full text in the source pipeline, a short AI-generated summary with a full structured digest behind it (what it is, setting and sample, key findings, argument, limitations, reported effect sizes), and a set of facet codings: review context, research construct, study design, bias subtype and review-process stage.

The two organising facets re-express the screening themes around the decomposition of judgment error into bias and noise (Kahneman, Sibony and Sunstein 2021): systematic error, random disagreement, validity and the interventions meant to reduce error. The grant subset is the default view everywhere on the site because the thesis is about grant review; the other contexts are the comparison.

IDENTIFICATIONSCREENINGELIGIBILITYINCLUDEDRecords identified fromdatabases and registersn = 6,906Embase 5,025 · MEDLINE 1,021 · PsycINFO 458 · OSF 402Records identified fromother methodsn = 2,877Citation chasing 2,709 · Consensus 116 · Organisational websites 52Duplicate records removedbefore screeningn = 541Records screened on title and abstractn = 9,242dual independent AI screen, Cohen's kappa 0.866Records excluded at title / abstractn = 7,615Reports sought for retrievaln = 1,6291,627 screening includes + 2 added during citation chasingRemoved before full-text assessmentDuplicate reports n = 69Not retrieved n = 176(OSF registration only 44; unavailable 94; recovered arm: unobtainable or not a real publication 38)Reports assessed for eligibility (full text)n = 1,384Reports excluded at full text n = 254wrong focus 143 · not peer review 68no substance 36 · analogue, not illustrative 7plus 7 reports with no readable textIncluded in the corpusn = 1,123
PRISMA-ScR flow of records from search to corpus (searches run 2026-06-29, final selection 2026-07-04). Title and abstract screening was a dual independent AI screen (two passes over every record, disagreements adjudicated by a third model): 96.0per cent raw agreement, Cohen’s kappa 0.866 (recovered arm 0.890), 670 records adjudicated. A full-text double screen on 134 overlapping reports reached kappa 0.742. Source: kappe-scoping data/funnel_counts.json (tools/reconcile_funnel.py), 2026-07-04. Provisional.

Methods behind the views

Charting and distributions(the overview) follow scoping-review charting practice (Arksey and O’Malley 2005; Tricco et al. 2018).

The evidence gap map crosses review-process stage with construct after Snilstveit et al. (2016) and White et al. (2020). Bubble colour is the share of controlled experiments in the cell: a description of study design, not a judgement of quality.

The reference network renders all 1,123 papers with standard bibliometric relations (Zupic and Cater 2015): 8,422 in-corpus citations, 29,030 bibliographic-coupling links (three or more shared references) and 15,756 co-citation links. Layouts are precomputed (force-directed placement and UMAP), communities via Louvain, inductive topics via BERTopic (Grootendorst 2022). References out of the corpus resolve through OpenAlex (17,533 of 19,108 works).

The reliability dataset locates every reliability and agreement coefficient the corpus papers report as their own result (1,960 measurements from 175 studies) and codes each to its exact form, sample, estimand and design, in the tradition of reliability generalisation (Vacha-Haase 1998) and extending Bornmann, Mutz and Daniel (2010) from journal to grant review. Each paper’s full text was coded independently by two AI models; their disagreements were adjudicated by a third model against the source text, and every value carries a verbatim quote and locator. Medians shown on the site are computed from the adjudicated table at build time, using only independent inter-rater agreement on submission scores; the pooled estimate comes from a random-effects model (metafor; Viechtbauer 2010) and is illustrative.

Why everything says provisional

The codings, summaries and estimates on this site are AI-assisted. Context, construct and study design are the majority vote of a five-model coding panel reading titles and abstracts, with book and monograph labels set from library metadata and panel ties adjudicated by hand; process stages and bias subtypes come from separate AI passes; and every paper summary is original AI-generated text (never the publisher’s abstract). The panel replaced the earlier single-pass coding, but the blind human validation of a sample of the corpus is still pending, so labels and counts may change.

The reliability measurements were dual-coded from full text and adjudicated as described above, and each row was cross-checked against published meta-analyses where possible, but the author’s own row-by-row verification against the primary sources has not yet been carried out. It is the next step on the project’s task list; until it is done, every measurement on the reliability page is published as unverified and the table’s verification column says so explicitly.

Treat every number as a well-founded estimate rather than a citable result. Self-authored studies (the thesis author’s own papers) are part of the corpus and are flagged wherever they appear. Before anything on this site enters the thesis or another publication, it is checked against the source’s full text.

The data

Every count, label, colour and median on the site derives from generated data files; nothing is hand-typed. The payloads are open to inspect:

Generated 2026-07-24 from the kappe-synthesis analysis pipeline, which remains the source of truth. If you reuse the data, carry the provisional flag with it and cite the primary sources, not the summaries.

Corrections and changes

Substantive changes to the data or methods behind this site are logged here, newest first. Each names the decision or commit that made it.

  • 2026-07-23 · data · DECISIONS.md D35, D36 and D37
    Duplicate reports linked; two full texts re-attributed

    Two preprint/published pairs inside the corpus (Bieri 2020/2021 and Simsek 2023/2024) are now linked as reports of the same study, following PRISMA 2020: the corpus counts 1,121 included studies described by 1,123 reports, and the preprints' coefficient rows were removed from the reliability dataset so no study is counted twice (now 1,960 coefficients from 175 studies). Separately, the stored full texts for two papers (Pina et al. 2015's eLife companion by Pina et al. 2021, and Carpenter et al. 2015) had been swapped, so their extracted coefficients were attributed to each other's papers; the attribution is corrected. All reliability rows remain provisional pending row-level human verification.

  • 2026-07-23 · data · DECISIONS.md D33 and D34
    Corpus corrected from 935 to 1,123 papers

    A de-duplication step in the scoping search had removed 188 eligible records without ever screening them: the record kept as the master of each near-duplicate cluster never entered the screen. Those 188 were recovered and admitted, so the corpus grows from 935 to 1,123. The PRISMA funnel is corrected with it: the pre-screening de-duplication box falls from 1,856 to 541 (1,315 of those removals were not duplicates) and every later box grows to match. The direction of the reliability findings is unchanged; the recovered set includes several highly cited papers, among them Wenneras and Wold 1997 and Mahoney 1977.

  • 2026-07-23 · display · grantscience site commits, 2026-07-23
    Reliability and screening headline refinements

    Removed the pooled single-reviewer ICC point estimate (0.42), which came from the superseded first-harvest analysis; the reliability page keeps its per-estimand median tiles instead. The screening-agreement figure now also reports the recovered arm's re-screen at Cohen's kappa 0.890 alongside the original 0.866.

Decisions on record

The synthesis runs on a dated methodological decision log. The ones most visible here: the grant-first default; colouring the gap map by study design rather than a quality label; publishing original AI-generated summaries instead of copyrighted abstracts; and the standard bibliometric definitions behind the network’s edge sets. The full log lives with the analysis pipeline and will accompany the thesis.

Colophon

Built by Jan-Ole Hesselberg (University of Oslo). The site is a static Next.js application; the network view uses sigma.js and graphology in 2D and Three.js in 3D; charts are hand-rendered SVG; type is set in your system’s fonts. No analytics, no cookies. Back to the overview.