grantscience.com · about
About this site
The corpus
The evidence base is a systematic scoping corpus assembled under PRISMA-ScR: 9,783 records identified, then 9,242 screened (dual AI screen, kappa 0.866; recovered arm 0.890), then 1,627 included, then 1,123 in the corpus after de-duplication and retrieval. Each paper carries its full text in the source pipeline, a short AI-generated summary with a full structured digest behind it (what it is, setting and sample, key findings, argument, limitations, reported effect sizes), and a set of facet codings: review context, research construct, study design, bias subtype and review-process stage.
The two organising facets re-express the screening themes around the decomposition of judgment error into bias and noise (Kahneman, Sibony and Sunstein 2021): systematic error, random disagreement, validity and the interventions meant to reduce error. The grant subset is the default view everywhere on the site because the thesis is about grant review; the other contexts are the comparison.
The search
All database and register searches were run on 2026-06-29; the final selection closed on 2026-07-04. Every search pairs a grant-focused backbone (grant review, proposal evaluation, research funding, funding decisions and related phrases) with a broad concept block covering reliability, agreement, variability, validity, bias, fairness and measurement terms. The generic phrase “peer review” counts only where it co-occurs with a grant, funding or proposal term, so the set does not fill with journal-review work; the journal literature enters through citation chasing instead. No language or publication-date restriction was applied, with two exceptions: forward citation chasing was floored at 2015 and the Consensus queries used recency slices.
The searches are reproducible from the strings below (the Consensus and organisational harvests are dated snapshots).
Embase · Ovid, UiO institutional access · 5,025 records
Non-MEDLINE records only (the Remove MEDLINE Records limit was on, since MEDLINE was searched separately). Concept block minus the polysemous 'variation'; 'variability' kept.
(("grant review*" or "grant peer review*" or "proposal review*" or "proposal evaluation*"
or "proposal assessment" or "grant application*" or "grant proposal*" or "research funding"
or "funding decision*" or "grant selection").tw.
or (("peer review"/ or (peer review* or peer-review* or referee or referees or refereeing).tw.)
and (grant or grants or funding or proposal or proposals or applicant or applicants).tw.))
and (reliab* or inter-rater or interrater or "inter rater" or inter-reviewer or agreement
or disagree* or concordance or discordance or consisten* or variability
or "intraclass correlation" or ICC or kappa or "average deviation" or validity
or "predictive validity" or bias* or fairness or "review quality").tw.MEDLINE · NCBI PubMed, E-utilities · 1,021 records
(
"grant review"[tiab] OR "grant peer review"[tiab] OR "proposal review"[tiab]
OR "proposal evaluation"[tiab] OR "proposal assessment"[tiab] OR "grant application"[tiab]
OR "grant applications"[tiab] OR "grant proposal"[tiab] OR "grant proposals"[tiab]
OR "research funding"[tiab] OR "funding decision"[tiab] OR "funding decisions"[tiab]
OR "grant selection"[tiab]
OR (
("peer review"[tiab] OR "Peer Review"[Mesh] OR referee[tiab] OR referees[tiab])
AND (grant[tiab] OR grants[tiab] OR funding[tiab] OR proposal[tiab] OR proposals[tiab]
OR applicant[tiab] OR applicants[tiab])
)
)
AND
(
"Reproducibility of Results"[Mesh] OR "Observer Variation"[Mesh]
OR reliability[tiab] OR reliable[tiab] OR unreliable[tiab] OR unreliability[tiab]
OR "inter-rater"[tiab] OR interrater[tiab] OR "inter rater"[tiab] OR "inter-reviewer"[tiab]
OR interreviewer[tiab] OR agreement[tiab] OR disagreement[tiab] OR disagree[tiab]
OR concordance[tiab] OR discordance[tiab] OR consistency[tiab] OR inconsistency[tiab]
OR variability[tiab] OR variation[tiab] OR "intraclass correlation"[tiab] OR ICC[tiab]
OR kappa[tiab] OR "average deviation"[tiab] OR validity[tiab] OR "predictive validity"[tiab]
OR bias[tiab] OR biases[tiab] OR biased[tiab] OR fairness[tiab]
OR "review quality"[tiab] OR "quality of peer review"[tiab]
)PsycINFO · Ovid, UiO institutional access · 458 records
459 exported; one within-library duplicate removed at import. Mapping to subject headings off; full concept block.
(("grant review*" or "grant peer review*" or "proposal review*" or "proposal evaluation*"
or "proposal assessment" or "grant application*" or "grant proposal*" or "research funding"
or "funding decision*" or "grant selection").tw.
or ((Peer Evaluation/ or (peer review* or peer-review* or referee or referees or refereeing).tw.)
and (grant or grants or funding or proposal or proposals or applicant or applicants).tw.))
and (reliab* or inter-rater or interrater or "inter rater" or inter-reviewer or agreement
or disagree* or concordance or discordance or consisten* or variability or variation
or "intraclass correlation" or ICC or kappa or "average deviation" or validity
or "predictive validity" or bias* or fairness or "review quality").tw.OSF · OSF API v2: registrations, preprints, projects · 402 records
The API matches case-insensitive substrings in titles and descriptions rather than tokenised phrases, so broad one and two-word terms were used.
Substring terms: peer review, grant peer review, inter-rater reliability, reviewer agreement, reviewer disagreement, funding decision
Citation chasing · OpenAlex: backward via reference lists, forward via citing works · 2,709 records
2,633 unique OpenAlex records plus 76 from the reference list of the Technopolis 'Review of Peer Review' report (mined manually, backward only). Forward chasing floored at 2015-01-01; backward chasing unrestricted.
Seed works (17): Lee et al. (2013) doi; Guthrie et al. (2017) doi; Bornmann et al. (2010) doi; Cicchetti (1991) doi; Marsh et al. (2008) doi; Mutz et al. (2012) doi; Jayasinghe et al. (2003) doi; Li and Agha (2015) doi; Fang et al. (2016) doi; Pina et al. (2021) doi; Seeber et al. (2021) doi; Erosheva et al. (2021) doi; Snell (2015) doi; Recio-Saucedo et al. (2022) doi; Hesselberg et al. (Paper 1) (2021) doi; Hesselberg et al. (Paper 2) (2023) doi; Hesselberg et al. (Paper 3) (2025) doi
Consensus · Consensus (Semantic Scholar, Scopus, PubMed, arXiv); operationalised via OpenAlex for DOI export · 116 records
Dated snapshot, not a rerunnable search: one focus per query, natural language.
- inter-rater reliability and agreement in grant peer review (from 2022)
- factors and predictors of reviewer disagreement in grant peer review (from 2022)
- validity and predictive value of grant peer review for research outcomes
- lottery and partial randomisation in research funding allocation
- artificial intelligence and large language models in grant peer review (from 2022)
- statistical methods for inter-rater agreement ICC kappa average deviation in evaluation
Organisational websites · Website harvesting, dated snapshot · 52 records
- Research on Research Institute (RoRI): working papers, reports, journal articles, insights
- UK Metascience Unit (UKRI publications)
Methods behind the views
Charting and distributions(the overview) follow scoping-review charting practice (Arksey and O’Malley 2005; Tricco et al. 2018).
The evidence gap map crosses review-process stage with construct after Snilstveit et al. (2016) and White et al. (2020). Bubble colour is the share of controlled experiments in the cell: a description of study design, not a judgement of quality.
The reference network renders all 1,123 papers with standard bibliometric relations (Zupic and Cater 2015): 8,422 in-corpus citations, 29,030 bibliographic-coupling links (three or more shared references) and 15,756 co-citation links. Layouts are precomputed (force-directed placement and UMAP), communities via Louvain, inductive topics via BERTopic (Grootendorst 2022). References out of the corpus resolve through OpenAlex (17,533 of 19,108 works).
The reliability dataset locates every reliability and agreement coefficient the corpus papers report as their own result (1,960 measurements from 175 studies) and codes each to its exact form, sample, estimand and design, in the tradition of reliability generalisation (Vacha-Haase 1998) and extending Bornmann, Mutz and Daniel (2010) from journal to grant review. Each paper’s full text was coded independently by two AI models; their disagreements were adjudicated by a third model against the source text, and every value carries a verbatim quote and locator. Medians shown on the site are computed from the adjudicated table at build time, using only independent inter-rater agreement on submission scores; the pooled estimate comes from a random-effects model (metafor; Viechtbauer 2010) and is illustrative.
Why everything says provisional
The codings, summaries and estimates on this site are AI-assisted. Context, construct and study design are the majority vote of a five-model coding panel reading titles and abstracts, with book and monograph labels set from library metadata and panel ties adjudicated by hand; process stages and bias subtypes come from separate AI passes; and every paper summary is original AI-generated text (never the publisher’s abstract). The panel replaced the earlier single-pass coding, but the blind human validation of a sample of the corpus is still pending, so labels and counts may change.
The reliability measurements were dual-coded from full text and adjudicated as described above, and each row was cross-checked against published meta-analyses where possible, but the author’s own row-by-row verification against the primary sources has not yet been carried out. It is the next step on the project’s task list; until it is done, every measurement on the reliability page is published as unverified and the table’s verification column says so explicitly.
Treat every number as a well-founded estimate rather than a citable result. Self-authored studies (the thesis author’s own papers) are part of the corpus and are flagged wherever they appear. Before anything on this site enters the thesis or another publication, it is checked against the source’s full text.
The data
Every count, label, colour and median on the site derives from generated data files; nothing is hand-typed. The payloads are open to inspect:
- /data/corpus.json (the corpus with facets and summaries)
- /data/network/core.json (network nodes)
- /data/network/edges.json (citation, coupling and co-citation edges)
- /data/network/layouts.json (the five layout coordinate sets)
- /data/network/layouts3d.json (the same five layouts in 3D)
- /data/reliability.json (the coded reliability coefficients)
/data/nodes/<key>.json(per-paper detail: coefficients, neighbours, external references)/data/digest/<key>.json(per-paper full digests)
Generated 2026-07-24 from the kappe-synthesis analysis pipeline, which remains the source of truth. If you reuse the data, carry the provisional flag with it and cite the primary sources, not the summaries.
Corrections and changes
Substantive changes to the data or methods behind this site are logged here, newest first. Each names the decision or commit that made it.
- 2026-07-23 · data · DECISIONS.md D35, D36 and D37Duplicate reports linked; two full texts re-attributed
Two preprint/published pairs inside the corpus (Bieri 2020/2021 and Simsek 2023/2024) are now linked as reports of the same study, following PRISMA 2020: the corpus counts 1,121 included studies described by 1,123 reports, and the preprints' coefficient rows were removed from the reliability dataset so no study is counted twice (now 1,960 coefficients from 175 studies). Separately, the stored full texts for two papers (Pina et al. 2015's eLife companion by Pina et al. 2021, and Carpenter et al. 2015) had been swapped, so their extracted coefficients were attributed to each other's papers; the attribution is corrected. All reliability rows remain provisional pending row-level human verification.
- 2026-07-23 · data · DECISIONS.md D33 and D34Corpus corrected from 935 to 1,123 papers
A de-duplication step in the scoping search had removed 188 eligible records without ever screening them: the record kept as the master of each near-duplicate cluster never entered the screen. Those 188 were recovered and admitted, so the corpus grows from 935 to 1,123. The PRISMA funnel is corrected with it: the pre-screening de-duplication box falls from 1,856 to 541 (1,315 of those removals were not duplicates) and every later box grows to match. The direction of the reliability findings is unchanged; the recovered set includes several highly cited papers, among them Wenneras and Wold 1997 and Mahoney 1977.
- 2026-07-23 · display · grantscience site commits, 2026-07-23Reliability and screening headline refinements
Removed the pooled single-reviewer ICC point estimate (0.42), which came from the superseded first-harvest analysis; the reliability page keeps its per-estimand median tiles instead. The screening-agreement figure now also reports the recovered arm's re-screen at Cohen's kappa 0.890 alongside the original 0.866.
Decisions on record
The synthesis runs on a dated methodological decision log. The ones most visible here: the grant-first default; colouring the gap map by study design rather than a quality label; publishing original AI-generated summaries instead of copyrighted abstracts; and the standard bibliometric definitions behind the network’s edge sets. The full log lives with the analysis pipeline and will accompany the thesis.
Colophon
Built by Jan-Ole Hesselberg (University of Oslo). The site is a static Next.js application; the network view uses sigma.js and graphology in 2D and Three.js in 3D; charts are hand-rendered SVG; type is set in your system’s fonts. No analytics, no cookies. Back to the overview.