{"rows":[{"key":"6754BVQM","au":"Andejeski, Yvonne","y":2002,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement (same rounded overall score) between the all-participant final score and the scientist-only score","estd":"percent agreement","v":0.762,"n":"2190","k":"","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1.0-5.0 one decimal, 1.0 = outstanding to 5.0 = acceptable","field":"biomedical (breast cancer research)","wr":"final all-member score vs scientist-only score on proposals","conf":"high","self":false,"doi":"10.1089/152460902317586010","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"For 2190 proposals, the final rounded score computed from all panel members' votes matched the score computed from scientists' votes only for 76.2 percent of proposals, showing consumer votes rarely changed the outcome.","vf":"unverified"},{"key":"6754BVQM","au":"Andejeski, Yvonne","y":2002,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson product-moment correlation coefficient","estd":"correlation","v":0.9492,"n":"2190","k":"","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1.0-5.0 one decimal, 1.0 = outstanding to 5.0 = acceptable","field":"biomedical (breast cancer research)","wr":"consumers and scientists on grant proposal scores","conf":"high","self":false,"doi":"10.1089/152460902317586010","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":false,"ms":"Across 2190 breast cancer research proposals, the mean score given by consumer panel members correlated very strongly (Pearson r of 0.9492) with the mean score given by scientist panel members, indicating the two groups ranked proposals almost identically. All members voted by anonymous ballot after shared panel discussion; scientists per proposal ranged from 4 to 21 and consumers from 1 to 2.","vf":"unverified"},{"key":"6754BVQM","au":"Andejeski, Yvonne","y":2002,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"percent of proposals with mean consumer score within 1 SD of mean scientist score (tolerance = 1 SD)","estd":"percent agreement","v":0.761,"n":"2190","k":"","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1.0-5.0 one decimal, 1.0 = outstanding to 5.0 = acceptable","field":"biomedical (breast cancer research)","wr":"consumers vs scientists on grant proposal scores","conf":"high","self":false,"doi":"10.1089/152460902317586010","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"For 76.1 percent of the 2190 proposals, the mean consumer score fell within one standard deviation of the mean scientist score, showing that the two assessor groups usually landed close together on the same proposals.","vf":"unverified"},{"key":"6754BVQM","au":"Andejeski, Yvonne","y":2002,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"percent of proposals with mean consumer score within 2 SD of mean scientist score (tolerance = 2 SD)","estd":"percent agreement","v":0.931,"n":"2190","k":"","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1.0-5.0 one decimal, 1.0 = outstanding to 5.0 = acceptable","field":"biomedical (breast cancer research)","wr":"consumers vs scientists on grant proposal scores","conf":"high","self":false,"doi":"10.1089/152460902317586010","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"For 93.1 percent of the 2190 proposals, the mean consumer score fell within two standard deviations of the mean scientist score, a wider tolerance version of the same closeness measure between the two assessor groups.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"other","form":"Mean of maximum difference between review scores per application","estd":"other","v":2.08,"n":"35","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external referee scores for LMS applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"External referees scored 35 advanced-career LMS applications. The mean maximum within-application score difference was 2.08 points.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"other","form":"Mean of standard deviation review scores per application","estd":"other","v":1.02,"n":"35","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external referee scores for LMS applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"External referees scored 35 advanced-career LMS applications. The mean within-application standard deviation was 1.02 points.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"other","form":"Mean of maximum difference between review scores per application","estd":"other","v":2.22,"n":"91","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external referee scores for grant applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Between two and six external referees scored each advanced-career application. The mean maximum within-application score difference was 2.22 points.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"other","form":"Mean of standard deviation review scores per application","estd":"other","v":1.06,"n":"91","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external referee scores for grant applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Between two and six external referees scored each advanced-career application. The mean within-application standard deviation was 1.06 points.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"other","form":"Mean of maximum difference between review scores per application","estd":"other","v":2.75,"n":"22","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external referee scores for SSH applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"External referees scored 22 advanced-career SSH applications. The mean maximum within-application score difference was 2.75 points.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"other","form":"Mean of standard deviation review scores per application","estd":"other","v":1.28,"n":"22","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external referee scores for SSH applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"External referees scored 22 advanced-career SSH applications. The mean within-application standard deviation was 1.28 points.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"other","form":"Mean of maximum difference between review scores per application","estd":"other","v":2.03,"n":"34","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external referee scores for STE applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"External referees scored 34 advanced-career STE applications. The mean maximum within-application score difference was 2.03 points.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"other","form":"Mean of standard deviation review scores per application","estd":"other","v":0.95,"n":"34","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external referee scores for STE applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"External referees scored 34 advanced-career STE applications. The mean within-application standard deviation was 0.95 points.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"correlation","form":"Kendall's tau","estd":"correlation","v":0.42,"n":"","k":"","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"committee scores before vs after interview (LMS)","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"The same LMS committee scored applicants before and after their interview; the two occasions correlate at Kendall's tau 0.42, a moderate stability of committee judgment across the interview phase.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"correlation","form":"Kendall's tau","estd":"correlation","v":0.62,"n":"","k":"","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"committee scores before vs after interview (SSH)","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"The same SSH committee scored applicants before and after their interview; the two occasions correlate at Kendall's tau 0.62, a fairly strong stability of committee judgment across the interview phase.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"correlation","form":"Kendall's tau","estd":"correlation","v":0.42,"n":"","k":"","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"committee scores before vs after interview (STE)","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"The same STE committee scored applicants before and after their interview; the two occasions correlate at Kendall's tau 0.42, a moderate stability of committee judgment across the interview phase.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"other","form":"Mean of maximum difference between review scores per application","estd":"other","v":1.45,"n":"161","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external referee scores for LMS applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"External referees scored 161 early-career LMS applications. The mean maximum within-application score difference was 1.45 points.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"other","form":"Mean of standard deviation review scores per application","estd":"other","v":0.94,"n":"161","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external referee scores for LMS applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"External referees scored 161 early-career LMS applications. The mean within-application standard deviation was 0.94 points.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"other","form":"Mean of maximum difference between review scores per application","estd":"other","v":1.57,"n":"453","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external referee scores for grant applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Between two and six external referees scored each early-career application. The mean maximum within-application score difference was 1.57 points.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"other","form":"Mean of standard deviation review scores per application","estd":"other","v":1.05,"n":"453","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external referee scores for grant applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Between two and six external referees scored each early-career application. The mean within-application standard deviation was 1.05 points.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"other","form":"Mean of maximum difference between review scores per application","estd":"other","v":1.6,"n":"141","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external referee scores for SSH applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"External referees scored 141 early-career SSH applications. The mean maximum within-application score difference was 1.60 points.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"other","form":"Mean of standard deviation review scores per application","estd":"other","v":1.13,"n":"141","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external referee scores for SSH applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"External referees scored 141 early-career SSH applications. The mean within-application standard deviation was 1.13 points.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"other","form":"Mean of maximum difference between review scores per application","estd":"other","v":1.68,"n":"151","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external referee scores for STE applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"External referees scored 151 early-career STE applications. The mean maximum within-application score difference was 1.68 points.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"other","form":"Mean of standard deviation review scores per application","estd":"other","v":1.09,"n":"151","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external referee scores for STE applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"External referees scored 151 early-career STE applications. The mean within-application standard deviation was 1.09 points.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"correlation","form":"Kendall's tau","estd":"correlation","v":0.36,"n":"","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external reviews vs committee research-impact score (LMS)","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"External review scores correlate with the LMS committee's research-impact scores at Kendall's tau 0.36, a weak-to-moderate agreement between external reviewers and the committee on this criterion.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"correlation","form":"Kendall's tau","estd":"correlation","v":0.22,"n":"","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external reviews vs committee research-impact score (SSH)","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"External review scores correlate with the SSH committee's research-impact scores at Kendall's tau 0.22, a weak agreement between external reviewers and the committee on this criterion.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"correlation","form":"Kendall's tau","estd":"correlation","v":0.29,"n":"","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external reviews vs committee research-impact score (STE)","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"External review scores correlate with the STE committee's research-impact scores at Kendall's tau 0.29, a weak agreement between external reviewers and the committee on this criterion.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"correlation","form":"Kendall's tau","estd":"correlation","v":0.64,"n":"","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external reviews vs committee quality-of-proposal score (LMS)","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"External review scores correlate with the LMS committee's quality-of-proposal scores at Kendall's tau 0.64, the strongest external-committee agreement, higher than for the other criteria.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"correlation","form":"Kendall's tau","estd":"correlation","v":0.55,"n":"","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external reviews vs committee quality-of-proposal score (SSH)","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"External review scores correlate with the SSH committee's quality-of-proposal scores at Kendall's tau 0.55, the strongest external-committee agreement among the criteria.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"correlation","form":"Kendall's tau","estd":"correlation","v":0.55,"n":"","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external reviews vs committee quality-of-proposal score (STE)","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"External review scores correlate with the STE committee's quality-of-proposal scores at Kendall's tau 0.55, the strongest external-committee agreement among the criteria.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"correlation","form":"Kendall's tau","estd":"correlation","v":0.32,"n":"","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external reviews vs committee quality-of-researcher score (LMS)","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"External review scores correlate with the LMS committee's quality-of-researcher scores at Kendall's tau 0.32, a weak agreement; only the LMS value is reported for this criterion.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"correlation","form":"Kendall's tau","estd":"correlation","v":0.53,"n":"","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external reviewers vs committee on grant applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"External referees and the committee each scored ACG career-grant applications; their scores correlate at Kendall's tau 0.53, a moderate agreement between the two assessor bodies.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"correlation","form":"Kendall's tau","estd":"correlation","v":0.53,"n":"","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external reviewers vs committee on grant applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"External referees and the review committee each scored ECG career-grant applications; their standardised scores correlate at Kendall's tau 0.53, a moderate agreement between the two assessor bodies.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"correlation","form":"Kendall's tau","estd":"correlation","v":0.52,"n":"","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external reviewers vs committee on grant applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":false,"ms":"External referees and the committee each scored ICG career-grant applications; their scores correlate at Kendall's tau 0.52, the value the authors foreground as the representative reviewer-committee agreement. No single pooled value across programmes is reported.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"other","form":"Mean of maximum difference between review scores per application","estd":"other","v":2.18,"n":"118","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external referee scores for LMS applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"External referees scored 118 intermediate-career LMS applications. The mean maximum within-application score difference was 2.18 points.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"other","form":"Mean of standard deviation review scores per application","estd":"other","v":1.16,"n":"118","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external referee scores for LMS applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"External referees scored 118 intermediate-career LMS applications. The mean within-application standard deviation was 1.16 points.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"other","form":"Mean of maximum difference between review scores per application","estd":"other","v":2.06,"n":"353","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external referee scores for grant applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Between two and six external referees scored each intermediate-career application. The mean maximum within-application score difference was 2.06 points.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"other","form":"Mean of standard deviation review scores per application","estd":"other","v":1.1,"n":"353","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external referee scores for grant applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Between two and six external referees scored each intermediate-career application. The mean within-application standard deviation was 1.10 points.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"other","form":"Mean of maximum difference between review scores per application","estd":"other","v":2.25,"n":"111","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external referee scores for SSH applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"External referees scored 111 intermediate-career SSH applications. The mean maximum within-application score difference was 2.25 points.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"other","form":"Mean of standard deviation review scores per application","estd":"other","v":1.21,"n":"111","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external referee scores for SSH applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"External referees scored 111 intermediate-career SSH applications. The mean within-application standard deviation was 1.21 points.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"other","form":"Mean of maximum difference between review scores per application","estd":"other","v":1.76,"n":"124","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external referee scores for STE applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"External referees scored 124 intermediate-career STE applications. The mean maximum within-application score difference was 1.76 points.","vf":"unverified"},{"key":"RVJKZWBR","au":"Arensbergen, P. van","y":2012,"cx":"Grant","ob":"fellowship","fam":"other","form":"Mean of standard deviation review scores per application","estd":"other","v":0.95,"n":"124","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (highest) to 6 (lowest)","field":"multi-field","wr":"external referee scores for STE applications","conf":"med","self":false,"doi":"10.1057/hep.2012.15","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"External referees scored 124 intermediate-career STE applications. The mean within-application standard deviation was 0.95 points.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Kendall's tau rank order correlation","estd":"correlation","v":0.59,"n":"","k":"","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"panel criteria scores against each other","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"The panel's own scores for quality of the proposal and quality of the researcher on the same applications correlated 0.59 after the interview in the LMS domain, so the two criteria measure related but distinct things.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Kendall's tau rank order correlation","estd":"correlation","v":0.44,"n":"","k":"","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"panel criteria scores against each other","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"The panel's own scores for quality of the proposal and quality of the researcher on the same applications correlated 0.44 before the interview in the LMS domain, so the two criteria measure related but distinct things.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Kendall's tau rank order correlation","estd":"correlation","v":0.57,"n":"","k":"","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"panel criteria scores against each other","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"The panel's own scores for quality of the proposal and quality of the researcher on the same applications correlated 0.57 after the interview in the SSH domain, so the two criteria measure related but distinct things.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Kendall's tau rank order correlation","estd":"correlation","v":0.5,"n":"","k":"","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"panel criteria scores against each other","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"The panel's own scores for quality of the proposal and quality of the researcher on the same applications correlated 0.50 before the interview in the SSH domain, so the two criteria measure related but distinct things.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Kendall's tau rank order correlation","estd":"correlation","v":0.37,"n":"","k":"","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"panel criteria scores against each other","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"The panel's own scores for research impact and quality of the researcher on the same applications correlated 0.37 after the interview in the SSH domain, so the two criteria measure related but distinct things.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Kendall's tau rank order correlation","estd":"correlation","v":0.33,"n":"","k":"","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"panel criteria scores against each other","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"The panel's own scores for research impact and quality of the researcher on the same applications correlated 0.33 before the interview in the SSH domain, so the two criteria measure related but distinct things.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Kendall's tau rank order correlation","estd":"correlation","v":0.53,"n":"91","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees versus panel on applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"For advanced career grant applications the standardised average score of the external referees (typically 4 per application) correlated 0.53 with the first panel score of the same applications, a moderate rank order agreement between the two review phases.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Kendall's tau rank order correlation","estd":"correlation","v":0.53,"n":"453","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees versus panel on applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"For early career grant applications the standardised average score of the external referees (typically 2 per application) correlated 0.53 with the first panel score of the same applications, a moderate rank order agreement between the two review phases.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Kendall's tau rank order correlation","estd":"correlation","v":0.52,"n":"353","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees versus panel on applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"For intermediate career grant applications the standardised average score of the external referees (typically 3 per application) correlated 0.52 with the first panel score of the same applications, a moderate rank order agreement between the two review phases.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Kendall's tau rank order correlation","estd":"correlation","v":0.64,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"referee scores versus panel criterion scores","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"In the LMS domain the standardised external referee scores correlated 0.64 with the first panel scores for the quality of the proposal criterion, showing how far referee judgments carry over to that particular panel criterion.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Kendall's tau rank order correlation","estd":"correlation","v":0.55,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"referee scores versus panel criterion scores","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"In the SSH domain the standardised external referee scores correlated 0.55 with the first panel scores for the quality of the proposal criterion, showing how far referee judgments carry over to that particular panel criterion.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Kendall's tau rank order correlation","estd":"correlation","v":0.55,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"referee scores versus panel criterion scores","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"In the STE domain the standardised external referee scores correlated 0.55 with the first panel scores for the quality of the proposal criterion, showing how far referee judgments carry over to that particular panel criterion.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Kendall's tau rank order correlation","estd":"correlation","v":0.32,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"referee scores versus panel criterion scores","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"In the LMS domain the standardised external referee scores correlated 0.32 with the first panel scores for the quality of the researcher criterion, showing how far referee judgments carry over to that particular panel criterion.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Kendall's tau rank order correlation","estd":"correlation","v":0.36,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"referee scores versus panel criterion scores","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"In the LMS domain the standardised external referee scores correlated 0.36 with the first panel scores for the research impact criterion, showing how far referee judgments carry over to that particular panel criterion.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Kendall's tau rank order correlation","estd":"correlation","v":0.22,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"referee scores versus panel criterion scores","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"In the SSH domain the standardised external referee scores correlated 0.22 with the first panel scores for the research impact criterion, showing how far referee judgments carry over to that particular panel criterion.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Kendall's tau rank order correlation","estd":"correlation","v":0.29,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"referee scores versus panel criterion scores","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"In the STE domain the standardised external referee scores correlated 0.29 with the first panel scores for the research impact criterion, showing how far referee judgments carry over to that particular panel criterion.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"percentage of grant decisions that would change under first panel scores","estd":"other","v":0.22,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"granted / rejected","field":"multi-field","wr":"first panel round versus final funding decisions","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"editor-decisions","rr":"restricted-top","pr":false,"he":false,"ms":"Had the panel scores from before the interview decided funding, 22 per cent of the grants would have gone to applicants who were in fact rejected, showing how much the second panel round changed the outcome.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"percentage of grant decisions that would change under referee scores","estd":"other","v":0.26,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"granted / rejected","field":"multi-field","wr":"referee ranking versus final funding decisions","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"restricted-other","pr":true,"he":false,"ms":"Twenty six per cent of the applicants would have had a different funding outcome if the external referee scores alone had decided, showing limited correspondence between referee judgments and the panel's allocation decisions.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"percentage of interview invitations that would change under referee scores","estd":"other","v":0.17,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"invited to interview / not invited","field":"multi-field","wr":"referee ranking versus panel interview decisions","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"restricted-other","pr":false,"he":false,"ms":"Seventeen per cent of the applicants actually invited to interview, 48 people, would not have been invited had the external referee scores decided, a measure of how far panel decisions depart from referee judgments.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"mean of maximum difference between review scores per application","estd":"other","v":2.22,"n":"91","k":"4","samp":"full-pool","blind":"single","agg":"single-rater","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees on career grant applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Typically 4 external referees scored each advanced career grant application on a six point scale in all domains, and the gap between the highest and lowest referee score per application averaged 2.22. Larger values mean more disagreement between referees.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"mean of maximum difference between review scores per application","estd":"other","v":2.08,"n":"35","k":"4","samp":"full-pool","blind":"single","agg":"single-rater","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees on career grant applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Typically 4 external referees scored each advanced career grant application on a six point scale in the LMS domain, and the gap between the highest and lowest referee score per application averaged 2.08. Larger values mean more disagreement between referees.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"mean of maximum difference between review scores per application","estd":"other","v":2.75,"n":"22","k":"4","samp":"full-pool","blind":"single","agg":"single-rater","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees on career grant applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Typically 4 external referees scored each advanced career grant application on a six point scale in the SSH domain, and the gap between the highest and lowest referee score per application averaged 2.75. Larger values mean more disagreement between referees.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"mean of maximum difference between review scores per application","estd":"other","v":2.03,"n":"34","k":"4","samp":"full-pool","blind":"single","agg":"single-rater","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees on career grant applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Typically 4 external referees scored each advanced career grant application on a six point scale in the STE domain, and the gap between the highest and lowest referee score per application averaged 2.03. Larger values mean more disagreement between referees.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"mean of maximum difference between review scores per application","estd":"other","v":1.57,"n":"453","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees on career grant applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Typically 2 external referees scored each early career grant application on a six point scale in all domains, and the gap between the highest and lowest referee score per application averaged 1.57. Larger values mean more disagreement between referees.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"mean of maximum difference between review scores per application","estd":"other","v":1.45,"n":"161","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees on career grant applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Typically 2 external referees scored each early career grant application on a six point scale in the LMS domain, and the gap between the highest and lowest referee score per application averaged 1.45. Larger values mean more disagreement between referees.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"mean of maximum difference between review scores per application","estd":"other","v":1.6,"n":"141","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees on career grant applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Typically 2 external referees scored each early career grant application on a six point scale in the SSH domain, and the gap between the highest and lowest referee score per application averaged 1.60. Larger values mean more disagreement between referees.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"mean of maximum difference between review scores per application","estd":"other","v":1.68,"n":"151","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees on career grant applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Typically 2 external referees scored each early career grant application on a six point scale in the STE domain, and the gap between the highest and lowest referee score per application averaged 1.68. Larger values mean more disagreement between referees.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"mean of maximum difference between review scores per application","estd":"other","v":2.06,"n":"353","k":"3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees on career grant applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Typically 3 external referees scored each intermediate career grant application on a six point scale in all domains, and the gap between the highest and lowest referee score per application averaged 2.06. Larger values mean more disagreement between referees.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"mean of maximum difference between review scores per application","estd":"other","v":2.18,"n":"118","k":"3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees on career grant applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Typically 3 external referees scored each intermediate career grant application on a six point scale in the LMS domain, and the gap between the highest and lowest referee score per application averaged 2.18. Larger values mean more disagreement between referees.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"mean of maximum difference between review scores per application","estd":"other","v":2.25,"n":"111","k":"3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees on career grant applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Typically 3 external referees scored each intermediate career grant application on a six point scale in the SSH domain, and the gap between the highest and lowest referee score per application averaged 2.25. Larger values mean more disagreement between referees.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"mean of maximum difference between review scores per application","estd":"other","v":1.76,"n":"124","k":"3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees on career grant applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Typically 3 external referees scored each intermediate career grant application on a six point scale in the STE domain, and the gap between the highest and lowest referee score per application averaged 1.76. Larger values mean more disagreement between referees.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"mean of standard deviation of review scores per application","estd":"other","v":1.06,"n":"91","k":"4","samp":"full-pool","blind":"single","agg":"single-rater","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees on career grant applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Typically 4 external referees scored each advanced career grant application on a six point scale in all domains, and the standard deviation of the referee scores within an application averaged 1.06. Larger values mean more disagreement between referees.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"mean of standard deviation of review scores per application","estd":"other","v":1.02,"n":"35","k":"4","samp":"full-pool","blind":"single","agg":"single-rater","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees on career grant applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Typically 4 external referees scored each advanced career grant application on a six point scale in the LMS domain, and the standard deviation of the referee scores within an application averaged 1.02. Larger values mean more disagreement between referees.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"mean of standard deviation of review scores per application","estd":"other","v":1.28,"n":"22","k":"4","samp":"full-pool","blind":"single","agg":"single-rater","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees on career grant applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Typically 4 external referees scored each advanced career grant application on a six point scale in the SSH domain, and the standard deviation of the referee scores within an application averaged 1.28. Larger values mean more disagreement between referees.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"mean of standard deviation of review scores per application","estd":"other","v":0.95,"n":"34","k":"4","samp":"full-pool","blind":"single","agg":"single-rater","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees on career grant applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Typically 4 external referees scored each advanced career grant application on a six point scale in the STE domain, and the standard deviation of the referee scores within an application averaged 0.95. Larger values mean more disagreement between referees.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"mean of standard deviation of review scores per application","estd":"other","v":1.05,"n":"453","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees on career grant applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Typically 2 external referees scored each early career grant application on a six point scale in all domains, and the standard deviation of the referee scores within an application averaged 1.05. Larger values mean more disagreement between referees.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"mean of standard deviation of review scores per application","estd":"other","v":0.94,"n":"161","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees on career grant applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Typically 2 external referees scored each early career grant application on a six point scale in the LMS domain, and the standard deviation of the referee scores within an application averaged 0.94. Larger values mean more disagreement between referees.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"mean of standard deviation of review scores per application","estd":"other","v":1.13,"n":"141","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees on career grant applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Typically 2 external referees scored each early career grant application on a six point scale in the SSH domain, and the standard deviation of the referee scores within an application averaged 1.13. Larger values mean more disagreement between referees.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"mean of standard deviation of review scores per application","estd":"other","v":1.09,"n":"151","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees on career grant applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Typically 2 external referees scored each early career grant application on a six point scale in the STE domain, and the standard deviation of the referee scores within an application averaged 1.09. Larger values mean more disagreement between referees.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"mean of standard deviation of review scores per application","estd":"other","v":1.1,"n":"353","k":"3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees on career grant applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Typically 3 external referees scored each intermediate career grant application on a six point scale in all domains, and the standard deviation of the referee scores within an application averaged 1.10. Larger values mean more disagreement between referees.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"mean of standard deviation of review scores per application","estd":"other","v":1.16,"n":"118","k":"3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees on career grant applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Typically 3 external referees scored each intermediate career grant application on a six point scale in the LMS domain, and the standard deviation of the referee scores within an application averaged 1.16. Larger values mean more disagreement between referees.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"mean of standard deviation of review scores per application","estd":"other","v":1.21,"n":"111","k":"3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees on career grant applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Typically 3 external referees scored each intermediate career grant application on a six point scale in the SSH domain, and the standard deviation of the referee scores within an application averaged 1.21. Larger values mean more disagreement between referees.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"mean of standard deviation of review scores per application","estd":"other","v":0.95,"n":"124","k":"3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees on career grant applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Typically 3 external referees scored each intermediate career grant application on a six point scale in the STE domain, and the standard deviation of the referee scores within an application averaged 0.95. Larger values mean more disagreement between referees.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Kendall's tau rank order correlation","estd":"correlation","v":0.52,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"external referees versus panel on applications","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":false,"ms":"The paper's conclusion gives a Kendall's tau of .52 between the standardised external review scores and the panel scores of the same applications as its headline figure for how far panel judgments track referee judgments.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Kendall's tau rank order correlation","estd":"correlation","v":0.42,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"same panel scoring applicants twice","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"In the LMS domain the same panel scored each interviewed applicant before and after the interview, and the two rounds of panel scores correlated 0.42, so rankings shifted noticeably between the rounds.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Kendall's tau rank order correlation","estd":"correlation","v":0.62,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"same panel scoring applicants twice","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"In the SSH domain the same panel scored each interviewed applicant before and after the interview, and the two rounds of panel scores correlated 0.62, so rankings shifted noticeably between the rounds.","vf":"unverified"},{"key":"W5SHTMNM","au":"Arensbergen, P. van","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Kendall's tau rank order correlation","estd":"correlation","v":0.42,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"six point scale, excellent (1) to poor (6)","field":"multi-field","wr":"same panel scoring applicants twice","conf":"med","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"In the STE domain the same panel scored each interviewed applicant before and after the interview, and the two rounds of panel scores correlated 0.42, so rankings shifted noticeably between the rounds.","vf":"unverified"},{"key":"7JG4HLD9","au":"Baethge, Christopher","y":2013,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"Cohen's kappa for condensed 2x2 cross table (accept vs reject, all pairs)","estd":"kappa","v":0.16,"n":"206","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"acceptance (with/without revision) / rejection","field":"general medicine","wr":"reviewers on manuscript accept/reject recommendations","conf":"high","self":false,"doi":"10.1371/journal.pone.0061401","ciLow":0.059,"ciHigh":0.252,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Cohen's kappa on the condensed accept-versus-reject table was 0.16, still in the poor range.","vf":"unverified"},{"key":"7JG4HLD9","au":"Baethge, Christopher","y":2013,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"Cohen's kappa for 3x3 cross table (all possible reviewer pairs)","estd":"kappa","v":0.059,"n":"206","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / accept after revision / reject","field":"general medicine","wr":"reviewers on manuscript accept/reject recommendations","conf":"high","self":false,"doi":"10.1371/journal.pone.0061401","ciLow":-0.016,"ciHigh":0.134,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Chance-corrected Cohen's kappa across all 529 reviewer pairs on the three-category scale was 0.059, indicating poor agreement.","vf":"unverified"},{"key":"7JG4HLD9","au":"Baethge, Christopher","y":2013,"cx":"Journal","ob":"journal-manuscript","fam":"Fleiss-kappa","form":"Fleiss' kappa (accept vs reject); handles >2 reviewers and varying reviewer subsets","estd":"kappa","v":0.16,"n":"206","k":"2.7","samp":"full-pool","blind":"single","agg":"single-rater","scale":"acceptance (with/without revision) / rejection","field":"general medicine","wr":"reviewers on manuscript accept/reject recommendations","conf":"high","self":false,"doi":"10.1371/journal.pone.0061401","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"A mean of 2.7 referees (range 2 to 7) rated each of 206 manuscripts; Fleiss' kappa across all reviewer recommendations was 0.16, indicating poor chance-corrected agreement.","vf":"unverified"},{"key":"7JG4HLD9","au":"Baethge, Christopher","y":2013,"cx":"Journal","ob":"journal-manuscript","fam":"Fleiss-kappa","form":"Fleiss' Kappa","estd":"kappa","v":0.15,"n":"206","k":"2.7","samp":"full-pool","blind":"single","agg":"single-rater","scale":"acceptance (with/without revision) / rejection","field":"general medicine","wr":"reviewers on manuscript accept/reject recommendations","conf":"high","self":false,"doi":"10.1371/journal.pone.0061401","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Table 4 reports Fleiss' kappa for this study's accept-versus-reject data as 0.15, slightly below the 0.16 stated in the text; the discrepancy is unexplained.","vf":"unverified"},{"key":"7JG4HLD9","au":"Baethge, Christopher","y":2013,"cx":"Journal","ob":"journal-manuscript","fam":"Gwet-AC","form":"alternative kappa value (AC) by Gwet on the accept vs reject table","estd":"Gwet AC","v":0.63,"n":"206","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"acceptance (with/without revision) / rejection","field":"general medicine","wr":"reviewers on manuscript accept/reject recommendations","conf":"high","self":false,"doi":"10.1371/journal.pone.0061401","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":true,"ms":"Gwet's AC1, the paper's advocated chance-corrected statistic, was 0.63 on the accept-versus-reject data, which the authors interpret as substantial reviewer agreement.","vf":"unverified"},{"key":"7JG4HLD9","au":"Baethge, Christopher","y":2013,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"proportion of negative (specific) agreement on rejection, exact agreement, Cicchetti-Feinstein","estd":"percent agreement","v":0.313,"n":"206","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"acceptance (with/without revision) / rejection","field":"general medicine","wr":"reviewers on manuscript accept/reject recommendations","conf":"high","self":false,"doi":"10.1371/journal.pone.0061401","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The specific negative agreement, the proportion of reviewer pairs agreeing on rejection given at least one reject, was only 0.313.","vf":"unverified"},{"key":"7JG4HLD9","au":"Baethge, Christopher","y":2013,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"percentage of concordant reviewer pairs, exact agreement on 2 categories, 2-reviewer manuscripts","estd":"percent agreement","v":0.752,"n":"93","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"acceptance (with/without revision) / rejection","field":"general medicine","wr":"reviewers on manuscript accept/reject recommendations","conf":"high","self":false,"doi":"10.1371/journal.pone.0061401","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For the 93 manuscripts with exactly two reviewers, 75.2% of reviewer pairs agreed on the condensed accept-versus-reject decision.","vf":"unverified"},{"key":"7JG4HLD9","au":"Baethge, Christopher","y":2013,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"overall percentage of concordant reviewer pairs, exact agreement on 2 categories (accept vs reject)","estd":"percent agreement","v":0.743,"n":"206","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"acceptance (with/without revision) / rejection","field":"general medicine","wr":"reviewers on manuscript accept/reject recommendations","conf":"high","self":false,"doi":"10.1371/journal.pone.0061401","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"When the three recommendation categories were condensed to accept versus reject, 74.3% of the 529 reviewer pairs agreed.","vf":"unverified"},{"key":"7JG4HLD9","au":"Baethge, Christopher","y":2013,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"percentage of concordant reviewer pairs, exact agreement on 3 categories, 2-reviewer manuscripts","estd":"percent agreement","v":0.624,"n":"93","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / accept after revision / reject","field":"general medicine","wr":"reviewers on manuscript accept/reject recommendations","conf":"high","self":false,"doi":"10.1371/journal.pone.0061401","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Restricting to the 93 manuscripts reviewed by exactly two referees, 62.4% of reviewer pairs gave concordant recommendations on the three-category scale.","vf":"unverified"},{"key":"7JG4HLD9","au":"Baethge, Christopher","y":2013,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"overall percentage of concordant reviewer pairs, exact agreement on 3 categories","estd":"percent agreement","v":0.609,"n":"206","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / accept after revision / reject","field":"general medicine","wr":"reviewers on manuscript accept/reject recommendations","conf":"high","self":false,"doi":"10.1371/journal.pone.0061401","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Across 529 reviewer pairs constructed from 554 reviews on 206 manuscripts, 60.9% of pairs gave concordant recommendations on the three-category scale (accept, accept after revision, reject).","vf":"unverified"},{"key":"7JG4HLD9","au":"Baethge, Christopher","y":2013,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"proportion of positive (specific) agreement on acceptance, exact agreement, Cicchetti-Feinstein","estd":"percent agreement","v":0.84,"n":"206","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"acceptance (with/without revision) / rejection","field":"general medicine","wr":"reviewers on manuscript accept/reject recommendations","conf":"high","self":false,"doi":"10.1371/journal.pone.0061401","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The specific positive agreement, the proportion of reviewer pairs agreeing on acceptance given at least one accept, was 0.84.","vf":"unverified"},{"key":"7JG4HLD9","au":"Baethge, Christopher","y":2013,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Spearman's correlation coefficient for the condensed 2x2 reviewer-pair table","estd":"correlation","v":0.16,"n":"206","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"acceptance (with/without revision) / rejection","field":"general medicine","wr":"reviewers on manuscript accept/reject recommendations","conf":"high","self":false,"doi":"10.1371/journal.pone.0061401","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"On the condensed accept-versus-reject scale, Spearman's rank correlation between paired reviewer recommendations was 0.16.","vf":"unverified"},{"key":"7JG4HLD9","au":"Baethge, Christopher","y":2013,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Spearman's correlation coefficient for the 3x3 reviewer-pair table","estd":"correlation","v":0.17,"n":"206","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / accept after revision / reject","field":"general medicine","wr":"reviewers on manuscript accept/reject recommendations","conf":"high","self":false,"doi":"10.1371/journal.pone.0061401","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Spearman's rank correlation between paired reviewer recommendations on the three-category scale was 0.17, indicating poor agreement.","vf":"unverified"},{"key":"ITEBWPJP","au":"Baimpos, Theodoros","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"all four scores within 0.5 of the median (within-0.5-point tolerance)","estd":"percent agreement","v":0.44,"n":"160","k":"4","samp":"funded-only","blind":"single","agg":"unspecified","scale":"1-5, 5 = highest merit; half-integer steps","field":"multi-field","wr":"four remote experts on funded proposal scores","conf":"high","self":false,"doi":"10.1093/reseval/rvz013","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Four independent remote experts scored each of 160 subsequently funded proposals. All four agreed within the paper's 0.5-point tolerance for 44% of them.","vf":"unverified"},{"key":"ITEBWPJP","au":"Baimpos, Theodoros","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"all four or three scores within 0.5 of the median (within-0.5-point tolerance)","estd":"percent agreement","v":0.91,"n":"160","k":"4","samp":"funded-only","blind":"single","agg":"unspecified","scale":"1-5, 5 = highest merit; half-integer steps","field":"multi-field","wr":"four remote experts on funded proposal scores","conf":"high","self":false,"doi":"10.1093/reseval/rvz013","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Four independent remote experts scored each funded proposal. Ninety-one per cent of the 160 funded proposals were in the all-four-agree or three-agree categories.","vf":"unverified"},{"key":"ITEBWPJP","au":"Baimpos, Theodoros","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"two agree or all disagree using the within-0.5-point tolerance","estd":"percent agreement","v":0.09,"n":"160","k":"4","samp":"funded-only","blind":"single","agg":"unspecified","scale":"1-5, 5 = highest merit; half-integer steps","field":"multi-field","wr":"four remote experts on funded proposal scores","conf":"high","self":false,"doi":"10.1093/reseval/rvz013","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Four independent remote experts scored each of 160 subsequently funded proposals. Nine per cent fell in the combined two-agree or all-disagree categories.","vf":"unverified"},{"key":"ITEBWPJP","au":"Baimpos, Theodoros","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"three scores within 0.5 of the median (within-0.5-point tolerance)","estd":"percent agreement","v":0.47,"n":"160","k":"4","samp":"funded-only","blind":"single","agg":"unspecified","scale":"1-5, 5 = highest merit; half-integer steps","field":"multi-field","wr":"four remote experts on funded proposal scores","conf":"high","self":false,"doi":"10.1093/reseval/rvz013","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Four independent remote experts scored each of 160 subsequently funded proposals. Three experts agreed within the paper's 0.5-point tolerance for 47% of them.","vf":"unverified"},{"key":"ITEBWPJP","au":"Baimpos, Theodoros","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"proportion of proposals with AD index below 0.7 (AD-threshold tolerance)","estd":"percent agreement","v":0.8,"n":"3764","k":"4","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1-5, 5 = highest merit; half-integer steps","field":"multi-field","wr":"four remote experts on grant proposal scores","conf":"high","self":false,"doi":"10.1093/reseval/rvz013","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Four independent remote experts scored each proposal. Eighty per cent of the 3,764 proposals had an Average Deviation index below 0.7, the paper's threshold for good agreement.","vf":"unverified"},{"key":"ITEBWPJP","au":"Baimpos, Theodoros","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"all four scores differ by more than 0.5 from the median (within-0.5-point tolerance)","estd":"percent agreement","v":0.09,"n":"3764","k":"4","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1-5, 5 = highest merit; half-integer steps","field":"multi-field","wr":"four remote experts on grant proposal scores","conf":"high","self":false,"doi":"10.1093/reseval/rvz013","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Four independent remote experts scored each of 3,764 proposals. The paper classified 9% of proposals in its all-disagree category under the 0.5-point rule.","vf":"unverified"},{"key":"ITEBWPJP","au":"Baimpos, Theodoros","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"all four scores within 0.5 of the median (within-0.5-point tolerance)","estd":"percent agreement","v":0.18,"n":"3764","k":"4","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1-5, 5 = highest merit; half-integer steps","field":"multi-field","wr":"four remote experts on grant proposal scores","conf":"high","self":false,"doi":"10.1093/reseval/rvz013","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Across 3,764 proposals each scored by four remote experts, all four agreed within 0.5 of the median on 18% of proposals; the full breakdown was all agree 18%, three agree 42%, two agree 31%, and all disagree 9%.","vf":"unverified"},{"key":"ITEBWPJP","au":"Baimpos, Theodoros","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"all four or three scores within 0.5 of the median (within-0.5-point tolerance)","estd":"percent agreement","v":0.6,"n":"3764","k":"4","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1-5, 5 = highest merit; half-integer steps","field":"multi-field","wr":"four remote experts on grant proposal scores","conf":"high","self":false,"doi":"10.1093/reseval/rvz013","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Four independent remote experts scored each proposal. All four or at least three agreed within the paper's 0.5-point tolerance for 60% of proposals.","vf":"unverified"},{"key":"ITEBWPJP","au":"Baimpos, Theodoros","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"Average Deviation (AD) index from the median (Burke, Finkelstein and Dusig 1999)","estd":"AD index","v":0.45,"n":"3764","k":"4","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1-5, 5 = highest merit; half-integer steps","field":"multi-field","wr":"four remote experts on grant proposal scores","conf":"high","self":false,"doi":"10.1093/reseval/rvz013","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Four external remote experts independently scored each of 3,764 interdisciplinary research proposals on a 1-5 scale; the median Average Deviation index across proposals was 0.45, indicating good inter-rater agreement, since smaller AD values mean closer agreement.","vf":"unverified"},{"key":"ITEBWPJP","au":"Baimpos, Theodoros","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"three scores within 0.5 of the median (within-0.5-point tolerance)","estd":"percent agreement","v":0.42,"n":"3764","k":"4","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1-5, 5 = highest merit; half-integer steps","field":"multi-field","wr":"four remote experts on grant proposal scores","conf":"high","self":false,"doi":"10.1093/reseval/rvz013","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Four independent remote experts scored each of 3,764 proposals. Three of the four experts agreed within the paper's 0.5-point rule for 42% of proposals.","vf":"unverified"},{"key":"ITEBWPJP","au":"Baimpos, Theodoros","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"two scores within 0.5 of the median (within-0.5-point tolerance)","estd":"percent agreement","v":0.31,"n":"3764","k":"4","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1-5, 5 = highest merit; half-integer steps","field":"multi-field","wr":"four remote experts on grant proposal scores","conf":"high","self":false,"doi":"10.1093/reseval/rvz013","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Four independent remote experts scored each of 3,764 proposals. Only two experts agreed within the paper's 0.5-point rule for 31% of proposals.","vf":"unverified"},{"key":"ITEBWPJP","au":"Baimpos, Theodoros","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"all four scores within 0.5 of the median (within-0.5-point tolerance)","estd":"percent agreement","v":0.44,"n":"70","k":"4","samp":"funded-only","blind":"single","agg":"unspecified","scale":"1-5, 5 = highest merit; half-integer steps","field":"multi-field","wr":"four remote experts on panel-selected proposal scores","conf":"high","self":false,"doi":"10.1093/reseval/rvz013","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Four independent remote experts scored each of 70 proposals that entered the funding list only after panel review. All four agreed within the paper's tolerance for 44%.","vf":"unverified"},{"key":"ITEBWPJP","au":"Baimpos, Theodoros","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"two agree or all disagree using the within-0.5-point tolerance","estd":"percent agreement","v":0.2,"n":"70","k":"4","samp":"funded-only","blind":"single","agg":"unspecified","scale":"1-5, 5 = highest merit; half-integer steps","field":"multi-field","wr":"four remote experts on panel-selected proposal scores","conf":"high","self":false,"doi":"10.1093/reseval/rvz013","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Four independent remote experts scored each of 70 proposals that entered the funding list only after panel review. Twenty per cent had low initial remote agreement.","vf":"unverified"},{"key":"ITEBWPJP","au":"Baimpos, Theodoros","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"three scores within 0.5 of the median (within-0.5-point tolerance)","estd":"percent agreement","v":0.36,"n":"70","k":"4","samp":"funded-only","blind":"single","agg":"unspecified","scale":"1-5, 5 = highest merit; half-integer steps","field":"multi-field","wr":"four remote experts on panel-selected proposal scores","conf":"high","self":false,"doi":"10.1093/reseval/rvz013","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Four independent remote experts scored each of 70 proposals that entered the funding list only after panel review. Three experts agreed within the paper's tolerance for 36%.","vf":"unverified"},{"key":"RJEGBDFX","au":"Bayindir, Esra Eren","y":2019,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC average measures, following Shrout & Fleiss (1979) and McGraw & Wong (1996)","estd":"ICC (average)","v":0.619,"n":"5254","k":"2","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"0-100","field":"regional development / public-sector projects","wr":"mean of two reviewers on development proposals","conf":"high","self":false,"doi":"10.1016/j.joi.2019.100981","ciLow":0.598,"ciHigh":0.639,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the same 5254 cutoff-65 proposals, the average-measures intraclass correlation of the two reviewers' mean score was 0.619, higher than the single-rater value as expected.","vf":"unverified"},{"key":"RJEGBDFX","au":"Bayindir, Esra Eren","y":2019,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC average measures, following Shrout & Fleiss (1979) and McGraw & Wong (1996)","estd":"ICC (average)","v":0.599,"n":"386","k":"2","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"0-100","field":"regional development / public-sector projects","wr":"mean of two reviewers on development proposals","conf":"high","self":false,"doi":"10.1016/j.joi.2019.100981","ciLow":0.51,"ciHigh":0.672,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 386 cutoff-75 proposals, the average-measures intraclass correlation of the two reviewers' mean score was 0.599.","vf":"unverified"},{"key":"RJEGBDFX","au":"Bayindir, Esra Eren","y":2019,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC single measures, following Shrout & Fleiss (1979) and McGraw & Wong (1996)","estd":"ICC (single/unspec)","v":0.448,"n":"5254","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-100","field":"regional development / public-sector projects","wr":"two reviewers scoring development project proposals","conf":"high","self":false,"doi":"10.1016/j.joi.2019.100981","ciLow":0.426,"ciHigh":0.469,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Two recruited reviewers independently scored development project proposals on a 0 to 100 scale; the single-measures intraclass correlation of 0.448 across 5254 proposals in cutoff-65 regions indicates modest reliability of one reviewer's score.","vf":"unverified"},{"key":"RJEGBDFX","au":"Bayindir, Esra Eren","y":2019,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC single measures, following Shrout & Fleiss (1979) and McGraw & Wong (1996)","estd":"ICC (single/unspec)","v":0.428,"n":"386","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-100","field":"regional development / public-sector projects","wr":"two reviewers scoring development project proposals","conf":"high","self":false,"doi":"10.1016/j.joi.2019.100981","ciLow":0.343,"ciHigh":0.506,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"In the single cutoff-75 region (386 proposals), the single-measures intraclass correlation between the first two reviewers was 0.428, similar to the cutoff-65 value.","vf":"unverified"},{"key":"RJEGBDFX","au":"Bayindir, Esra Eren","y":2019,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC average measures, following Shrout & Fleiss (1979) and McGraw & Wong (1996)","estd":"ICC (average)","v":-1.539,"n":"1774","k":"2","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"0-100","field":"regional development / public-sector projects","wr":"first two reviewers on disagreement-subset proposals","conf":"high","self":false,"doi":"10.1016/j.joi.2019.100981","ciLow":-1.786,"ciHigh":-1.313,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For the same 1774 cutoff-65 disagreement projects, the average-measures intraclass correlation of the two reviewers was -1.539, an out-of-range value reflecting systematic disagreement.","vf":"unverified"},{"key":"RJEGBDFX","au":"Bayindir, Esra Eren","y":2019,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC average measures, following Shrout & Fleiss (1979) and McGraw & Wong (1996)","estd":"ICC (average)","v":-2.207,"n":"154","k":"2","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"0-100","field":"regional development / public-sector projects","wr":"first two reviewers on disagreement-subset proposals","conf":"high","self":false,"doi":"10.1016/j.joi.2019.100981","ciLow":-3.405,"ciHigh":-1.334,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For the 154 cutoff-75 disagreement projects, the average-measures intraclass correlation of the two reviewers was -2.207, an out-of-range value reflecting systematic disagreement.","vf":"unverified"},{"key":"RJEGBDFX","au":"Bayindir, Esra Eren","y":2019,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC single measures, following Shrout & Fleiss (1979) and McGraw & Wong (1996)","estd":"ICC (single/unspec)","v":-0.435,"n":"1774","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-100","field":"regional development / public-sector projects","wr":"first two reviewers on disagreement-subset proposals","conf":"high","self":false,"doi":"10.1016/j.joi.2019.100981","ciLow":-0.472,"ciHigh":-0.396,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Restricting to the 1774 cutoff-65 projects where the first two reviewers disagreed enough to trigger a third reviewer, their single-measures intraclass correlation was strongly negative (-0.435).","vf":"unverified"},{"key":"RJEGBDFX","au":"Bayindir, Esra Eren","y":2019,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC single measures, following Shrout & Fleiss (1979) and McGraw & Wong (1996)","estd":"ICC (single/unspec)","v":-0.525,"n":"154","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-100","field":"regional development / public-sector projects","wr":"first two reviewers on disagreement-subset proposals","conf":"high","self":false,"doi":"10.1016/j.joi.2019.100981","ciLow":-0.63,"ciHigh":-0.4,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For the 154 cutoff-75 projects where the first two reviewers disagreed enough to trigger a third reviewer, their single-measures intraclass correlation was -0.525.","vf":"unverified"},{"key":"RJEGBDFX","au":"Bayindir, Esra Eren","y":2019,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC average measures, following Shrout & Fleiss (1979) and McGraw & Wong (1996)","estd":"ICC (average)","v":-0.215,"n":"1774","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"0-100","field":"regional development / public-sector projects","wr":"mean of three reviewers on development proposals","conf":"high","self":false,"doi":"10.1016/j.joi.2019.100981","ciLow":-0.316,"ciHigh":-0.12,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For the same 1774 cutoff-65 third-reviewer projects, the average-measures intraclass correlation across three reviewers was -0.215.","vf":"unverified"},{"key":"RJEGBDFX","au":"Bayindir, Esra Eren","y":2019,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC average measures, following Shrout & Fleiss (1979) and McGraw & Wong (1996)","estd":"ICC (average)","v":-0.354,"n":"154","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"0-100","field":"regional development / public-sector projects","wr":"mean of three reviewers on development proposals","conf":"high","self":false,"doi":"10.1016/j.joi.2019.100981","ciLow":-0.771,"ciHigh":-0.022,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For the 154 cutoff-75 third-reviewer projects, the average-measures intraclass correlation across three reviewers was -0.354.","vf":"unverified"},{"key":"RJEGBDFX","au":"Bayindir, Esra Eren","y":2019,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC single measures, following Shrout & Fleiss (1979) and McGraw & Wong (1996)","estd":"ICC (single/unspec)","v":-0.063,"n":"1774","k":"3","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-100","field":"regional development / public-sector projects","wr":"three reviewers scoring development project proposals","conf":"high","self":false,"doi":"10.1016/j.joi.2019.100981","ciLow":-0.087,"ciHigh":-0.037,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Across the 1774 cutoff-65 projects that required a third reviewer, the single-measures intraclass correlation among all three reviewers was slightly negative (-0.063), indicating essentially no agreement.","vf":"unverified"},{"key":"RJEGBDFX","au":"Bayindir, Esra Eren","y":2019,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC single measures, following Shrout & Fleiss (1979) and McGraw & Wong (1996)","estd":"ICC (single/unspec)","v":-0.095,"n":"154","k":"3","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-100","field":"regional development / public-sector projects","wr":"three reviewers scoring development project proposals","conf":"high","self":false,"doi":"10.1016/j.joi.2019.100981","ciLow":-0.17,"ciHigh":-0.007,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Across the 154 cutoff-75 projects that required a third reviewer, the single-measures intraclass correlation among all three reviewers was -0.095.","vf":"unverified"},{"key":"BSI9PZ5P","au":"Bharti, Prabhat Kumar","y":2026,"cx":"Journal","ob":"review-report","fam":"other","form":"'Alignment Score' column (undefined; distinct from the Fleiss kappa in same row)","estd":"other","v":0.833,"n":"","k":"7","samp":"special","blind":"unclear","agg":"unspecified","scale":"exhaustive / trivial","field":"machine learning (NLP)","wr":"pooled LLM-human alignment on exhaustive/trivial review labels","conf":"med","self":false,"doi":"10.1007/s11192-025-05435-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Table 2 reports an 'Alignment Score' of 0.833 for all models combined, a separate agreement figure the paper does not clearly define, distinct from the same row's Fleiss kappa of 0.225.","vf":"unverified"},{"key":"BSI9PZ5P","au":"Bharti, Prabhat Kumar","y":2026,"cx":"Journal","ob":"review-report","fam":"Fleiss-kappa","form":"Fleiss' kappa across all LLMs and human annotators","estd":"kappa","v":0.225,"n":"","k":"7","samp":"special","blind":"unclear","agg":"single-rater","scale":"exhaustive / trivial","field":"machine learning (NLP)","wr":"four LLMs vs three human annotators labelling reviews exhaustive/trivial","conf":"med","self":false,"doi":"10.1007/s11192-025-05435-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":true,"he":false,"ms":"Four large language models and three human annotators each classified peer reviews as exhaustive or trivial; a pooled Fleiss kappa of 0.225 indicates only fair chance-corrected agreement across the full model and annotator pool.","vf":"unverified"},{"key":"BSI9PZ5P","au":"Bharti, Prabhat Kumar","y":2026,"cx":"Journal","ob":"review-report","fam":"other","form":"'Alignment Score' column (undefined; distinct from the Cohen kappa in same row)","estd":"other","v":0.881,"n":"","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"exhaustive / trivial","field":"machine learning (NLP)","wr":"Gemma2-9b alignment with human labels","conf":"med","self":false,"doi":"10.1007/s11192-025-05435-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Table 2 reports an 'Alignment Score' of 0.881 for Gemma2-9b, a separate agreement figure alongside its Cohen kappa of 0.747; the paper does not clearly define what the alignment score measures.","vf":"unverified"},{"key":"BSI9PZ5P","au":"Bharti, Prabhat Kumar","y":2026,"cx":"Journal","ob":"review-report","fam":"kappa","form":"Cohen's kappa (LLM vs majority-vote human consensus)","estd":"kappa","v":0.747,"n":"","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"exhaustive / trivial","field":"machine learning (NLP)","wr":"Gemma2-9b vs human consensus on exhaustive/trivial review labels","conf":"med","self":false,"doi":"10.1007/s11192-025-05435-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Gemma2-9b's exhaustive/trivial classifications were compared with the human-consensus labels; a Cohen kappa of 0.747 indicates substantial chance-corrected agreement.","vf":"unverified"},{"key":"BSI9PZ5P","au":"Bharti, Prabhat Kumar","y":2026,"cx":"Journal","ob":"review-report","fam":"other","form":"'Alignment Score' column (undefined; distinct from the Cohen kappa in same row)","estd":"other","v":0.641,"n":"","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"exhaustive / trivial","field":"machine learning (NLP)","wr":"GPT-4 alignment with human labels","conf":"med","self":false,"doi":"10.1007/s11192-025-05435-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Table 2 reports an 'Alignment Score' of 0.641 for GPT-4, a separate agreement figure alongside its Cohen kappa of 0.340; the paper does not clearly define what the alignment score measures.","vf":"unverified"},{"key":"BSI9PZ5P","au":"Bharti, Prabhat Kumar","y":2026,"cx":"Journal","ob":"review-report","fam":"kappa","form":"Cohen's kappa (LLM vs majority-vote human consensus)","estd":"kappa","v":0.34,"n":"","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"exhaustive / trivial","field":"machine learning (NLP)","wr":"GPT-4 vs human consensus on exhaustive/trivial review labels","conf":"med","self":false,"doi":"10.1007/s11192-025-05435-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"GPT-4's exhaustive/trivial classifications were compared with the human-consensus labels; a Cohen kappa of 0.340 indicates only fair chance-corrected agreement, the lowest among the four models.","vf":"unverified"},{"key":"BSI9PZ5P","au":"Bharti, Prabhat Kumar","y":2026,"cx":"Journal","ob":"review-report","fam":"other","form":"'Alignment Score' column (undefined; distinct from the Cohen kappa in same row)","estd":"other","v":0.897,"n":"","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"exhaustive / trivial","field":"machine learning (NLP)","wr":"Llama-3.1-70b alignment with human labels","conf":"med","self":false,"doi":"10.1007/s11192-025-05435-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Table 2 reports an 'Alignment Score' of 0.897 for Llama-3.1-70b, a separate agreement figure alongside its Cohen kappa of 0.759; the paper does not clearly define what the alignment score measures.","vf":"unverified"},{"key":"BSI9PZ5P","au":"Bharti, Prabhat Kumar","y":2026,"cx":"Journal","ob":"review-report","fam":"kappa","form":"Cohen's kappa (LLM vs majority-vote human consensus)","estd":"kappa","v":0.759,"n":"","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"exhaustive / trivial","field":"machine learning (NLP)","wr":"Llama-3.1-70b vs human consensus on exhaustive/trivial review labels","conf":"med","self":false,"doi":"10.1007/s11192-025-05435-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Llama-3.1-70b's exhaustive/trivial classifications were compared with the majority-vote consensus of three human annotators; a Cohen kappa of 0.759 indicates substantial chance-corrected agreement.","vf":"unverified"},{"key":"BSI9PZ5P","au":"Bharti, Prabhat Kumar","y":2026,"cx":"Journal","ob":"review-report","fam":"other","form":"'Alignment Score' column (undefined; distinct from the Cohen kappa in same row)","estd":"other","v":0.867,"n":"","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"exhaustive / trivial","field":"machine learning (NLP)","wr":"Mixtral-8x7b alignment with human labels","conf":"med","self":false,"doi":"10.1007/s11192-025-05435-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Table 2 reports an 'Alignment Score' of 0.867 for Mixtral-8x7b, a separate agreement figure alongside its Cohen kappa of 0.551; the paper does not clearly define what the alignment score measures.","vf":"unverified"},{"key":"BSI9PZ5P","au":"Bharti, Prabhat Kumar","y":2026,"cx":"Journal","ob":"review-report","fam":"kappa","form":"Cohen's kappa (LLM vs majority-vote human consensus)","estd":"kappa","v":0.551,"n":"","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"exhaustive / trivial","field":"machine learning (NLP)","wr":"Mixtral-8x7b vs human consensus on exhaustive/trivial review labels","conf":"med","self":false,"doi":"10.1007/s11192-025-05435-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Mixtral-8x7b's exhaustive/trivial classifications were compared with the human-consensus labels; a Cohen kappa of 0.551 indicates moderate chance-corrected agreement.","vf":"unverified"},{"key":"VPW3P5QI","au":"Bieri, Marco","y":2021,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"simple agreement, exact match on binary funded/not-funded outcome; bootstrap 95% CI","estd":"percent agreement","v":0.9,"n":"40","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"funded / not funded","field":"multi-field","wr":"funding decisions on Biology fellowship applications","conf":"high","self":false,"doi":"10.1136/bmjopen-2020-047386","ciLow":0.8,"ciHigh":0.975,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For the 40 Biology applications, the simulated procedure using only the two panel members' scores matched the official funding decision in 90.0% of cases.","vf":"unverified"},{"key":"VPW3P5QI","au":"Bieri, Marco","y":2021,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"simple agreement, exact match on binary funded/not-funded outcome; bootstrap 95% CI","estd":"percent agreement","v":0.75,"n":"16","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"funded / not funded","field":"multi-field","wr":"funding decisions on Humanities fellowship applications","conf":"high","self":false,"doi":"10.1136/bmjopen-2020-047386","ciLow":0.5,"ciHigh":0.938,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For the 16 Humanities applications, the simulated procedure using only the two panel members' scores matched the official funding decision in 75.0% of cases, the lowest agreement of any panel.","vf":"unverified"},{"key":"VPW3P5QI","au":"Bieri, Marco","y":2021,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"simple agreement, exact match on binary funded/not-funded outcome; bootstrap 95% CI","estd":"percent agreement","v":0.9,"n":"20","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"funded / not funded","field":"multi-field","wr":"funding decisions on Medicine fellowship applications","conf":"high","self":false,"doi":"10.1136/bmjopen-2020-047386","ciLow":0.75,"ciHigh":1,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For the 20 Medicine applications, the simulated procedure using only the two panel members' scores matched the official funding decision in 90.0% of cases.","vf":"unverified"},{"key":"VPW3P5QI","au":"Bieri, Marco","y":2021,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"simple agreement, exact match on binary funded/not-funded outcome; bootstrap 95% CI","estd":"percent agreement","v":0.86,"n":"86","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"funded / not funded","field":"multi-field","wr":"funding decisions on applications from men","conf":"high","self":false,"doi":"10.1136/bmjopen-2020-047386","ciLow":0.779,"ciHigh":0.93,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For the 86 applications submitted by men, the simulated procedure using only the two panel members' scores matched the official funding decision in 86.0% of cases.","vf":"unverified"},{"key":"VPW3P5QI","au":"Bieri, Marco","y":2021,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"simple agreement, exact match on binary funded/not-funded outcome; bootstrap 95% CI","estd":"percent agreement","v":0.866,"n":"134","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"funded / not funded","field":"multi-field","wr":"funding decisions on fellowship applications","conf":"high","self":false,"doi":"10.1136/bmjopen-2020-047386","ciLow":0.806,"ciHigh":0.918,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"Using only the two panel members' independent 6-point scores for the same 134 fellowship applications, the simulated funding outcome matched the official panel-based decision for 86.6% of applications.","vf":"unverified"},{"key":"VPW3P5QI","au":"Bieri, Marco","y":2021,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"simple agreement, exact match on binary funded/not-funded outcome; bootstrap 95% CI","estd":"percent agreement","v":0.913,"n":"23","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"funded / not funded","field":"multi-field","wr":"funding decisions on Social Sciences fellowship applications","conf":"high","self":false,"doi":"10.1136/bmjopen-2020-047386","ciLow":0.783,"ciHigh":1,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For the 23 Social Sciences applications, the simulated procedure using only the two panel members' scores matched the official funding decision in 91.3% of cases, the highest agreement of any panel.","vf":"unverified"},{"key":"VPW3P5QI","au":"Bieri, Marco","y":2021,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"simple agreement, exact match on binary funded/not-funded outcome; bootstrap 95% CI","estd":"percent agreement","v":0.829,"n":"35","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"funded / not funded","field":"multi-field","wr":"funding decisions on STEM fellowship applications","conf":"high","self":false,"doi":"10.1136/bmjopen-2020-047386","ciLow":0.686,"ciHigh":0.943,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For the 35 STEM applications, the simulated procedure using only the two panel members' scores matched the official funding decision in 82.9% of cases.","vf":"unverified"},{"key":"VPW3P5QI","au":"Bieri, Marco","y":2021,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"simple agreement, exact match on binary funded/not-funded outcome; bootstrap 95% CI","estd":"percent agreement","v":0.731,"n":"67","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"funded / not funded","field":"multi-field","wr":"funding decisions on applications discussed in panel","conf":"high","self":false,"doi":"10.1136/bmjopen-2020-047386","ciLow":0.627,"ciHigh":0.836,"mt":"inter-rater","tgt":"editor-decisions","rr":"restricted-other","pr":false,"he":false,"ms":"For the 67 middle-ranked applications discussed in the official panel meeting, the simulated procedure using only the two panel members' scores matched the funding decision in 73.1% of cases.","vf":"unverified"},{"key":"VPW3P5QI","au":"Bieri, Marco","y":2021,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"simple agreement, exact match on binary funded/not-funded outcome; no CI reported","estd":"percent agreement","v":1,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"funded / not funded","field":"multi-field","wr":"funding decisions on directly funded applications","conf":"high","self":false,"doi":"10.1136/bmjopen-2020-047386","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"restricted-top","pr":false,"he":false,"ms":"For the 37 applications the official triage sent straight to funding, the simulated procedure using only the two panel members' scores reproduced the funding decision in every case.","vf":"unverified"},{"key":"VPW3P5QI","au":"Bieri, Marco","y":2021,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"simple agreement, exact match on binary funded/not-funded outcome; no CI reported","estd":"percent agreement","v":1,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"funded / not funded","field":"multi-field","wr":"funding decisions on directly rejected applications","conf":"high","self":false,"doi":"10.1136/bmjopen-2020-047386","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"restricted-other","pr":false,"he":false,"ms":"For the 30 applications the official triage rejected without discussion, the simulated procedure using only the two panel members' scores reproduced the rejection in every case.","vf":"unverified"},{"key":"VPW3P5QI","au":"Bieri, Marco","y":2021,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"simple agreement, exact match on binary funded/not-funded outcome; bootstrap 95% CI","estd":"percent agreement","v":0.875,"n":"48","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"funded / not funded","field":"multi-field","wr":"funding decisions on applications from women","conf":"high","self":false,"doi":"10.1136/bmjopen-2020-047386","ciLow":0.771,"ciHigh":0.958,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For the 48 applications submitted by women, the simulated procedure using only the two panel members' scores matched the official funding decision in 87.5% of cases.","vf":"unverified"},{"key":"VPW3P5QI","au":"Bieri, Marco","y":2021,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"simple agreement, exact match on binary funded/not-funded outcome; bootstrap 95% CI","estd":"percent agreement","v":0.85,"n":"40","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"funded / not funded","field":"multi-field","wr":"funding decisions on Biology fellowship applications","conf":"high","self":false,"doi":"10.1136/bmjopen-2020-047386","ciLow":0.725,"ciHigh":0.95,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For the 40 applications handled by the Biology panel, the simulated three-reviewer procedure reached the same funding decision as the official panel process in 85.0% of cases.","vf":"unverified"},{"key":"VPW3P5QI","au":"Bieri, Marco","y":2021,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"simple agreement, exact match on binary funded/not-funded outcome; bootstrap 95% CI","estd":"percent agreement","v":0.75,"n":"16","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"funded / not funded","field":"multi-field","wr":"funding decisions on Humanities fellowship applications","conf":"high","self":false,"doi":"10.1136/bmjopen-2020-047386","ciLow":0.5,"ciHigh":0.938,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For the 16 applications handled by the Humanities panel, the simulated three-reviewer procedure reached the same funding decision as the official panel process in 75.0% of cases.","vf":"unverified"},{"key":"VPW3P5QI","au":"Bieri, Marco","y":2021,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"simple agreement, exact match on binary funded/not-funded outcome; bootstrap 95% CI","estd":"percent agreement","v":0.9,"n":"20","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"funded / not funded","field":"multi-field","wr":"funding decisions on Medicine fellowship applications","conf":"high","self":false,"doi":"10.1136/bmjopen-2020-047386","ciLow":0.75,"ciHigh":1,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For the 20 applications handled by the Medicine panel, the simulated three-reviewer procedure agreed with the official funding decision in 90.0% of cases, the highest agreement of any panel.","vf":"unverified"},{"key":"VPW3P5QI","au":"Bieri, Marco","y":2021,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"simple agreement, exact match on binary funded/not-funded outcome; bootstrap 95% CI","estd":"percent agreement","v":0.791,"n":"86","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"funded / not funded","field":"multi-field","wr":"funding decisions on applications from men","conf":"high","self":false,"doi":"10.1136/bmjopen-2020-047386","ciLow":0.698,"ciHigh":0.872,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For the 86 applications submitted by men, the simulated three-reviewer procedure reached the same funding decision as the official panel process in 79.1% of cases.","vf":"unverified"},{"key":"VPW3P5QI","au":"Bieri, Marco","y":2021,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"simple agreement, exact match on binary funded/not-funded outcome; bootstrap 95% CI","estd":"percent agreement","v":0.806,"n":"134","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"funded / not funded","field":"multi-field","wr":"funding decisions on fellowship applications","conf":"high","self":false,"doi":"10.1136/bmjopen-2020-047386","ciLow":0.739,"ciHigh":0.873,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":true,"he":false,"ms":"Two panel members and one additional external expert each scored 134 Postdoc.Mobility fellowship applications on a 6-point scale. Funding outcomes simulated from the mean of these three reviews matched the official panel-based funding decision for 80.6% of applications.","vf":"unverified"},{"key":"VPW3P5QI","au":"Bieri, Marco","y":2021,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"simple agreement, exact match on binary funded/not-funded outcome; bootstrap 95% CI","estd":"percent agreement","v":0.739,"n":"23","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"funded / not funded","field":"multi-field","wr":"funding decisions on Social Sciences fellowship applications","conf":"high","self":false,"doi":"10.1136/bmjopen-2020-047386","ciLow":0.565,"ciHigh":0.913,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For the 23 applications handled by the Social Sciences panel, the simulated three-reviewer procedure agreed with the official funding decision in 73.9% of cases, the lowest agreement of any panel.","vf":"unverified"},{"key":"VPW3P5QI","au":"Bieri, Marco","y":2021,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"simple agreement, exact match on binary funded/not-funded outcome; bootstrap 95% CI","estd":"percent agreement","v":0.771,"n":"35","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"funded / not funded","field":"multi-field","wr":"funding decisions on STEM fellowship applications","conf":"high","self":false,"doi":"10.1136/bmjopen-2020-047386","ciLow":0.629,"ciHigh":0.914,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For the 35 applications handled by the STEM panel, the simulated three-reviewer procedure reached the same funding decision as the official panel process in 77.1% of cases.","vf":"unverified"},{"key":"VPW3P5QI","au":"Bieri, Marco","y":2021,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"simple agreement, exact match on binary funded/not-funded outcome; bootstrap 95% CI","estd":"percent agreement","v":0.642,"n":"67","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"funded / not funded","field":"multi-field","wr":"funding decisions on applications discussed in panel","conf":"high","self":false,"doi":"10.1136/bmjopen-2020-047386","ciLow":0.522,"ciHigh":0.761,"mt":"inter-rater","tgt":"editor-decisions","rr":"restricted-other","pr":false,"he":false,"ms":"For the 67 middle-ranked applications that the official process discussed in the panel meeting, the simulated three-reviewer procedure gave the same funding decision in only 64.2% of cases.","vf":"unverified"},{"key":"VPW3P5QI","au":"Bieri, Marco","y":2021,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"simple agreement, exact match on binary funded/not-funded outcome; bootstrap 95% CI","estd":"percent agreement","v":0.973,"n":"","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"funded / not funded","field":"multi-field","wr":"funding decisions on directly funded applications","conf":"high","self":false,"doi":"10.1136/bmjopen-2020-047386","ciLow":0.919,"ciHigh":1,"mt":"inter-rater","tgt":"editor-decisions","rr":"restricted-top","pr":false,"he":false,"ms":"For the 37 applications the official triage sent straight to funding without discussion, the simulated three-reviewer procedure gave the same funding decision in 97.3% of cases.","vf":"unverified"},{"key":"VPW3P5QI","au":"Bieri, Marco","y":2021,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"simple agreement, exact match on binary funded/not-funded outcome; bootstrap 95% CI","estd":"percent agreement","v":0.967,"n":"","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"funded / not funded","field":"multi-field","wr":"funding decisions on directly rejected applications","conf":"high","self":false,"doi":"10.1136/bmjopen-2020-047386","ciLow":0.9,"ciHigh":1,"mt":"inter-rater","tgt":"editor-decisions","rr":"restricted-other","pr":false,"he":false,"ms":"For the 30 applications the official triage rejected without discussion, the simulated three-reviewer procedure gave the same funding decision in 96.7% of cases.","vf":"unverified"},{"key":"VPW3P5QI","au":"Bieri, Marco","y":2021,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"simple agreement, exact match on binary funded/not-funded outcome; bootstrap 95% CI","estd":"percent agreement","v":0.833,"n":"48","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"funded / not funded","field":"multi-field","wr":"funding decisions on applications from women","conf":"high","self":false,"doi":"10.1136/bmjopen-2020-047386","ciLow":0.729,"ciHigh":0.938,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For the 48 applications submitted by women, the simulated three-reviewer procedure reached the same funding decision as the official panel process in 83.3% of cases.","vf":"unverified"},{"key":"HPC6SYSG","au":"Blackburn, Jessica L.","y":2006,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC1 (Shrout & Fleiss, 1979)","estd":"ICC (single/unspec)","v":0.29,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"5-point scales, averaged over 5 criteria","field":"industrial-organisational psychology","wr":"reviewers on conference poster ratings","conf":"med","self":false,"doi":"10.1111/j.1467-9280.2006.01715.x","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Separate one-way ICCs were computed for posters with differing numbers of reviewers; the highest of these individual-rating reliabilities was 0.29.","vf":"unverified"},{"key":"HPC6SYSG","au":"Blackburn, Jessica L.","y":2006,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC1 (Shrout & Fleiss, 1979)","estd":"ICC (single/unspec)","v":-0.07,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"5-point scales, averaged over 5 criteria","field":"industrial-organisational psychology","wr":"reviewers on conference poster ratings","conf":"med","self":false,"doi":"10.1111/j.1467-9280.2006.01715.x","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Separate one-way ICCs were computed for posters with differing numbers of reviewers; the lowest of these individual-rating reliabilities was -0.07.","vf":"unverified"},{"key":"HPC6SYSG","au":"Blackburn, Jessica L.","y":2006,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC1 (Shrout & Fleiss, 1979), computed on within-reviewer z scores","estd":"ICC (single/unspec)","v":0.21,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"within-reviewer z scores","field":"industrial-organisational psychology","wr":"reviewers on z-standardised poster ratings","conf":"med","self":false,"doi":"10.1111/j.1467-9280.2006.01715.x","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"For z-standardised ratings, separate one-way ICCs by number of reviewers ranged up to a maximum of 0.21.","vf":"unverified"},{"key":"HPC6SYSG","au":"Blackburn, Jessica L.","y":2006,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC1 (Shrout & Fleiss, 1979), computed on within-reviewer z scores","estd":"ICC (single/unspec)","v":0.14,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"within-reviewer z scores","field":"industrial-organisational psychology","wr":"reviewers on z-standardised poster ratings","conf":"med","self":false,"doi":"10.1111/j.1467-9280.2006.01715.x","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"For z-standardised ratings, separate one-way ICCs by number of reviewers ranged downward to a minimum of 0.14.","vf":"unverified"},{"key":"HPC6SYSG","au":"Blackburn, Jessica L.","y":2006,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC1 (Shrout & Fleiss, 1979), computed on within-reviewer z scores","estd":"ICC (single/unspec)","v":0.16,"n":"1983","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"within-reviewer z scores","field":"industrial-organisational psychology","wr":"reviewers on z-standardised poster ratings","conf":"med","self":false,"doi":"10.1111/j.1467-9280.2006.01715.x","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"After converting each reviewer's ratings to within-reviewer z scores, the weighted-average one-way ICC was 0.16, showing that standardisation did not improve the reliability of the poster ratings.","vf":"unverified"},{"key":"HPC6SYSG","au":"Blackburn, Jessica L.","y":2006,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC1 (Shrout & Fleiss, 1979)","estd":"ICC (single/unspec)","v":0.18,"n":"1983","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"5-point scales, averaged over 5 criteria","field":"industrial-organisational psychology","wr":"reviewers on conference poster ratings","conf":"med","self":false,"doi":"10.1111/j.1467-9280.2006.01715.x","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":false,"ms":"A pool of 1,557 conference reviewers produced 7,383 ratings of 1,983 poster submissions on 5-point scales. The weighted-average one-way ICC for individual ratings was 0.18, indicating low inter-rater reliability.","vf":"unverified"},{"key":"PSBEMHHH","au":"Bollen, Johan","y":2017,"cx":"Grant","ob":"other","fam":"correlation","form":"Pearson R","estd":"correlation","v":0.2683,"n":"65610","k":"","samp":"special","blind":"unclear","agg":"unspecified","scale":"award amount in US Dollars","field":"multi-field","wr":"simulated and actual funding per scientist","conf":"high","self":false,"doi":"10.1007/s11192-016-2110-3","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"unclear","pr":true,"he":false,"ms":"For 65,610 matched scientists, funding allocated by a simulated crowd-based system was compared with their actual NSF and NIH funding for 2000 to 2010. Pearson R was 0.2683.","vf":"unverified"},{"key":"PSBEMHHH","au":"Bollen, Johan","y":2017,"cx":"Grant","ob":"other","fam":"correlation","form":"Spearman ρ","estd":"correlation","v":0.2999,"n":"65610","k":"","samp":"special","blind":"unclear","agg":"unspecified","scale":"award amount in US Dollars","field":"multi-field","wr":"simulated and actual funding per scientist","conf":"high","self":false,"doi":"10.1007/s11192-016-2110-3","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"unclear","pr":false,"he":false,"ms":"For 65,610 matched scientists, ranks of funding allocated by a simulated crowd-based system were compared with ranks of their actual NSF and NIH funding for 2000 to 2010. Spearman rho was 0.2999.","vf":"unverified"},{"key":"FDNSXEWP","au":"Bornmann, Lutz","y":2005,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"agreement defined as decision reached in the first Board round (indirect proxy, not pairwise exact agreement)","estd":"percent agreement","v":0.76,"n":"2524","k":"7","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"A / A- / A-B and below","field":"biomedical","wr":"Board members on fellowship approve/reject decisions","conf":"med","self":false,"doi":"10.1007/s11192-005-0214-2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":true,"he":false,"ms":"The seven-member Board of Trustees decided on 2,524 fellowship applications after discussing each one; 76% were settled in the first of up to three decision rounds, which the authors treat as the share of decisions characterised by agreement among the trustees.","vf":"unverified"},{"key":"5CTCSBI3","au":"Bornmann, Lutz","y":2006,"cx":"Grant","ob":"fellowship","fam":"correlation","form":"Cramer's V","estd":"correlation","v":0.36,"n":"1003","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"reviewer 1 award/2 possible/3 no; board approved/rejected","field":"biomedical","wr":"external reviewer rating vs Board funding decision","conf":"med","self":false,"doi":"10.1016/j.joi.2006.09.005","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":true,"he":false,"ms":"For 1003 doctoral fellowship applications, the association between each external reviewer's recommendation and the Board of Trustees' approve or reject decision was Cramer's V = 0.36, a moderate concordance between the two assessment levels.","vf":"unverified"},{"key":"5CTCSBI3","au":"Bornmann, Lutz","y":2006,"cx":"Grant","ob":"fellowship","fam":"correlation","form":"Cramer's V","estd":"correlation","v":0.27,"n":"326","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"reviewer 1 award/2 possible/3 no; board approved/rejected","field":"biomedical","wr":"external reviewer rating vs Board funding decision","conf":"med","self":false,"doi":"10.1016/j.joi.2006.09.005","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For 326 post-doctoral fellowship applications, the association between each external reviewer's recommendation and the Board of Trustees' approve or reject decision was Cramer's V = 0.27, a weak-to-moderate concordance between the two assessment levels.","vf":"unverified"},{"key":"49ZXXL9X","au":"Bornmann, Lutz","y":2008,"cx":"Grant","ob":"fellowship","fam":"other","form":"a kind of test-retest reliability, defined as 1 minus proportion of measurement error of change (latent non-stationary Markov model)","estd":"other","v":0.8,"n":"1474","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"award / possible award / no award","field":"biomedical","wr":"doctoral fellowship application ratings","conf":"med","self":false,"doi":"10.1016/j.joi.2008.05.003","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"External reviewers, staff members and the Board of Trustees assessed 1474 doctoral fellowship applications over three dependent stages. The latent Markov reliability for this group was approximately 0.80.","vf":"unverified"},{"key":"49ZXXL9X","au":"Bornmann, Lutz","y":2008,"cx":"Grant","ob":"fellowship","fam":"other","form":"true stability (measurement-error-corrected stability, latent Markov model)","estd":"other","v":0.2,"n":"1474","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"award / possible award / no award","field":"biomedical","wr":"doctoral fellowship application ratings","conf":"med","self":false,"doi":"10.1016/j.joi.2008.05.003","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"The latent Markov model estimated the true (measurement-error-corrected) proportion of doctoral fellowship applications keeping the same rating across all three assessment stages as 0.20.","vf":"unverified"},{"key":"49ZXXL9X","au":"Bornmann, Lutz","y":2008,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"exact agreement (same category) across three evaluation stages, manifest data","estd":"percent agreement","v":0.24,"n":"1474","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"award / possible award / no award","field":"biomedical","wr":"doctoral fellowship application ratings","conf":"med","self":false,"doi":"10.1016/j.joi.2008.05.003","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"The three evaluation stages gave the same rating category to a doctoral fellowship application in 24% of cases in the observed data.","vf":"unverified"},{"key":"49ZXXL9X","au":"Bornmann, Lutz","y":2008,"cx":"Grant","ob":"fellowship","fam":"other","form":"a kind of test-retest reliability, defined as 1 minus proportion of measurement error of change (latent non-stationary Markov model)","estd":"other","v":0.8,"n":"1954","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"award / possible award / no award","field":"biomedical","wr":"three-stage assessors on fellowship applications","conf":"med","self":false,"doi":"10.1016/j.joi.2008.05.003","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":false,"ms":"Across 1954 Boehringer Ingelheim Fonds fellowship applications, categorical judgments at three sequential evaluation stages (external reviewer, staff member, Board of Trustees) were modelled with a latent non-stationary Markov model. The reliability of about 0.80, defined as one minus the measurement error of change, indicates the multi-stage process separates applications sufficiently reliably. The stage ratings are dependent rather than independent.","vf":"unverified"},{"key":"49ZXXL9X","au":"Bornmann, Lutz","y":2008,"cx":"Grant","ob":"fellowship","fam":"other","form":"a kind of test-retest reliability, defined as 1 minus proportion of measurement error of change (latent non-stationary Markov model)","estd":"other","v":0.8,"n":"480","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"award / possible award / no award","field":"biomedical","wr":"post-doctoral fellowship application ratings","conf":"med","self":false,"doi":"10.1016/j.joi.2008.05.003","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"External reviewers, staff members and the Board of Trustees assessed 480 post-doctoral fellowship applications over three dependent stages. The latent Markov reliability for this group was approximately 0.80.","vf":"unverified"},{"key":"49ZXXL9X","au":"Bornmann, Lutz","y":2008,"cx":"Grant","ob":"fellowship","fam":"other","form":"true stability (measurement-error-corrected stability, latent Markov model)","estd":"other","v":0.19,"n":"480","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"award / possible award / no award","field":"biomedical","wr":"post-doctoral fellowship application ratings","conf":"med","self":false,"doi":"10.1016/j.joi.2008.05.003","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"The latent Markov model estimated the true (measurement-error-corrected) proportion of post-doctoral fellowship applications keeping the same rating across all three assessment stages as 0.19.","vf":"unverified"},{"key":"49ZXXL9X","au":"Bornmann, Lutz","y":2008,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"exact agreement (same category) across three evaluation stages, manifest data","estd":"percent agreement","v":0.22,"n":"480","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"award / possible award / no award","field":"biomedical","wr":"post-doctoral fellowship application ratings","conf":"med","self":false,"doi":"10.1016/j.joi.2008.05.003","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"The three evaluation stages gave the same rating category to a post-doctoral fellowship application in 22% of cases in the observed data.","vf":"unverified"},{"key":"MAPYT9ZR","au":"Bornmann, Lutz","y":2008,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"unweighted kappa (k, u), subgroup of two-review Communications","estd":"kappa","v":0.27,"n":"718","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"Yes, without alterations / Yes, minor / Yes, major / No","field":"chemistry","wr":"referees' accept recommendations on Communications","conf":"high","self":false,"doi":"10.1002/anie.200800513","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"In the restricted subgroup of 718 Communications with exactly two complete reviews, the unweighted kappa was 0.27, somewhat higher than in the whole group.","vf":"unverified"},{"key":"MAPYT9ZR","au":"Bornmann, Lutz","y":2008,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"weighted kappa (k, w), subgroup of two-review Communications","estd":"weighted kappa","v":0.43,"n":"718","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"Yes, without alterations / Yes, minor / Yes, major / No","field":"chemistry","wr":"referees' accept recommendations on Communications","conf":"high","self":false,"doi":"10.1002/anie.200800513","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"In the same 718-Communication subgroup, a weighted kappa crediting near-agreement gave 0.43, higher than for the whole group but still moderate.","vf":"unverified"},{"key":"MAPYT9ZR","au":"Bornmann, Lutz","y":2008,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"unweighted kappa (k, u), three referees","estd":"kappa","v":0.1,"n":"535","k":"3","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"Yes, without alterations / Yes, minor / Yes, major / No","field":"chemistry","wr":"referees' accept recommendations on Communications","conf":"high","self":false,"doi":"10.1002/anie.200800513","ciLow":0.07,"ciHigh":0.14,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For 535 Communications each reviewed by three independent external referees, the unweighted kappa was 0.10, indicating low agreement among referees.","vf":"unverified"},{"key":"MAPYT9ZR","au":"Bornmann, Lutz","y":2008,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"observed exact agreement between two referees across the four response categories (percent)","estd":"percent agreement","v":0.418,"n":"952","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"Yes, without alterations / Yes, minor / Yes, major / No","field":"chemistry","wr":"referees' accept recommendations on Communications","conf":"high","self":false,"doi":"10.1002/anie.200800513","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Among 952 two-referee Communications, the two referees gave exactly the same recommendation for 41.8 percent of manuscripts, against 31.8 percent expected by chance.","vf":"unverified"},{"key":"MAPYT9ZR","au":"Bornmann, Lutz","y":2008,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"weighted observed agreement: partial agreement with 0.6667 and 0.3333 category-distance weights","estd":"percent agreement","v":0.691,"n":"952","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"Yes, without alterations / Yes, minor / Yes, major / No","field":"chemistry","wr":"referees' accept recommendations on Communications","conf":"high","self":false,"doi":"10.1002/anie.200800513","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the same 952 two-referee Communications, the weighted observed agreement crediting adjacent-category near-agreement was 69.1 percent, against 61.2 percent expected by chance.","vf":"unverified"},{"key":"MAPYT9ZR","au":"Bornmann, Lutz","y":2008,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"unweighted kappa (k, u)","estd":"kappa","v":0.15,"n":"952","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"Yes, without alterations / Yes, minor / Yes, major / No","field":"chemistry","wr":"referees' accept recommendations on Communications","conf":"high","self":false,"doi":"10.1002/anie.200800513","ciLow":0.1,"ciHigh":0.19,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For 952 Communications reviewed by two independent external referees each, agreement on the acceptance recommendation was low, with an unweighted kappa of 0.15.","vf":"unverified"},{"key":"MAPYT9ZR","au":"Bornmann, Lutz","y":2008,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"weighted kappa (k, w); weights 1/0.6667/0.3333/0 for full/two-thirds/one-third/no agreement","estd":"weighted kappa","v":0.21,"n":"952","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"Yes, without alterations / Yes, minor / Yes, major / No","field":"chemistry","wr":"referees' accept recommendations on Communications","conf":"high","self":false,"doi":"10.1002/anie.200800513","ciLow":0.16,"ciHigh":0.25,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the same 952 two-referee Communications, a weighted kappa crediting adjacent-category near-agreement gave 0.21, still indicating low agreement between referees.","vf":"unverified"},{"key":"MAPYT9ZR","au":"Bornmann, Lutz","y":2008,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"unweighted kappa (k, u), two to five referees","estd":"kappa","v":0.12,"n":"1507","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"Yes, without alterations / Yes, minor / Yes, major / No","field":"chemistry","wr":"referees' accept recommendations on Communications","conf":"high","self":false,"doi":"10.1002/anie.200800513","ciLow":0.09,"ciHigh":0.15,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Across all 1507 Communications with between two and five independent referees, the overall unweighted kappa was 0.12, the study's headline result of low inter-referee agreement.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"conditional κ (‘fair’ and ‘poor’)","estd":"kappa","v":0.22,"n":"356","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript presentation-quality ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Conditional kappa for presentation quality among the 356 two-reviewer manuscripts, with the two unfavourable categories combined, was 0.22.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"conditional κ (‘excellent’ and ‘good’)","estd":"kappa","v":0.29,"n":"356","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript presentation-quality ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Conditional kappa for presentation quality among the 356 two-reviewer manuscripts, with the two favourable categories combined, was 0.29.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"conditional κ (‘yes, after major alterations’ and ‘no’)","estd":"kappa","v":0.25,"n":"356","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: yes(no alt.), yes(minor), yes(major), no","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript publication recommendation","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Conditional kappa for the publication recommendation among the 356 two-reviewer manuscripts, with the two major-alterations and rejection categories combined, was 0.25.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"conditional κ (‘yes, without alterations’ and ‘yes, after minor alterations’)","estd":"kappa","v":0.29,"n":"356","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: yes(no alt.), yes(minor), yes(major), no","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript publication recommendation","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Conditional kappa for the publication recommendation among the 356 two-reviewer manuscripts, with the two acceptance categories combined, was 0.29.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"conditional κ (‘fair’ and ‘poor’)","estd":"kappa","v":0.2,"n":"356","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript scientific-quality ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Conditional kappa for scientific quality among the 356 two-reviewer manuscripts, with the two unfavourable categories combined, was 0.2.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"conditional κ (‘excellent’ and ‘good’)","estd":"kappa","v":0.24,"n":"356","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript scientific-quality ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Conditional kappa for scientific quality among the 356 two-reviewer manuscripts, with the two favourable categories combined, was 0.24.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"conditional κ (‘fair’ and ‘poor’)","estd":"kappa","v":0.36,"n":"356","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript scientific-significance ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Conditional kappa for scientific significance among the 356 two-reviewer manuscripts, with the two unfavourable categories combined, was 0.36.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"conditional κ (‘excellent’ and ‘good’)","estd":"kappa","v":0.28,"n":"356","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript scientific-significance ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Conditional kappa for scientific significance among the 356 two-reviewer manuscripts, with the two favourable categories combined, was 0.28.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC (random-effects ANOVA model)","estd":"ICC (single/unspec)","v":0.31,"n":"465","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript presentation-quality ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"A one-way random effects ANOVA ICC for presentation quality across 465 manuscripts (1,058 reviews) was 0.31.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC (random-effects ANOVA model)","estd":"ICC (single/unspec)","v":0.26,"n":"465","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: yes(no alt.), yes(minor), yes(major), no","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript publication recommendation","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"A one-way random effects ANOVA ICC for the publication recommendation across 465 manuscripts (1,058 reviews) was 0.26.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC (random-effects ANOVA model)","estd":"ICC (single/unspec)","v":0.26,"n":"465","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript scientific-quality ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"A one-way random effects ANOVA ICC for scientific quality across 465 manuscripts (1,058 reviews) was 0.26.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC (random-effects ANOVA model)","estd":"ICC (single/unspec)","v":0.34,"n":"465","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript scientific-significance ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"A one-way random effects ANOVA ICC for scientific significance across 465 manuscripts (1,058 reviews) was 0.34. This is a reasonable level of agreement by Hargens and Herting's benchmark for manuscript review.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"unconditional ICC (random-intercept proportional odds model)","estd":"ICC (single/unspec)","v":0.36,"n":"465","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript presentation-quality ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The unconditional ICC from a random-intercept proportional odds model for presentation quality across 465 manuscripts (1,058 reviews) was 0.36.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"unconditional ICC (random-intercept proportional odds model)","estd":"ICC (single/unspec)","v":0.51,"n":"15","k":"4","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: yes(no alt.), yes(minor), yes(major), no","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript publication recommendation","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 15 manuscripts with four reviews each, the proportional-odds unconditional ICC for the publication recommendation was 0.51, the highest ICC stratum, included as the maximum of the per-reviewer-count strata.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"unconditional ICC (random-intercept proportional odds model)","estd":"ICC (single/unspec)","v":0.33,"n":"465","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: yes(no alt.), yes(minor), yes(major), no","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript publication recommendation","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The unconditional ICC from a random-intercept proportional odds model for the publication recommendation across 465 manuscripts (1,058 reviews) was 0.33.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"unconditional ICC (random-intercept proportional odds model)","estd":"ICC (single/unspec)","v":0.01,"n":"15","k":"4","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript scientific-quality ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 15 manuscripts with four reviews each, the proportional-odds unconditional ICC for scientific quality was 0.01, the lowest ICC in Table 1, included as the minimum of the per-reviewer-count strata.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"unconditional ICC (random-intercept proportional odds model)","estd":"ICC (single/unspec)","v":0.3,"n":"465","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript scientific-quality ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The unconditional ICC from a random-intercept proportional odds model for scientific quality across 465 manuscripts (1,058 reviews) was 0.3.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"unconditional ICC (random-intercept proportional odds model)","estd":"ICC (single/unspec)","v":0.41,"n":"465","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript scientific-significance ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The unconditional ICC from a random-intercept proportional odds model for scientific significance across 465 manuscripts (1,058 reviews) was 0.41.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"observed agreement (%); exact agreement","estd":"percent agreement","v":0.492,"n":"356","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript presentation-quality ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two reviewers gave identical presentation quality ratings for 49.2% of the 356 two-reviewer manuscripts (observed exact agreement, before chance correction).","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"observed agreement (%); exact agreement","estd":"percent agreement","v":0.587,"n":"356","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: yes(no alt.), yes(minor), yes(major), no","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript publication recommendation","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two reviewers gave the identical publication recommendation for 58.7% of the 356 two-reviewer manuscripts (observed exact agreement, before chance correction).","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"observed agreement (%); exact agreement","estd":"percent agreement","v":0.461,"n":"356","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript scientific-quality ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two reviewers gave identical scientific quality ratings for 46.1% of the 356 two-reviewer manuscripts (observed exact agreement, before chance correction).","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"observed agreement (%); exact agreement","estd":"percent agreement","v":0.551,"n":"356","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript scientific-significance ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two reviewers gave identical scientific significance ratings for 55.1% of the 356 two-reviewer manuscripts (observed exact agreement, before chance correction).","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"unweighted κ","estd":"kappa","v":0.15,"n":"465","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript presentation-quality ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"At the open-access journal ACP, external reviewers rating presentation quality across 465 manuscripts (1,058 reviews, 2-5 reviewers each) showed an unweighted kappa of 0.15, indicating only slight chance-corrected agreement.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"unweighted κ","estd":"kappa","v":0.25,"n":"15","k":"4","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: yes(no alt.), yes(minor), yes(major), no","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript publication recommendation","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 15 manuscripts with four reviews each, unweighted kappa for the publication recommendation was 0.25, the highest unweighted-kappa stratum, included as the maximum of the per-reviewer-count strata.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"unweighted κ","estd":"kappa","v":0.2,"n":"465","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: yes(no alt.), yes(minor), yes(major), no","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript publication recommendation","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"At ACP, external reviewers' publication recommendations on 465 manuscripts (1,058 reviews, 2-5 reviewers each) showed an unweighted kappa of 0.20, only slight chance-corrected agreement; selected as the bottom-line reliability result among four parallel dimensions the paper reports without privileging one.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"unweighted κ","estd":"kappa","v":-0.01,"n":"15","k":"4","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript scientific-quality ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 15 manuscripts with four reviews each (60 reviews), unweighted kappa for scientific quality was -0.01, the lowest kappa in Table 1, included as the minimum of the per-reviewer-count strata.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"unweighted κ","estd":"kappa","v":0.12,"n":"465","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript scientific-quality ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"At the open-access journal ACP, external reviewers rating scientific quality across 465 manuscripts (1,058 reviews, 2-5 reviewers each) showed an unweighted kappa of 0.12, indicating only slight chance-corrected agreement.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"unweighted κ","estd":"kappa","v":0.19,"n":"465","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript scientific-significance ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"At the open-access journal ACP, external reviewers rating scientific significance across 465 manuscripts (1,058 reviews, 2-5 reviewers each) showed an unweighted kappa of 0.19, indicating only slight chance-corrected agreement.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"weighted κ","estd":"weighted kappa","v":0.24,"n":"356","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript presentation-quality ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 356 two-reviewer manuscripts, a linearly weighted kappa for presentation quality was 0.24, giving partial credit for adjacent categories.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"weighted κ","estd":"weighted kappa","v":0.22,"n":"356","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: yes(no alt.), yes(minor), yes(major), no","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript publication recommendation","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 356 two-reviewer manuscripts, a linearly weighted kappa for the publication recommendation was 0.22, giving partial credit for adjacent categories.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"weighted κ","estd":"weighted kappa","v":0.2,"n":"356","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript scientific-quality ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 356 two-reviewer manuscripts, a linearly weighted kappa for scientific quality was 0.2, giving partial credit for adjacent categories.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"weighted κ","estd":"weighted kappa","v":0.27,"n":"356","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript scientific-significance ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 356 two-reviewer manuscripts, a linearly weighted kappa for scientific significance was 0.27, giving partial credit for adjacent categories.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"observed agreement (%); weighted: exact=1, neighbouring=0.6667, two apart=0.3333, contrary=0","estd":"percent agreement","v":0.814,"n":"356","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript presentation-quality ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 356 two-reviewer manuscripts, weighted observed agreement for presentation quality was 81.4%, giving partial credit for neighbouring categories, before chance correction.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"observed agreement (%); weighted: exact=1, neighbouring=0.6667, two apart=0.3333, contrary=0","estd":"percent agreement","v":0.857,"n":"356","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: yes(no alt.), yes(minor), yes(major), no","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript publication recommendation","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 356 two-reviewer manuscripts, weighted observed agreement for the publication recommendation was 85.7%, giving partial credit for neighbouring categories, before chance correction.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"observed agreement (%); weighted: exact=1, neighbouring=0.6667, two apart=0.3333, contrary=0","estd":"percent agreement","v":0.788,"n":"356","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript scientific-quality ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 356 two-reviewer manuscripts, weighted observed agreement for scientific quality was 78.8%, giving partial credit for neighbouring categories, before chance correction.","vf":"unverified"},{"key":"SHHLHEPD","au":"Bornmann, Lutz","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"observed agreement (%); weighted: exact=1, neighbouring=0.6667, two apart=0.3333, contrary=0","estd":"percent agreement","v":0.83,"n":"356","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-4: excellent, good, fair, poor","field":"atmospheric sciences (geosciences)","wr":"reviewers on manuscript scientific-significance ratings","conf":"high","self":false,"doi":"10.1087/20100207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 356 two-reviewer manuscripts, weighted observed agreement for scientific significance was 83.0%, giving partial credit for neighbouring categories, before chance correction.","vf":"unverified"},{"key":"8AAUI9F9","au":"Bornmann, Lutz","y":2015,"cx":"General","ob":"journal-manuscript","fam":"ICC","form":"ICC via Stata icc command, model unspecified","estd":"ICC (single/unspec)","v":0.11,"n":"72","k":"8","samp":"funded-only","blind":"open","agg":"unspecified","scale":"Good (1), Very good (2), Exceptional (3)","field":"biomedical","wr":"Faculty members on published papers' recommendation scores","conf":"high","self":false,"doi":"10.1002/asi.23334","ciLow":0.05,"ciHigh":0.2,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"For 72 papers each recommended by eight Faculty members, an ICC of 0.11 indicates low inter-rater reliability.","vf":"unverified"},{"key":"8AAUI9F9","au":"Bornmann, Lutz","y":2015,"cx":"General","ob":"journal-manuscript","fam":"ICC","form":"ICC via Stata icc command, model unspecified","estd":"ICC (single/unspec)","v":0.15,"n":"637","k":"5","samp":"funded-only","blind":"open","agg":"unspecified","scale":"Good (1), Very good (2), Exceptional (3)","field":"biomedical","wr":"Faculty members on published papers' recommendation scores","conf":"high","self":false,"doi":"10.1002/asi.23334","ciLow":0.12,"ciHigh":0.18,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"For 637 papers each recommended by five Faculty members, an ICC of 0.15 indicates low inter-rater reliability.","vf":"unverified"},{"key":"8AAUI9F9","au":"Bornmann, Lutz","y":2015,"cx":"General","ob":"journal-manuscript","fam":"ICC","form":"ICC via Stata icc command, model unspecified","estd":"ICC (single/unspec)","v":0.19,"n":"1602","k":"4","samp":"funded-only","blind":"open","agg":"unspecified","scale":"Good (1), Very good (2), Exceptional (3)","field":"biomedical","wr":"Faculty members on published papers' recommendation scores","conf":"high","self":false,"doi":"10.1002/asi.23334","ciLow":0.17,"ciHigh":0.22,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"For 1,602 papers each recommended by four Faculty members, an ICC of 0.19 indicates low inter-rater reliability.","vf":"unverified"},{"key":"8AAUI9F9","au":"Bornmann, Lutz","y":2015,"cx":"General","ob":"journal-manuscript","fam":"ICC","form":"ICC via Stata icc command, model unspecified","estd":"ICC (single/unspec)","v":0.17,"n":"49","k":"9","samp":"funded-only","blind":"open","agg":"unspecified","scale":"Good (1), Very good (2), Exceptional (3)","field":"biomedical","wr":"Faculty members on published papers' recommendation scores","conf":"high","self":false,"doi":"10.1002/asi.23334","ciLow":0.1,"ciHigh":0.28,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"For 49 papers each recommended by nine Faculty members, an ICC of 0.17 indicates low inter-rater reliability.","vf":"unverified"},{"key":"8AAUI9F9","au":"Bornmann, Lutz","y":2015,"cx":"General","ob":"journal-manuscript","fam":"ICC","form":"ICC via Stata icc command, model unspecified","estd":"ICC (single/unspec)","v":0.13,"n":"138","k":"7","samp":"funded-only","blind":"open","agg":"unspecified","scale":"Good (1), Very good (2), Exceptional (3)","field":"biomedical","wr":"Faculty members on published papers' recommendation scores","conf":"high","self":false,"doi":"10.1002/asi.23334","ciLow":0.08,"ciHigh":0.2,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"For 138 papers each recommended by seven Faculty members, an ICC of 0.13 indicates low inter-rater reliability.","vf":"unverified"},{"key":"8AAUI9F9","au":"Bornmann, Lutz","y":2015,"cx":"General","ob":"journal-manuscript","fam":"ICC","form":"ICC via Stata icc command, model unspecified","estd":"ICC (single/unspec)","v":0.18,"n":"266","k":"6","samp":"funded-only","blind":"open","agg":"unspecified","scale":"Good (1), Very good (2), Exceptional (3)","field":"biomedical","wr":"Faculty members on published papers' recommendation scores","conf":"high","self":false,"doi":"10.1002/asi.23334","ciLow":0.13,"ciHigh":0.23,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"For 266 papers each recommended by six Faculty members, an ICC of 0.18 indicates low inter-rater reliability.","vf":"unverified"},{"key":"8AAUI9F9","au":"Bornmann, Lutz","y":2015,"cx":"General","ob":"journal-manuscript","fam":"ICC","form":"ICC via Stata icc command, model unspecified","estd":"ICC (single/unspec)","v":0.21,"n":"4225","k":"3","samp":"funded-only","blind":"open","agg":"unspecified","scale":"Good (1), Very good (2), Exceptional (3)","field":"biomedical","wr":"Faculty members on published papers' recommendation scores","conf":"high","self":false,"doi":"10.1002/asi.23334","ciLow":0.19,"ciHigh":0.23,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"For 4,225 papers each recommended by three Faculty members, an ICC of 0.21 indicates low inter-rater reliability.","vf":"unverified"},{"key":"8AAUI9F9","au":"Bornmann, Lutz","y":2015,"cx":"General","ob":"journal-manuscript","fam":"ICC","form":"ICC via Stata icc command, model unspecified","estd":"ICC (single/unspec)","v":0.21,"n":"14476","k":"2","samp":"funded-only","blind":"open","agg":"unspecified","scale":"Good (1), Very good (2), Exceptional (3)","field":"biomedical","wr":"Faculty members on published papers' recommendation scores","conf":"high","self":false,"doi":"10.1002/asi.23334","ciLow":0.2,"ciHigh":0.23,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":true,"he":false,"ms":"Faculty members recommended published biomedical papers on a three-point scale; for 14,476 papers each rated by two members, an ICC of 0.21 indicates low inter-rater reliability.","vf":"unverified"},{"key":"8AAUI9F9","au":"Bornmann, Lutz","y":2015,"cx":"General","ob":"journal-manuscript","fam":"kappa","form":"Kappa via Stata kappa command","estd":"kappa","v":0.35,"n":"223","k":"2","samp":"funded-only","blind":"open","agg":"single-rater","scale":"Good (1), Very good (2), Exceptional (3)","field":"biomedical","wr":"two Faculty members who both tagged a paper as Confirmation","conf":"high","self":false,"doi":"10.1002/asi.23334","ciLow":0.21,"ciHigh":0.47,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"For 223 papers where two Faculty members both applied the Confirmation tag, a kappa of 0.35 indicates fair chance-corrected agreement on the recommendation score.","vf":"unverified"},{"key":"8AAUI9F9","au":"Bornmann, Lutz","y":2015,"cx":"General","ob":"journal-manuscript","fam":"kappa","form":"Kappa via Stata kappa command","estd":"kappa","v":0.17,"n":"78","k":"2","samp":"funded-only","blind":"open","agg":"single-rater","scale":"Good (1), Very good (2), Exceptional (3)","field":"biomedical","wr":"two Faculty members who both tagged a paper as Hypothesis","conf":"high","self":false,"doi":"10.1002/asi.23334","ciLow":-0.02,"ciHigh":0.34,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"For 78 papers where two Faculty members both applied the Interesting Hypothesis tag, a kappa of 0.17 indicates only slight chance-corrected agreement on the recommendation score.","vf":"unverified"},{"key":"8AAUI9F9","au":"Bornmann, Lutz","y":2015,"cx":"General","ob":"journal-manuscript","fam":"kappa","form":"Kappa via Stata kappa command","estd":"kappa","v":0.14,"n":"3634","k":"2","samp":"funded-only","blind":"open","agg":"single-rater","scale":"Good (1), Very good (2), Exceptional (3)","field":"biomedical","wr":"two Faculty members who both tagged a paper as New finding","conf":"high","self":false,"doi":"10.1002/asi.23334","ciLow":0.11,"ciHigh":0.17,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"For 3,634 papers where two Faculty members both applied the New Finding tag, a kappa of 0.14 indicates only slight chance-corrected agreement on the recommendation score.","vf":"unverified"},{"key":"8AAUI9F9","au":"Bornmann, Lutz","y":2015,"cx":"General","ob":"journal-manuscript","fam":"kappa","form":"Kappa via Stata kappa command","estd":"kappa","v":0.2,"n":"353","k":"2","samp":"funded-only","blind":"open","agg":"single-rater","scale":"Good (1), Very good (2), Exceptional (3)","field":"biomedical","wr":"two Faculty members who both tagged a paper as Technical Advance","conf":"high","self":false,"doi":"10.1002/asi.23334","ciLow":0.09,"ciHigh":0.29,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"For 353 papers where two Faculty members both applied the Technical Advance tag, a kappa of 0.20 indicates only slight chance-corrected agreement on the recommendation score.","vf":"unverified"},{"key":"8AAUI9F9","au":"Bornmann, Lutz","y":2015,"cx":"General","ob":"journal-manuscript","fam":"kappa","form":"Kappa via Stata kappa command","estd":"kappa","v":0.14,"n":"72","k":"8","samp":"funded-only","blind":"open","agg":"single-rater","scale":"Good (1), Very good (2), Exceptional (3)","field":"biomedical","wr":"Faculty members on published papers' recommendation scores","conf":"high","self":false,"doi":"10.1002/asi.23334","ciLow":0.01,"ciHigh":0.31,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"For 72 papers each recommended by eight Faculty members, a kappa of 0.14 indicates only slight chance-corrected agreement.","vf":"unverified"},{"key":"8AAUI9F9","au":"Bornmann, Lutz","y":2015,"cx":"General","ob":"journal-manuscript","fam":"kappa","form":"Kappa via Stata kappa command","estd":"kappa","v":0.09,"n":"637","k":"5","samp":"funded-only","blind":"open","agg":"single-rater","scale":"Good (1), Very good (2), Exceptional (3)","field":"biomedical","wr":"Faculty members on published papers' recommendation scores","conf":"high","self":false,"doi":"10.1002/asi.23334","ciLow":0,"ciHigh":0.13,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"For 637 papers each recommended by five Faculty members, a kappa of 0.09 indicates only slight chance-corrected agreement.","vf":"unverified"},{"key":"8AAUI9F9","au":"Bornmann, Lutz","y":2015,"cx":"General","ob":"journal-manuscript","fam":"kappa","form":"Kappa via Stata kappa command","estd":"kappa","v":0.12,"n":"1602","k":"4","samp":"funded-only","blind":"open","agg":"single-rater","scale":"Good (1), Very good (2), Exceptional (3)","field":"biomedical","wr":"Faculty members on published papers' recommendation scores","conf":"high","self":false,"doi":"10.1002/asi.23334","ciLow":0.07,"ciHigh":0.15,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"For 1,602 papers each recommended by four Faculty members, a kappa of 0.12 indicates only slight chance-corrected agreement.","vf":"unverified"},{"key":"8AAUI9F9","au":"Bornmann, Lutz","y":2015,"cx":"General","ob":"journal-manuscript","fam":"kappa","form":"Kappa via Stata kappa command","estd":"kappa","v":0.25,"n":"49","k":"9","samp":"funded-only","blind":"open","agg":"single-rater","scale":"Good (1), Very good (2), Exceptional (3)","field":"biomedical","wr":"Faculty members on published papers' recommendation scores","conf":"high","self":false,"doi":"10.1002/asi.23334","ciLow":0.09,"ciHigh":0.54,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"For 49 papers each recommended by nine Faculty members, a kappa of 0.25 indicates fair chance-corrected agreement, the only value above slight in the table.","vf":"unverified"},{"key":"8AAUI9F9","au":"Bornmann, Lutz","y":2015,"cx":"General","ob":"journal-manuscript","fam":"kappa","form":"Kappa via Stata kappa command","estd":"kappa","v":0,"n":"138","k":"7","samp":"funded-only","blind":"open","agg":"single-rater","scale":"Good (1), Very good (2), Exceptional (3)","field":"biomedical","wr":"Faculty members on published papers' recommendation scores","conf":"high","self":false,"doi":"10.1002/asi.23334","ciLow":-0.13,"ciHigh":0.1,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"For 138 papers each recommended by seven Faculty members, a kappa of 0.00 indicates no chance-corrected agreement.","vf":"unverified"},{"key":"8AAUI9F9","au":"Bornmann, Lutz","y":2015,"cx":"General","ob":"journal-manuscript","fam":"kappa","form":"Kappa via Stata kappa command","estd":"kappa","v":0.06,"n":"266","k":"6","samp":"funded-only","blind":"open","agg":"single-rater","scale":"Good (1), Very good (2), Exceptional (3)","field":"biomedical","wr":"Faculty members on published papers' recommendation scores","conf":"high","self":false,"doi":"10.1002/asi.23334","ciLow":-0.04,"ciHigh":0.13,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"For 266 papers each recommended by six Faculty members, a kappa of 0.06 indicates essentially no chance-corrected agreement.","vf":"unverified"},{"key":"8AAUI9F9","au":"Bornmann, Lutz","y":2015,"cx":"General","ob":"journal-manuscript","fam":"kappa","form":"Kappa via Stata kappa command","estd":"kappa","v":0.14,"n":"4225","k":"3","samp":"funded-only","blind":"open","agg":"single-rater","scale":"Good (1), Very good (2), Exceptional (3)","field":"biomedical","wr":"Faculty members on published papers' recommendation scores","conf":"high","self":false,"doi":"10.1002/asi.23334","ciLow":0.12,"ciHigh":0.17,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"For 4,225 papers each recommended by three Faculty members, a kappa of 0.14 indicates only slight chance-corrected agreement.","vf":"unverified"},{"key":"8AAUI9F9","au":"Bornmann, Lutz","y":2015,"cx":"General","ob":"journal-manuscript","fam":"kappa","form":"Kappa via Stata kappa command","estd":"kappa","v":0.14,"n":"14476","k":"2","samp":"funded-only","blind":"open","agg":"single-rater","scale":"Good (1), Very good (2), Exceptional (3)","field":"biomedical","wr":"Faculty members on published papers' recommendation scores","conf":"high","self":false,"doi":"10.1002/asi.23334","ciLow":0.13,"ciHigh":0.15,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"For 14,476 papers each recommended by two Faculty members, a kappa of 0.14 indicates only slight chance-corrected agreement.","vf":"unverified"},{"key":"8AAUI9F9","au":"Bornmann, Lutz","y":2015,"cx":"General","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement (both members gave the identical recommendation score)","estd":"percent agreement","v":0.51,"n":"14476","k":"2","samp":"funded-only","blind":"open","agg":"single-rater","scale":"Good (1), Very good (2), Exceptional (3)","field":"biomedical","wr":"Faculty members on published papers' recommendation scores","conf":"high","self":false,"doi":"10.1002/asi.23334","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Of 14,476 papers each recommended by two Faculty members, the two members gave the identical recommendation score for about 51 per cent (raw, non-chance-corrected agreement).","vf":"unverified"},{"key":"TE7UT3JG","au":"Boudreau, Kevin J","y":2016,"cx":"Grant","ob":"other","fam":"percent-agreement","form":"exact agreement on binary fund / do-not-fund decision (both fund or both do not fund)","estd":"percent agreement","v":0.59,"n":"120","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"binary: fund / do not fund","field":"theatre arts","wr":"crowd (Kickstarter) vs expert judge, fund/not-fund on theatre projects","conf":"med","self":false,"doi":"10.1287/mnsc.2015.2207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":true,"he":false,"ms":"For 120 theatre projects, the crowd's Kickstarter funding outcomes and expert judges' hypothetical funding decisions agreed on about 59% of cases (range 56.7% to 65% across three funding screens and 1,024 judge combinations), above chance.","vf":"unverified"},{"key":"TE7UT3JG","au":"Boudreau, Kevin J","y":2016,"cx":"Grant","ob":"other","fam":"weighted-kappa","form":"kappa with squared weighting scheme (Cohen, 1968)","estd":"weighted kappa","v":0.45,"n":"60","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1 to 5, 5 being 'Strongly Agree'","field":"theatre arts","wr":"expert judges on theatre project funding proposals","conf":"med","self":false,"doi":"10.1287/mnsc.2015.2207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Same measurement as the text kappa: the Table 1 note reports inter-rater reliability for the doubly rated projects as 0.45, conflicting with the 0.44 given in the text.","vf":"unverified"},{"key":"TE7UT3JG","au":"Boudreau, Kevin J","y":2016,"cx":"Grant","ob":"other","fam":"weighted-kappa","form":"kappa with squared weighting scheme (Cohen, 1968)","estd":"weighted kappa","v":0.44,"n":"60","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1 to 5, 5 being 'Strongly Agree'","field":"theatre arts","wr":"expert judges on theatre project funding proposals","conf":"med","self":false,"doi":"10.1287/mnsc.2015.2207","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Sixty theatre crowdfunding projects were each scored independently by two recruited expert judges on 1 to 5 scales; a squared-weighted kappa of 0.44 indicates moderate chance-corrected agreement.","vf":"unverified"},{"key":"PB3TYFEA","au":"Bowen, Donald D.","y":1972,"cx":"Journal","ob":"conference-abstract","fam":"Kendall-W","form":"Kendall's coefficient of concordance (W)","estd":"Kendall W","v":0.24,"n":"8","k":"3","samp":"funded-only","blind":"unclear","agg":"unspecified","scale":"ranking, 1 to 8","field":"psychology (consumer)","wr":"3 least-consensual judges ranking 8 papers","conf":"high","self":false,"doi":"10.1037/h0033849","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"The three judges whose rankings were furthest from the consensus agreed among themselves with a Kendall's W of 0.240 across the eight papers, indicating low agreement.","vf":"unverified"},{"key":"PB3TYFEA","au":"Bowen, Donald D.","y":1972,"cx":"Journal","ob":"conference-abstract","fam":"Kendall-W","form":"Kendall's coefficient of concordance (W)","estd":"Kendall W","v":0.661,"n":"8","k":"3","samp":"funded-only","blind":"unclear","agg":"unspecified","scale":"ranking, 1 to 8","field":"psychology (consumer)","wr":"3 most-consensual judges ranking 8 papers","conf":"high","self":false,"doi":"10.1037/h0033849","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"The three judges whose rankings were closest to the consensus reached only a minimal level of agreement among themselves, with a Kendall's W of 0.661 across the eight papers.","vf":"unverified"},{"key":"PB3TYFEA","au":"Bowen, Donald D.","y":1972,"cx":"Journal","ob":"conference-abstract","fam":"Kendall-W","form":"Kendall's coefficient of concordance (W)","estd":"Kendall W","v":0.173,"n":"6","k":"10","samp":"funded-only","blind":"unclear","agg":"unspecified","scale":"ranking, 1 to 8","field":"psychology (consumer)","wr":"10 judges ranking the best and worst three papers","conf":"high","self":false,"doi":"10.1037/h0033849","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Restricting the analysis to the three highest and three lowest ranked papers did not improve agreement among the ten judges, giving a Kendall's W of 0.173.","vf":"unverified"},{"key":"PB3TYFEA","au":"Bowen, Donald D.","y":1972,"cx":"Journal","ob":"conference-abstract","fam":"Kendall-W","form":"Kendall's coefficient of concordance (W)","estd":"Kendall W","v":0.106,"n":"8","k":"10","samp":"funded-only","blind":"unclear","agg":"unspecified","scale":"ranking, 1 to 8","field":"psychology (consumer)","wr":"past presidents ranking 8 contest papers","conf":"high","self":false,"doi":"10.1037/h0033849","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":true,"he":true,"ms":"Ten past Division 23 presidents independently ranked the eight programme-selected papers from 1 to 8; Kendall's W of 0.106 shows very low agreement among the judges.","vf":"unverified"},{"key":"73TNGMX8","au":"Bravo, Giangiacomo","y":2018,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"number of referee recommendations that should be changed to reach a perfect agreement among the referees, divided by the number of referees","estd":"other","v":0.19,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"accept / minor revisions / major revisions / reject","field":"computer science","wr":"referees' recommendations on accepted manuscripts","conf":"med","self":false,"doi":"10.1016/j.joi.2017.12.002","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"External referees recommended outcomes for journal manuscripts. Among manuscripts the editor accepted, the mean disagreement score was 0.19, where 0 means perfect agreement and higher values mean greater disagreement.","vf":"unverified"},{"key":"73TNGMX8","au":"Bravo, Giangiacomo","y":2018,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"number of referee recommendations that should be changed to reach a perfect agreement among the referees, divided by the number of referees","estd":"other","v":0.32,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"accept / minor revisions / major revisions / reject","field":"computer science","wr":"referees' recommendations on major-revision manuscripts","conf":"med","self":false,"doi":"10.1016/j.joi.2017.12.002","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"External referees recommended outcomes for journal manuscripts. Among manuscripts receiving a major-revisions decision, the mean disagreement score was 0.32, the highest of the four decision groups.","vf":"unverified"},{"key":"73TNGMX8","au":"Bravo, Giangiacomo","y":2018,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"number of referee recommendations that should be changed to reach a perfect agreement among the referees, divided by the number of referees","estd":"other","v":0.24,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"accept / minor revisions / major revisions / reject","field":"computer science","wr":"referees' recommendations on minor-revision manuscripts","conf":"med","self":false,"doi":"10.1016/j.joi.2017.12.002","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"External referees recommended outcomes for journal manuscripts. Among manuscripts receiving a minor-revisions decision, the mean disagreement score was 0.24, where 0 means perfect agreement.","vf":"unverified"},{"key":"73TNGMX8","au":"Bravo, Giangiacomo","y":2018,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"number of referee recommendations that should be changed to reach a perfect agreement among the referees, divided by the number of referees","estd":"other","v":0.21,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"accept / minor revisions / major revisions / reject","field":"computer science","wr":"referees' recommendations on rejected manuscripts","conf":"med","self":false,"doi":"10.1016/j.joi.2017.12.002","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"External referees recommended outcomes for journal manuscripts. Among manuscripts the editor rejected, the mean disagreement score was 0.21, where 0 means perfect agreement.","vf":"unverified"},{"key":"BGW9Q5PA","au":"Butters, Anna","y":2026,"cx":"Grant","ob":"grant-proposal","fam":"weighted-kappa","form":"weighted kappa, quadratic weights; mean weighted kappa of 41 proposals","estd":"weighted kappa","v":0.38,"n":"","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"5pt, 'constructive' to 'non-constructive'","field":"multi-field (interdisciplinary)","wr":"co-applicants rating constructiveness of reviews received","conf":"high","self":false,"doi":"10.31222/osf.io/t2p56_v2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":true,"he":false,"ms":"Co-applicants on the same proposal each rated the constructiveness of the distributed-peer-review comments their proposal received, via an optional feedback survey. The mean quadratic-weighted kappa across 41 proposals (two of 43 excluded for lack of variance) was 0.38, only fair agreement.","vf":"unverified"},{"key":"BGW9Q5PA","au":"Butters, Anna","y":2026,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"Intraclass Correlation Coefficient (ICC) was calculated for random effects; model: Score ~ 1 + (1|proposal) + (1|reviewer)","estd":"ICC (single/unspec)","v":0.2004,"n":"140","k":"9.92","samp":"full-pool","blind":"double","agg":"unspecified","scale":"A+ ('outstanding', 9) to C- ('unsuitable', 1), 9pt","field":"multi-field (interdisciplinary)","wr":"reviewer-effect share of DPR proposal-score variance","conf":"high","self":false,"doi":"10.31222/osf.io/t2p56_v2","ciLow":null,"ciHigh":null,"mt":"other","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"In the same crossed mixed model of distributed-peer-review scores, 20.04% of overall-score variance was attributable to systematic between-reviewer differences, exceeding the between-proposal share and signalling high reviewer inconsistency. This is a variance share, not an agreement coefficient.","vf":"unverified"},{"key":"BGW9Q5PA","au":"Butters, Anna","y":2026,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"average proportion in same quartile (exact agreement), one DPR vs one panel reviewer, 500 bootstrapped samples","estd":"percent agreement","v":0.39,"n":"70","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"quartiles 1-4 of proposal ranking","field":"multi-field (interdisciplinary)","wr":"DPR vs panel reviewer quartile placement of proposals","conf":"high","self":false,"doi":"10.31222/osf.io/t2p56_v2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Bootstrapped comparison of one DPR reviewer and one panel reviewer placing 70 shortlisted proposals into ranking quartiles; they agreed on top-quartile placement 39% of the time, above the 25% chance baseline.","vf":"unverified"},{"key":"BGW9Q5PA","au":"Butters, Anna","y":2026,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"average proportion in same quartile (exact agreement), one DPR vs one panel reviewer, 500 bootstrapped samples","estd":"percent agreement","v":0.27,"n":"70","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"quartiles 1-4 of proposal ranking","field":"multi-field (interdisciplinary)","wr":"DPR vs panel reviewer quartile placement of proposals","conf":"high","self":false,"doi":"10.31222/osf.io/t2p56_v2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Bootstrapped comparison of one DPR reviewer and one panel reviewer placing 70 shortlisted proposals into ranking quartiles; they agreed on second-quartile placement 27% of the time, near the 25% chance baseline.","vf":"unverified"},{"key":"BGW9Q5PA","au":"Butters, Anna","y":2026,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"average proportion in same quartile (exact agreement), one DPR vs one panel reviewer, 500 bootstrapped samples","estd":"percent agreement","v":0.28,"n":"70","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"quartiles 1-4 of proposal ranking","field":"multi-field (interdisciplinary)","wr":"DPR vs panel reviewer quartile placement of proposals","conf":"high","self":false,"doi":"10.31222/osf.io/t2p56_v2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Bootstrapped comparison of one DPR reviewer and one panel reviewer placing 70 shortlisted proposals into ranking quartiles; they agreed on third-quartile placement 28% of the time, near the 25% chance baseline.","vf":"unverified"},{"key":"BGW9Q5PA","au":"Butters, Anna","y":2026,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"average proportion in same quartile (exact agreement), one DPR vs one panel reviewer, 500 bootstrapped samples","estd":"percent agreement","v":0.31,"n":"70","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"quartiles 1-4 of proposal ranking","field":"multi-field (interdisciplinary)","wr":"DPR vs panel reviewer quartile placement of proposals","conf":"high","self":false,"doi":"10.31222/osf.io/t2p56_v2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Bootstrapped comparison of one DPR reviewer and one panel reviewer placing 70 shortlisted proposals into ranking quartiles; they agreed on bottom-quartile placement 31% of the time, above the 25% chance baseline.","vf":"unverified"},{"key":"BGW9Q5PA","au":"Butters, Anna","y":2026,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"Intraclass Correlation Coefficient (ICC) was calculated for random effects; model: Score ~ 1 + (1|proposal) + (1|reviewer)","estd":"ICC (single/unspec)","v":0.0887,"n":"140","k":"9.92","samp":"full-pool","blind":"double","agg":"single-rater","scale":"A+ ('outstanding', 9) to C- ('unsuitable', 1), 9pt","field":"multi-field (interdisciplinary)","wr":"applicant-reviewers on grant proposal overall scores","conf":"high","self":false,"doi":"10.31222/osf.io/t2p56_v2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"In distributed peer review, 323 applicant-reviewers independently scored 140 interdisciplinary proposals (mean 9.92 reviews each, 1387 reviews). A crossed mixed model attributed only 8.87% of overall-score variance to between-proposal differences, indicating very low single-reviewer reliability.","vf":"unverified"},{"key":"SA7AMBC5","au":"Callaham, M L","y":1998,"cx":"Journal","ob":"review-report","fam":"percent-agreement","form":"exact agreement","estd":"percent agreement","v":0.5,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"acceptance, rejection, or revision","field":"biomedical","wr":"reviewer recommendations and editorial decisions","conf":"high","self":false,"doi":"10.1001/jama.280.3.229","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"Across the studied reviews by regular reviewers, the reviewer's recommendation exactly matched the editor's final decision in 50% of reviews.","vf":"unverified"},{"key":"SA7AMBC5","au":"Callaham, M L","y":1998,"cx":"Journal","ob":"review-report","fam":"percent-agreement","form":"exact agreement","estd":"percent agreement","v":0.21,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"acceptance, rejection, or revision","field":"biomedical","wr":"decisions for reviews rated poor","conf":"high","self":false,"doi":"10.1001/jama.280.3.229","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For reviews given a quality rating of 1, reviewer recommendations exactly matched editorial decisions in 21% of cases.","vf":"unverified"},{"key":"SA7AMBC5","au":"Callaham, M L","y":1998,"cx":"Journal","ob":"review-report","fam":"percent-agreement","form":"exact agreement","estd":"percent agreement","v":0.27,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"acceptance, rejection, or revision","field":"biomedical","wr":"decisions for reviews rated 2","conf":"high","self":false,"doi":"10.1001/jama.280.3.229","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For reviews given a quality rating of 2, reviewer recommendations exactly matched editorial decisions in 27% of cases.","vf":"unverified"},{"key":"SA7AMBC5","au":"Callaham, M L","y":1998,"cx":"Journal","ob":"review-report","fam":"percent-agreement","form":"exact agreement","estd":"percent agreement","v":0.41,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"acceptance, rejection, or revision","field":"biomedical","wr":"decisions for reviews rated 3","conf":"high","self":false,"doi":"10.1001/jama.280.3.229","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For reviews given a quality rating of 3, reviewer recommendations exactly matched editorial decisions in 41% of cases.","vf":"unverified"},{"key":"SA7AMBC5","au":"Callaham, M L","y":1998,"cx":"Journal","ob":"review-report","fam":"percent-agreement","form":"exact agreement","estd":"percent agreement","v":0.55,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"acceptance, rejection, or revision","field":"biomedical","wr":"decisions for reviews rated 4","conf":"high","self":false,"doi":"10.1001/jama.280.3.229","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For reviews given a quality rating of 4, reviewer recommendations exactly matched editorial decisions in 55% of cases.","vf":"unverified"},{"key":"SA7AMBC5","au":"Callaham, M L","y":1998,"cx":"Journal","ob":"review-report","fam":"percent-agreement","form":"exact agreement","estd":"percent agreement","v":0.62,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"acceptance, rejection, or revision","field":"biomedical","wr":"decisions for reviews rated excellent","conf":"high","self":false,"doi":"10.1001/jama.280.3.229","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For reviews given a quality rating of 5, reviewer recommendations exactly matched editorial decisions in 62% of cases.","vf":"unverified"},{"key":"SA7AMBC5","au":"Callaham, M L","y":1998,"cx":"Journal","ob":"review-report","fam":"weighted-kappa","form":"weighted k","estd":"weighted kappa","v":0.034,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"acceptance, rejection, or revision","field":"biomedical","wr":"reviewer recommendations and editorial decisions","conf":"high","self":false,"doi":"10.1001/jama.280.3.229","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":true,"he":false,"ms":"Individual reviewers' accept, reject or revise recommendations were compared with editors' final decisions on the same manuscripts. The weighted kappa of 0.034 indicates essentially no chance-corrected agreement between recommendation and decision.","vf":"unverified"},{"key":"SA7AMBC5","au":"Callaham, M L","y":1998,"cx":"Journal","ob":"review-report","fam":"ICC","form":"intraclass correlation for editor","estd":"ICC (single/unspec)","v":0.24,"n":"2686","k":"1","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"1 (poor) to 5 (excellent)","field":"biomedical","wr":"editors rating quality of peer reviews","conf":"high","self":false,"doi":"10.1001/jama.280.3.229","ciLow":null,"ciHigh":null,"mt":"other","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"In the same variance decomposition of single editorial ratings of 2686 reviews, the editor intraclass correlation of 0.24 describes clustering of review-quality ratings by the rating editor, about 6% of the variance.","vf":"unverified"},{"key":"SA7AMBC5","au":"Callaham, M L","y":1998,"cx":"Journal","ob":"review-report","fam":"ICC","form":"intraclass correlation for manuscript","estd":"ICC (single/unspec)","v":0.12,"n":"2686","k":"1","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"1 (poor) to 5 (excellent)","field":"biomedical","wr":"editors rating quality of peer reviews","conf":"high","self":false,"doi":"10.1001/jama.280.3.229","ciLow":null,"ciHigh":null,"mt":"other","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"In the same variance decomposition, the manuscript intraclass correlation of 0.12 describes clustering of review-quality ratings by the manuscript reviewed (973 manuscripts), about 1% of the variance.","vf":"unverified"},{"key":"SA7AMBC5","au":"Callaham, M L","y":1998,"cx":"Journal","ob":"review-report","fam":"ICC","form":"within-reviewer intraclass correlation","estd":"ICC (single/unspec)","v":0.44,"n":"2686","k":"1","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"1 (poor) to 5 (excellent)","field":"biomedical","wr":"editors rating quality of peer reviews","conf":"high","self":false,"doi":"10.1001/jama.280.3.229","ciLow":null,"ciHigh":null,"mt":"other","tgt":"review-quality","rr":"none","pr":true,"he":false,"ms":"Thirty-six editors each gave one quality rating to 2686 reviews written by 395 reviewers. The within-reviewer intraclass correlation of 0.44 means about 20% of the variance in review-quality ratings was attributable to the reviewer, the paper's headline reliability result.","vf":"unverified"},{"key":"W2VF7UBP","au":"Campos‐Arceiz, Ahimsa","y":2015,"cx":"Journal","ob":"journal-manuscript","fam":"Gwet-AC","form":"Gwet's kappa (AC1)","estd":"Gwet AC","v":0.28,"n":"108","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"'accept', 'minor revision', 'major revision', or 'reject'","field":"conservation biology","wr":"two reviewers' recommendations on manuscripts from China","conf":"high","self":false,"doi":"10.1016/j.biocon.2015.02.025","ciLow":0.16,"ciHigh":0.41,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Gwet's AC1 for the two reviewers' recommendations on the 108 manuscripts from China was 0.28, with a reported interval of 0.16 to 0.41.","vf":"unverified"},{"key":"W2VF7UBP","au":"Campos‐Arceiz, Ahimsa","y":2015,"cx":"Journal","ob":"journal-manuscript","fam":"Gwet-AC","form":"Gwet's kappa (AC1)","estd":"Gwet AC","v":0.17,"n":"1985","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"'accept', 'minor revision', 'major revision', or 'reject'","field":"conservation biology","wr":"two reviewers' recommendations on English-speaking-country manuscripts","conf":"high","self":false,"doi":"10.1016/j.biocon.2015.02.025","ciLow":0.14,"ciHigh":0.2,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Gwet's AC1 for the two reviewers' recommendations on the 1985 manuscripts from English-speaking countries was 0.17, with a reported interval of 0.14 to 0.20.","vf":"unverified"},{"key":"W2VF7UBP","au":"Campos‐Arceiz, Ahimsa","y":2015,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC1k, because reviewers are not the same across manuscripts (ICC1) and average reliability (k) is of interest","estd":"ICC (average)","v":0.55,"n":"108","k":"2","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"'accept', 'minor revision', 'major revision', or 'reject'","field":"conservation biology","wr":"two reviewers' recommendations on manuscripts from China","conf":"high","self":false,"doi":"10.1016/j.biocon.2015.02.025","ciLow":0.34,"ciHigh":0.69,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The ICC1k for the average of the two reviewers on the 108 manuscripts from China was 0.55, with a reported interval of 0.34 to 0.69, higher than for English-speaking countries because reviewers often agreed on reject or major revision.","vf":"unverified"},{"key":"W2VF7UBP","au":"Campos‐Arceiz, Ahimsa","y":2015,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC1k, because reviewers are not the same across manuscripts (ICC1) and average reliability (k) is of interest","estd":"ICC (average)","v":0.25,"n":"1985","k":"2","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"'accept', 'minor revision', 'major revision', or 'reject'","field":"conservation biology","wr":"two reviewers' recommendations on English-speaking-country manuscripts","conf":"high","self":false,"doi":"10.1016/j.biocon.2015.02.025","ciLow":0.18,"ciHigh":0.3,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The ICC1k for the average of the two reviewers on the 1985 manuscripts from English-speaking countries was 0.25, with a reported interval of 0.18 to 0.30.","vf":"unverified"},{"key":"W2VF7UBP","au":"Campos‐Arceiz, Ahimsa","y":2015,"cx":"Journal","ob":"journal-manuscript","fam":"Gwet-AC","form":"Gwet's kappa (AC1)","estd":"Gwet AC","v":0.17,"n":"2093","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"'accept', 'minor revision', 'major revision', or 'reject'","field":"conservation biology","wr":"two reviewers' recommendations on submitted manuscripts","conf":"high","self":false,"doi":"10.1016/j.biocon.2015.02.025","ciLow":0.15,"ciHigh":0.2,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Gwet's AC1 for the two reviewers' recommendations across all 2093 reviewed manuscripts was 0.17, with a reported interval of 0.15 to 0.20, again indicating poor agreement.","vf":"unverified"},{"key":"W2VF7UBP","au":"Campos‐Arceiz, Ahimsa","y":2015,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC1k, because reviewers are not the same across manuscripts (ICC1) and average reliability (k) is of interest","estd":"ICC (average)","v":0.27,"n":"2093","k":"2","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"'accept', 'minor revision', 'major revision', or 'reject'","field":"conservation biology","wr":"two reviewers' recommendations on submitted manuscripts","conf":"high","self":false,"doi":"10.1016/j.biocon.2015.02.025","ciLow":0.2,"ciHigh":0.33,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":true,"ms":"A one-way random intraclass correlation for the average of the two reviewers (ICC1k) across all 2093 reviewed manuscripts was 0.27, with a reported interval of 0.20 to 0.33. This is the study's headline estimate of reviewer reliability.","vf":"unverified"},{"key":"W2VF7UBP","au":"Campos‐Arceiz, Ahimsa","y":2015,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement, both reviewers giving the same recommendation category","estd":"percent agreement","v":0.364,"n":"2093","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"'accept', 'minor revision', 'major revision', or 'reject'","field":"conservation biology","wr":"two reviewers' recommendations on submitted manuscripts","conf":"high","self":false,"doi":"10.1016/j.biocon.2015.02.025","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For 2093 manuscripts each reviewed by two external reviewers, the two reviewers picked exactly the same recommendation category in 36.4% of cases. The authors note that 25% agreement would be expected by chance.","vf":"unverified"},{"key":"W2VF7UBP","au":"Campos‐Arceiz, Ahimsa","y":2015,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"weighted kappa","estd":"weighted kappa","v":0.16,"n":"2093","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"'accept', 'minor revision', 'major revision', or 'reject'","field":"conservation biology","wr":"two reviewers' recommendations on submitted manuscripts","conf":"high","self":false,"doi":"10.1016/j.biocon.2015.02.025","ciLow":0.11,"ciHigh":0.2,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Weighted kappa for the two reviewers' four-category recommendations across all 2093 reviewed manuscripts was 0.16 with a reported interval of 0.11 to 0.20, which the authors describe as poor agreement.","vf":"unverified"},{"key":"W2VF7UBP","au":"Campos‐Arceiz, Ahimsa","y":2015,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement, both reviewers giving the same recommendation category","estd":"percent agreement","v":0.444,"n":"108","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"'accept', 'minor revision', 'major revision', or 'reject'","field":"conservation biology","wr":"two reviewers' recommendations on manuscripts from China","conf":"high","self":false,"doi":"10.1016/j.biocon.2015.02.025","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For the 108 reviewed manuscripts whose corresponding author was at a Chinese institution, the two reviewers gave exactly the same recommendation in 44.4% of cases.","vf":"unverified"},{"key":"W2VF7UBP","au":"Campos‐Arceiz, Ahimsa","y":2015,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement, both reviewers giving the same recommendation category","estd":"percent agreement","v":0.36,"n":"1985","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"'accept', 'minor revision', 'major revision', or 'reject'","field":"conservation biology","wr":"two reviewers' recommendations on English-speaking-country manuscripts","conf":"high","self":false,"doi":"10.1016/j.biocon.2015.02.025","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For the 1985 reviewed manuscripts from English-speaking countries, the two reviewers gave exactly the same recommendation in 36.0% of cases.","vf":"unverified"},{"key":"W2VF7UBP","au":"Campos‐Arceiz, Ahimsa","y":2015,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"weighted kappa","estd":"weighted kappa","v":0.38,"n":"108","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"'accept', 'minor revision', 'major revision', or 'reject'","field":"conservation biology","wr":"two reviewers' recommendations on manuscripts from China","conf":"high","self":false,"doi":"10.1016/j.biocon.2015.02.025","ciLow":0.22,"ciHigh":0.53,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Weighted kappa for the two reviewers' recommendations on the 108 manuscripts from China was 0.38, with a reported interval of 0.22 to 0.53, which the authors call fair agreement.","vf":"unverified"},{"key":"W2VF7UBP","au":"Campos‐Arceiz, Ahimsa","y":2015,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"weighted kappa","estd":"weighted kappa","v":0.14,"n":"1985","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"'accept', 'minor revision', 'major revision', or 'reject'","field":"conservation biology","wr":"two reviewers' recommendations on English-speaking-country manuscripts","conf":"high","self":false,"doi":"10.1016/j.biocon.2015.02.025","ciLow":0.1,"ciHigh":0.18,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Weighted kappa for the two reviewers' recommendations on the 1985 manuscripts from English-speaking countries was 0.14, with a reported interval of 0.10 to 0.18.","vf":"unverified"},{"key":"KNJE6VS9","au":"Cao, Jing","y":2010,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Kendall's tau (suitable for ordinal data with ties)","estd":"correlation","v":0.56,"n":"39","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1-3; 3 definitely fund, 1 not worthy of funding","field":"multi-field","wr":"panel judges on grant proposal scores","conf":"high","self":false,"doi":"10.3102/1076998609353116","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":false,"ms":"Ten faculty panel judges each scored 39 grant proposals on a 1 to 3 scale; the highest pairwise Kendall tau between any two judges was 0.56, so even the most concordant pair agreed only modestly.","vf":"unverified"},{"key":"KNJE6VS9","au":"Cao, Jing","y":2010,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Kendall's tau (suitable for ordinal data with ties)","estd":"correlation","v":-0.25,"n":"39","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1-3; 3 definitely fund, 1 not worthy of funding","field":"multi-field","wr":"panel judges on grant proposal scores","conf":"high","self":false,"doi":"10.3102/1076998609353116","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Among the same ten judges scoring 39 proposals, six pairwise Kendall tau correlations were negative, the lowest being -0.25, showing some judge pairs ordered proposals in nearly opposite ways.","vf":"unverified"},{"key":"TZXYKKCC","au":"Carpenter, Afton S","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"R²","estd":"correlation","v":0.74,"n":"260","k":"","samp":"special","blind":"unclear","agg":"unspecified","scale":"1.0–5.0, 1 highest merit, 5 lowest merit","field":"biomedical","wr":"premeeting and final grant merit scores","conf":"med","self":false,"doi":"10.1136/bmjopen-2015-009138","ciLow":null,"ciHigh":null,"mt":"other","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":false,"ms":"Two assigned reviewers provided premeeting scores and voting panel members provided final scores for 260 applications. The regression yielded R²=0.74 in face-to-face panels.","vf":"unverified"},{"key":"TZXYKKCC","au":"Carpenter, Afton S","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"R²","estd":"correlation","v":0.82,"n":"212","k":"","samp":"special","blind":"unclear","agg":"unspecified","scale":"1.0–5.0, 1 highest merit, 5 lowest merit","field":"biomedical","wr":"premeeting and final grant merit scores","conf":"med","self":false,"doi":"10.1136/bmjopen-2015-009138","ciLow":null,"ciHigh":null,"mt":"other","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Two assigned reviewers provided premeeting scores and voting panel members provided final scores for 212 applications. The regression yielded R²=0.82 in teleconference panels.","vf":"unverified"},{"key":"TZXYKKCC","au":"Carpenter, Afton S","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement between primary and secondary reviewer scores","estd":"percent agreement","v":0.277,"n":"260","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"1.0–5.0, 1 highest merit, 5 lowest merit","field":"biomedical","wr":"assigned reviewers on grant merit","conf":"med","self":false,"doi":"10.1136/bmjopen-2015-009138","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Primary and secondary reviewers rescored 260 face-to-face applications after discussion. They gave exactly the same postdiscussion score to 27.7% of applications.","vf":"unverified"},{"key":"TZXYKKCC","au":"Carpenter, Afton S","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement between primary and secondary reviewer scores","estd":"percent agreement","v":0.208,"n":"212","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"1.0–5.0, 1 highest merit, 5 lowest merit","field":"biomedical","wr":"assigned reviewers on grant merit","conf":"med","self":false,"doi":"10.1136/bmjopen-2015-009138","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Primary and secondary reviewers rescored 212 teleconference applications after discussion. They gave exactly the same postdiscussion score to 20.8% of applications.","vf":"unverified"},{"key":"TZXYKKCC","au":"Carpenter, Afton S","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement between primary and secondary reviewer scores","estd":"percent agreement","v":0.123,"n":"260","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"1.0–5.0, 1 highest merit, 5 lowest merit","field":"biomedical","wr":"assigned reviewers on grant merit","conf":"med","self":false,"doi":"10.1136/bmjopen-2015-009138","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Primary and secondary reviewers independently scored 260 face-to-face applications before discussion. They gave exactly the same score to 12.3% of applications.","vf":"unverified"},{"key":"TZXYKKCC","au":"Carpenter, Afton S","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement between primary and secondary reviewer scores","estd":"percent agreement","v":0.104,"n":"212","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"1.0–5.0, 1 highest merit, 5 lowest merit","field":"biomedical","wr":"assigned reviewers on grant merit","conf":"med","self":false,"doi":"10.1136/bmjopen-2015-009138","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Primary and secondary reviewers independently scored 212 teleconference applications before discussion. They gave exactly the same score to 10.4% of applications.","vf":"unverified"},{"key":"TZXYKKCC","au":"Carpenter, Afton S","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement between premeeting and postdiscussion scores","estd":"percent agreement","v":0.388,"n":"260","k":"1","samp":"special","blind":"unclear","agg":"single-rater","scale":"1.0–5.0, 1 highest merit, 5 lowest merit","field":"biomedical","wr":"primary reviewers on grant merit","conf":"med","self":false,"doi":"10.1136/bmjopen-2015-009138","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Primary reviewers scored 260 face-to-face applications before and after panel discussion. Their score was exactly unchanged for 38.8% of applications.","vf":"unverified"},{"key":"TZXYKKCC","au":"Carpenter, Afton S","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement between premeeting and postdiscussion scores","estd":"percent agreement","v":0.557,"n":"212","k":"1","samp":"special","blind":"unclear","agg":"single-rater","scale":"1.0–5.0, 1 highest merit, 5 lowest merit","field":"biomedical","wr":"primary reviewers on grant merit","conf":"med","self":false,"doi":"10.1136/bmjopen-2015-009138","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Primary reviewers scored 212 teleconference applications before and after panel discussion. Their score was exactly unchanged for 55.7% of applications.","vf":"unverified"},{"key":"TZXYKKCC","au":"Carpenter, Afton S","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement between premeeting and postdiscussion scores","estd":"percent agreement","v":0.415,"n":"260","k":"1","samp":"special","blind":"unclear","agg":"single-rater","scale":"1.0–5.0, 1 highest merit, 5 lowest merit","field":"biomedical","wr":"secondary reviewers on grant merit","conf":"med","self":false,"doi":"10.1136/bmjopen-2015-009138","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Secondary reviewers scored 260 face-to-face applications before and after panel discussion. Their score was exactly unchanged for 41.5% of applications.","vf":"unverified"},{"key":"TZXYKKCC","au":"Carpenter, Afton S","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement between premeeting and postdiscussion scores","estd":"percent agreement","v":0.519,"n":"212","k":"1","samp":"special","blind":"unclear","agg":"single-rater","scale":"1.0–5.0, 1 highest merit, 5 lowest merit","field":"biomedical","wr":"secondary reviewers on grant merit","conf":"med","self":false,"doi":"10.1136/bmjopen-2015-009138","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Secondary reviewers scored 212 teleconference applications before and after panel discussion. Their score was exactly unchanged for 51.9% of applications.","vf":"unverified"},{"key":"Q3AG2M2K","au":"Chen, Shiping","y":2025,"cx":"Journal","ob":"review-report","fam":"Krippendorff","form":"Krippendorff's alpha","estd":"Krippendorff","v":0.502,"n":"48","k":"3","samp":"special","blind":"unclear","agg":"single-rater","scale":"5-point Likert scale","field":"human-computer interaction","wr":"researchers rating quality of participant-written reviews","conf":"high","self":false,"doi":"10.1145/3729176.3729196","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":true,"he":false,"ms":"Three experienced HCI researchers independently scored the accuracy of all 48 participant-written reviews on a 5-point scale; a Krippendorff's alpha of 0.502 indicates moderate agreement. n_ratings_total derived from stated complete crossing (3 raters x 48 reviews).","vf":"unverified"},{"key":"Q3AG2M2K","au":"Chen, Shiping","y":2025,"cx":"Journal","ob":"review-report","fam":"Krippendorff","form":"Krippendorff's alpha","estd":"Krippendorff","v":0.502,"n":"48","k":"3","samp":"special","blind":"unclear","agg":"single-rater","scale":"5-point Likert scale","field":"human-computer interaction","wr":"researchers rating quality of participant-written reviews","conf":"high","self":false,"doi":"10.1145/3729176.3729196","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Three experienced HCI researchers independently scored the clarity and coherence of all 48 participant-written reviews on a 5-point scale; a Krippendorff's alpha of 0.502 indicates moderate agreement. n_ratings_total derived from stated complete crossing (3 raters x 48 reviews).","vf":"unverified"},{"key":"Q3AG2M2K","au":"Chen, Shiping","y":2025,"cx":"Journal","ob":"review-report","fam":"Krippendorff","form":"Krippendorff's alpha","estd":"Krippendorff","v":0.519,"n":"48","k":"3","samp":"special","blind":"unclear","agg":"single-rater","scale":"5-point Likert scale","field":"human-computer interaction","wr":"researchers rating quality of participant-written reviews","conf":"high","self":false,"doi":"10.1145/3729176.3729196","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Three experienced HCI researchers independently scored the completeness of all 48 participant-written reviews on a 5-point scale; a Krippendorff's alpha of 0.519 indicates moderate agreement. n_ratings_total derived from stated complete crossing (3 raters x 48 reviews).","vf":"unverified"},{"key":"Q3AG2M2K","au":"Chen, Shiping","y":2025,"cx":"Journal","ob":"review-report","fam":"Krippendorff","form":"Krippendorff's alpha","estd":"Krippendorff","v":0.503,"n":"48","k":"3","samp":"special","blind":"unclear","agg":"single-rater","scale":"5-point Likert scale","field":"human-computer interaction","wr":"researchers rating quality of participant-written reviews","conf":"high","self":false,"doi":"10.1145/3729176.3729196","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Three experienced HCI researchers independently scored the constructiveness of all 48 participant-written reviews on a 5-point scale; a Krippendorff's alpha of 0.503 indicates moderate agreement. n_ratings_total derived from stated complete crossing (3 raters x 48 reviews).","vf":"unverified"},{"key":"Q3AG2M2K","au":"Chen, Shiping","y":2025,"cx":"Journal","ob":"review-report","fam":"Krippendorff","form":"Krippendorff's alpha","estd":"Krippendorff","v":0.545,"n":"48","k":"3","samp":"special","blind":"unclear","agg":"single-rater","scale":"5-point Likert scale","field":"human-computer interaction","wr":"researchers rating quality of participant-written reviews","conf":"high","self":false,"doi":"10.1145/3729176.3729196","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Three experienced HCI researchers independently scored the tone of all 48 participant-written reviews on a 5-point scale; a Krippendorff's alpha of 0.545 indicates moderate agreement. n_ratings_total derived from stated complete crossing (3 raters x 48 reviews).","vf":"unverified"},{"key":"UYGWPDNF","au":"Chibnik, Michael","y":2016,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"mean agreement score, percent better than expected by chance","estd":"other","v":0.08,"n":"418","k":"3","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"accept / revise and resubmit / reject","field":"anthropology","wr":"three reviewers' recommendations on journal manuscripts","conf":"high","self":false,"doi":"10.1111/aman.12523","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The observed mean agreement score of 1.90 exceeded the chance expectation of 1.76 by only 8 percent. The paper uses this to argue that reviewers agree little more than chance.","vf":"unverified"},{"key":"UYGWPDNF","au":"Chibnik, Michael","y":2016,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"deviation from chance as percent of maximum possible","estd":"other","v":0.123,"n":"418","k":"3","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"accept / revise and resubmit / reject","field":"anthropology","wr":"three reviewers' recommendations on journal manuscripts","conf":"high","self":false,"doi":"10.1111/aman.12523","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":true,"ms":"The gain of the observed mean agreement score over chance was only 12.3 percent of the maximum possible gain. This is the paper's culminating chance-corrected summary of reviewer agreement.","vf":"unverified"},{"key":"UYGWPDNF","au":"Chibnik, Michael","y":2016,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"mean per-manuscript agreement score (perfect=3, good=2, poor=1) averaged across manuscripts; 1.76 expected by chance","estd":"other","v":1.9,"n":"418","k":"3","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"accept / revise and resubmit / reject","field":"anthropology","wr":"three reviewers' recommendations on journal manuscripts","conf":"high","self":false,"doi":"10.1111/aman.12523","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Each manuscript's three recommendations were scored 3 for perfect, 2 for good, or 1 for poor agreement and averaged across the 418 manuscripts; the observed mean of 1.90 compares with 1.76 expected by chance.","vf":"unverified"},{"key":"UYGWPDNF","au":"Chibnik, Michael","y":2016,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement: all three reviewers gave the identical recommendation","estd":"percent agreement","v":0.165,"n":"418","k":"3","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"accept / revise and resubmit / reject","field":"anthropology","wr":"three reviewers' recommendations on journal manuscripts","conf":"high","self":false,"doi":"10.1111/aman.12523","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For each of 418 American Anthropologist manuscripts, three reviewers independently recommended accept, revise and resubmit, or reject; all three gave the identical recommendation for 16.5 percent of manuscripts, against 12.1 percent expected by chance.","vf":"unverified"},{"key":"UYGWPDNF","au":"Chibnik, Michael","y":2016,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"within-one-category agreement: perfect or good agreement combined (a majority agree and any third reviewer differs by only one adjacent category)","estd":"percent agreement","v":0.739,"n":"418","k":"3","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"accept / revise and resubmit / reject","field":"anthropology","wr":"three reviewers' recommendations on journal manuscripts","conf":"high","self":false,"doi":"10.1111/aman.12523","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Combining perfect agreement with good agreement, where a majority agree and the third reviewer differs by one adjacent category, reviewers agreed on 73.9 percent of the 418 manuscripts. The paper stresses this figure is misleading because much agreement is expected by chance.","vf":"unverified"},{"key":"UYGWPDNF","au":"Chibnik, Michael","y":2016,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"correlation (type unspecified) between summed reviewer recommendation score and editor's initial decision","estd":"correlation","v":0.67,"n":"418","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"reviewer recommendation score 0-6 vs initial decision 0-2","field":"anthropology","wr":"reviewers' recommendations vs editor's initial decision on manuscripts","conf":"high","self":false,"doi":"10.1111/aman.12523","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"restricted-other","pr":true,"he":false,"ms":"Across 418 manuscripts, the summed reviewer recommendation score (0 to 6) correlated .67 with the editor's initial decision, scored 2 for accept, 1 for revise and resubmit, 0 for reject. Editors decided after reading the reviews, so this indexes reviewer influence on editors rather than independent inter-rater agreement.","vf":"unverified"},{"key":"UYGWPDNF","au":"Chibnik, Michael","y":2016,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"correlation (type unspecified) between summed reviewer recommendation score and editor's ultimate accept/reject decision","estd":"correlation","v":0.59,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"reviewer recommendation score 0-6 vs ultimate decision accept/reject","field":"anthropology","wr":"reviewers' recommendations vs editor's ultimate accept/reject decision","conf":"high","self":false,"doi":"10.1111/aman.12523","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"restricted-other","pr":false,"he":false,"ms":"The summed reviewer recommendation score correlated .59 with the editor's ultimate accept or reject decision, excluding revise-and-resubmit manuscripts never resubmitted. Editors had read the reviews before deciding, so this reflects influence rather than independent agreement.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"specific-category observed agreement, exact (2-category, no partial credit)","estd":"percent agreement","v":0.7077,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"reject / possibly accept","field":"psychology","wr":"observed agreement on the dichotomised possibly-accept category","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"In the two-category schema, observed agreement between the two reviews on the possibly accept category was 0.7077.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"specific-category observed agreement, exact (2-category, no partial credit)","estd":"percent agreement","v":0.8257,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"reject / possibly accept","field":"psychology","wr":"observed agreement on the dichotomised reject category","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"In the two-category schema, observed agreement between the two reviews on the reject category was 0.8257.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"specific-category observed agreement (PO), weights 1, .5, 0","estd":"percent agreement","v":0.6364,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"reject / resubmit / accept (collapsed from 5)","field":"psychology","wr":"observed agreement on the collapsed accept recommendation","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"With recommendations collapsed to three categories, observed agreement between the two reviews on the accept category was 0.6364.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"specific-category observed agreement (PO), weights 1, .5, 0","estd":"percent agreement","v":0.8624,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"reject / resubmit / accept (collapsed from 5)","field":"psychology","wr":"observed agreement on the collapsed reject recommendation","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"With recommendations collapsed to three categories, observed agreement between the two reviews on the reject category was 0.8624.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"specific-category observed agreement (PO), weights 1, .5, 0","estd":"percent agreement","v":0.8437,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"reject / resubmit / accept (collapsed from 5)","field":"psychology","wr":"observed agreement on the resubmit recommendation","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"With recommendations collapsed to three categories, observed agreement between the two reviews on the resubmit category was 0.8437.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"specific-category observed agreement (PO), linear partial-credit weighting","estd":"percent agreement","v":0.7222,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"1 = reject to 5 = accept in present form","field":"psychology","wr":"observed agreement on the accept-as-is recommendation","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Within the five-category schema, observed agreement between the two reviews on the accept in present form category was 0.7222.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"specific-category observed agreement (PO), linear partial-credit weighting","estd":"percent agreement","v":0.6,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"1 = reject to 5 = accept in present form","field":"psychology","wr":"observed agreement on the accept-with-minor-revision recommendation","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Within the five-category schema, observed agreement between the two reviews on the accept subject to minor revision category was 0.6000.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"specific-category observed agreement (PO), linear partial-credit weighting","estd":"percent agreement","v":0.8621,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"1 = reject to 5 = accept in present form","field":"psychology","wr":"observed agreement on the reject recommendation","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Within the five-category schema, observed agreement between the two reviews on the reject category was 0.8621.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"specific-category observed agreement (PO), linear partial-credit weighting","estd":"percent agreement","v":0.8672,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"1 = reject to 5 = accept in present form","field":"psychology","wr":"observed agreement on the reject/resubmit recommendation","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Within the five-category schema, observed agreement between the two reviews on the reject and resubmit category was 0.8672.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"specific-category observed agreement (PO), linear partial-credit weighting","estd":"percent agreement","v":0.8523,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"1 = reject to 5 = accept in present form","field":"psychology","wr":"observed agreement on reject/submit-to-technical-journal recommendation","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Within the five-category schema, observed agreement between the two reviews on the reject and submit to a technical journal category was 0.8523.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"specific-category chance-corrected agreement (Cicchetti, Lee, Fontana, & Dowds, 1978)","estd":"kappa","v":0.53,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"reject / possibly accept","field":"psychology","wr":"reviewer agreement on the dichotomised possibly-accept category","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"In the two-category schema, chance-corrected agreement between the two reviews on the possibly accept category was 0.53, identical to the overall dichotomous level by degrees-of-freedom restriction.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"specific-category chance-corrected agreement (Cicchetti, Lee, Fontana, & Dowds, 1978)","estd":"kappa","v":0.53,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"reject / possibly accept","field":"psychology","wr":"reviewer agreement on the dichotomised reject category","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"In the two-category schema, chance-corrected agreement between the two reviews on the reject category was 0.53, identical to the overall dichotomous level by degrees-of-freedom restriction.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"specific-category chance-corrected agreement (Cicchetti, Lee, Fontana, & Dowds, 1978)","estd":"kappa","v":0.5,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"reject / resubmit / accept (collapsed from 5)","field":"psychology","wr":"reviewer agreement on the collapsed accept recommendation","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"With recommendations collapsed to three categories, chance-corrected agreement between the two reviews on the accept category was 0.50.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"specific-category chance-corrected agreement (Cicchetti, Lee, Fontana, & Dowds, 1978)","estd":"kappa","v":0.51,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"reject / resubmit / accept (collapsed from 5)","field":"psychology","wr":"reviewer agreement on the collapsed reject recommendation","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"With recommendations collapsed to three categories, chance-corrected agreement between the two reviews on the reject category was 0.51.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"specific-category chance-corrected agreement (Cicchetti, Lee, Fontana, & Dowds, 1978)","estd":"kappa","v":0.62,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"reject / resubmit / accept (collapsed from 5)","field":"psychology","wr":"reviewer agreement on the resubmit recommendation","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"With recommendations collapsed to three categories, chance-corrected agreement between the two reviews on the resubmit category was 0.62, the highest.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"specific-category chance-corrected agreement (Cicchetti, Lee, Fontana, & Dowds, 1978)","estd":"kappa","v":0.61,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"1 = reject to 5 = accept in present form","field":"psychology","wr":"reviewer agreement on the accept-as-is recommendation","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Within the five-category schema, chance-corrected agreement between the two reviews on the accept in present form category was 0.61.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"specific-category chance-corrected agreement (Cicchetti, Lee, Fontana, & Dowds, 1978)","estd":"kappa","v":0.23,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"1 = reject to 5 = accept in present form","field":"psychology","wr":"reviewer agreement on the accept-with-minor-revision recommendation","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Within the five-category schema, chance-corrected agreement between the two reviews on the accept subject to minor revision category was only 0.23, the least reliable category.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"specific-category chance-corrected agreement (Cicchetti, Lee, Fontana, & Dowds, 1978)","estd":"kappa","v":0.53,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"1 = reject to 5 = accept in present form","field":"psychology","wr":"reviewer agreement on the reject recommendation","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Within the five-category schema, chance-corrected agreement between the two reviews on the reject category was 0.53.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"specific-category chance-corrected agreement (Cicchetti, Lee, Fontana, & Dowds, 1978)","estd":"kappa","v":0.63,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"1 = reject to 5 = accept in present form","field":"psychology","wr":"reviewer agreement on the reject/resubmit recommendation","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Within the five-category schema, chance-corrected agreement between the two reviews on the reject and resubmit category was 0.63, the most reliable category.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"specific-category chance-corrected agreement (Cicchetti, Lee, Fontana, & Dowds, 1978)","estd":"kappa","v":0.49,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"1 = reject to 5 = accept in present form","field":"psychology","wr":"reviewer agreement on reject/submit-to-technical-journal recommendation","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Within the five-category schema, chance-corrected agreement between the two reviews on the reject and submit to a technical journal category was 0.49.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (Ri); different manuscripts evaluated by different reviewers","estd":"ICC (single/unspec)","v":0.53,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"reject / possibly accept","field":"psychology","wr":"reviewers' recommendations on journal manuscripts","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"With recommendations collapsed to a reject or possibly accept dichotomy, the one-way intraclass correlation between the two independent reviews of 87 manuscripts was 0.53.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (Ri); different manuscripts evaluated by different reviewers","estd":"ICC (single/unspec)","v":0.5,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"reject / resubmit / accept (collapsed from 5)","field":"psychology","wr":"reviewers' recommendations on journal manuscripts","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"With the recommendation categories collapsed to three, the one-way intraclass correlation between the two independent reviews of 87 manuscripts was 0.50.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (Ri); different manuscripts evaluated by different reviewers","estd":"ICC (single/unspec)","v":0.54,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"1 = reject to 5 = accept in present form","field":"psychology","wr":"reviewers' recommendations on journal manuscripts","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Two independent reviewers each rated 87 manuscripts on a five-category recommendation scale, with different reviewer pairs across manuscripts. A one-way intraclass correlation of 0.54 indicates moderate agreement.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"kappa (Cohen, 1960), 2-category, no partial agreement weights possible","estd":"kappa","v":0.53,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"reject / possibly accept","field":"psychology","wr":"reviewers' recommendations on journal manuscripts","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"With recommendations dichotomised to reject versus possibly accept, an unweighted Cohen kappa of 0.53 gives the chance-corrected agreement between the two independent reviews.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"observed proportion agreement, exact (2-category, no partial credit)","estd":"percent agreement","v":0.78,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"reject / possibly accept","field":"psychology","wr":"reviewers' dichotomous recommendations on journal manuscripts","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"In the two-category reject or possibly accept schema, the two independent reviews of 87 manuscripts gave the same recommendation for 78 percent of manuscripts.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"weighted observed agreement (PO), weights 1, .5, 0","estd":"percent agreement","v":0.8161,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"reject / resubmit / accept (collapsed from 5)","field":"psychology","wr":"reviewers' three-category recommendations on journal manuscripts","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the three-category schema, the weighted observed agreement between the two independent reviews of 87 manuscripts was 0.8161, reported as .82 in the text.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"weighted observed agreement (PO), linear partial-credit weights by category distance","estd":"percent agreement","v":0.8247,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"1 = reject to 5 = accept in present form","field":"psychology","wr":"reviewers' five-category recommendations on journal manuscripts","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the five-category schema, the linearly weighted observed agreement between the two independent reviews of 87 manuscripts was 0.8247, reported as .82 in the text.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"observed agreement within one category (5-category scale)","estd":"percent agreement","v":0.79,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"1 = reject to 5 = accept in present form","field":"psychology","wr":"reviewers' five-category recommendations on journal manuscripts","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"On the full five-category scale, the two independent reviews of 87 manuscripts agreed exactly or within one category on 79 percent of manuscripts.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"weighted kappa, linear weights (1, .5, 0)","estd":"weighted kappa","v":0.53,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"reject / resubmit / accept (collapsed from 5)","field":"psychology","wr":"reviewers' recommendations on journal manuscripts","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"With recommendations collapsed to three categories, a linearly weighted kappa of 0.53 gives the chance-corrected agreement between the two independent reviews.","vf":"unverified"},{"key":"RRJ5XCFM","au":"Cicchetti, Domenic V.","y":1980,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"weighted kappa, linear weights (Cohen, 1968; Cicchetti, 1976 formulas)","estd":"weighted kappa","v":0.52,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"1 = reject to 5 = accept in present form","field":"psychology","wr":"reviewers' recommendations on journal manuscripts","conf":"high","self":false,"doi":"10.1037/0003-066x.35.3.300","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the full five-category recommendation, a linearly weighted kappa of 0.52 gives the chance-corrected agreement between the two independent reviews of 87 manuscripts.","vf":"unverified"},{"key":"QQRFQCNU","au":"Cicchetti, Domenic V.","y":1985,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Finn's r (rF)","estd":"other","v":0.56,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"possibly accept / reject (five categories collapsed to two)","field":"psychology","wr":"reviewers' recommendations on journal manuscripts","conf":"high","self":false,"doi":"10.1037/0003-066x.40.5.563","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"For the same 87 American Psychologist manuscripts collapsed to accept versus reject, Finn's r, an alternative agreement statistic, was 0.56.","vf":"unverified"},{"key":"QQRFQCNU","au":"Cicchetti, Domenic V.","y":1985,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (rI)","estd":"ICC (single/unspec)","v":0.53,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"possibly accept / reject (five categories collapsed to two)","field":"psychology","wr":"reviewers' recommendations on journal manuscripts","conf":"high","self":false,"doi":"10.1037/0003-066x.40.5.563","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":true,"he":false,"ms":"Two reviewers' recommendations on 87 American Psychologist manuscripts, collapsed to possibly accept versus reject, gave an intraclass correlation of 0.53, the paper's own headline actual-case result.","vf":"unverified"},{"key":"QQRFQCNU","au":"Cicchetti, Domenic V.","y":1985,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"observed exact agreement (proportion) on the acceptance category","estd":"percent agreement","v":0.7077,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"possibly accept / reject (five categories collapsed to two)","field":"psychology","wr":"reviewers' recommendations on journal manuscripts","conf":"high","self":false,"doi":"10.1037/0003-066x.40.5.563","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"On the acceptance category of the collapsed American Psychologist data, observed exact reviewer agreement was 70.77 per cent.","vf":"unverified"},{"key":"QQRFQCNU","au":"Cicchetti, Domenic V.","y":1985,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"observed exact agreement (proportion), collapsed accept/reject","estd":"percent agreement","v":0.7816,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"possibly accept / reject (five categories collapsed to two)","field":"psychology","wr":"reviewers' recommendations on journal manuscripts","conf":"high","self":false,"doi":"10.1037/0003-066x.40.5.563","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"Across the 87 American Psychologist manuscripts collapsed to accept versus reject, the two reviewers agreed exactly on 78.16 per cent of cases.","vf":"unverified"},{"key":"QQRFQCNU","au":"Cicchetti, Domenic V.","y":1985,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"observed exact agreement (proportion) on the rejection category","estd":"percent agreement","v":0.8257,"n":"87","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"possibly accept / reject (five categories collapsed to two)","field":"psychology","wr":"reviewers' recommendations on journal manuscripts","conf":"high","self":false,"doi":"10.1037/0003-066x.40.5.563","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"On the rejection category of the collapsed American Psychologist data, observed exact reviewer agreement was 82.57 per cent.","vf":"unverified"},{"key":"QQRFQCNU","au":"Cicchetti, Domenic V.","y":1985,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Finn's r (rF)","estd":"other","v":0.43,"n":"73","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"possibly accept / possibly reject (four categories collapsed to two)","field":"psychology","wr":"reviewers' recommendations on journal manuscripts","conf":"high","self":false,"doi":"10.1037/0003-066x.40.5.563","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"For the collapsed accept-reject coding of the 73 Developmental Review and Merrill-Palmer Quarterly manuscripts, Finn's r was 0.43.","vf":"unverified"},{"key":"QQRFQCNU","au":"Cicchetti, Domenic V.","y":1985,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (rI)","estd":"ICC (single/unspec)","v":0.27,"n":"73","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"possibly accept / possibly reject (four categories collapsed to two)","field":"psychology","wr":"reviewers' recommendations on journal manuscripts","conf":"high","self":false,"doi":"10.1037/0003-066x.40.5.563","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"Two reviewers' recommendations on 73 Developmental Review and Merrill-Palmer Quarterly manuscripts, collapsed to possibly accept versus possibly reject, gave an intraclass correlation of 0.27.","vf":"unverified"},{"key":"QQRFQCNU","au":"Cicchetti, Domenic V.","y":1985,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"observed exact agreement (proportion) on the acceptance category","estd":"percent agreement","v":0.4615,"n":"73","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"possibly accept / possibly reject (four categories collapsed to two)","field":"psychology","wr":"reviewers' recommendations on journal manuscripts","conf":"high","self":false,"doi":"10.1037/0003-066x.40.5.563","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"On the acceptance category of the collapsed Developmental Review and Merrill-Palmer Quarterly data, observed exact reviewer agreement was 46.15 per cent.","vf":"unverified"},{"key":"QQRFQCNU","au":"Cicchetti, Domenic V.","y":1985,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"observed exact agreement (proportion), collapsed accept/reject","estd":"percent agreement","v":0.7123,"n":"73","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"possibly accept / possibly reject (four categories collapsed to two)","field":"psychology","wr":"reviewers' recommendations on journal manuscripts","conf":"high","self":false,"doi":"10.1037/0003-066x.40.5.563","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"Across the 73 Developmental Review and Merrill-Palmer Quarterly manuscripts collapsed to accept versus reject, the two reviewers agreed exactly on 71.23 per cent of cases.","vf":"unverified"},{"key":"QQRFQCNU","au":"Cicchetti, Domenic V.","y":1985,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"observed exact agreement (proportion) on the rejection category","estd":"percent agreement","v":0.8037,"n":"73","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"possibly accept / possibly reject (four categories collapsed to two)","field":"psychology","wr":"reviewers' recommendations on journal manuscripts","conf":"high","self":false,"doi":"10.1037/0003-066x.40.5.563","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"On the rejection category of the collapsed Developmental Review and Merrill-Palmer Quarterly data, observed exact reviewer agreement was 80.37 per cent.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass correlation coefficient (R_i, Model III), based on the average number of reviews per grant proposal","estd":"ICC (single/unspec)","v":0.18,"n":"50","k":"4.24","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"10 (lowest score possible) to 50 (highest score possible)","field":"multi-field","wr":"reviewers scoring grant proposals","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":true,"ms":"Reviewers independently scored 50 chemical dynamics grant proposals (COSPUP blind, mean 4.24 reviews each); an intraclass correlation (R_i, Model III) of 0.18 indicates low chance-corrected agreement between reviewers.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass correlation coefficient (R_i, Model III), based on the average number of reviews per grant proposal","estd":"ICC (single/unspec)","v":0.37,"n":"49","k":"4.04","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"10 (lowest score possible) to 50 (highest score possible)","field":"multi-field","wr":"reviewers scoring grant proposals","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":true,"ms":"Reviewers independently scored 49 economics grant proposals (COSPUP blind, mean 4.04 reviews each); an intraclass correlation (R_i, Model III) of 0.37 indicates low chance-corrected agreement between reviewers.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass correlation coefficient (R_i, Model III), based on the average number of reviews per grant proposal","estd":"ICC (single/unspec)","v":0.33,"n":"50","k":"4.06","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"10 (lowest score possible) to 50 (highest score possible)","field":"multi-field","wr":"reviewers scoring grant proposals","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":true,"ms":"Reviewers independently scored 50 solid state physics grant proposals (COSPUP blind, mean 4.06 reviews each); an intraclass correlation (R_i, Model III) of 0.33 indicates low chance-corrected agreement between reviewers.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass correlation coefficient (R_i, Model III), based on the average number of reviews per grant proposal","estd":"ICC (single/unspec)","v":0.32,"n":"50","k":"4.26","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"10 (lowest score possible) to 50 (highest score possible)","field":"multi-field","wr":"reviewers scoring grant proposals","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":true,"ms":"Reviewers independently scored 50 chemical dynamics grant proposals (COSPUP open, mean 4.26 reviews each); an intraclass correlation (R_i, Model III) of 0.32 indicates low chance-corrected agreement between reviewers.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass correlation coefficient (R_i, Model III), based on the average number of reviews per grant proposal","estd":"ICC (single/unspec)","v":0.36,"n":"49","k":"3.69","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"10 (lowest score possible) to 50 (highest score possible)","field":"multi-field","wr":"reviewers scoring grant proposals","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":true,"ms":"Reviewers independently scored 49 economics grant proposals (COSPUP open, mean 3.69 reviews each); an intraclass correlation (R_i, Model III) of 0.36 indicates low chance-corrected agreement between reviewers.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass correlation coefficient (R_i, Model III), based on the average number of reviews per grant proposal","estd":"ICC (single/unspec)","v":0.34,"n":"49","k":"3.86","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"10 (lowest score possible) to 50 (highest score possible)","field":"multi-field","wr":"reviewers scoring grant proposals","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":true,"ms":"Reviewers independently scored 49 solid state physics grant proposals (COSPUP open, mean 3.86 reviews each); an intraclass correlation (R_i, Model III) of 0.34 indicates low chance-corrected agreement between reviewers.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass correlation coefficient (R_i, Model III), based on the average number of reviews per grant proposal","estd":"ICC (single/unspec)","v":0.25,"n":"50","k":"4.84","samp":"unclear","blind":"single","agg":"single-rater","scale":"10 (lowest score possible) to 50 (highest score possible)","field":"multi-field","wr":"reviewers scoring grant proposals","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":true,"ms":"Reviewers independently scored 50 chemical dynamics grant proposals (NSF open, mean 4.84 reviews each); an intraclass correlation (R_i, Model III) of 0.25 indicates low chance-corrected agreement between reviewers.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass correlation coefficient (R_i, Model III), based on the average number of reviews per grant proposal","estd":"ICC (single/unspec)","v":0.37,"n":"42","k":"3.69","samp":"unclear","blind":"single","agg":"single-rater","scale":"10 (lowest score possible) to 50 (highest score possible)","field":"multi-field","wr":"reviewers scoring grant proposals","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":true,"ms":"Reviewers independently scored 42 economics grant proposals (NSF open, mean 3.69 reviews each); an intraclass correlation (R_i, Model III) of 0.37 indicates low chance-corrected agreement between reviewers.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass correlation coefficient (R_i, Model III), based on the average number of reviews per grant proposal","estd":"ICC (single/unspec)","v":0.32,"n":"50","k":"3.84","samp":"unclear","blind":"single","agg":"single-rater","scale":"10 (lowest score possible) to 50 (highest score possible)","field":"multi-field","wr":"reviewers scoring grant proposals","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":true,"ms":"Reviewers independently scored 50 solid state physics grant proposals (NSF open, mean 3.84 reviews each); an intraclass correlation (R_i, Model III) of 0.32 indicates low chance-corrected agreement between reviewers.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"exact agreement on 'accept' recommendations","estd":"percent agreement","v":0.66,"n":"62","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"dichotomised: accept vs reject","field":"multi-field","wr":"reviewers agreeing to accept manuscripts","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Among 62 manuscripts one reviewer recommended for acceptance at American Psychologist, the second reviewer agreed 66% of the time, showing lower agreement on acceptance.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"overall exact agreement on dichotomised accept/reject","estd":"percent agreement","v":0.74,"n":"159","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"dichotomised: accept vs reject","field":"multi-field","wr":"reviewers' overall accept/reject agreement","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Overall, two independent reviewers gave the same accept/reject recommendation on 74% of 159 manuscripts submitted to American Psychologist.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"exact agreement on 'reject' recommendations","estd":"percent agreement","v":0.78,"n":"97","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"dichotomised: accept vs reject","field":"multi-field","wr":"reviewers agreeing to reject manuscripts","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Among 97 manuscripts one reviewer recommended for rejection at American Psychologist, the second reviewer agreed 78% of the time, higher than agreement on acceptance.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"R_i on dichotomised accept/reject (Methods: Model I for manuscript-review designs)","estd":"ICC (single/unspec)","v":0.45,"n":"159","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"dichotomised: accept vs reject","field":"multi-field","wr":"reviewers' accept/reject on manuscripts","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers' dichotomised accept/reject recommendations on 159 manuscripts submitted to American Psychologist yielded a chance-corrected agreement (R_i) of 0.45, indicating low reliability.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"exact agreement on 'accept' recommendations","estd":"percent agreement","v":0.52,"n":"25","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"dichotomised: accept vs reject","field":"multi-field","wr":"reviewers agreeing to accept manuscripts","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Among 25 manuscripts one reviewer recommended for acceptance at Developmental Review, the second reviewer agreed 52% of the time, showing lower agreement on acceptance.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"overall exact agreement on dichotomised accept/reject","estd":"percent agreement","v":0.67,"n":"72","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"dichotomised: accept vs reject","field":"multi-field","wr":"reviewers' overall accept/reject agreement","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Overall, two independent reviewers gave the same accept/reject recommendation on 67% of 72 manuscripts submitted to Developmental Review.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"exact agreement on 'reject' recommendations","estd":"percent agreement","v":0.74,"n":"47","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"dichotomised: accept vs reject","field":"multi-field","wr":"reviewers agreeing to reject manuscripts","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Among 47 manuscripts one reviewer recommended for rejection at Developmental Review, the second reviewer agreed 74% of the time, higher than agreement on acceptance.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"R_i on dichotomised accept/reject (Methods: Model I for manuscript-review designs)","estd":"ICC (single/unspec)","v":0.27,"n":"72","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"dichotomised: accept vs reject","field":"multi-field","wr":"reviewers' accept/reject on manuscripts","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers' dichotomised accept/reject recommendations on 72 manuscripts submitted to Developmental Review yielded a chance-corrected agreement (R_i) of 0.27, indicating low reliability.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"exact agreement on 'accept' recommendations","estd":"percent agreement","v":0.44,"n":"462","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"dichotomised: accept vs reject","field":"multi-field","wr":"reviewers agreeing to accept manuscripts","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Among 462 manuscripts one reviewer recommended for acceptance at Journal of Abnormal Psychology, the second reviewer agreed 44% of the time, showing lower agreement on acceptance.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"overall exact agreement on dichotomised accept/reject","estd":"percent agreement","v":0.61,"n":"1319","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"dichotomised: accept vs reject","field":"multi-field","wr":"reviewers' overall accept/reject agreement","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Overall, two independent reviewers gave the same accept/reject recommendation on 61% of 1319 manuscripts submitted to Journal of Abnormal Psychology.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"exact agreement on 'reject' recommendations","estd":"percent agreement","v":0.7,"n":"857","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"dichotomised: accept vs reject","field":"multi-field","wr":"reviewers agreeing to reject manuscripts","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Among 857 manuscripts one reviewer recommended for rejection at Journal of Abnormal Psychology, the second reviewer agreed 70% of the time, higher than agreement on acceptance.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"R_i on dichotomised accept/reject (Methods: Model I for manuscript-review designs)","estd":"ICC (single/unspec)","v":0.14,"n":"1319","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"dichotomised: accept vs reject","field":"multi-field","wr":"reviewers' accept/reject on manuscripts","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers' dichotomised accept/reject recommendations on 1319 manuscripts submitted to Journal of Abnormal Psychology yielded a chance-corrected agreement (R_i) of 0.14, indicating low reliability.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"exact agreement on 'accept' recommendations","estd":"percent agreement","v":0.5,"n":"289","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"dichotomised: accept vs reject","field":"multi-field","wr":"reviewers agreeing to accept manuscripts","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Among 289 manuscripts one reviewer recommended for acceptance at Untitled Medical Specialty Journal, the second reviewer agreed 50% of the time, showing lower agreement on acceptance.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"overall exact agreement on dichotomised accept/reject","estd":"percent agreement","v":0.67,"n":"866","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"dichotomised: accept vs reject","field":"multi-field","wr":"reviewers' overall accept/reject agreement","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Overall, two independent reviewers gave the same accept/reject recommendation on 67% of 866 manuscripts submitted to Untitled Medical Specialty Journal.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"exact agreement on 'reject' recommendations","estd":"percent agreement","v":0.76,"n":"577","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"dichotomised: accept vs reject","field":"multi-field","wr":"reviewers agreeing to reject manuscripts","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Among 577 manuscripts one reviewer recommended for rejection at Untitled Medical Specialty Journal, the second reviewer agreed 76% of the time, higher than agreement on acceptance.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"R_i on dichotomised accept/reject (Methods: Model I for manuscript-review designs)","estd":"ICC (single/unspec)","v":0.26,"n":"866","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"dichotomised: accept vs reject","field":"multi-field","wr":"reviewers' accept/reject on manuscripts","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers' dichotomised accept/reject recommendations on 866 manuscripts submitted to Untitled Medical Specialty Journal yielded a chance-corrected agreement (R_i) of 0.26, indicating low reliability.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"overall exact agreement on dichotomised high/low ratings","estd":"percent agreement","v":0.6,"n":"50","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"dichotomised: high (40-50) vs low (10-39)","field":"multi-field","wr":"reviewers' overall high/low agreement","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":true,"ms":"Reviewers gave the same high/low classification on 60% of the 50 chemical dynamics grant proposals overall.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"exact agreement on high (40-50) rated proposals","estd":"percent agreement","v":0.41,"n":"17","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"dichotomised: high (40-50) vs low (10-39)","field":"multi-field","wr":"reviewers agreeing on high-rated proposals","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Reviewers agreed on 41% of the 17 highly rated chemical dynamics grant proposals, lower than agreement on low-rated proposals.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"exact agreement on low (10-39) rated proposals","estd":"percent agreement","v":0.7,"n":"33","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"dichotomised: high (40-50) vs low (10-39)","field":"multi-field","wr":"reviewers agreeing on low-rated proposals","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Reviewers agreed on 70% of the 33 low-rated chemical dynamics grant proposals, higher than agreement on high-rated proposals.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"chance-corrected agreement R_i on dichotomised high/low grant ratings","estd":"ICC (single/unspec)","v":0.16,"n":"50","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"dichotomised: high (40-50) vs low (10-39)","field":"multi-field","wr":"reviewers' high/low ratings on grant proposals","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":true,"ms":"NSF and COSPUP reviewers' dichotomised high/low ratings on 50 chemical dynamics grant proposals gave a chance-corrected agreement (R_i) of 0.16 (not significant).","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"overall exact agreement on dichotomised high/low ratings","estd":"percent agreement","v":0.68,"n":"150","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"dichotomised: high (40-50) vs low (10-39)","field":"multi-field","wr":"reviewers' overall high/low agreement","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":true,"ms":"Reviewers gave the same high/low classification on 68% of the 150 Combined (3 areas) grant proposals overall.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"exact agreement on high (40-50) rated proposals","estd":"percent agreement","v":0.54,"n":"52","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"dichotomised: high (40-50) vs low (10-39)","field":"multi-field","wr":"reviewers agreeing on high-rated proposals","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Reviewers agreed on 54% of the 52 highly rated Combined (3 areas) grant proposals, lower than agreement on low-rated proposals.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"exact agreement on low (10-39) rated proposals","estd":"percent agreement","v":0.76,"n":"98","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"dichotomised: high (40-50) vs low (10-39)","field":"multi-field","wr":"reviewers agreeing on low-rated proposals","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Reviewers agreed on 76% of the 98 low-rated Combined (3 areas) grant proposals, higher than agreement on high-rated proposals.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"chance-corrected agreement R_i on dichotomised high/low grant ratings","estd":"ICC (single/unspec)","v":0.32,"n":"150","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"dichotomised: high (40-50) vs low (10-39)","field":"multi-field","wr":"reviewers' high/low ratings on grant proposals","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":true,"he":true,"ms":"Across three research areas, NSF and COSPUP reviewers' dichotomised high/low ratings on 150 grant proposals gave a chance-corrected agreement (R_i) of 0.32. This is the study's overall inter-reviewer agreement result.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"overall exact agreement on dichotomised high/low ratings","estd":"percent agreement","v":0.76,"n":"50","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"dichotomised: high (40-50) vs low (10-39)","field":"multi-field","wr":"reviewers' overall high/low agreement","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":true,"ms":"Reviewers gave the same high/low classification on 76% of the 50 economics grant proposals overall.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"exact agreement on high (40-50) rated proposals","estd":"percent agreement","v":0.6,"n":"15","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"dichotomised: high (40-50) vs low (10-39)","field":"multi-field","wr":"reviewers agreeing on high-rated proposals","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Reviewers agreed on 60% of the 15 highly rated economics grant proposals, lower than agreement on low-rated proposals.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"exact agreement on low (10-39) rated proposals","estd":"percent agreement","v":0.83,"n":"35","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"dichotomised: high (40-50) vs low (10-39)","field":"multi-field","wr":"reviewers agreeing on low-rated proposals","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Reviewers agreed on 83% of the 35 low-rated economics grant proposals, higher than agreement on high-rated proposals.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"chance-corrected agreement R_i on dichotomised high/low grant ratings","estd":"ICC (single/unspec)","v":0.44,"n":"50","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"dichotomised: high (40-50) vs low (10-39)","field":"multi-field","wr":"reviewers' high/low ratings on grant proposals","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":true,"ms":"NSF and COSPUP reviewers' dichotomised high/low ratings on 50 economics grant proposals gave a chance-corrected agreement (R_i) of 0.44.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"overall exact agreement on dichotomised high/low ratings","estd":"percent agreement","v":0.68,"n":"50","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"dichotomised: high (40-50) vs low (10-39)","field":"multi-field","wr":"reviewers' overall high/low agreement","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":true,"ms":"Reviewers gave the same high/low classification on 68% of the 50 solid state physics grant proposals overall.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"exact agreement on high (40-50) rated proposals","estd":"percent agreement","v":0.6,"n":"20","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"dichotomised: high (40-50) vs low (10-39)","field":"multi-field","wr":"reviewers agreeing on high-rated proposals","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Reviewers agreed on 60% of the 20 highly rated solid state physics grant proposals, lower than agreement on low-rated proposals.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"exact agreement on low (10-39) rated proposals","estd":"percent agreement","v":0.73,"n":"30","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"dichotomised: high (40-50) vs low (10-39)","field":"multi-field","wr":"reviewers agreeing on low-rated proposals","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Reviewers agreed on 73% of the 30 low-rated solid state physics grant proposals, higher than agreement on high-rated proposals.","vf":"unverified"},{"key":"HHVUQFXP","au":"Cicchetti, Domenic V.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"chance-corrected agreement R_i on dichotomised high/low grant ratings","estd":"ICC (single/unspec)","v":0.34,"n":"50","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"dichotomised: high (40-50) vs low (10-39)","field":"multi-field","wr":"reviewers' high/low ratings on grant proposals","conf":"med","self":false,"doi":"10.1017/S0140525X00065675","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":true,"ms":"NSF and COSPUP reviewers' dichotomised high/low ratings on 50 solid state physics grant proposals gave a chance-corrected agreement (R_i) of 0.34.","vf":"unverified"},{"key":"SBVJDN7K","au":"Clarke, Philip","y":2016,"cx":"Grant","ob":"fellowship","fam":"Gwet-AC","form":"Gwet's statistic (chance-adjusted agreement)","estd":"Gwet AC","v":0.75,"n":"60","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"panel-consensus","scale":"funded / not funded","field":"health and medical research","wr":"two panels' funding decisions, chance-adjusted","conf":"high","self":false,"doi":"10.1016/j.jclinepi.2015.04.010","ciLow":0.58,"ciHigh":0.92,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For the same 60 applications judged by two independent panels, Gwet's chance-adjusted agreement on the funding decision was 0.75, indicating substantial agreement after correcting for expected chance agreement.","vf":"unverified"},{"key":"SBVJDN7K","au":"Clarke, Philip","y":2016,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"overall percent agreement in funding, exact agreement (funded by both or not funded by both)","estd":"percent agreement","v":0.83,"n":"60","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"panel-consensus","scale":"funded / not funded","field":"health and medical research","wr":"two panels' fund/not-fund decisions on fellowship applications","conf":"high","self":false,"doi":"10.1016/j.jclinepi.2015.04.010","ciLow":0.73,"ciHigh":0.92,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":true,"he":false,"ms":"Two independent NHMRC panels, each of four reviewers, judged the same 60 duplicated Early Career Fellowship applications; they reached the same fund or not-fund outcome for 83% of applications, a raw agreement that still includes chance.","vf":"unverified"},{"key":"WU4XSEWV","au":"Cobo, E","y":2011,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient","estd":"ICC (single/unspec)","v":0.46,"n":"","k":"3","samp":"funded-only","blind":"unclear","agg":"single-rater","scale":"5 point Likert scale from 1 (low) to 5 (high)","field":"biomedical","wr":"three statisticians rating submitted manuscripts' quality","conf":"med","self":false,"doi":"10.1136/bmj.d6783","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":true,"he":true,"ms":"Three junior statisticians independently scored the reporting quality of manuscripts conditionally accepted by Medicina Clinica, using a 37 item instrument on a 1 to 5 Likert scale, before discussing consensus scores. The intraclass correlation coefficient of 0.46 reflects the reliability of their individual ratings; the paper does not state how many manuscripts or which rating occasion entered this coefficient.","vf":"unverified"},{"key":"N8M2AUCU","au":"Cohen, Ira Todd","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"kappa = N (PMS-EMS)/{N.PMS + (k-1) RMS+(N-1)(k-1) EMS}","estd":"kappa","v":0.21,"n":"43","k":"6","samp":"full-pool","blind":"double","agg":"unspecified","scale":"1-4: rejection, possible rejection, possible acceptance, acceptance","field":"anesthesiology","wr":"reviewers on conference abstract acceptance scores","conf":"high","self":false,"doi":"10.46374/volvii-issue2-cohen2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":false,"ms":"Six society reviewers scored 43 conference abstracts on a four-point acceptance scale; an ANOVA-based kappa of 0.21 indicates fair agreement between reviewers.","vf":"unverified"},{"key":"N8M2AUCU","au":"Cohen, Ira Todd","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"kappa = N (PMS-EMS)/{N.PMS + (k-1) RMS+(N-1)(k-1) EMS}","estd":"kappa","v":0.39,"n":"44","k":"5","samp":"full-pool","blind":"double","agg":"unspecified","scale":"1-4: rejection, possible rejection, possible acceptance, acceptance","field":"anesthesiology","wr":"reviewers on conference abstract acceptance scores","conf":"high","self":false,"doi":"10.46374/volvii-issue2-cohen2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Five society reviewers scored 44 conference abstracts on a four-point acceptance scale; an ANOVA-based kappa of 0.39 indicates fair agreement between reviewers.","vf":"unverified"},{"key":"JN8MZ48Y","au":"Cohen, Ira Todd","y":2006,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"Kappa via two-way ANOVA mean squares (PMS, RMS, EMS) per Fleiss (1989)","estd":"kappa","v":0.02,"n":"39","k":"","samp":"full-pool","blind":"double","agg":"single-rater","scale":"70, 73, 77, 80 (Nominal)","field":"biomedical (anesthesiology)","wr":"reviewers on conference abstract scores","conf":"med","self":false,"doi":"10.1213/01.ane.0000200314.73035.4d","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Nine reviewers scored 39 Society A anesthesiology abstracts on a four-category nominal acceptance scale; a kappa of 0.02 indicates essentially chance-level agreement.","vf":"unverified"},{"key":"JN8MZ48Y","au":"Cohen, Ira Todd","y":2006,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"Kappa via two-way ANOVA mean squares (PMS, RMS, EMS) per Fleiss (1989)","estd":"kappa","v":0.14,"n":"31","k":"","samp":"full-pool","blind":"double","agg":"single-rater","scale":"70, 73, 77, 80 (Nominal)","field":"biomedical (anesthesiology)","wr":"reviewers on conference abstract scores","conf":"med","self":false,"doi":"10.1213/01.ane.0000200314.73035.4d","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Five reviewers scored 31 Society B anesthesiology abstracts on a four-category nominal acceptance scale; a kappa of 0.14 indicates poor chance-corrected agreement.","vf":"unverified"},{"key":"JN8MZ48Y","au":"Cohen, Ira Todd","y":2006,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"Kappa via two-way ANOVA mean squares (PMS, RMS, EMS) per Fleiss (1989)","estd":"kappa","v":0.17,"n":"53","k":"","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 - 5 (Interval)","field":"biomedical (anesthesiology)","wr":"reviewers on conference abstract scores","conf":"med","self":false,"doi":"10.1213/01.ane.0000200314.73035.4d","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Nine reviewers scored 53 Society C anesthesiology abstracts on a 1 to 5 interval scale; a kappa of 0.17 indicates poor chance-corrected agreement.","vf":"unverified"},{"key":"JN8MZ48Y","au":"Cohen, Ira Todd","y":2006,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"Kappa via two-way ANOVA mean squares (PMS, RMS, EMS) per Fleiss (1989)","estd":"kappa","v":0.09,"n":"","k":"","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 - 4 (Nominal)","field":"biomedical (anesthesiology)","wr":"reviewers on conference abstract scores","conf":"med","self":false,"doi":"10.1213/01.ane.0000200314.73035.4d","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Society D divided its 146 abstracts into subgroups; for one subgroup reviewers scored on a 1 to 4 nominal scale with a kappa of 0.09. Subgroup size and rater count were not reported.","vf":"unverified"},{"key":"JN8MZ48Y","au":"Cohen, Ira Todd","y":2006,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"Kappa via two-way ANOVA mean squares (PMS, RMS, EMS) per Fleiss (1989)","estd":"kappa","v":0.24,"n":"","k":"","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 - 4 (Nominal)","field":"biomedical (anesthesiology)","wr":"reviewers on conference abstract scores","conf":"med","self":false,"doi":"10.1213/01.ane.0000200314.73035.4d","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Society D divided its 146 abstracts into subgroups; for a second subgroup reviewers scored on a 1 to 4 nominal scale with a kappa of 0.24. Subgroup size and rater count were not reported.","vf":"unverified"},{"key":"JN8MZ48Y","au":"Cohen, Ira Todd","y":2006,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"Kappa via two-way ANOVA mean squares (PMS, RMS, EMS) per Fleiss (1989)","estd":"kappa","v":0.27,"n":"","k":"","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 - 4 (Nominal)","field":"biomedical (anesthesiology)","wr":"reviewers on conference abstract scores","conf":"med","self":false,"doi":"10.1213/01.ane.0000200314.73035.4d","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Society D divided its 146 abstracts into subgroups; for a third subgroup reviewers scored on a 1 to 4 nominal scale with a kappa of 0.27. Subgroup size and rater count were not reported.","vf":"unverified"},{"key":"JN8MZ48Y","au":"Cohen, Ira Todd","y":2006,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"Kappa via two-way ANOVA mean squares (PMS, RMS, EMS) per Fleiss (1989)","estd":"kappa","v":0.21,"n":"","k":"","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 - 4 (Nominal)","field":"biomedical (anesthesiology)","wr":"reviewers on conference abstract scores","conf":"med","self":false,"doi":"10.1213/01.ane.0000200314.73035.4d","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Society E divided its 87 abstracts into subgroups; for one subgroup reviewers scored on a 1 to 4 nominal scale with a kappa of 0.21. Subgroup size and rater count were not reported.","vf":"unverified"},{"key":"JN8MZ48Y","au":"Cohen, Ira Todd","y":2006,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"Kappa via two-way ANOVA mean squares (PMS, RMS, EMS) per Fleiss (1989)","estd":"kappa","v":0.39,"n":"","k":"","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 - 4 (Nominal)","field":"biomedical (anesthesiology)","wr":"reviewers on conference abstract scores","conf":"med","self":false,"doi":"10.1213/01.ane.0000200314.73035.4d","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Society E divided its 87 abstracts into subgroups; for a second subgroup reviewers scored on a 1 to 4 nominal scale with a kappa of 0.39. Subgroup size and rater count were not reported.","vf":"unverified"},{"key":"JN8MZ48Y","au":"Cohen, Ira Todd","y":2006,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"Kappa via two-way ANOVA mean squares (PMS, RMS, EMS) per Fleiss (1989)","estd":"kappa","v":0.4,"n":"99","k":"","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 - 4 (Nominal)","field":"biomedical (anesthesiology)","wr":"reviewers on conference abstract scores","conf":"med","self":false,"doi":"10.1213/01.ane.0000200314.73035.4d","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":false,"ms":"Five reviewers scored 99 Society F anesthesiology abstracts on a 1 to 4 nominal scale; a kappa of 0.40 indicates fair agreement. No overall kappa was reported, so this largest single-society undivided sample carries the primary flag.","vf":"unverified"},{"key":"JN8MZ48Y","au":"Cohen, Ira Todd","y":2006,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"Kappa via two-way ANOVA mean squares (PMS, RMS, EMS) per Fleiss (1989)","estd":"kappa","v":0.2,"n":"","k":"","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 - 5 (Interval)","field":"biomedical (anesthesiology)","wr":"reviewers on conference abstract scores","conf":"med","self":false,"doi":"10.1213/01.ane.0000200314.73035.4d","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Society G divided its 166 abstracts into subgroups; for one subgroup reviewers scored on a 1 to 5 interval scale with a kappa of 0.20. Subgroup size and rater count were not reported.","vf":"unverified"},{"key":"JN8MZ48Y","au":"Cohen, Ira Todd","y":2006,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"Kappa via two-way ANOVA mean squares (PMS, RMS, EMS) per Fleiss (1989)","estd":"kappa","v":0.22,"n":"","k":"","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 - 5 (Interval)","field":"biomedical (anesthesiology)","wr":"reviewers on conference abstract scores","conf":"med","self":false,"doi":"10.1213/01.ane.0000200314.73035.4d","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Society G divided its 166 abstracts into subgroups; for a second subgroup reviewers scored on a 1 to 5 interval scale with a kappa of 0.22. Subgroup size and rater count were not reported.","vf":"unverified"},{"key":"JN8MZ48Y","au":"Cohen, Ira Todd","y":2006,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"Kappa via two-way ANOVA mean squares (PMS, RMS, EMS) per Fleiss (1989)","estd":"kappa","v":0.28,"n":"","k":"","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 - 5 (Interval)","field":"biomedical (anesthesiology)","wr":"reviewers on conference abstract scores","conf":"med","self":false,"doi":"10.1213/01.ane.0000200314.73035.4d","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Society G divided its 166 abstracts into subgroups; for a third subgroup reviewers scored on a 1 to 5 interval scale with a kappa of 0.28. Subgroup size and rater count were not reported.","vf":"unverified"},{"key":"JN8MZ48Y","au":"Cohen, Ira Todd","y":2006,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"Kappa via two-way ANOVA mean squares (PMS, RMS, EMS) per Fleiss (1989)","estd":"kappa","v":0.43,"n":"","k":"","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 - 5 (Interval)","field":"biomedical (anesthesiology)","wr":"reviewers on conference abstract scores","conf":"med","self":false,"doi":"10.1213/01.ane.0000200314.73035.4d","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Society G divided its 166 abstracts into subgroups; for a fourth subgroup reviewers scored on a 1 to 5 interval scale with a kappa of 0.43, moderate agreement. Subgroup size and rater count were not reported.","vf":"unverified"},{"key":"JN8MZ48Y","au":"Cohen, Ira Todd","y":2006,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"Kappa via two-way ANOVA mean squares (PMS, RMS, EMS) per Fleiss (1989)","estd":"kappa","v":0.52,"n":"42","k":"","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 - 10 (Interval)","field":"biomedical (anesthesiology)","wr":"reviewers on conference abstract scores","conf":"med","self":false,"doi":"10.1213/01.ane.0000200314.73035.4d","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Nine reviewers scored 42 Society H anesthesiology abstracts on a 1 to 10 interval scale; a kappa of 0.52 indicates moderate agreement, the highest of the eight societies.","vf":"unverified"},{"key":"II5SPEMP","au":"Cole, S","y":1981,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"correlation coefficient between mean NSF rating and mean COSPUP rating per proposal","estd":"correlation","v":0.595,"n":"50","k":"","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"excellent/very good/good/fair/poor, scored 50-10","field":"multi-field (chemical dynamics, economics, solid-state physics)","wr":"two reviewer panels on grant proposal scores","conf":"high","self":false,"doi":"10.1126/science.7302566","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For 50 chemical-dynamics NSF proposals, an independent set of COSPUP reviewers re-scored each proposal; the correlation of 0.60 between the two panels' mean ratings shows moderate agreement between independently selected reviewer groups.","vf":"unverified"},{"key":"II5SPEMP","au":"Cole, S","y":1981,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"correlation coefficient between mean NSF rating and mean COSPUP rating per proposal","estd":"correlation","v":0.659,"n":"50","k":"","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"excellent/very good/good/fair/poor, scored 50-10","field":"multi-field (chemical dynamics, economics, solid-state physics)","wr":"two reviewer panels on grant proposal scores","conf":"high","self":false,"doi":"10.1126/science.7302566","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For 50 economics NSF proposals, independent COSPUP reviewers re-scored each proposal; the correlation of 0.66 between the two panels' mean ratings shows moderate agreement between independently selected reviewer groups.","vf":"unverified"},{"key":"II5SPEMP","au":"Cole, S","y":1981,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"correlation coefficient between mean NSF rating and mean COSPUP rating per proposal","estd":"correlation","v":0.623,"n":"50","k":"","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"excellent/very good/good/fair/poor, scored 50-10","field":"multi-field (chemical dynamics, economics, solid-state physics)","wr":"two reviewer panels on grant proposal scores","conf":"high","self":false,"doi":"10.1126/science.7302566","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"For 50 solid-state-physics NSF proposals, independent COSPUP reviewers re-scored each proposal; the correlation of 0.62 between the two panels' mean ratings is the paper's worked-example agreement result, showing only moderate agreement between independently selected reviewer groups.","vf":"unverified"},{"key":"II5SPEMP","au":"Cole, S","y":1981,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"reversal rate: proportion of proposals shifting between the top-25 (fund) and bottom-25 (decline) halves under independent COSPUP re-review versus the actual NSF funding decision (this is a disagreement rate, higher = more disagreement)","estd":"other","v":0.3,"n":"50","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"fund / decline (top 25 vs bottom 25)","field":"multi-field (chemical dynamics, economics, solid-state physics)","wr":"fund/decline decisions on grant proposals","conf":"high","self":false,"doi":"10.1126/science.7302566","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"Comparing the COSPUP re-review ranking with the actual NSF funding decision on 50 chemical-dynamics proposals, 30 percent of proposals would have moved across the fund/decline line, a decision reversal (disagreement) rate.","vf":"unverified"},{"key":"II5SPEMP","au":"Cole, S","y":1981,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"reversal rate: proportion of proposals shifting between the top-25 and bottom-25 halves under COSPUP re-review versus the ranking implied by NSF reviewers' mean ratings (disagreement rate, higher = more disagreement)","estd":"other","v":0.3,"n":"50","k":"","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"top 25 / bottom 25 by mean rating","field":"multi-field (chemical dynamics, economics, solid-state physics)","wr":"fund/decline decisions on grant proposals","conf":"high","self":false,"doi":"10.1126/science.7302566","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"Comparing the COSPUP re-review ranking with the ranking implied by NSF reviewers' mean ratings on 50 chemical-dynamics proposals, 30 percent of proposals would have moved across the fund/decline line, a reversal (disagreement) rate.","vf":"unverified"},{"key":"II5SPEMP","au":"Cole, S","y":1981,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"reversal rate: proportion of proposals shifting between the top-25 (fund) and bottom-25 (decline) halves under independent COSPUP re-review versus the actual NSF funding decision (this is a disagreement rate, higher = more disagreement)","estd":"other","v":0.24,"n":"50","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"fund / decline (top 25 vs bottom 25)","field":"multi-field (chemical dynamics, economics, solid-state physics)","wr":"fund/decline decisions on grant proposals","conf":"high","self":false,"doi":"10.1126/science.7302566","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"Comparing the COSPUP re-review ranking with the actual NSF funding decision on 50 economics proposals, 24 percent of proposals would have moved across the fund/decline line, a decision reversal (disagreement) rate.","vf":"unverified"},{"key":"II5SPEMP","au":"Cole, S","y":1981,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"reversal rate: proportion of proposals shifting between the top-25 and bottom-25 halves under COSPUP re-review versus the ranking implied by NSF reviewers' mean ratings (disagreement rate, higher = more disagreement)","estd":"other","v":0.28,"n":"50","k":"","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"top 25 / bottom 25 by mean rating","field":"multi-field (chemical dynamics, economics, solid-state physics)","wr":"fund/decline decisions on grant proposals","conf":"high","self":false,"doi":"10.1126/science.7302566","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"Comparing the COSPUP re-review ranking with the ranking implied by NSF reviewers' mean ratings on 50 economics proposals, 28 percent of proposals would have moved across the fund/decline line, a reversal (disagreement) rate.","vf":"unverified"},{"key":"II5SPEMP","au":"Cole, S","y":1981,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"overall reversal rate, about 25 percent: proportion of fund/decline outcomes reversed under independent COSPUP re-review across the three programmes (disagreement rate, higher = more disagreement)","estd":"other","v":0.25,"n":"150","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"fund / decline","field":"multi-field (chemical dynamics, economics, solid-state physics)","wr":"fund/decline decisions on grant proposals","conf":"high","self":false,"doi":"10.1126/science.7302566","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":true,"he":false,"ms":"Across the 150 NSF proposals re-reviewed by independent COSPUP reviewers, about 25 percent of fund/decline outcomes would have been reversed; the paper's headline conclusion is that a grant's fate is roughly half determined by the luck of the reviewer draw.","vf":"unverified"},{"key":"II5SPEMP","au":"Cole, S","y":1981,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"reversal rate within COSPUP rank-order quintile: proportion shifting between top-25 and bottom-25 halves, mean-rating comparison (maximum stratum of 30 quintile cells)","estd":"other","v":0.6,"n":"","k":"","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"top 25 / bottom 25 by mean rating","field":"multi-field (chemical dynamics, economics, solid-state physics)","wr":"fund/decline decisions on grant proposals","conf":"high","self":false,"doi":"10.1126/science.7302566","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"restricted-other","pr":false,"he":false,"ms":"Within the middle quintile of the COSPUP rank order for chemical dynamics, 60 percent of proposals were classified on opposite sides of the fund/decline line by the NSF and COSPUP mean ratings; this is the maximum quintile reversal rate reported.","vf":"unverified"},{"key":"II5SPEMP","au":"Cole, S","y":1981,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"reversal rate within COSPUP rank-order quintile: proportion shifting between top-25 and bottom-25 halves, mean-rating comparison (minimum stratum of 30 quintile cells)","estd":"other","v":0,"n":"","k":"","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"top 25 / bottom 25 by mean rating","field":"multi-field (chemical dynamics, economics, solid-state physics)","wr":"fund/decline decisions on grant proposals","conf":"high","self":false,"doi":"10.1126/science.7302566","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"restricted-other","pr":false,"he":false,"ms":"Within the lowest quintile of the COSPUP rank order for economics, no proposals were classified on opposite sides of the fund/decline line by the NSF and COSPUP mean ratings; this is the minimum quintile reversal rate reported.","vf":"unverified"},{"key":"II5SPEMP","au":"Cole, S","y":1981,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"reversal rate: proportion of proposals shifting between the top-25 (fund) and bottom-25 (decline) halves under independent COSPUP re-review versus the actual NSF funding decision (this is a disagreement rate, higher = more disagreement)","estd":"other","v":0.25,"n":"50","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"fund / decline (top 25 vs bottom 25)","field":"multi-field (chemical dynamics, economics, solid-state physics)","wr":"fund/decline decisions on grant proposals","conf":"high","self":false,"doi":"10.1126/science.7302566","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"Comparing the COSPUP re-review ranking with the actual NSF funding decision on 50 solid-state-physics proposals, 25 percent of proposals would have moved across the fund/decline line, a decision reversal (disagreement) rate.","vf":"unverified"},{"key":"II5SPEMP","au":"Cole, S","y":1981,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"reversal rate: proportion of proposals shifting between the top-25 and bottom-25 halves under COSPUP re-review versus the ranking implied by NSF reviewers' mean ratings (disagreement rate, higher = more disagreement)","estd":"other","v":0.27,"n":"50","k":"","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"top 25 / bottom 25 by mean rating","field":"multi-field (chemical dynamics, economics, solid-state physics)","wr":"fund/decline decisions on grant proposals","conf":"high","self":false,"doi":"10.1126/science.7302566","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"Comparing the COSPUP re-review ranking with the ranking implied by NSF reviewers' mean ratings on 50 solid-state-physics proposals, 27 percent of proposals would have moved across the fund/decline line, a reversal (disagreement) rate.","vf":"unverified"},{"key":"II5SPEMP","au":"Cole, S","y":1981,"cx":"Grant","ob":"grant-proposal","fam":"G-theory","form":"percent of total rating variance due to within-proposal differences among reviewers (one-way ANOVA); complement of the intraclass reliability","estd":"G-theory","v":0.53,"n":"50","k":"","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"excellent/very good/good/fair/poor, scored 50-10","field":"multi-field (chemical dynamics, economics, solid-state physics)","wr":"external reviewers on grant proposal scores","conf":"high","self":false,"doi":"10.1126/science.7302566","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"A one-way analysis of variance on the COSPUP reviews of 50 chemical-dynamics proposals attributed 53 percent of the total rating variance to disagreement among reviewers of the same proposal.","vf":"unverified"},{"key":"II5SPEMP","au":"Cole, S","y":1981,"cx":"Grant","ob":"grant-proposal","fam":"G-theory","form":"percent of total rating variance due to within-proposal differences among reviewers (one-way ANOVA); complement of the intraclass reliability","estd":"G-theory","v":0.6,"n":"50","k":"","samp":"special","blind":"single","agg":"single-rater","scale":"excellent/very good/good/fair/poor, scored 50-10","field":"multi-field (chemical dynamics, economics, solid-state physics)","wr":"external reviewers on grant proposal scores","conf":"high","self":false,"doi":"10.1126/science.7302566","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"A one-way analysis of variance on the NSF reviews of 50 chemical-dynamics proposals attributed 60 percent of the total rating variance to disagreement among reviewers of the same proposal, meaning most of the variance was reviewer disagreement rather than differences between proposals.","vf":"unverified"},{"key":"II5SPEMP","au":"Cole, S","y":1981,"cx":"Grant","ob":"grant-proposal","fam":"G-theory","form":"percent of total rating variance due to within-proposal differences among reviewers (one-way ANOVA); complement of the intraclass reliability","estd":"G-theory","v":0.49,"n":"50","k":"","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"excellent/very good/good/fair/poor, scored 50-10","field":"multi-field (chemical dynamics, economics, solid-state physics)","wr":"external reviewers on grant proposal scores","conf":"high","self":false,"doi":"10.1126/science.7302566","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"A one-way analysis of variance on the COSPUP reviews of 50 economics proposals attributed 49 percent of the total rating variance to disagreement among reviewers of the same proposal.","vf":"unverified"},{"key":"II5SPEMP","au":"Cole, S","y":1981,"cx":"Grant","ob":"grant-proposal","fam":"G-theory","form":"percent of total rating variance due to within-proposal differences among reviewers (one-way ANOVA); complement of the intraclass reliability","estd":"G-theory","v":0.51,"n":"50","k":"","samp":"special","blind":"single","agg":"single-rater","scale":"excellent/very good/good/fair/poor, scored 50-10","field":"multi-field (chemical dynamics, economics, solid-state physics)","wr":"external reviewers on grant proposal scores","conf":"high","self":false,"doi":"10.1126/science.7302566","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"A one-way analysis of variance on the NSF reviews of 50 economics proposals attributed 51 percent of the total rating variance to disagreement among reviewers of the same proposal.","vf":"unverified"},{"key":"II5SPEMP","au":"Cole, S","y":1981,"cx":"Grant","ob":"grant-proposal","fam":"G-theory","form":"percent of total rating variance due to within-proposal differences among reviewers (one-way ANOVA), reported as a 35-63 percent range across ten programs","estd":"G-theory","v":0.63,"n":"","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"excellent/very good/good/fair/poor, scored 50-10","field":"multi-field (chemical dynamics, economics, solid-state physics)","wr":"external reviewers on grant proposal scores","conf":"high","self":false,"doi":"10.1126/science.7302566","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Replicating the one-way analysis of variance across the ten programmes studied in phase I, the share of total variance due to within-proposal reviewer disagreement ranged from 35 to 63 percent; this row records the maximum of that reported range.","vf":"unverified"},{"key":"II5SPEMP","au":"Cole, S","y":1981,"cx":"Grant","ob":"grant-proposal","fam":"G-theory","form":"percent of total rating variance due to within-proposal differences among reviewers (one-way ANOVA), reported as a 35-63 percent range across ten programs","estd":"G-theory","v":0.35,"n":"","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"excellent/very good/good/fair/poor, scored 50-10","field":"multi-field (chemical dynamics, economics, solid-state physics)","wr":"external reviewers on grant proposal scores","conf":"high","self":false,"doi":"10.1126/science.7302566","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Replicating the one-way analysis of variance across the ten programmes studied in phase I, the share of total variance due to within-proposal reviewer disagreement ranged from 35 to 63 percent; this row records the minimum of that reported range.","vf":"unverified"},{"key":"II5SPEMP","au":"Cole, S","y":1981,"cx":"Grant","ob":"grant-proposal","fam":"G-theory","form":"percent of total rating variance due to within-proposal differences among reviewers (one-way ANOVA); complement of the intraclass reliability","estd":"G-theory","v":0.47,"n":"50","k":"","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"excellent/very good/good/fair/poor, scored 50-10","field":"multi-field (chemical dynamics, economics, solid-state physics)","wr":"external reviewers on grant proposal scores","conf":"high","self":false,"doi":"10.1126/science.7302566","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"A one-way analysis of variance on the COSPUP reviews of 50 solid-state-physics proposals attributed 47 percent of the total rating variance to disagreement among reviewers of the same proposal.","vf":"unverified"},{"key":"II5SPEMP","au":"Cole, S","y":1981,"cx":"Grant","ob":"grant-proposal","fam":"G-theory","form":"percent of total rating variance due to within-proposal differences among reviewers (one-way ANOVA); complement of the intraclass reliability","estd":"G-theory","v":0.43,"n":"50","k":"","samp":"special","blind":"single","agg":"single-rater","scale":"excellent/very good/good/fair/poor, scored 50-10","field":"multi-field (chemical dynamics, economics, solid-state physics)","wr":"external reviewers on grant proposal scores","conf":"high","self":false,"doi":"10.1126/science.7302566","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"A one-way analysis of variance on the NSF reviews of 50 solid-state-physics proposals attributed 43 percent of the total rating variance to disagreement among reviewers of the same proposal.","vf":"unverified"},{"key":"G6BW4RFI","au":"Conix, Stijn","y":2025,"cx":"General","ob":"other","fam":"other","form":"Brennan-Prediger coefficient, binary ratings","estd":"other","v":0.18,"n":"","k":"","samp":"funded-only","blind":"unclear","agg":"single-rater","scale":"relevant / not relevant (fund vs. do not fund)","field":"humanities","wr":"early-career researchers rating published abstracts' societal relevance","conf":"med","self":false,"doi":"10.1162/qss.a.19","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"restricted-top","pr":true,"he":false,"ms":"In Group 1, early-career humanities researchers independently judged whether each published abstract described societally relevant research; a Brennan-Prediger coefficient of 0.18 indicates low chance-corrected agreement on the binary relevance decision.","vf":"unverified"},{"key":"G6BW4RFI","au":"Conix, Stijn","y":2025,"cx":"General","ob":"other","fam":"other","form":"Brennan-Prediger coefficient, binary ratings","estd":"other","v":0.22,"n":"","k":"","samp":"funded-only","blind":"unclear","agg":"single-rater","scale":"relevant / not relevant (fund vs. do not fund)","field":"humanities","wr":"early-career researchers rating published abstracts' societal relevance","conf":"med","self":false,"doi":"10.1162/qss.a.19","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"In Group 2, early-career humanities researchers independently judged whether each published abstract described societally relevant research; a Brennan-Prediger coefficient of 0.22 indicates low chance-corrected agreement on the binary relevance decision.","vf":"unverified"},{"key":"G6BW4RFI","au":"Conix, Stijn","y":2025,"cx":"General","ob":"other","fam":"other","form":"Brennan-Prediger coefficient, binary ratings","estd":"other","v":0.17,"n":"","k":"","samp":"funded-only","blind":"unclear","agg":"single-rater","scale":"relevant / not relevant (fund vs. do not fund)","field":"humanities","wr":"early-career researchers (incl. non-humanities) rating relevance","conf":"med","self":false,"doi":"10.1162/qss.a.19","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"In Group 3, which added non-humanities raters, researchers independently judged whether each published abstract described societally relevant research; a Brennan-Prediger coefficient of 0.17 indicates low chance-corrected agreement on the binary relevance decision.","vf":"unverified"},{"key":"G6BW4RFI","au":"Conix, Stijn","y":2025,"cx":"General","ob":"other","fam":"percent-agreement","form":"percent full consensus (all raters in exact agreement) on binary score","estd":"percent agreement","v":0.0893,"n":"","k":"","samp":"funded-only","blind":"unclear","agg":"unspecified","scale":"relevant / not relevant (fund vs. do not fund)","field":"humanities","wr":"early-career researchers rating published abstracts' relevance","conf":"med","self":false,"doi":"10.1162/qss.a.19","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"Across all abstracts in the three groups, raters reached full consensus on the binary relevance decision for only 8.93 percent, an exact-agreement measure of how rarely every rater agreed.","vf":"unverified"},{"key":"G6BW4RFI","au":"Conix, Stijn","y":2025,"cx":"General","ob":"other","fam":"percent-agreement","form":"percent of cases with identical rankings of a five-abstract set (exact agreement)","estd":"percent agreement","v":0.2296,"n":"","k":"","samp":"funded-only","blind":"unclear","agg":"unspecified","scale":"rank 1-5 within a set of five abstracts","field":"humanities","wr":"early-career researchers ranking abstracts by relevance","conf":"med","self":false,"doi":"10.1162/qss.a.19","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"Raters assigned exactly identical rankings to the same set of five published abstracts in 22.96 percent of cases, an exact-agreement measure of how often their full rank orders matched.","vf":"unverified"},{"key":"G6BW4RFI","au":"Conix, Stijn","y":2025,"cx":"General","ob":"other","fam":"Kendall-W","form":"Kendall's W for rank scores, averaged over sets within group","estd":"Kendall W","v":0.38,"n":"","k":"","samp":"funded-only","blind":"unclear","agg":"unspecified","scale":"rank 1-5 within a set of five abstracts","field":"humanities","wr":"early-career researchers ranking abstracts by relevance","conf":"med","self":false,"doi":"10.1162/qss.a.19","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"In Group 1, raters independently ranked each set of five published abstracts by societal relevance; a Kendall's W of 0.38, averaged over sets, indicates low to moderate concordance among raters' rankings.","vf":"unverified"},{"key":"G6BW4RFI","au":"Conix, Stijn","y":2025,"cx":"General","ob":"other","fam":"Kendall-W","form":"Kendall's W for rank scores, averaged over sets within group","estd":"Kendall W","v":0.35,"n":"","k":"","samp":"funded-only","blind":"unclear","agg":"unspecified","scale":"rank 1-5 within a set of five abstracts","field":"humanities","wr":"early-career researchers ranking abstracts by relevance","conf":"med","self":false,"doi":"10.1162/qss.a.19","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"In Group 2, raters independently ranked each set of five published abstracts by societal relevance; a Kendall's W of 0.35, averaged over sets, indicates low to moderate concordance among raters' rankings.","vf":"unverified"},{"key":"G6BW4RFI","au":"Conix, Stijn","y":2025,"cx":"General","ob":"other","fam":"Kendall-W","form":"Kendall's W for rank scores, averaged over sets within group","estd":"Kendall W","v":0.4,"n":"","k":"","samp":"funded-only","blind":"unclear","agg":"unspecified","scale":"rank 1-5 within a set of five abstracts","field":"humanities","wr":"early-career researchers ranking abstracts by relevance","conf":"med","self":false,"doi":"10.1162/qss.a.19","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"In Group 3, raters independently ranked each set of five published abstracts by societal relevance; a Kendall's W of 0.40, averaged over sets, indicates low to moderate concordance among raters' rankings.","vf":"unverified"},{"key":"G6BW4RFI","au":"Conix, Stijn","y":2025,"cx":"General","ob":"other","fam":"other","form":"percent of abstracts on which more than one rater disagreed","estd":"other","v":0.7561,"n":"","k":"","samp":"funded-only","blind":"unclear","agg":"unspecified","scale":"relevant / not relevant (fund vs. do not fund)","field":"humanities","wr":"early-career researchers rating published abstracts' relevance","conf":"med","self":false,"doi":"10.1162/qss.a.19","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"Across all abstracts in the three groups, more than one rater dissented from the others' binary relevance decision for 75.61 percent of abstracts, a disagreement-prevalence measure showing how common substantial dissent was.","vf":"unverified"},{"key":"RJLFSWWX","au":"Crable, Erika L","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha","estd":"Krippendorff","v":0.84,"n":"30","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-3, 3 = all criteria fully met","field":"biomedical implementation and improvement science","wr":"two coders scoring proposals on the care/quality gap item","conf":"high","self":false,"doi":"10.1186/s13012-018-0770-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the care or quality gap item, two coders independently scored 30 grant proposals; a Krippendorff's alpha of 0.84 indicates strong agreement on this item.","vf":"unverified"},{"key":"RJLFSWWX","au":"Crable, Erika L","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha","estd":"Krippendorff","v":0.99,"n":"30","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-3, 3 = all criteria fully met","field":"biomedical implementation and improvement science","wr":"two coders scoring proposals on the conceptual model item","conf":"high","self":false,"doi":"10.1186/s13012-018-0770-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the conceptual model and theoretical justification item, two coders independently scored 30 grant proposals; a Krippendorff's alpha of 0.99 indicates near-perfect agreement on this item.","vf":"unverified"},{"key":"RJLFSWWX","au":"Crable, Erika L","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha","estd":"Krippendorff","v":0.77,"n":"30","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-3, 3 = all criteria fully met","field":"biomedical implementation and improvement science","wr":"two coders scoring proposals on the evidence-based treatment item","conf":"high","self":false,"doi":"10.1186/s13012-018-0770-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the evidence-based treatment item, two coders independently scored 30 grant proposals; a Krippendorff's alpha of 0.77 indicates acceptable agreement on this item.","vf":"unverified"},{"key":"RJLFSWWX","au":"Crable, Erika L","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha","estd":"Krippendorff","v":0.84,"n":"30","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-3, 3 = all criteria fully met","field":"biomedical implementation and improvement science","wr":"two coders scoring proposals on the feasibility item","conf":"high","self":false,"doi":"10.1186/s13012-018-0770-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the feasibility of research design item, two coders independently scored 30 grant proposals; a Krippendorff's alpha of 0.84 indicates strong agreement on this item.","vf":"unverified"},{"key":"RJLFSWWX","au":"Crable, Erika L","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha","estd":"Krippendorff","v":0.84,"n":"30","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-3, 3 = all criteria fully met","field":"biomedical implementation and improvement science","wr":"two coders scoring proposals on the implementation strategy item","conf":"high","self":false,"doi":"10.1186/s13012-018-0770-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the implementation strategy/process item, two coders independently scored 30 grant proposals; a Krippendorff's alpha of 0.84 indicates strong agreement on this item.","vf":"unverified"},{"key":"RJLFSWWX","au":"Crable, Erika L","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha","estd":"Krippendorff","v":0.78,"n":"30","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-3, 3 = all criteria fully met","field":"biomedical implementation and improvement science","wr":"two coders scoring proposals on the measurement and analysis item","conf":"high","self":false,"doi":"10.1186/s13012-018-0770-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the measurement and analysis item, two coders independently scored 30 grant proposals; a Krippendorff's alpha of 0.78 indicates acceptable agreement on this item.","vf":"unverified"},{"key":"RJLFSWWX","au":"Crable, Erika L","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha","estd":"Krippendorff","v":0.77,"n":"30","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-3, 3 = all criteria fully met","field":"biomedical implementation and improvement science","wr":"two coders scoring proposals on the policy/funding environment item","conf":"high","self":false,"doi":"10.1186/s13012-018-0770-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the policy/funding environment item, two coders independently scored 30 grant proposals; a Krippendorff's alpha of 0.77 indicates acceptable agreement on this item.","vf":"unverified"},{"key":"RJLFSWWX","au":"Crable, Erika L","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha","estd":"Krippendorff","v":0.96,"n":"30","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-3, 3 = all criteria fully met","field":"biomedical implementation and improvement science","wr":"two coders scoring proposals on the setting readiness item","conf":"high","self":false,"doi":"10.1186/s13012-018-0770-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the setting's readiness item, two coders independently scored 30 grant proposals; a Krippendorff's alpha of 0.96 indicates excellent agreement on this item.","vf":"unverified"},{"key":"RJLFSWWX","au":"Crable, Erika L","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha","estd":"Krippendorff","v":0.88,"n":"30","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-3, 3 = all criteria fully met","field":"biomedical implementation and improvement science","wr":"two coders scoring proposals on the stakeholder priorities item","conf":"high","self":false,"doi":"10.1186/s13012-018-0770-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the stakeholder priorities and engagement item, two coders independently scored 30 grant proposals; a Krippendorff's alpha of 0.88 indicates strong agreement on this item.","vf":"unverified"},{"key":"RJLFSWWX","au":"Crable, Erika L","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha","estd":"Krippendorff","v":0.96,"n":"30","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-3, 3 = all criteria fully met","field":"biomedical implementation and improvement science","wr":"two coders scoring proposals on the team experience item","conf":"high","self":false,"doi":"10.1186/s13012-018-0770-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the team experience item, two coders independently scored 30 grant proposals; a Krippendorff's alpha of 0.96 indicates excellent agreement on this item.","vf":"unverified"},{"key":"RJLFSWWX","au":"Crable, Erika L","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha","estd":"Krippendorff","v":0.88,"n":"30","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-3, 3 = all criteria fully met","field":"biomedical implementation and improvement science","wr":"two coders applying INSPECT to grant proposals","conf":"high","self":false,"doi":"10.1186/s13012-018-0770-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Two research team members independently scored 30 implementation-science pilot grant proposals with the 10-item INSPECT criteria; a Krippendorff's alpha of 0.88 indicates excellent overall agreement between the two coders.","vf":"unverified"},{"key":"QWNRA5FG","au":"Crowe, Michael","y":2011,"cx":"General","ob":"other","fam":"G-theory","form":"absolute G coefficient","estd":"G-theory","v":0.6,"n":"5","k":"1","samp":"special","blind":"unclear","agg":"single-rater","scale":"categories scored 0-5; total as percent (0-100)","field":"health research","wr":"projected reliability of CCAT scores","conf":"high","self":false,"doi":"10.1111/j.1744-1609.2011.00237.x","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"The D-study, fed by the trial's own ratings, projected an absolute G coefficient of 0.60 for one CCAT rater scoring the five papers.","vf":"unverified"},{"key":"QWNRA5FG","au":"Crowe, Michael","y":2011,"cx":"General","ob":"other","fam":"G-theory","form":"absolute G coefficient","estd":"G-theory","v":0.94,"n":"5","k":"10","samp":"special","blind":"unclear","agg":"average-of-k","scale":"categories scored 0-5; total as percent (0-100)","field":"health research","wr":"projected reliability of CCAT scores","conf":"high","self":false,"doi":"10.1111/j.1744-1609.2011.00237.x","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"The empirical D-study projected an absolute G coefficient of 0.94 for ten CCAT raters scoring five papers.","vf":"unverified"},{"key":"QWNRA5FG","au":"Crowe, Michael","y":2011,"cx":"General","ob":"other","fam":"G-theory","form":"absolute G coefficient","estd":"G-theory","v":0.88,"n":"5","k":"5","samp":"special","blind":"unclear","agg":"average-of-k","scale":"categories scored 0-5; total as percent (0-100)","field":"health research","wr":"reliability of CCAT scores","conf":"high","self":false,"doi":"10.1111/j.1744-1609.2011.00237.x","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the observed five-rater CCAT design, the absolute G coefficient was 0.88 across five papers, matching the absolute-agreement ICC. Total ratings derived from stated complete crossing.","vf":"unverified"},{"key":"QWNRA5FG","au":"Crowe, Michael","y":2011,"cx":"General","ob":"other","fam":"G-theory","form":"relative G coefficient","estd":"G-theory","v":0.62,"n":"5","k":"1","samp":"special","blind":"unclear","agg":"single-rater","scale":"categories scored 0-5; total as percent (0-100)","field":"health research","wr":"projected reliability of CCAT scores","conf":"high","self":false,"doi":"10.1111/j.1744-1609.2011.00237.x","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"The D-study, fed by the trial's own ratings, projected a relative G coefficient of 0.62 for one CCAT rater scoring the five papers.","vf":"unverified"},{"key":"QWNRA5FG","au":"Crowe, Michael","y":2011,"cx":"General","ob":"other","fam":"G-theory","form":"relative G coefficient","estd":"G-theory","v":0.94,"n":"5","k":"10","samp":"special","blind":"unclear","agg":"average-of-k","scale":"categories scored 0-5; total as percent (0-100)","field":"health research","wr":"projected reliability of CCAT scores","conf":"high","self":false,"doi":"10.1111/j.1744-1609.2011.00237.x","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"The empirical D-study projected a relative G coefficient of 0.94 for ten CCAT raters scoring five papers.","vf":"unverified"},{"key":"QWNRA5FG","au":"Crowe, Michael","y":2011,"cx":"General","ob":"other","fam":"G-theory","form":"relative G coefficient","estd":"G-theory","v":0.89,"n":"5","k":"5","samp":"special","blind":"unclear","agg":"average-of-k","scale":"categories scored 0-5; total as percent (0-100)","field":"health research","wr":"reliability of CCAT scores","conf":"high","self":false,"doi":"10.1111/j.1744-1609.2011.00237.x","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the observed five-rater CCAT design, the relative G coefficient was 0.89 across five papers, matching the consistency ICC. Total ratings derived from stated complete crossing.","vf":"unverified"},{"key":"QWNRA5FG","au":"Crowe, Michael","y":2011,"cx":"General","ob":"other","fam":"ICC","form":"ICC for multiple raters, absolute agreement","estd":"ICC (average)","v":0.88,"n":"5","k":"5","samp":"special","blind":"unclear","agg":"average-of-k","scale":"categories scored 0-5; total as percent (0-100)","field":"health research","wr":"5 raters scoring 5 published research papers","conf":"high","self":false,"doi":"10.1111/j.1744-1609.2011.00237.x","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":true,"he":false,"ms":"Five intervention-group raters using the Crowe Critical Appraisal Tool independently scored five published papers. The absolute-agreement ICC of 0.88 was the study's featured reliability result. Total ratings derived from stated complete crossing.","vf":"unverified"},{"key":"QWNRA5FG","au":"Crowe, Michael","y":2011,"cx":"General","ob":"other","fam":"ICC","form":"ICC for multiple raters, consistency","estd":"ICC (average)","v":0.89,"n":"5","k":"5","samp":"special","blind":"unclear","agg":"average-of-k","scale":"categories scored 0-5; total as percent (0-100)","field":"health research","wr":"5 raters scoring 5 published research papers","conf":"high","self":false,"doi":"10.1111/j.1744-1609.2011.00237.x","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Five intervention-group raters using the Crowe Critical Appraisal Tool independently scored five published papers. The consistency ICC of 0.89 estimates agreement in rank ordering across the five raters. Total ratings derived from stated complete crossing.","vf":"unverified"},{"key":"QWNRA5FG","au":"Crowe, Michael","y":2011,"cx":"General","ob":"other","fam":"G-theory","form":"absolute G coefficient","estd":"G-theory","v":0.38,"n":"5","k":"1","samp":"special","blind":"unclear","agg":"single-rater","scale":"0 (lowest) to 10 (highest), converted to percentage","field":"health research","wr":"projected reliability of informal appraisal scores","conf":"high","self":false,"doi":"10.1111/j.1744-1609.2011.00237.x","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"The D-study, fed by the trial's own ratings, projected an absolute G coefficient of 0.38 for one informal-appraisal rater scoring the five papers.","vf":"unverified"},{"key":"QWNRA5FG","au":"Crowe, Michael","y":2011,"cx":"General","ob":"other","fam":"G-theory","form":"absolute G coefficient","estd":"G-theory","v":0.86,"n":"5","k":"10","samp":"special","blind":"unclear","agg":"average-of-k","scale":"0 (lowest) to 10 (highest), converted to percentage","field":"health research","wr":"projected reliability of informal appraisal scores","conf":"high","self":false,"doi":"10.1111/j.1744-1609.2011.00237.x","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"The empirical D-study projected an absolute G coefficient of 0.86 for ten informal-appraisal raters scoring five papers.","vf":"unverified"},{"key":"QWNRA5FG","au":"Crowe, Michael","y":2011,"cx":"General","ob":"other","fam":"G-theory","form":"absolute G coefficient","estd":"G-theory","v":0.76,"n":"5","k":"5","samp":"special","blind":"unclear","agg":"average-of-k","scale":"0 (lowest) to 10 (highest), converted to percentage","field":"health research","wr":"reliability of informal appraisal scores","conf":"high","self":false,"doi":"10.1111/j.1744-1609.2011.00237.x","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the observed five-rater informal-appraisal design, the absolute G coefficient was 0.76 across five papers, matching the absolute-agreement ICC. Total ratings derived from stated complete crossing.","vf":"unverified"},{"key":"QWNRA5FG","au":"Crowe, Michael","y":2011,"cx":"General","ob":"other","fam":"G-theory","form":"relative G coefficient","estd":"G-theory","v":0.51,"n":"5","k":"1","samp":"special","blind":"unclear","agg":"single-rater","scale":"0 (lowest) to 10 (highest), converted to percentage","field":"health research","wr":"projected reliability of informal appraisal scores","conf":"high","self":false,"doi":"10.1111/j.1744-1609.2011.00237.x","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"The D-study, fed by the trial's own ratings, projected a relative G coefficient of 0.51 for one informal-appraisal rater scoring the five papers.","vf":"unverified"},{"key":"QWNRA5FG","au":"Crowe, Michael","y":2011,"cx":"General","ob":"other","fam":"G-theory","form":"relative G coefficient","estd":"G-theory","v":0.91,"n":"5","k":"10","samp":"special","blind":"unclear","agg":"average-of-k","scale":"0 (lowest) to 10 (highest), converted to percentage","field":"health research","wr":"projected reliability of informal appraisal scores","conf":"high","self":false,"doi":"10.1111/j.1744-1609.2011.00237.x","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"The empirical D-study projected a relative G coefficient of 0.91 for ten informal-appraisal raters scoring five papers.","vf":"unverified"},{"key":"QWNRA5FG","au":"Crowe, Michael","y":2011,"cx":"General","ob":"other","fam":"G-theory","form":"relative G coefficient","estd":"G-theory","v":0.84,"n":"5","k":"5","samp":"special","blind":"unclear","agg":"average-of-k","scale":"0 (lowest) to 10 (highest), converted to percentage","field":"health research","wr":"reliability of informal appraisal scores","conf":"high","self":false,"doi":"10.1111/j.1744-1609.2011.00237.x","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the observed five-rater informal-appraisal design, the relative G coefficient was 0.84 across five papers, matching the consistency ICC. Total ratings derived from stated complete crossing.","vf":"unverified"},{"key":"QWNRA5FG","au":"Crowe, Michael","y":2011,"cx":"General","ob":"other","fam":"ICC","form":"ICC for multiple raters, absolute agreement","estd":"ICC (average)","v":0.76,"n":"5","k":"5","samp":"special","blind":"unclear","agg":"average-of-k","scale":"0 (lowest) to 10 (highest), converted to percentage","field":"health research","wr":"5 raters scoring 5 published research papers","conf":"high","self":false,"doi":"10.1111/j.1744-1609.2011.00237.x","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Five control-group raters using informal appraisal independently scored five published health research papers. The absolute-agreement ICC of 0.76 estimates agreement across the five raters. Total ratings derived from stated complete crossing.","vf":"unverified"},{"key":"QWNRA5FG","au":"Crowe, Michael","y":2011,"cx":"General","ob":"other","fam":"ICC","form":"ICC for multiple raters, consistency","estd":"ICC (average)","v":0.84,"n":"5","k":"5","samp":"special","blind":"unclear","agg":"average-of-k","scale":"0 (lowest) to 10 (highest), converted to percentage","field":"health research","wr":"5 raters scoring 5 published research papers","conf":"high","self":false,"doi":"10.1111/j.1744-1609.2011.00237.x","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Five control-group raters using informal appraisal independently scored five published health research papers. The consistency ICC of 0.84 estimates agreement in rank ordering across the five raters. Total ratings derived from stated complete crossing.","vf":"unverified"},{"key":"YWD3GJ8R","au":"Daniel, Hans‐Dieter","y":1993,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"category-specific agreement coefficient","estd":"other","v":0.28,"n":"392","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"Yes, without alterations / minor alterations / major alterations / No","field":"chemistry","wr":"referees' use of the \"No\" (reject) recommendation","conf":"med","self":false,"doi":"10.1002/anie.199302341","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the single response category \"No\", meaning reject, the category-specific agreement coefficient between the two referees was 0.28. This was the highest of the four category-specific values, so referees agreed most readily about rejection.","vf":"unverified"},{"key":"YWD3GJ8R","au":"Daniel, Hans‐Dieter","y":1993,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"category-specific agreement coefficient","estd":"other","v":0.09,"n":"392","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"Yes, without alterations / minor alterations / major alterations / No","field":"chemistry","wr":"referees' use of the \"Yes, after major alterations\" recommendation","conf":"med","self":false,"doi":"10.1002/anie.199302341","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the response category \"Yes, but only after major alterations\", the category-specific agreement coefficient between the two referees was only 0.09, again showing very weak convergence.","vf":"unverified"},{"key":"YWD3GJ8R","au":"Daniel, Hans‐Dieter","y":1993,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"category-specific agreement coefficient","estd":"other","v":0.1,"n":"392","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"Yes, without alterations / minor alterations / major alterations / No","field":"chemistry","wr":"referees' use of the \"Yes, after minor alterations\" recommendation","conf":"med","self":false,"doi":"10.1002/anie.199302341","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the response category \"Yes, after minor alterations\", the category-specific agreement coefficient between the two referees was only 0.10, indicating very weak convergence on this middle recommendation.","vf":"unverified"},{"key":"YWD3GJ8R","au":"Daniel, Hans‐Dieter","y":1993,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"category-specific agreement coefficient","estd":"other","v":0.07,"n":"392","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"Yes, without alterations / minor alterations / major alterations / No","field":"chemistry","wr":"referees' use of the \"Yes, without alterations\" recommendation","conf":"med","self":false,"doi":"10.1002/anie.199302341","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the response category \"Yes, without alterations\", the category-specific agreement coefficient between the two referees was only 0.07, so referees almost never converged on unconditional acceptance of the same communication.","vf":"unverified"},{"key":"YWD3GJ8R","au":"Daniel, Hans‐Dieter","y":1993,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient","estd":"ICC (single/unspec)","v":0.25,"n":"392","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"Yes, without alterations / minor alterations / major alterations / No","field":"chemistry","wr":"referees' acceptance recommendations on submitted communications","conf":"med","self":false,"doi":"10.1002/anie.199302341","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Treating the four-category recommendation as a numeric score, the intraclass correlation between the two referees' judgments of the same communications was 0.25. No ICC model, unit or agreement definition is stated.","vf":"unverified"},{"key":"YWD3GJ8R","au":"Daniel, Hans‐Dieter","y":1993,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"kappa coefficient","estd":"kappa","v":0.14,"n":"392","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"Yes, without alterations / minor alterations / major alterations / No","field":"chemistry","wr":"referees' acceptance recommendations on submitted communications","conf":"med","self":false,"doi":"10.1002/anie.199302341","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Two independent external referees each rated 392 communications submitted to Angewandte Chemie in 1984 on a four-category acceptance recommendation. A kappa of 0.14 indicates very low chance-corrected agreement between the paired referees.","vf":"unverified"},{"key":"YWD3GJ8R","au":"Daniel, Hans‐Dieter","y":1993,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"weighted kappa coefficient","estd":"weighted kappa","v":0.2,"n":"392","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"Yes, without alterations / minor alterations / major alterations / No","field":"chemistry","wr":"referees' acceptance recommendations on submitted communications","conf":"med","self":false,"doi":"10.1002/anie.199302341","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The same paired referee recommendations on the four ordered acceptance categories gave a weighted kappa of 0.20. The paper does not state whether linear or quadratic weights were used.","vf":"unverified"},{"key":"YWD3GJ8R","au":"Daniel, Hans‐Dieter","y":1993,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"kappa with scores computed as agreement if within one point","estd":"kappa","v":0.67,"n":"392","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"Yes, without alterations / minor alterations / major alterations / No","field":"chemistry","wr":"referees' acceptance recommendations on submitted communications","conf":"med","self":false,"doi":"10.1002/anie.199302341","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Following Crandall, recommendations differing by only a single category were counted as full agreement. The resulting kappa for the same paired referee recommendations was 0.67, which the author reads as substantial reviewer agreement.","vf":"unverified"},{"key":"FSW6MJXT","au":"Das, N K","y":1985,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"reliability coefficient r(x) = 1 - (S2w/S2b), Winer 1962 (within vs total rating variance)","estd":"ICC (single/unspec)","v":0.97,"n":"","k":"13","samp":"full-pool","blind":"single","agg":"unspecified","scale":"priority rating 1.0 best to 5.0 worst","field":"biomedical (allergy and immunology)","wr":"committee reviewers on P01/P50 grant priority scores","conf":"med","self":false,"doi":"10.1007/BF00929456","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"NIH AITRC committee members scored an RFA group of P01/P50 grant applications on priority; an intraclass reliability coefficient of 0.97 with 13 reviewers indicates most score variance reflected real between-application quality differences rather than reviewer error.","vf":"unverified"},{"key":"FSW6MJXT","au":"Das, N K","y":1985,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"reliability coefficient r(x) = 1 - (S2w/S2b), Winer 1962 (within vs total rating variance)","estd":"ICC (single/unspec)","v":0.98,"n":"","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"priority rating 1.0 best to 5.0 worst","field":"biomedical (allergy and immunology)","wr":"committee reviewers on P01/P50 grant priority scores","conf":"med","self":false,"doi":"10.1007/BF00929456","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"NIH AITRC committee members scored an RFA group of P01/P50 grant applications on priority; an intraclass reliability coefficient of 0.98 indicates most score variance reflected real between-application quality differences (reviewer count OCR-garbled).","vf":"unverified"},{"key":"FSW6MJXT","au":"Das, N K","y":1985,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"reliability coefficient r(x) = 1 - (S2w/S2b), Winer 1962 (within vs total rating variance)","estd":"ICC (single/unspec)","v":0.98,"n":"","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"priority rating 1.0 best to 5.0 worst","field":"biomedical (allergy and immunology)","wr":"committee reviewers on P01/P50 grant priority scores","conf":"med","self":false,"doi":"10.1007/BF00929456","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"NIH AITRC committee members scored an RFA group of P01/P50 grant applications on priority; an intraclass reliability coefficient of 0.98 indicates most score variance reflected real between-application quality differences (reviewer count OCR-garbled).","vf":"unverified"},{"key":"FSW6MJXT","au":"Das, N K","y":1985,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"reliability coefficient r(x) = 1 - (S2w/S2b), Winer 1962 (within vs total rating variance)","estd":"ICC (single/unspec)","v":0.99,"n":"","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"priority rating 1.0 best to 5.0 worst","field":"biomedical (allergy and immunology)","wr":"committee reviewers on P01/P50 grant priority scores","conf":"med","self":false,"doi":"10.1007/BF00929456","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"NIH AITRC committee members scored an RFA group of P01/P50 grant applications on priority; an intraclass reliability coefficient of 0.99 indicates almost all score variance reflected real between-application quality differences (reviewer count OCR-garbled).","vf":"unverified"},{"key":"FSW6MJXT","au":"Das, N K","y":1985,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"mean of RFA-group reliability coefficients r = 1 - (S2w/S2b), Winer 1962","estd":"ICC (single/unspec)","v":0.98,"n":"18","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"priority rating 1.0 best to 5.0 worst","field":"biomedical (allergy and immunology)","wr":"committee reviewers on P01/P50 grant priority scores","conf":"med","self":false,"doi":"10.1007/BF00929456","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Across the FY1980-1981 RFA groups (18 P01/P50 applications), the mean intraclass reliability coefficient of committee priority scores was 0.98, indicating high consensus on relative application quality.","vf":"unverified"},{"key":"FSW6MJXT","au":"Das, N K","y":1985,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"reliability coefficient r(x) = 1 - (S2w/S2b), Winer 1962 (within vs total rating variance)","estd":"ICC (single/unspec)","v":0.98,"n":"","k":"8","samp":"full-pool","blind":"single","agg":"unspecified","scale":"priority rating 1.0 best to 5.0 worst","field":"biomedical (allergy and immunology)","wr":"committee reviewers on P01/P50 grant priority scores","conf":"med","self":false,"doi":"10.1007/BF00929456","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"NIH AITRC committee members scored an RFA group of P01/P50 grant applications on priority; an intraclass reliability coefficient of 0.98 with 8 reviewers indicates most score variance reflected real between-application quality differences.","vf":"unverified"},{"key":"FSW6MJXT","au":"Das, N K","y":1985,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"reliability coefficient r(x) = 1 - (S2w/S2b), Winer 1962 (within vs total rating variance)","estd":"ICC (single/unspec)","v":0.97,"n":"","k":"13","samp":"full-pool","blind":"single","agg":"unspecified","scale":"priority rating 1.0 best to 5.0 worst","field":"biomedical (allergy and immunology)","wr":"committee reviewers on P01/P50 grant priority scores","conf":"med","self":false,"doi":"10.1007/BF00929456","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"NIH AITRC committee members scored an RFA group of P01/P50 grant applications on priority; an intraclass reliability coefficient of 0.97 with 13 reviewers indicates most score variance reflected real between-application quality differences.","vf":"unverified"},{"key":"FSW6MJXT","au":"Das, N K","y":1985,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"reliability coefficient r(x) = 1 - (S2w/S2b), Winer 1962 (within vs total rating variance)","estd":"ICC (single/unspec)","v":0.94,"n":"","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"priority rating 1.0 best to 5.0 worst","field":"biomedical (allergy and immunology)","wr":"committee reviewers on P01/P50 grant priority scores","conf":"med","self":false,"doi":"10.1007/BF00929456","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"NIH AITRC committee members scored an RFA group of P01/P50 grant applications on priority; an intraclass reliability coefficient of 0.94 indicates most score variance reflected real between-application quality differences (reviewer count OCR-garbled).","vf":"unverified"},{"key":"FSW6MJXT","au":"Das, N K","y":1985,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"reliability coefficient r(x) = 1 - (S2w/S2b), Winer 1962 (within vs total rating variance)","estd":"ICC (single/unspec)","v":0.98,"n":"","k":"13","samp":"full-pool","blind":"single","agg":"unspecified","scale":"priority rating 1.0 best to 5.0 worst","field":"biomedical (allergy and immunology)","wr":"committee reviewers on P01/P50 grant priority scores","conf":"med","self":false,"doi":"10.1007/BF00929456","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"NIH AITRC committee members scored an RFA group of P01/P50 grant applications on priority; an intraclass reliability coefficient of 0.98 with 13 reviewers indicates most score variance reflected real between-application quality differences.","vf":"unverified"},{"key":"FSW6MJXT","au":"Das, N K","y":1985,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"mean of RFA-group reliability coefficients r = 1 - (S2w/S2b), Winer 1962","estd":"ICC (single/unspec)","v":0.97,"n":"26","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"priority rating 1.0 best to 5.0 worst","field":"biomedical (allergy and immunology)","wr":"committee reviewers on P01/P50 grant priority scores","conf":"med","self":false,"doi":"10.1007/BF00929456","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":false,"ms":"Across the FY1982 RFA groups (26 P01/P50 applications, the largest yearly sample), the mean intraclass reliability coefficient of committee priority scores was 0.97, indicating high consensus on relative application quality; marked primary as no single grand-overall value is reported.","vf":"unverified"},{"key":"FSW6MJXT","au":"Das, N K","y":1985,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"reliability coefficient r(x) = 1 - (S2w/S2b), Winer 1962 (within vs total rating variance)","estd":"ICC (single/unspec)","v":0.98,"n":"","k":"16","samp":"full-pool","blind":"single","agg":"unspecified","scale":"priority rating 1.0 best to 5.0 worst","field":"biomedical (allergy and immunology)","wr":"committee reviewers on P01/P50 grant priority scores","conf":"med","self":false,"doi":"10.1007/BF00929456","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"NIH AITRC committee members scored an RFA group of P01/P50 grant applications on priority; an intraclass reliability coefficient of 0.98 with 16 reviewers indicates most score variance reflected real between-application quality differences.","vf":"unverified"},{"key":"FSW6MJXT","au":"Das, N K","y":1985,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"reliability coefficient r(x) = 1 - (S2w/S2b), Winer 1962 (within vs total rating variance)","estd":"ICC (single/unspec)","v":0.99,"n":"","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"priority rating 1.0 best to 5.0 worst","field":"biomedical (allergy and immunology)","wr":"committee reviewers on P01/P50 grant priority scores","conf":"med","self":false,"doi":"10.1007/BF00929456","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"NIH AITRC committee members scored an RFA group of P01/P50 grant applications on priority; an intraclass reliability coefficient of 0.99 indicates almost all score variance reflected real between-application quality differences (reviewer count OCR-garbled).","vf":"unverified"},{"key":"FSW6MJXT","au":"Das, N K","y":1985,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"reliability coefficient r(x) = 1 - (S2w/S2b), Winer 1962 (within vs total rating variance)","estd":"ICC (single/unspec)","v":0.93,"n":"","k":"9","samp":"full-pool","blind":"single","agg":"unspecified","scale":"priority rating 1.0 best to 5.0 worst","field":"biomedical (allergy and immunology)","wr":"committee reviewers on P01/P50 grant priority scores","conf":"med","self":false,"doi":"10.1007/BF00929456","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"NIH AITRC committee members scored an RFA group of P01/P50 grant applications on priority; an intraclass reliability coefficient of 0.93 with 9 reviewers indicates most score variance reflected real between-application quality differences.","vf":"unverified"},{"key":"FSW6MJXT","au":"Das, N K","y":1985,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"reliability coefficient r(x) = 1 - (S2w/S2b), Winer 1962 (within vs total rating variance)","estd":"ICC (single/unspec)","v":0.99,"n":"","k":"19","samp":"full-pool","blind":"single","agg":"unspecified","scale":"priority rating 1.0 best to 5.0 worst","field":"biomedical (allergy and immunology)","wr":"committee reviewers on P01/P50 grant priority scores","conf":"med","self":false,"doi":"10.1007/BF00929456","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"NIH AITRC committee members scored an RFA group of P01/P50 grant applications on priority; an intraclass reliability coefficient of 0.99 with 19 reviewers indicates almost all score variance reflected real between-application quality differences.","vf":"unverified"},{"key":"FSW6MJXT","au":"Das, N K","y":1985,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"mean of RFA-group reliability coefficients r = 1 - (S2w/S2b), Winer 1962","estd":"ICC (single/unspec)","v":0.97,"n":"21","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"priority rating 1.0 best to 5.0 worst","field":"biomedical (allergy and immunology)","wr":"committee reviewers on P01/P50 grant priority scores","conf":"med","self":false,"doi":"10.1007/BF00929456","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Across the FY1983 RFA groups (21 P01/P50 applications), the mean intraclass reliability coefficient of committee priority scores was 0.97, indicating high consensus on relative application quality.","vf":"unverified"},{"key":"4BPNLHNQ","au":"Doi, Suhail A.R.","y":2015,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"weighted kappa","estd":"weighted kappa","v":0.37,"n":"16","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"weak vs intermediate/strong; accept vs reject","field":"biomedical","wr":"CoRE score-class vs final editorial decision, manuscripts","conf":"high","self":false,"doi":"10.1080/08989621.2014.1002835","ciLow":0.02,"ciHigh":0.73,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the 16 manuscripts that reached a final editorial decision, the CoRE score-based classification (weak versus intermediate/strong) was compared with the editorial decision (accept versus reject). A weighted kappa of 0.37 indicates only fair concordance between the instrument classification and editorial decisions.","vf":"unverified"},{"key":"4BPNLHNQ","au":"Doi, Suhail A.R.","y":2015,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"quadratically weighted kappa","estd":"weighted kappa","v":0.79,"n":"17","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"weak/intermediate/strong vs reject/revise/accept","field":"biomedical","wr":"CoRE score-class vs reviewer recommendation, manuscripts","conf":"high","self":false,"doi":"10.1080/08989621.2014.1002835","ciLow":0.51,"ciHigh":1,"mt":"other","tgt":"other","rr":"none","pr":true,"he":false,"ms":"For 17 consecutively reviewed manuscripts at one medical journal, each manuscript's CoRE score-based class (weak, intermediate, or strong) was compared with the first responding reviewer's own recommendation (reject, revise, or accept). A quadratically weighted kappa of 0.79 indicates strong concordance between the instrument's classification and reviewer recommendations.","vf":"unverified"},{"key":"HEYTAC8Z","au":"Dun, Chen","y":2025,"cx":"General","ob":"review-report","fam":"kappa","form":"Cohen's kappa on binarised yes/no ratings","estd":"kappa","v":0.76,"n":"110","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"yes/no; 1-2 = no, 3-4 = yes","field":"health economics","wr":"reviewers rating binarised LLM assessment quality per item","conf":"med","self":false,"doi":"10.36469/jheor.2025.145214","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"When reviewers' 0-4 quality ratings were binarised to yes or no, the highest per-item Cohen's kappa between paired reviewers rose to 0.76, with eight of twenty-six items reaching moderate to substantial agreement.","vf":"unverified"},{"key":"HEYTAC8Z","au":"Dun, Chen","y":2025,"cx":"General","ob":"review-report","fam":"kappa","form":"Cohen's kappa","estd":"kappa","v":0.43,"n":"110","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"0-4 LLM-performance scale, 4 = accurate and supported","field":"health economics","wr":"reviewers rating LLM CHEERS-assessment quality per item","conf":"med","self":false,"doi":"10.36469/jheor.2025.145214","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":true,"he":false,"ms":"Eight trained informatics students and faculty rated the quality of an LLM's per-item CHEERS assessments across 110 papers, two per paper; per-item Cohen's kappa reached at most 0.43, and the headline interrater agreement is reported only as the range -0.07 to 0.43.","vf":"unverified"},{"key":"HEYTAC8Z","au":"Dun, Chen","y":2025,"cx":"General","ob":"review-report","fam":"kappa","form":"Cohen's kappa","estd":"kappa","v":-0.07,"n":"110","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"0-4 LLM-performance scale, 4 = accurate and supported","field":"health economics","wr":"reviewers rating LLM CHEERS-assessment quality per item","conf":"med","self":false,"doi":"10.36469/jheor.2025.145214","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Eight trained informatics students and faculty rated the quality of an LLM's per-item CHEERS assessments across 110 papers, two reviewers per paper; the lowest per-item Cohen's kappa was -0.07, worse than chance.","vf":"unverified"},{"key":"HEYTAC8Z","au":"Dun, Chen","y":2025,"cx":"General","ob":"review-report","fam":"kappa","form":"pairwise Cohen's kappa","estd":"kappa","v":0.425,"n":"","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"0-4 LLM-performance scale, 4 = accurate and supported","field":"health economics","wr":"highest pairwise agreement between two named reviewers","conf":"med","self":false,"doi":"10.36469/jheor.2025.145214","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"The highest agreement between any two of the eight reviewers was a pairwise Cohen's kappa of 0.425, between the pseudonymous reviewers apple pie and milkshake.","vf":"unverified"},{"key":"HEYTAC8Z","au":"Dun, Chen","y":2025,"cx":"General","ob":"review-report","fam":"kappa","form":"pairwise Cohen's kappa","estd":"kappa","v":-0.049,"n":"","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"0-4 LLM-performance scale, 4 = accurate and supported","field":"health economics","wr":"lowest pairwise agreement between two reviewers","conf":"med","self":false,"doi":"10.36469/jheor.2025.145214","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"The lowest agreement between any two of the eight reviewers was a pairwise Cohen's kappa of -0.049, reported for reviewers coconut and jellybean, indicating worse-than-chance agreement.","vf":"unverified"},{"key":"HEYTAC8Z","au":"Dun, Chen","y":2025,"cx":"General","ob":"review-report","fam":"percent-agreement","form":"percent agreement, exact match on binarised yes/no ratings","estd":"percent agreement","v":0.991,"n":"110","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"yes/no; 1-2 = no, 3-4 = yes","field":"health economics","wr":"reviewers rating binarised LLM assessment quality per item","conf":"med","self":false,"doi":"10.36469/jheor.2025.145214","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"With ratings binarised to yes or no, the highest per-item exact agreement between the two human reviewers was 99.1 percent.","vf":"unverified"},{"key":"HEYTAC8Z","au":"Dun, Chen","y":2025,"cx":"General","ob":"review-report","fam":"percent-agreement","form":"percent agreement, exact match on binarised yes/no ratings","estd":"percent agreement","v":0.527,"n":"110","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"yes/no; 1-2 = no, 3-4 = yes","field":"health economics","wr":"reviewers rating binarised LLM assessment quality per item","conf":"med","self":false,"doi":"10.36469/jheor.2025.145214","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"With ratings binarised to yes or no, the lowest per-item exact agreement between the two human reviewers was 52.7 percent.","vf":"unverified"},{"key":"HEYTAC8Z","au":"Dun, Chen","y":2025,"cx":"General","ob":"review-report","fam":"percent-agreement","form":"raw percent agreement, exact match on 0-4 ratings","estd":"percent agreement","v":0.955,"n":"110","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"0-4 LLM-performance scale, 4 = accurate and supported","field":"health economics","wr":"reviewers rating LLM assessment quality per item","conf":"med","self":false,"doi":"10.36469/jheor.2025.145214","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"The highest per-item raw exact agreement between the two human reviewers rating the LLM's assessment quality was 95.5 percent.","vf":"unverified"},{"key":"HEYTAC8Z","au":"Dun, Chen","y":2025,"cx":"General","ob":"review-report","fam":"percent-agreement","form":"raw percent agreement, exact match on 0-4 ratings","estd":"percent agreement","v":0.282,"n":"110","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"0-4 LLM-performance scale, 4 = accurate and supported","field":"health economics","wr":"reviewers rating LLM assessment quality per item","conf":"med","self":false,"doi":"10.36469/jheor.2025.145214","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"The lowest per-item raw exact agreement between the two human reviewers rating the LLM's assessment quality was 28.2 percent.","vf":"unverified"},{"key":"HEYTAC8Z","au":"Dun, Chen","y":2025,"cx":"General","ob":"review-report","fam":"kappa","form":"Cohen's kappa, mean of pairwise values","estd":"kappa","v":0.214,"n":"","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"0-4 LLM-performance scale, 4 = accurate and supported","field":"health economics","wr":"reviewer apple pie's mean pairwise agreement with peers","conf":"med","self":false,"doi":"10.36469/jheor.2025.145214","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Reviewer apple pie had a mean pairwise Cohen's kappa of 0.214 with the other reviewers (median 0.213, range 0.027 to 0.425).","vf":"unverified"},{"key":"HEYTAC8Z","au":"Dun, Chen","y":2025,"cx":"General","ob":"review-report","fam":"kappa","form":"Cohen's kappa, mean of pairwise values","estd":"kappa","v":0.067,"n":"","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"0-4 LLM-performance scale, 4 = accurate and supported","field":"health economics","wr":"reviewer caramel's mean pairwise agreement with peers","conf":"med","self":false,"doi":"10.36469/jheor.2025.145214","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Reviewer caramel had the lowest mean pairwise Cohen's kappa with the other reviewers at 0.067 (median 0.075, range -0.014 to 0.121).","vf":"unverified"},{"key":"HEYTAC8Z","au":"Dun, Chen","y":2025,"cx":"General","ob":"review-report","fam":"kappa","form":"Cohen's kappa, mean of pairwise values","estd":"kappa","v":0.139,"n":"","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"0-4 LLM-performance scale, 4 = accurate and supported","field":"health economics","wr":"reviewer cherry pop's mean pairwise agreement with peers","conf":"med","self":false,"doi":"10.36469/jheor.2025.145214","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Reviewer cherry pop had a mean pairwise Cohen's kappa of 0.139 with the other reviewers (median 0.109, range 0.027 to 0.409).","vf":"unverified"},{"key":"HEYTAC8Z","au":"Dun, Chen","y":2025,"cx":"General","ob":"review-report","fam":"kappa","form":"Cohen's kappa, mean of pairwise values","estd":"kappa","v":0.155,"n":"","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"0-4 LLM-performance scale, 4 = accurate and supported","field":"health economics","wr":"reviewer coconut's mean pairwise agreement with peers","conf":"med","self":false,"doi":"10.36469/jheor.2025.145214","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Reviewer coconut had a mean pairwise Cohen's kappa of 0.155 with the other reviewers (median 0.188, range -0.049 to 0.265), including worse-than-chance agreement with some peers.","vf":"unverified"},{"key":"HEYTAC8Z","au":"Dun, Chen","y":2025,"cx":"General","ob":"review-report","fam":"kappa","form":"Cohen's kappa, mean of pairwise values","estd":"kappa","v":0.094,"n":"","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"0-4 LLM-performance scale, 4 = accurate and supported","field":"health economics","wr":"reviewer jellybean's mean pairwise agreement with peers","conf":"med","self":false,"doi":"10.36469/jheor.2025.145214","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Reviewer jellybean had a mean pairwise Cohen's kappa of 0.094 with the other reviewers (median 0.111, range -0.049 to 0.231), including worse-than-chance agreement with some peers.","vf":"unverified"},{"key":"HEYTAC8Z","au":"Dun, Chen","y":2025,"cx":"General","ob":"review-report","fam":"kappa","form":"Cohen's kappa, mean of pairwise values","estd":"kappa","v":0.24,"n":"","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"0-4 LLM-performance scale, 4 = accurate and supported","field":"health economics","wr":"reviewer lollipop's mean pairwise agreement with peers","conf":"med","self":false,"doi":"10.36469/jheor.2025.145214","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Reviewer lollipop had the highest mean pairwise Cohen's kappa with the other seven reviewers at 0.24 (median 0.226, range 0.015 to 0.409).","vf":"unverified"},{"key":"HEYTAC8Z","au":"Dun, Chen","y":2025,"cx":"General","ob":"review-report","fam":"kappa","form":"Cohen's kappa, mean of pairwise values","estd":"kappa","v":0.181,"n":"","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"0-4 LLM-performance scale, 4 = accurate and supported","field":"health economics","wr":"reviewer milkshake's mean pairwise agreement with peers","conf":"med","self":false,"doi":"10.36469/jheor.2025.145214","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Reviewer milkshake had a mean pairwise Cohen's kappa of 0.181 with the other reviewers (median 0.108, range 0.039 to 0.425).","vf":"unverified"},{"key":"HEYTAC8Z","au":"Dun, Chen","y":2025,"cx":"General","ob":"review-report","fam":"kappa","form":"Cohen's kappa, mean of pairwise values","estd":"kappa","v":0.163,"n":"","k":"2","samp":"special","blind":"unclear","agg":"single-rater","scale":"0-4 LLM-performance scale, 4 = accurate and supported","field":"health economics","wr":"reviewer waffle cone's mean pairwise agreement with peers","conf":"med","self":false,"doi":"10.36469/jheor.2025.145214","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Reviewer waffle cone had a mean pairwise Cohen's kappa of 0.163 with the other reviewers (median 0.133, range 0.074 to 0.265).","vf":"unverified"},{"key":"2NTZHLW6","au":"Eaton, Warren O","y":1983,"cx":"General","ob":"other","fam":"kappa","form":"kappa for individual rating category, unweighted","estd":"kappa","v":0.11,"n":"111","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"acceptable / unclear (can not determine acceptability) / unacceptable","field":"psychology (research ethics)","wr":"reviewers on the 'acceptable' rating category","conf":"high","self":false,"doi":"10.1037/h0080684","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Category-specific Cohen's kappa for the 'acceptable' rating was 0.11, the highest of the three categories but still negligible chance-corrected agreement.","vf":"unverified"},{"key":"2NTZHLW6","au":"Eaton, Warren O","y":1983,"cx":"General","ob":"other","fam":"kappa","form":"kappa for individual rating category, unweighted","estd":"kappa","v":-0.05,"n":"111","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"acceptable / unclear (can not determine acceptability) / unacceptable","field":"psychology (research ethics)","wr":"reviewers on the 'unacceptable' rating category","conf":"high","self":false,"doi":"10.1037/h0080684","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Category-specific Cohen's kappa for the 'unacceptable' rating was -0.05, slightly below chance, the poorest agreement of the three categories.","vf":"unverified"},{"key":"2NTZHLW6","au":"Eaton, Warren O","y":1983,"cx":"General","ob":"other","fam":"kappa","form":"kappa for individual rating category, unweighted","estd":"kappa","v":0.07,"n":"111","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"acceptable / unclear (can not determine acceptability) / unacceptable","field":"psychology (research ethics)","wr":"reviewers on the 'unclear' rating category","conf":"high","self":false,"doi":"10.1037/h0080684","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Category-specific Cohen's kappa for the 'unclear' rating was 0.07, indicating essentially no chance-corrected agreement on that category.","vf":"unverified"},{"key":"2NTZHLW6","au":"Eaton, Warren O","y":1983,"cx":"General","ob":"other","fam":"kappa","form":"kappa statistic (Cohen, 1960), unweighted, entire category set","estd":"kappa","v":0.08,"n":"111","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"acceptable / unclear (can not determine acceptability) / unacceptable","field":"psychology (research ethics)","wr":"reviewers on research proposals' ethical acceptability","conf":"high","self":false,"doi":"10.1037/h0080684","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Pairs drawn from four ethics committee members independently rated the ethical acceptability of 111 research proposals into three categories; an overall Cohen's kappa of 0.08 indicates almost no chance-corrected agreement.","vf":"unverified"},{"key":"2NTZHLW6","au":"Eaton, Warren O","y":1983,"cx":"General","ob":"other","fam":"percent-agreement","form":"exact (perfect) agreement, both reviewers same category, entire category set","estd":"percent agreement","v":0.676,"n":"111","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"acceptable / unclear (can not determine acceptability) / unacceptable","field":"psychology (research ethics)","wr":"reviewers on research proposals' ethical acceptability","conf":"high","self":false,"doi":"10.1037/h0080684","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Pairs of committee reviewers agreed exactly on 75 of 111 proposals (67.6%); this uncorrected percent agreement overstates reliability given the skewed category use (abstract reports 67.7%).","vf":"unverified"},{"key":"2NTZHLW6","au":"Eaton, Warren O","y":1983,"cx":"General","ob":"other","fam":"percent-agreement","form":"obtained (exact) agreement for the 'acceptable' category","estd":"percent agreement","v":0.811,"n":"111","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"acceptable / unclear (can not determine acceptability) / unacceptable","field":"psychology (research ethics)","wr":"reviewers on the 'acceptable' rating category","conf":"high","self":false,"doi":"10.1037/h0080684","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Obtained (uncorrected) agreement for the 'acceptable' category was 0.811, the highest of the three categories.","vf":"unverified"},{"key":"2NTZHLW6","au":"Eaton, Warren O","y":1983,"cx":"General","ob":"other","fam":"percent-agreement","form":"obtained (exact) agreement for the 'unacceptable' category","estd":"percent agreement","v":0,"n":"111","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"acceptable / unclear (can not determine acceptability) / unacceptable","field":"psychology (research ethics)","wr":"reviewers on the 'unacceptable' rating category","conf":"high","self":false,"doi":"10.1037/h0080684","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Obtained (uncorrected) agreement for the 'unacceptable' category was 0.000; no proposal was rated unacceptable by both reviewers of a pair.","vf":"unverified"},{"key":"2NTZHLW6","au":"Eaton, Warren O","y":1983,"cx":"General","ob":"other","fam":"percent-agreement","form":"obtained (exact) agreement for the 'unclear' category","estd":"percent agreement","v":0.222,"n":"111","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"acceptable / unclear (can not determine acceptability) / unacceptable","field":"psychology (research ethics)","wr":"reviewers on the 'unclear' rating category","conf":"high","self":false,"doi":"10.1037/h0080684","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Obtained (uncorrected) agreement for the 'unclear' category was 0.222.","vf":"unverified"},{"key":"V5LUUSHS","au":"Erosheva, Elena A.","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater IRR (IRR_1), one-way random-effects/mixed model, REML via lme4, as in Pier et al. 2018","estd":"ICC (single/unspec)","v":0.37,"n":"72","k":"3","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 (best) to 5 (worst), 0.1 increments","field":"biomedical","wr":"reviewers scoring grant proposals on overall scientific merit","conf":"high","self":false,"doi":"10.1111/rssa.12681","ciLow":0.22,"ciHigh":0.52,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three expertise-matched reviewers independently scored each of 72 AIBS biomedical grant applications for overall scientific merit, with no panel discussion; across the full submission range the single-rater intraclass correlation was 0.37 (95% CI 0.22 to 0.52), a fair level of agreement between individual reviewers, with the Bayesian estimate almost identical.","vf":"unverified"},{"key":"V5LUUSHS","au":"Erosheva, Elena A.","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"multiple-rater IRR_n with n=3, one-way random-effects/mixed model, REML","estd":"ICC (average)","v":0.64,"n":"72","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1 (best) to 5 (worst), 0.1 increments","field":"biomedical","wr":"reviewers scoring grant proposals on overall scientific merit","conf":"high","self":false,"doi":"10.1111/rssa.12681","ciLow":0.46,"ciHigh":0.76,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the same 72 AIBS applications, the reliability of the mean of three reviewers' overall merit scores was 0.64 (95% CI 0.46 to 0.76), indicating good agreement when funding decisions rest on the three-reviewer average rather than a single reviewer.","vf":"unverified"},{"key":"V5LUUSHS","au":"Erosheva, Elena A.","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater IRR (IRR_1), one-way random-effects/mixed model, REML via lme4, as in Pier et al. 2018","estd":"ICC (single/unspec)","v":0.34,"n":"2076","k":"2.79","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 (best) to 9 (worst), whole numbers","field":"biomedical","wr":"reviewers scoring grant proposals on preliminary overall impact","conf":"high","self":false,"doi":"10.1111/rssa.12681","ciLow":0.31,"ciHigh":0.37,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Assigned NIH study-section reviewers independently scored a random sample of 2076 R01 applications (mean 2.79 reviewers each, range 1-5) on Preliminary Overall Impact before panel discussion; across the full submission range the single-rater intraclass correlation was 0.34 (95% CI 0.31 to 0.37), identical under REML and Bayesian estimation. This is the paper's headline result that full-range single-rater IRR is well above the previously reported estimate of zero.","vf":"unverified"},{"key":"V5LUUSHS","au":"Erosheva, Elena A.","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"multiple-rater IRR_n with n=3, one-way random-effects/mixed model, REML","estd":"ICC (average)","v":0.61,"n":"2076","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1 (best) to 9 (worst), whole numbers","field":"biomedical","wr":"reviewers scoring grant proposals on preliminary overall impact","conf":"high","self":false,"doi":"10.1111/rssa.12681","ciLow":0.58,"ciHigh":0.64,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the same NIH sample (mean 2.79 reviewers per application), the reliability of a three-reviewer average score was 0.61 (95% CI 0.58 to 0.64), indicating good agreement across the complete range of submissions even though single-reviewer reliability was only fair.","vf":"unverified"},{"key":"SA84R68P","au":"Etzel, Franka Tabitha","y":2025,"cx":"General","ob":"journal-manuscript","fam":"Fleiss-kappa","form":"Fleiss Kappa (Fleiss, 1971)","estd":"kappa","v":-0.02,"n":"52","k":"3","samp":"special","blind":"single","agg":"single-rater","scale":"single-choice or multiple-choice criterion","field":"psychology","wr":"students on formal modelling of a theory","conf":"high","self":false,"doi":"10.31234/osf.io/4w7rb_v2","ciLow":-0.18,"ciHigh":0.14,"mt":"inter-rater","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"Three psychology students independently rated nominated papers on the criterion Formal modeling of a theory; a Fleiss kappa of -0.02 shows chance-level agreement. Tied minimum of the 20 Study 1 criterion kappas (capped extraction).","vf":"unverified"},{"key":"SA84R68P","au":"Etzel, Franka Tabitha","y":2025,"cx":"General","ob":"journal-manuscript","fam":"Fleiss-kappa","form":"Fleiss Kappa (Fleiss, 1971)","estd":"kappa","v":0.88,"n":"52","k":"3","samp":"special","blind":"single","agg":"single-rater","scale":"single-choice or multiple-choice criterion","field":"psychology","wr":"students on presence of preregistration","conf":"high","self":false,"doi":"10.31234/osf.io/4w7rb_v2","ciLow":0.74,"ciHigh":1,"mt":"inter-rater","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"Three psychology students independently rated nominated papers on the criterion Preregistration; a Fleiss kappa of 0.88 indicates almost perfect agreement. Maximum of the 20 Study 1 criterion kappas (capped extraction).","vf":"unverified"},{"key":"SA84R68P","au":"Etzel, Franka Tabitha","y":2025,"cx":"General","ob":"journal-manuscript","fam":"Fleiss-kappa","form":"Fleiss Kappa (Fleiss, 1971)","estd":"kappa","v":-0.02,"n":"52","k":"3","samp":"special","blind":"single","agg":"single-rater","scale":"single-choice or multiple-choice criterion","field":"psychology","wr":"students on version control of open scripts","conf":"high","self":false,"doi":"10.31234/osf.io/4w7rb_v2","ciLow":-0.19,"ciHigh":0.16,"mt":"inter-rater","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"Three psychology students independently rated nominated papers on the criterion FAIR format of open scripts: version control; a Fleiss kappa of -0.02 shows chance-level agreement. Tied minimum of the 20 Study 1 criterion kappas (capped extraction).","vf":"unverified"},{"key":"SA84R68P","au":"Etzel, Franka Tabitha","y":2025,"cx":"General","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,1)","estd":"ICC (single/unspec)","v":0.94,"n":"21","k":"3","samp":"special","blind":"single","agg":"single-rater","scale":"Relative Rigor Score averaged across up to three papers","field":"psychology","wr":"students on author-level average rigour score","conf":"high","self":false,"doi":"10.31234/osf.io/4w7rb_v2","ciLow":0.88,"ciHigh":0.97,"mt":"inter-rater","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"Relative Rigor Scores were averaged across up to three papers for each of the 21 nominating researchers; the author-level single-rater reliability was ICC(1,1) = 0.94.","vf":"unverified"},{"key":"SA84R68P","au":"Etzel, Franka Tabitha","y":2025,"cx":"General","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,1)","estd":"ICC (single/unspec)","v":0.91,"n":"52","k":"3","samp":"special","blind":"single","agg":"single-rater","scale":"Relative Rigor Score across applicable criteria","field":"psychology","wr":"students on overall paper rigour score","conf":"high","self":false,"doi":"10.31234/osf.io/4w7rb_v2","ciLow":0.86,"ciHigh":0.94,"mt":"inter-rater","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"Three student raters independently scored each paper on all applicable rigour criteria; the composite Relative Rigor Score reached ICC(1,1) = 0.91, excellent single-rater reliability and Study 1's overall result.","vf":"unverified"},{"key":"SA84R68P","au":"Etzel, Franka Tabitha","y":2025,"cx":"General","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,1), Type 1 (Shrout & Fleiss, 1979)","estd":"ICC (single/unspec)","v":-0.03,"n":"110","k":"3","samp":"special","blind":"single","agg":"single-rater","scale":"does not apply (1) / partly applies (2) / does apply (3)","field":"psychology","wr":"students on challenging consensus research goals","conf":"high","self":false,"doi":"10.31234/osf.io/4w7rb_v2","ciLow":-0.05,"ciHigh":0.01,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Three of nine untrained student raters scored 110 psychology papers on whether they challenge consensus research goals; ICC(1,1) = -0.03 is the single-rater reliability and the minimum criterion-level estimate (capped extraction).","vf":"unverified"},{"key":"SA84R68P","au":"Etzel, Franka Tabitha","y":2025,"cx":"General","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,3), Type 1 (Shrout & Fleiss, 1979)","estd":"ICC (average)","v":-0.08,"n":"110","k":"3","samp":"special","blind":"single","agg":"average-of-k","scale":"does not apply (1) / partly applies (2) / does apply (3)","field":"psychology","wr":"students on challenging consensus research goals","conf":"high","self":false,"doi":"10.31234/osf.io/4w7rb_v2","ciLow":-0.16,"ciHigh":0.02,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the same 110 papers, ICC(1,3) = -0.08 is the reliability of the mean of three raters' judgments on this criterion and the minimum criterion-level average-rating estimate (capped extraction).","vf":"unverified"},{"key":"SA84R68P","au":"Etzel, Franka Tabitha","y":2025,"cx":"General","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,1), Type 1 (Shrout & Fleiss, 1979)","estd":"ICC (single/unspec)","v":0.89,"n":"110","k":"3","samp":"special","blind":"single","agg":"single-rater","scale":"does not apply (1) / partly applies (2) / does apply (3)","field":"psychology","wr":"students on availability of open code","conf":"high","self":false,"doi":"10.31234/osf.io/4w7rb_v2","ciLow":0.87,"ciHigh":0.92,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Three of nine untrained student raters scored 110 psychology papers on open code availability; ICC(1,1) = 0.89 is the single-rater reliability and the maximum criterion-level estimate (capped extraction).","vf":"unverified"},{"key":"SA84R68P","au":"Etzel, Franka Tabitha","y":2025,"cx":"General","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,3), Type 1 (Shrout & Fleiss, 1979)","estd":"ICC (average)","v":0.96,"n":"110","k":"3","samp":"special","blind":"single","agg":"average-of-k","scale":"does not apply (1) / partly applies (2) / does apply (3)","field":"psychology","wr":"students on availability of open code","conf":"high","self":false,"doi":"10.31234/osf.io/4w7rb_v2","ciLow":0.95,"ciHigh":0.97,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the same 110 papers, ICC(1,3) = 0.96 is the reliability of the mean of three raters' judgments on open code and the maximum criterion-level average-rating estimate (capped extraction).","vf":"unverified"},{"key":"SA84R68P","au":"Etzel, Franka Tabitha","y":2025,"cx":"General","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,1), Type 1 (Shrout & Fleiss, 1979)","estd":"ICC (single/unspec)","v":-0.02,"n":"110","k":"3","samp":"special","blind":"single","agg":"single-rater","scale":"overall score across consensus criteria (continuous composite)","field":"psychology","wr":"students on overall consensus-criteria score","conf":"high","self":false,"doi":"10.31234/osf.io/4w7rb_v2","ciLow":-0.04,"ciHigh":0.02,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Three of nine untrained student raters scored 110 psychology papers; the overall score across consensus-building criteria showed a single-rater reliability of ICC(1,1) = -0.02, effectively zero.","vf":"unverified"},{"key":"SA84R68P","au":"Etzel, Franka Tabitha","y":2025,"cx":"General","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,3), Type 1 (Shrout & Fleiss, 1979)","estd":"ICC (average)","v":-0.05,"n":"110","k":"3","samp":"special","blind":"single","agg":"average-of-k","scale":"overall score across consensus criteria (continuous composite)","field":"psychology","wr":"students on overall consensus-criteria score","conf":"high","self":false,"doi":"10.31234/osf.io/4w7rb_v2","ciLow":-0.13,"ciHigh":0.05,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the same 110 papers, the reliability of the mean of three raters' overall consensus-criteria scores was ICC(1,3) = -0.05, effectively zero.","vf":"unverified"},{"key":"SA84R68P","au":"Etzel, Franka Tabitha","y":2025,"cx":"General","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,1), Type 1 (Shrout & Fleiss, 1979)","estd":"ICC (single/unspec)","v":0.74,"n":"110","k":"3","samp":"special","blind":"single","agg":"single-rater","scale":"overall score across methodological rigour criteria (continuous composite)","field":"psychology","wr":"students on overall methodological rigour score","conf":"high","self":false,"doi":"10.31234/osf.io/4w7rb_v2","ciLow":0.69,"ciHigh":0.8,"mt":"inter-rater","tgt":"other","rr":"none","pr":true,"he":false,"ms":"Three of nine untrained student raters scored 110 psychology papers; the overall methodological rigour score reached ICC(1,1) = 0.74, the preregistered main endpoint of Study 2 and the paper's headline single-rater result.","vf":"unverified"},{"key":"SA84R68P","au":"Etzel, Franka Tabitha","y":2025,"cx":"General","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,3), Type 1 (Shrout & Fleiss, 1979)","estd":"ICC (average)","v":0.9,"n":"110","k":"3","samp":"special","blind":"single","agg":"average-of-k","scale":"overall score across methodological rigour criteria (continuous composite)","field":"psychology","wr":"students on overall methodological rigour score","conf":"high","self":false,"doi":"10.31234/osf.io/4w7rb_v2","ciLow":0.87,"ciHigh":0.92,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the same 110 papers, the reliability of the mean of three raters' overall methodological rigour scores was ICC(1,3) = 0.90, good to excellent.","vf":"unverified"},{"key":"3MVDPR4T","au":"Feliciani, Thomas","y":2022,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"average normalized Hamming distance between reviewers' topic-criteria mappings for one criterion (dissimilarity; higher = more disagreement)","estd":"other","v":0.359,"n":"","k":"2","samp":"special","blind":"single","agg":"unspecified","scale":"binary link: present / absent","field":"multi-field","wr":"reviewers' topic mappings for the 'applicant' criterion","conf":"med","self":false,"doi":"10.1162/qss_a_00207","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the 'applicant' evaluation criterion alone, the average normalised Hamming distance between the 261 reviewers' topic mappings was 0.359, the lowest of the three criteria, indicating slightly more shared interpretation for this criterion.","vf":"unverified"},{"key":"3MVDPR4T","au":"Feliciani, Thomas","y":2022,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"average normalized Hamming distance between reviewers' topic-criteria mappings for one criterion (dissimilarity; higher = more disagreement)","estd":"other","v":0.389,"n":"","k":"2","samp":"special","blind":"single","agg":"unspecified","scale":"binary link: present / absent","field":"multi-field","wr":"reviewers' topic mappings for the 'potential for impact' criterion","conf":"med","self":false,"doi":"10.1162/qss_a_00207","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the 'potential for impact' evaluation criterion, the average normalised Hamming distance between the 261 reviewers' topic mappings was 0.389, the highest of the three criteria, indicating the least shared interpretation among reviewers.","vf":"unverified"},{"key":"3MVDPR4T","au":"Feliciani, Thomas","y":2022,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"average normalized Hamming distance between reviewers' topic-criteria mappings for one criterion (dissimilarity; higher = more disagreement)","estd":"other","v":0.36,"n":"","k":"2","samp":"special","blind":"single","agg":"unspecified","scale":"binary link: present / absent","field":"multi-field","wr":"reviewers' topic mappings for the 'proposed research' criterion","conf":"med","self":false,"doi":"10.1162/qss_a_00207","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the 'proposed research' evaluation criterion, the average normalised Hamming distance between the 261 reviewers' topic mappings was 0.36, indicating a similar degree of interpretive heterogeneity to the 'applicant' criterion.","vf":"unverified"},{"key":"3MVDPR4T","au":"Feliciani, Thomas","y":2022,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"average normalized Hamming distance between reviewers' topic-criteria mapping networks (dissimilarity; higher = more disagreement)","estd":"other","v":0.37,"n":"","k":"2","samp":"special","blind":"single","agg":"unspecified","scale":"binary link: present / absent","field":"multi-field","wr":"reviewers' topic-to-criterion mapping choices","conf":"med","self":false,"doi":"10.1162/qss_a_00207","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":true,"he":false,"ms":"In a survey, 261 Science Foundation Ireland grant reviewers reported which review topics they associate with each evaluation criterion; the average normalised Hamming distance of about 0.37 between reviewers' mapping networks indicates moderate heterogeneity in how reviewers interpret the criteria, not agreement on proposal scores.","vf":"unverified"},{"key":"GNW32DYJ","au":"Fiske, Donald W.","y":1990,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlations","estd":"ICC (single/unspec)","v":0.68,"n":"","k":"","samp":"special","blind":"unclear","agg":"unspecified","scale":"15-point recommendation scale","field":"psychology","wr":"reviewers' overall publication recommendations for manuscripts","conf":"high","self":false,"doi":"10.1037/0003-066x.45.5.591","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The highest of the 12 per-editor-journal intraclass correlations for reviewers' recommendations was 0.68, the strongest agreement observed in any single sample.","vf":"unverified"},{"key":"GNW32DYJ","au":"Fiske, Donald W.","y":1990,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlations","estd":"ICC (single/unspec)","v":0.2,"n":"","k":"","samp":"special","blind":"unclear","agg":"unspecified","scale":"15-point recommendation scale","field":"psychology","wr":"reviewers' overall publication recommendations for manuscripts","conf":"high","self":false,"doi":"10.1037/0003-066x.45.5.591","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Reviewers' overall recommendations for manuscripts submitted to seven APA journals were coded onto a 15-point scale. The mean intraclass correlation across the 12 editor-journal samples was 0.20, indicating very low agreement between reviewers.","vf":"unverified"},{"key":"GNW32DYJ","au":"Fiske, Donald W.","y":1990,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlations","estd":"ICC (single/unspec)","v":-0.23,"n":"","k":"","samp":"special","blind":"unclear","agg":"unspecified","scale":"15-point recommendation scale","field":"psychology","wr":"reviewers' overall publication recommendations for manuscripts","conf":"high","self":false,"doi":"10.1037/0003-066x.45.5.591","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The lowest of the 12 per-editor-journal intraclass correlations for reviewers' recommendations was -0.23, showing no agreement between reviewers in that sample.","vf":"unverified"},{"key":"Y5JCHZCD","au":"Fleurence, Rachael","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"proportion of scores that did not change by 10 or more points from prediscussion to postdiscussion (tolerance: less than 10 points)","estd":"percent agreement","v":0.58,"n":"","k":"1","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1 (exceptional) to 9 (poor), multiplied by 10","field":"biomedical","wr":"lead reviewers rescoring grant proposals after discussion","conf":"high","self":false,"doi":"10.7326/m13-2412","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":false,"ms":"Each phase 2 lead reviewer gave the same application a score before and after the in-person panel discussion. Across 377 reviewer-application observations, 58% of postdiscussion scores changed by less than 10 points from the reviewer's own prediscussion score on the 10 to 90 scale.","vf":"unverified"},{"key":"Y5JCHZCD","au":"Fleurence, Rachael","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"proportion of scores that did not change by 10 or more points from prediscussion to postdiscussion (tolerance: less than 10 points)","estd":"percent agreement","v":0.53,"n":"","k":"1","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1 (exceptional) to 9 (poor), multiplied by 10","field":"biomedical","wr":"patient reviewers rescoring grant proposals after discussion","conf":"high","self":false,"doi":"10.7326/m13-2412","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Patient lead reviewers rescored their applications after the panel meeting. In 98 reviewer-application observations, 53% of postdiscussion scores changed by less than 10 points from the same patient's prediscussion score, the least stable of the three reviewer types.","vf":"unverified"},{"key":"Y5JCHZCD","au":"Fleurence, Rachael","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"proportion of scores that did not change by 10 or more points from prediscussion to postdiscussion (tolerance: less than 10 points)","estd":"percent agreement","v":0.6,"n":"","k":"1","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1 (exceptional) to 9 (poor), multiplied by 10","field":"biomedical","wr":"scientist reviewers rescoring grant proposals after discussion","conf":"high","self":false,"doi":"10.7326/m13-2412","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Scientist lead reviewers rescored their applications after the panel meeting. In 188 reviewer-application observations, 60% of postdiscussion scores changed by less than 10 points from the same scientist's prediscussion score.","vf":"unverified"},{"key":"Y5JCHZCD","au":"Fleurence, Rachael","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"proportion of scores that did not change by 10 or more points from prediscussion to postdiscussion (tolerance: less than 10 points)","estd":"percent agreement","v":0.62,"n":"","k":"1","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1 (exceptional) to 9 (poor), multiplied by 10","field":"biomedical","wr":"stakeholder reviewers rescoring grant proposals after discussion","conf":"high","self":false,"doi":"10.7326/m13-2412","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Stakeholder lead reviewers rescored their applications after the panel meeting. In 91 reviewer-application observations, 62% of postdiscussion scores changed by less than 10 points from the same stakeholder's prediscussion score.","vf":"unverified"},{"key":"JCNY992P","au":"Fogelholm, Mikael","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"kappa","form":"Cohen kappa, dichotomised at cutoff score 5 (scores 5-6 vs 1-4)","estd":"kappa","v":0.17,"n":"65","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"scores 1-4 vs 5-6 (low vs high funding likelihood)","field":"biomedical","wr":"two reviewers on funded-vs-not classification, panel A","conf":"high","self":false,"doi":"10.1016/j.jclinepi.2011.05.001","ciLow":-0.08,"ciHigh":0.42,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Within panel A, the two reviewers' preliminary scores were dichotomised at the funding cutoff; their Cohen kappa on the 65 proposals was 0.17, indicating slight chance-corrected agreement.","vf":"unverified"},{"key":"JCNY992P","au":"Fogelholm, Mikael","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"kappa","form":"Cohen kappa, dichotomised at cutoff score 5 (scores 5-6 vs 1-4)","estd":"kappa","v":0.08,"n":"65","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"scores 1-4 vs 5-6 (low vs high funding likelihood)","field":"biomedical","wr":"two reviewers on funded-vs-not classification, panel B","conf":"high","self":false,"doi":"10.1016/j.jclinepi.2011.05.001","ciLow":-0.18,"ciHigh":0.33,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Within panel B, the two reviewers' preliminary scores were dichotomised at the funding cutoff; their Cohen kappa on the 65 proposals was 0.08, indicating near-zero chance-corrected agreement.","vf":"unverified"},{"key":"JCNY992P","au":"Fogelholm, Mikael","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement on dichotomised classification (funded 5-6 vs 1-4)","estd":"percent agreement","v":0.625,"n":"65","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"scores 1-4 vs 5-6 (low vs high funding likelihood)","field":"biomedical","wr":"two reviewers, exact agreement on funding class, panel A","conf":"high","self":false,"doi":"10.1016/j.jclinepi.2011.05.001","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The two panel A reviewers classified 62.5% of the 65 proposals identically as likely funded or not, an observed exact-agreement proportion uncorrected for chance.","vf":"unverified"},{"key":"JCNY992P","au":"Fogelholm, Mikael","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement on dichotomised classification (funded 5-6 vs 1-4)","estd":"percent agreement","v":0.574,"n":"65","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"scores 1-4 vs 5-6 (low vs high funding likelihood)","field":"biomedical","wr":"two reviewers, exact agreement on funding class, panel B","conf":"high","self":false,"doi":"10.1016/j.jclinepi.2011.05.001","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The two panel B reviewers classified 57.4% of the 65 proposals identically as likely funded or not, an observed exact-agreement proportion uncorrected for chance.","vf":"unverified"},{"key":"JCNY992P","au":"Fogelholm, Mikael","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Spearman rank correlation","estd":"correlation","v":0.05,"n":"65","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1-6, 1 = weak to 6 = outstanding","field":"biomedical","wr":"two reviewers on grant proposal scores, panel A","conf":"high","self":false,"doi":"10.1016/j.jclinepi.2011.05.001","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The Spearman rank correlation between the two panel A reviewers' preliminary scores on the 65 proposals was 0.05, indicating essentially no inter-reviewer association.","vf":"unverified"},{"key":"JCNY992P","au":"Fogelholm, Mikael","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Spearman rank correlation","estd":"correlation","v":0.26,"n":"65","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1-6, 1 = weak to 6 = outstanding","field":"biomedical","wr":"two reviewers on grant proposal scores, panel B","conf":"high","self":false,"doi":"10.1016/j.jclinepi.2011.05.001","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The Spearman rank correlation between the two panel B reviewers' preliminary scores on the 65 proposals was 0.26, indicating weak inter-reviewer association.","vf":"unverified"},{"key":"JCNY992P","au":"Fogelholm, Mikael","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"kappa","form":"Cohen kappa, dichotomised at cutoff score 5 (scores 5-6 vs 1-4)","estd":"kappa","v":0.12,"n":"65","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"panel-consensus","scale":"scores 1-4 vs 5-6 (low vs high funding likelihood)","field":"biomedical","wr":"two panels on funded-vs-not classification (consensus)","conf":"high","self":false,"doi":"10.1016/j.jclinepi.2011.05.001","ciLow":-0.14,"ciHigh":0.36,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Comparing the two panels' post-discussion consensus scores on the dichotomised funding classification of the 65 proposals gave a Cohen kappa of 0.12, indicating slight interpanel agreement.","vf":"unverified"},{"key":"JCNY992P","au":"Fogelholm, Mikael","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"kappa","form":"Cohen kappa, dichotomised at cutoff score 5, half scores rounded up","estd":"kappa","v":0.32,"n":"65","k":"4","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"mean scores dichotomised 1-4 vs 5-6, halves rounded up","field":"biomedical","wr":"two panels on funded-vs-not classification (mean scores)","conf":"high","self":false,"doi":"10.1016/j.jclinepi.2011.05.001","ciLow":0.08,"ciHigh":0.56,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Comparing the two panels' mean preliminary reviewer scores on the dichotomised funding classification of the 65 proposals gave a Cohen kappa of 0.32, the highest interpanel agreement observed.","vf":"unverified"},{"key":"JCNY992P","au":"Fogelholm, Mikael","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement on dichotomised classification (funded 5-6 vs 1-4)","estd":"percent agreement","v":0.646,"n":"65","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"panel-consensus","scale":"scores 1-4 vs 5-6 (low vs high funding likelihood)","field":"biomedical","wr":"two panels, exact agreement on funding class (consensus)","conf":"high","self":false,"doi":"10.1016/j.jclinepi.2011.05.001","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"The two panels' post-discussion consensus scores classified 64.6% of the 65 proposals identically as likely funded or not, an observed exact-agreement proportion uncorrected for chance.","vf":"unverified"},{"key":"JCNY992P","au":"Fogelholm, Mikael","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement on dichotomised classification, half scores rounded up","estd":"percent agreement","v":0.692,"n":"65","k":"4","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"mean scores dichotomised 1-4 vs 5-6, halves rounded up","field":"biomedical","wr":"two panels, exact agreement on funding class (mean scores)","conf":"high","self":false,"doi":"10.1016/j.jclinepi.2011.05.001","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The two panels' mean preliminary reviewer scores classified 69.2% of the 65 proposals identically as likely funded or not, the highest observed exact-agreement proportion.","vf":"unverified"},{"key":"JCNY992P","au":"Fogelholm, Mikael","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Spearman rank correlation","estd":"correlation","v":0.3,"n":"65","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"panel-consensus","scale":"1-6, 1 = weak to 6 = outstanding","field":"biomedical","wr":"two panels' consensus scores on grant proposals","conf":"high","self":false,"doi":"10.1016/j.jclinepi.2011.05.001","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"The Spearman rank correlation between the two panels' post-discussion consensus scores on the 65 proposals was 0.30, indicating low interpanel association.","vf":"unverified"},{"key":"JCNY992P","au":"Fogelholm, Mikael","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Spearman rank correlation","estd":"correlation","v":0.41,"n":"65","k":"4","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"mean of two 1-6 reviewer scores","field":"biomedical","wr":"two panels' mean reviewer scores on grant proposals","conf":"high","self":false,"doi":"10.1016/j.jclinepi.2011.05.001","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The Spearman rank correlation between the two panels' mean preliminary reviewer scores on the 65 proposals was 0.41, the strongest interpanel association observed.","vf":"unverified"},{"key":"JCNY992P","au":"Fogelholm, Mikael","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"weighted-kappa","form":"Cohen's kappa coefficients with linear weighting","estd":"weighted kappa","v":0.23,"n":"65","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"panel-consensus","scale":"1-6, 1 = weak to 6 = outstanding","field":"biomedical","wr":"two panels' consensus scores on grant proposals","conf":"high","self":false,"doi":"10.1016/j.jclinepi.2011.05.001","ciLow":0.08,"ciHigh":0.39,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":false,"ms":"Two randomised 15-member panels each assigned post-discussion consensus scores to the same 65 grant proposals; the linearly weighted Cohen kappa between the panels was 0.23, indicating low interpanel agreement.","vf":"unverified"},{"key":"JCNY992P","au":"Fogelholm, Mikael","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"weighted-kappa","form":"kappa coefficients with linear weighting","estd":"weighted kappa","v":0.23,"n":"65","k":"4","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"mean of two reviewer scores, rounded to 1-6","field":"biomedical","wr":"two panels' mean reviewer scores on grant proposals","conf":"high","self":false,"doi":"10.1016/j.jclinepi.2011.05.001","ciLow":0,"ciHigh":0.46,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"When each panel's score was replaced by the mean of its two reviewers' preliminary independent scores, the interpanel linear-weighted kappa was again 0.23, showing discussion did not improve agreement.","vf":"unverified"},{"key":"JCNY992P","au":"Fogelholm, Mikael","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"proportion of score pairs differing by at least two points","estd":"other","v":0.26,"n":"65","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"panel-consensus","scale":"1-6, 1 = weak to 6 = outstanding","field":"biomedical","wr":"two panels' consensus scores, large differences","conf":"high","self":false,"doi":"10.1016/j.jclinepi.2011.05.001","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"The two panels' post-discussion consensus scores on the 65 proposals differed by at least two points on the six-point scale for 26% of proposals, a disagreement proportion.","vf":"unverified"},{"key":"JCNY992P","au":"Fogelholm, Mikael","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"proportion of score pairs differing by at least two points","estd":"other","v":0.14,"n":"65","k":"4","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"mean of two 1-6 reviewer scores","field":"biomedical","wr":"two panels' mean reviewer scores, large differences","conf":"high","self":false,"doi":"10.1016/j.jclinepi.2011.05.001","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The two panels' mean preliminary reviewer scores on the 65 proposals differed by at least two points for 14% of proposals, the smallest disagreement proportion observed.","vf":"unverified"},{"key":"JCNY992P","au":"Fogelholm, Mikael","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"proportion of score pairs differing by at least two points","estd":"other","v":0.4,"n":"65","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1-6, 1 = weak to 6 = outstanding","field":"biomedical","wr":"two reviewers' preliminary scores, large differences, panel A","conf":"high","self":false,"doi":"10.1016/j.jclinepi.2011.05.001","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Within panel A, the two reviewers' preliminary scores on the 65 proposals differed by at least two points on the six-point scale for 40% of proposals, a disagreement proportion.","vf":"unverified"},{"key":"JCNY992P","au":"Fogelholm, Mikael","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"proportion of score pairs differing by at least two points","estd":"other","v":0.36,"n":"65","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1-6, 1 = weak to 6 = outstanding","field":"biomedical","wr":"two reviewers' preliminary scores, large differences, panel B","conf":"high","self":false,"doi":"10.1016/j.jclinepi.2011.05.001","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Within panel B, the two reviewers' preliminary scores on the 65 proposals differed by at least two points on the six-point scale for 36% of proposals, a disagreement proportion.","vf":"unverified"},{"key":"JCNY992P","au":"Fogelholm, Mikael","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Spearman rank correlation","estd":"correlation","v":0.94,"n":"","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"1-6, 1 = weak to 6 = outstanding","field":"biomedical","wr":"individual reviewer vs own panel's consensus scores","conf":"high","self":false,"doi":"10.1016/j.jclinepi.2011.05.001","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Each reviewer's preliminary scores were correlated with the panel's final consensus scores on that reviewer's proposals; across 30 reviewers the Spearman correlations ranged from 0.03 to 0.94, this row records the maximum.","vf":"unverified"},{"key":"JCNY992P","au":"Fogelholm, Mikael","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Spearman rank correlation","estd":"correlation","v":0.03,"n":"","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"1-6, 1 = weak to 6 = outstanding","field":"biomedical","wr":"individual reviewer vs own panel's consensus scores","conf":"high","self":false,"doi":"10.1016/j.jclinepi.2011.05.001","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Each reviewer's preliminary scores were correlated with the panel's final consensus scores on that reviewer's proposals; across 30 reviewers the Spearman correlations ranged from 0.03 to 0.94, this row records the minimum.","vf":"unverified"},{"key":"A42HFS8G","au":"Forscher, Patrick S.","y":2019,"cx":"Grant","ob":"grant-proposal","fam":"G-theory","form":"reliability","estd":"G-theory","v":0.5,"n":"48","k":"86","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"number of words in each of 9 categories","field":"biomedical","wr":"word-category use in proposal critiques","conf":"med","self":false,"doi":"10.31234/osf.io/483zj","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"restricted-other","pr":true,"he":false,"ms":"Reviewers independently wrote critiques of 48 proposals, analysed across nine word categories. The authors forecast that at least 86 critiques per proposal would be needed for the aggregate word counts to reach a reliability of 0.50; at three reviewers the reliabilities are essentially zero.","vf":"unverified"},{"key":"A42HFS8G","au":"Forscher, Patrick S.","y":2019,"cx":"Grant","ob":"grant-proposal","fam":"G-theory","form":"absolute reliability","estd":"G-theory","v":0.29,"n":"48","k":"3","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"1 (exceptional) to 9 (poor)","field":"biomedical","wr":"reviewers on grant proposal Investigator scores","conf":"med","self":false,"doi":"10.31234/osf.io/483zj","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For the Investigator dimension, an aggregate of three primary reviewers yields an absolute Generalizability Theory reliability of 0.29, the highest of any dimension but still low.","vf":"unverified"},{"key":"A42HFS8G","au":"Forscher, Patrick S.","y":2019,"cx":"Grant","ob":"grant-proposal","fam":"G-theory","form":"relative reliability","estd":"G-theory","v":0.35,"n":"48","k":"3","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"1 (exceptional) to 9 (poor)","field":"biomedical","wr":"reviewers on grant proposal Investigator scores","conf":"med","self":false,"doi":"10.31234/osf.io/483zj","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For the Investigator dimension, an aggregate of three primary reviewers yields a relative Generalizability Theory reliability of 0.35, the best rank-order reliability observed across dimensions but still modest.","vf":"unverified"},{"key":"A42HFS8G","au":"Forscher, Patrick S.","y":2019,"cx":"Grant","ob":"grant-proposal","fam":"G-theory","form":"reliability","estd":"G-theory","v":0.2,"n":"48","k":"3","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"1 (exceptional) to 9 (poor)","field":"biomedical","wr":"reviewers on grant proposal Overall Impact scores","conf":"med","self":false,"doi":"10.31234/osf.io/483zj","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":true,"ms":"412 scientists each independently re-reviewed three of 48 NIH R01 proposals. Using Generalizability Theory, the authors estimate that an aggregate of three reviewers yields a reliability of about 0.2 for Overall Impact, the most heavily weighted dimension, indicating low agreement.","vf":"unverified"},{"key":"A42HFS8G","au":"Forscher, Patrick S.","y":2019,"cx":"Grant","ob":"grant-proposal","fam":"G-theory","form":"reliability","estd":"G-theory","v":0.5,"n":"48","k":"12","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"1 (exceptional) to 9 (poor)","field":"biomedical","wr":"reviewers on grant proposal Overall Impact scores","conf":"med","self":false,"doi":"10.31234/osf.io/483zj","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"From the same variance components, the authors forecast that averaging 12 reviewers per proposal would be needed to raise the reliability of Overall Impact scores to 0.50.","vf":"unverified"},{"key":"A42HFS8G","au":"Forscher, Patrick S.","y":2019,"cx":"Grant","ob":"grant-proposal","fam":"G-theory","form":"absolute reliability","estd":"G-theory","v":0.5,"n":"48","k":"21","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"1 (exceptional) to 9 (poor)","field":"biomedical","wr":"reviewers on grant proposal Significance scores","conf":"med","self":false,"doi":"10.31234/osf.io/483zj","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For the Significance dimension, the authors forecast that 21 reviewers per proposal would be required to reach an absolute reliability of 0.50.","vf":"unverified"},{"key":"A42HFS8G","au":"Forscher, Patrick S.","y":2019,"cx":"Grant","ob":"grant-proposal","fam":"G-theory","form":"relative reliability","estd":"G-theory","v":0.5,"n":"48","k":"17","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"1 (exceptional) to 9 (poor)","field":"biomedical","wr":"reviewers on grant proposal Significance scores","conf":"med","self":false,"doi":"10.31234/osf.io/483zj","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For the Significance dimension, the authors forecast that 17 reviewers per proposal would be required to reach a relative reliability of 0.50.","vf":"unverified"},{"key":"NNIEFSDW","au":"Forsythe, Laura P","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson correlation","estd":"correlation","v":0.15,"n":"1312","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 (Exceptional) to 9 (Poor)","field":"health/comparative-effectiveness research","wr":"patients and stakeholders on patient-centredness","conf":"med","self":false,"doi":"10.1016/j.jval.2018.03.017","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Patient and stakeholder reviewers independently scored the same 1312 applications on patient-centredness before panel discussion. Their preliminary scores correlated at 0.15.","vf":"unverified"},{"key":"NNIEFSDW","au":"Forsythe, Laura P","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson correlation","estd":"correlation","v":0.22,"n":"1312","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 (Exceptional) to 9 (Poor)","field":"health/comparative-effectiveness research","wr":"patients and stakeholders on engagement","conf":"med","self":false,"doi":"10.1016/j.jval.2018.03.017","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Patient and stakeholder reviewers independently scored the same 1312 applications on patient and stakeholder engagement before panel discussion. Their preliminary scores correlated at 0.22.","vf":"unverified"},{"key":"NNIEFSDW","au":"Forsythe, Laura P","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson correlation","estd":"correlation","v":0.12,"n":"1312","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 (Exceptional) to 9 (Poor)","field":"health/comparative-effectiveness research","wr":"patients and stakeholders on improvement potential","conf":"med","self":false,"doi":"10.1016/j.jval.2018.03.017","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Patient and stakeholder reviewers independently scored the same 1312 applications on potential to improve health care and outcomes before panel discussion. Their preliminary scores correlated at 0.12, the lowest cross-type correlation.","vf":"unverified"},{"key":"NNIEFSDW","au":"Forsythe, Laura P","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson correlation","estd":"correlation","v":0.16,"n":"1312","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 (Exceptional) to 9 (Poor)","field":"health/comparative-effectiveness research","wr":"patients and stakeholders on preliminary overall scores","conf":"med","self":false,"doi":"10.1016/j.jval.2018.03.017","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Patient and stakeholder reviewers independently assigned preliminary overall scores to the same 1312 applications before panel discussion. Their scores correlated at 0.16.","vf":"unverified"},{"key":"NNIEFSDW","au":"Forsythe, Laura P","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson correlation","estd":"correlation","v":0.28,"n":"1312","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 (Exceptional) to 9 (Poor)","field":"health/comparative-effectiveness research","wr":"scientists and patients on patient-centredness","conf":"med","self":false,"doi":"10.1016/j.jval.2018.03.017","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Scientist and patient reviewers independently scored the same 1312 applications on patient-centredness before panel discussion. Their preliminary scores correlated at 0.28.","vf":"unverified"},{"key":"NNIEFSDW","au":"Forsythe, Laura P","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson correlation","estd":"correlation","v":0.33,"n":"1312","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 (Exceptional) to 9 (Poor)","field":"health/comparative-effectiveness research","wr":"scientists and patients on engagement","conf":"med","self":false,"doi":"10.1016/j.jval.2018.03.017","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Scientist and patient reviewers independently scored the same 1312 applications on patient and stakeholder engagement before panel discussion. Their preliminary scores correlated at 0.33.","vf":"unverified"},{"key":"NNIEFSDW","au":"Forsythe, Laura P","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson correlation","estd":"correlation","v":0.18,"n":"1312","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 (Exceptional) to 9 (Poor)","field":"health/comparative-effectiveness research","wr":"scientists and patients on improvement potential","conf":"med","self":false,"doi":"10.1016/j.jval.2018.03.017","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Scientist and patient reviewers independently scored the same 1312 applications on potential to improve health care and outcomes before panel discussion. Their preliminary scores correlated at 0.18.","vf":"unverified"},{"key":"NNIEFSDW","au":"Forsythe, Laura P","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson correlation","estd":"correlation","v":0.2,"n":"1312","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 (Exceptional) to 9 (Poor)","field":"health/comparative-effectiveness research","wr":"scientists and patients on preliminary overall scores","conf":"med","self":false,"doi":"10.1016/j.jval.2018.03.017","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Scientist and patient reviewers independently assigned preliminary overall scores to the same 1312 applications before panel discussion. Their scores correlated at 0.20.","vf":"unverified"},{"key":"NNIEFSDW","au":"Forsythe, Laura P","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson correlation","estd":"correlation","v":0.29,"n":"1312","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 (Exceptional) to 9 (Poor)","field":"health/comparative-effectiveness research","wr":"scientists and stakeholders on patient-centredness","conf":"med","self":false,"doi":"10.1016/j.jval.2018.03.017","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Scientist and stakeholder reviewers independently scored the same 1312 applications on patient-centredness before panel discussion. Their preliminary scores correlated at 0.29.","vf":"unverified"},{"key":"NNIEFSDW","au":"Forsythe, Laura P","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson correlation","estd":"correlation","v":0.36,"n":"1312","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 (Exceptional) to 9 (Poor)","field":"health/comparative-effectiveness research","wr":"scientists and stakeholders on engagement","conf":"med","self":false,"doi":"10.1016/j.jval.2018.03.017","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Scientist and stakeholder reviewers independently scored the same 1312 applications on patient and stakeholder engagement before panel discussion. Their preliminary scores correlated at 0.36, the highest cross-type correlation.","vf":"unverified"},{"key":"NNIEFSDW","au":"Forsythe, Laura P","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson correlation","estd":"correlation","v":0.21,"n":"1312","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 (Exceptional) to 9 (Poor)","field":"health/comparative-effectiveness research","wr":"scientists and stakeholders on improvement potential","conf":"med","self":false,"doi":"10.1016/j.jval.2018.03.017","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Scientist and stakeholder reviewers independently scored the same 1312 applications on potential to improve health care and outcomes before panel discussion. Their preliminary scores correlated at 0.21.","vf":"unverified"},{"key":"NNIEFSDW","au":"Forsythe, Laura P","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson correlation","estd":"correlation","v":0.21,"n":"1312","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 (Exceptional) to 9 (Poor)","field":"health/comparative-effectiveness research","wr":"scientists and stakeholders on preliminary overall scores","conf":"med","self":false,"doi":"10.1016/j.jval.2018.03.017","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Scientist and stakeholder reviewers independently assigned preliminary overall scores to the same 1312 applications before panel discussion. Their scores correlated at 0.21.","vf":"unverified"},{"key":"NNIEFSDW","au":"Forsythe, Laura P","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"proportion of scientists' top-20% applications also in the top 20% for both patients and stakeholders (exact top-20% membership)","estd":"percent agreement","v":0.57,"n":"123","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"top 20% vs not (from 1-9 final scores)","field":"health/comparative-effectiveness research","wr":"all three reviewer types on top-20% ranking","conf":"med","self":false,"doi":"10.1016/j.jval.2018.03.017","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":true,"he":false,"ms":"After in-person panel discussion, applications were ranked by each reviewer type's average final overall score. Of the 123 applications in the scientists' top 20%, 57% were also in the top 20% for both patients and stakeholders, the paper's headline post-discussion agreement result.","vf":"unverified"},{"key":"NNIEFSDW","au":"Forsythe, Laura P","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"conditional overlap: proportion of scientists' top-20% applications also in patients' top-20% (exact top-20% membership)","estd":"percent agreement","v":0.66,"n":"123","k":"2","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"top 20% vs not (from 1-9 final scores)","field":"health/comparative-effectiveness research","wr":"scientist and patient reviewer types on top-20% ranking","conf":"med","self":false,"doi":"10.1016/j.jval.2018.03.017","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"After in-person panel discussion, applications were ranked by each reviewer type's average final overall score. Of the 123 applications in the scientists' top 20%, 66% were also in the patients' top 20%.","vf":"unverified"},{"key":"NNIEFSDW","au":"Forsythe, Laura P","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"conditional overlap: proportion of scientists' top-20% applications also in stakeholders' top-20% (exact top-20% membership)","estd":"percent agreement","v":0.74,"n":"123","k":"2","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"top 20% vs not (from 1-9 final scores)","field":"health/comparative-effectiveness research","wr":"scientist and stakeholder reviewer types on top-20% ranking","conf":"med","self":false,"doi":"10.1016/j.jval.2018.03.017","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"After in-person panel discussion, applications were ranked by each reviewer type's average final overall score. Of the 123 applications in the scientists' top 20%, 74% were also in the stakeholders' top 20%.","vf":"unverified"},{"key":"NNIEFSDW","au":"Forsythe, Laura P","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson correlation between the two preliminary scientist reviewers, same criterion","estd":"correlation","v":0.4,"n":"1312","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 (Exceptional) to 9 (Poor)","field":"health/comparative-effectiveness research","wr":"two scientist reviewers on application criterion scores","conf":"med","self":false,"doi":"10.1016/j.jval.2018.03.017","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two scientist reviewers independently scored each of 1312 funding applications online before panel discussion. Their patient-and-stakeholder-engagement scores correlated at 0.40, the highest of the five per-criterion correlations.","vf":"unverified"},{"key":"NNIEFSDW","au":"Forsythe, Laura P","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson correlation between the two preliminary scientist reviewers, same criterion","estd":"correlation","v":0.19,"n":"1312","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 (Exceptional) to 9 (Poor)","field":"health/comparative-effectiveness research","wr":"two scientist reviewers on application criterion scores","conf":"med","self":false,"doi":"10.1016/j.jval.2018.03.017","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two scientist reviewers independently scored each of 1312 funding applications online before panel discussion. Their impact-of-condition scores correlated at 0.19, the lowest of the five per-criterion correlations.","vf":"unverified"},{"key":"WWP3Q9HN","au":"François, Olivier","y":2015,"cx":"Journal","ob":"conference-abstract","fam":"other","form":"arbitrariness: conditional probability for an accepted submission to get rejected if examined by a second committee","estd":"other","v":0.6,"n":"166","k":"2","samp":"re-reviewed-subset","blind":"double","agg":"panel-consensus","scale":"accept / reject","field":"computer science (machine learning)","wr":"two program committees' accept or reject decisions","conf":"high","self":false,"doi":"10.48550/arxiv.1507.06411","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"unclear","pr":false,"he":false,"ms":"Two independent program committees each decided on the same 166 NIPS 2014 submissions, and 60 per cent of the submissions accepted by one committee were rejected by the other. This observed disagreement rate is the NIPS organisers' reported figure that the paper reanalyses.","vf":"unverified"},{"key":"WWP3Q9HN","au":"François, Olivier","y":2015,"cx":"Journal","ob":"conference-abstract","fam":"other","form":"Bayesian estimate of arbitrariness from approximate Bayesian computation in the RAFC model with alpha = 5%","estd":"other","v":0.57,"n":"166","k":"2","samp":"re-reviewed-subset","blind":"double","agg":"panel-consensus","scale":"accept / reject","field":"computer science (machine learning)","wr":"two program committees' accept or reject decisions","conf":"high","self":false,"doi":"10.48550/arxiv.1507.06411","ciLow":0.41,"ciHigh":0.63,"mt":"inter-rater","tgt":"editor-decisions","rr":"unclear","pr":false,"he":false,"ms":"Under the alternative Reject, Accept or Flip a Coin model, which assumes 5 per cent of submissions are accepted by any committee, the same NIPS data gave an arbitrariness estimate of 57 per cent, with a 95 per cent credibility interval from 0.41 to 0.63.","vf":"unverified"},{"key":"WWP3Q9HN","au":"François, Olivier","y":2015,"cx":"Journal","ob":"conference-abstract","fam":"other","form":"Bayesian Monte Carlo estimate of the arbitrariness parameter, a = 1 - y, in the RFC model","estd":"other","v":0.61,"n":"166","k":"2","samp":"re-reviewed-subset","blind":"double","agg":"panel-consensus","scale":"accept / reject","field":"computer science (machine learning)","wr":"two program committees' accept or reject decisions","conf":"high","self":false,"doi":"10.48550/arxiv.1507.06411","ciLow":0.43,"ciHigh":0.73,"mt":"inter-rater","tgt":"editor-decisions","rr":"unclear","pr":true,"he":false,"ms":"Fitting the Reject or Flip a Coin model to the NIPS experiment data for 166 twice reviewed submissions gave a posterior estimate of arbitrariness of 61 per cent, with a 95 per cent credibility interval from 0.43 to 0.73. This is the paper's own headline estimate of how often two independent committees would reach opposite decisions.","vf":"unverified"},{"key":"T67A4S8E","au":"Galbraith, Craig S.","y":2010,"cx":"Grant","ob":"grant-proposal","fam":"Kendall-W","form":"Kendall's W (coefficient of concordance)","estd":"Kendall W","v":0.55,"n":"5","k":"6","samp":"full-pool","blind":"open","agg":"unspecified","scale":"11-point Likert per item, six items summed (0-66)","field":"technology commercialisation (multi-field)","wr":"panel members' scores of technology grant proposals","conf":"high","self":false,"doi":"10.1007/s10961-009-9122-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"The same six panel members re-scored the five proposals after presentation, Q&A and discussion; Kendall's W rose to 0.550, indicating higher concordance among their informed rankings.","vf":"unverified"},{"key":"T67A4S8E","au":"Galbraith, Craig S.","y":2010,"cx":"Grant","ob":"grant-proposal","fam":"Kendall-W","form":"Kendall's W (coefficient of concordance)","estd":"Kendall W","v":0.287,"n":"5","k":"6","samp":"full-pool","blind":"open","agg":"unspecified","scale":"11-point Likert per item, six items summed (0-66)","field":"technology commercialisation (multi-field)","wr":"panel members' scores of technology grant proposals","conf":"high","self":false,"doi":"10.1007/s10961-009-9122-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":true,"ms":"Six panel members independently scored five technology grant proposals before presentation and discussion; a Kendall's W of 0.287 indicates low concordance among their rankings.","vf":"unverified"},{"key":"T67A4S8E","au":"Galbraith, Craig S.","y":2010,"cx":"Grant","ob":"grant-proposal","fam":"Kendall-W","form":"Kendall's W (coefficient of concordance)","estd":"Kendall W","v":0.615,"n":"4","k":"4","samp":"full-pool","blind":"open","agg":"unspecified","scale":"11-point Likert per item, six items summed (0-66)","field":"technology commercialisation (multi-field)","wr":"panel members' scores of technology grant proposals","conf":"high","self":false,"doi":"10.1007/s10961-009-9122-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"The same four panel members re-scored the four proposals after presentation, Q&A and discussion; Kendall's W rose to 0.615, indicating substantially higher concordance among their informed rankings.","vf":"unverified"},{"key":"T67A4S8E","au":"Galbraith, Craig S.","y":2010,"cx":"Grant","ob":"grant-proposal","fam":"Kendall-W","form":"Kendall's W (coefficient of concordance)","estd":"Kendall W","v":0.186,"n":"4","k":"4","samp":"full-pool","blind":"open","agg":"unspecified","scale":"11-point Likert per item, six items summed (0-66)","field":"technology commercialisation (multi-field)","wr":"panel members' scores of technology grant proposals","conf":"high","self":false,"doi":"10.1007/s10961-009-9122-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Four panel members independently scored four electronics technology proposals before discussion; a Kendall's W of 0.186 indicates low concordance among their rankings.","vf":"unverified"},{"key":"T67A4S8E","au":"Galbraith, Craig S.","y":2010,"cx":"Grant","ob":"grant-proposal","fam":"Kendall-W","form":"Kendall's W (coefficient of concordance)","estd":"Kendall W","v":0.453,"n":"3","k":"5","samp":"full-pool","blind":"open","agg":"unspecified","scale":"11-point Likert per item, six items summed (0-66)","field":"technology commercialisation (multi-field)","wr":"panel members' scores of technology grant proposals","conf":"high","self":false,"doi":"10.1007/s10961-009-9122-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"The same five panel members re-scored the three proposals after presentation, Q&A and discussion; Kendall's W rose to 0.453, indicating higher concordance among their informed rankings.","vf":"unverified"},{"key":"T67A4S8E","au":"Galbraith, Craig S.","y":2010,"cx":"Grant","ob":"grant-proposal","fam":"Kendall-W","form":"Kendall's W (coefficient of concordance)","estd":"Kendall W","v":0.044,"n":"3","k":"5","samp":"full-pool","blind":"open","agg":"unspecified","scale":"11-point Likert per item, six items summed (0-66)","field":"technology commercialisation (multi-field)","wr":"panel members' scores of technology grant proposals","conf":"high","self":false,"doi":"10.1007/s10961-009-9122-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Five panel members independently scored three electronics technology proposals before discussion; a Kendall's W of 0.044 indicates essentially no concordance among their rankings.","vf":"unverified"},{"key":"T67A4S8E","au":"Galbraith, Craig S.","y":2010,"cx":"Grant","ob":"grant-proposal","fam":"Kendall-W","form":"Kendall's W (coefficient of concordance)","estd":"Kendall W","v":0.217,"n":"4","k":"4","samp":"full-pool","blind":"open","agg":"unspecified","scale":"11-point Likert per item, six items summed (0-66)","field":"technology commercialisation (multi-field)","wr":"panel members' scores of technology grant proposals","conf":"high","self":false,"doi":"10.1007/s10961-009-9122-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"The same four panel members re-scored the four proposals after presentation, Q&A and discussion; Kendall's W rose to 0.217, still indicating weak concordance among their informed rankings.","vf":"unverified"},{"key":"T67A4S8E","au":"Galbraith, Craig S.","y":2010,"cx":"Grant","ob":"grant-proposal","fam":"Kendall-W","form":"Kendall's W (coefficient of concordance)","estd":"Kendall W","v":0.075,"n":"4","k":"4","samp":"full-pool","blind":"open","agg":"unspecified","scale":"11-point Likert per item, six items summed (0-66)","field":"technology commercialisation (multi-field)","wr":"panel members' scores of technology grant proposals","conf":"high","self":false,"doi":"10.1007/s10961-009-9122-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Four panel members independently scored four software technology proposals before discussion; a Kendall's W of 0.075 indicates almost no concordance among their rankings.","vf":"unverified"},{"key":"8JXS4WJA","au":"Gallo, Stephen A","y":2013,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraclass correlation described by Cicchetti, R i Model III, single-rater (intra-application correlation), based on average 10 reviews per application","estd":"ICC (single/unspec)","v":0.87,"n":"","k":"10","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1.0 to 5.0, 1 highest and 5 lowest","field":"biomedical","wr":"panel members on grant application merit scores","conf":"high","self":false,"doi":"10.1371/journal.pone.0071693","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"Across 2009 to 2012, panels averaging ten reviewers scored each grant application after discussion; the highest annual single-rater ICC was 0.87, indicating high agreement between individual reviewers on a given application.","vf":"unverified"},{"key":"8JXS4WJA","au":"Gallo, Stephen A","y":2013,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraclass correlation described by Cicchetti, R i Model III, single-rater (intra-application correlation), based on average 10 reviews per application","estd":"ICC (single/unspec)","v":0.84,"n":"","k":"10","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1.0 to 5.0, 1 highest and 5 lowest","field":"biomedical","wr":"panel members on grant application merit scores","conf":"high","self":false,"doi":"10.1371/journal.pone.0071693","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"Across 2009 to 2012, panels averaging ten reviewers scored each grant application after discussion; the lowest annual single-rater ICC was 0.84, indicating high agreement between individual reviewers on a given application.","vf":"unverified"},{"key":"8JXS4WJA","au":"Gallo, Stephen A","y":2013,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"reliability of the mean application rating (IRR), estimated from the ICC using the Spearman-Brown formula","estd":"ICC (average)","v":0.98,"n":"1600","k":"10","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1.0 to 5.0, 1 highest and 5 lowest","field":"biomedical","wr":"panel members on grant application merit scores","conf":"high","self":false,"doi":"10.1371/journal.pone.0071693","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":false,"ms":"Approximately 1600 grant applications from 2009 to 2012 were each scored by an average of ten panel members after discussion. The reliability of the mean application rating, derived from the single-rater ICC via the Spearman-Brown formula, was 0.98 in every year, indicating very high reliability of the averaged panel score.","vf":"unverified"},{"key":"9EBX8PWV","au":"Gallo, Stephen A","y":2016,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraproposal correlation coefficient rho = sigma2_mu / (sigma2_mu + sigma2_eps), the correlation between two ratings of the same application","estd":"ICC (single/unspec)","v":0.23,"n":"725","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1.0-5.0, 1 = highest merit, 5 = lowest","field":"biomedical","wr":"single reviewer on grant proposal scientific-merit scores","conf":"high","self":false,"doi":"10.1371/journal.pone.0165147","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The correlation between two reviewers' scientific-merit scores of the same application, 0.23, reflects the reliability of a single reviewer's judgment and indicates poor agreement.","vf":"unverified"},{"key":"9EBX8PWV","au":"Gallo, Stephen A","y":2016,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"inter-rater reliability (IRR), the reliability of the average rating, calculated using rho and the Spearman-Brown formula","estd":"ICC (average)","v":0.37,"n":"725","k":"2","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1.0-5.0, 1 = highest merit, 5 = lowest","field":"biomedical","wr":"two reviewers on grant proposal scientific-merit scores","conf":"high","self":false,"doi":"10.1371/journal.pone.0165147","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Two independent reviewers scored the scientific merit of 725 biomedical funding applications on a 1.0-5.0 scale; the reliability of the two-reviewer average, 0.37, indicates poor inter-rater reliability.","vf":"unverified"},{"key":"9EBX8PWV","au":"Gallo, Stephen A","y":2016,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"proportion of applications where both reviewers agreed on fundability status (exact agreement on the SM < 2.0 binary)","estd":"percent agreement","v":0.33,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"fundable / unfundable (SM < 2.0 threshold)","field":"biomedical","wr":"two reviewers agreeing on fundability of top-15% applications","conf":"high","self":false,"doi":"10.1371/journal.pone.0165147","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"For higher-expertise reviewer pairs, the two reviewers agreed on the hypothetical fundability status of only 33% of top-15% (fundable) applications, showing low agreement on the best proposals.","vf":"unverified"},{"key":"9EBX8PWV","au":"Gallo, Stephen A","y":2016,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"proportion of applications where both reviewers agreed on fundability status (exact agreement on the SM < 2.0 binary)","estd":"percent agreement","v":0.82,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"fundable / unfundable (SM < 2.0 threshold)","field":"biomedical","wr":"two reviewers agreeing on fundability of bottom-85% applications","conf":"high","self":false,"doi":"10.1371/journal.pone.0165147","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For higher-expertise reviewer pairs, the two reviewers agreed on the hypothetical fundability status of 82% of bottom-85% (unfundable) applications, showing much higher agreement on weaker proposals.","vf":"unverified"},{"key":"9EBX8PWV","au":"Gallo, Stephen A","y":2016,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"proportion of applications where both reviewers agreed on fundability status (exact agreement on the SM < 2.0 binary)","estd":"percent agreement","v":0.35,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"fundable / unfundable (SM < 2.0 threshold)","field":"biomedical","wr":"two reviewers agreeing on fundability of top-15% applications","conf":"high","self":false,"doi":"10.1371/journal.pone.0165147","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"For lower-expertise reviewer pairs, the two reviewers agreed on the hypothetical fundability status of 35% of top-15% (fundable) applications, similarly low agreement on the best proposals.","vf":"unverified"},{"key":"9EBX8PWV","au":"Gallo, Stephen A","y":2016,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"proportion of applications where both reviewers agreed on fundability status (exact agreement on the SM < 2.0 binary)","estd":"percent agreement","v":0.81,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"fundable / unfundable (SM < 2.0 threshold)","field":"biomedical","wr":"two reviewers agreeing on fundability of bottom-85% applications","conf":"high","self":false,"doi":"10.1371/journal.pone.0165147","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For lower-expertise reviewer pairs, the two reviewers agreed on the hypothetical fundability status of 81% of bottom-85% (unfundable) applications, again much higher agreement on weaker proposals.","vf":"unverified"},{"key":"9EBX8PWV","au":"Gallo, Stephen A","y":2016,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"Intra-proposal Correlation (correlation between two ratings of one application) from the covariate-adjusted random-intercept model, Table 4","estd":"ICC (single/unspec)","v":0.21,"n":"725","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1.0-5.0, 1 = highest merit, 5 = lowest","field":"biomedical","wr":"single reviewer on grant proposal scientific-merit scores (adjusted)","conf":"high","self":false,"doi":"10.1371/journal.pone.0165147","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"In the covariate-adjusted model, the correlation between two reviewers' scores of the same application is 0.21, the single-reviewer reliability, again indicating poor agreement.","vf":"unverified"},{"key":"9EBX8PWV","au":"Gallo, Stephen A","y":2016,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"Inter-Rater Reliability (reliability of the average rating) from the covariate-adjusted random-intercept model, Table 4","estd":"ICC (average)","v":0.35,"n":"725","k":"2","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1.0-5.0, 1 = highest merit, 5 = lowest","field":"biomedical","wr":"two reviewers on grant proposal scientific-merit scores (adjusted)","conf":"high","self":false,"doi":"10.1371/journal.pone.0165147","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"In the covariate-adjusted model controlling for expertise, seniority and sector, the reliability of the two-reviewer average score is 0.35, still indicating poor inter-rater reliability.","vf":"unverified"},{"key":"9EBX8PWV","au":"Gallo, Stephen A","y":2016,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"average scoring differences (absolute differences) in SM score between the two assigned reviewers","estd":"AD index","v":0.66,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1.0-5.0, 1 = highest merit, 5 = lowest","field":"biomedical","wr":"mean absolute score difference between two reviewers, fundable apps","conf":"high","self":false,"doi":"10.1371/journal.pone.0165147","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"For higher-expertise reviewer pairs, the two reviewers' scientific-merit scores differed on average by 0.66 points on the 1.0-5.0 scale for top-15% (fundable) applications.","vf":"unverified"},{"key":"9EBX8PWV","au":"Gallo, Stephen A","y":2016,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"average scoring differences (absolute differences) in SM score between the two assigned reviewers","estd":"AD index","v":1.09,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1.0-5.0, 1 = highest merit, 5 = lowest","field":"biomedical","wr":"mean absolute score difference between two reviewers, unfundable apps","conf":"high","self":false,"doi":"10.1371/journal.pone.0165147","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For higher-expertise reviewer pairs, the two reviewers' scores differed on average by 1.09 points on the 1.0-5.0 scale for bottom-85% (unfundable) applications, a larger gap than for fundable ones.","vf":"unverified"},{"key":"9EBX8PWV","au":"Gallo, Stephen A","y":2016,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"average scoring differences (absolute differences) in SM score between the two assigned reviewers","estd":"AD index","v":0.57,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1.0-5.0, 1 = highest merit, 5 = lowest","field":"biomedical","wr":"mean absolute score difference between two reviewers, fundable apps","conf":"high","self":false,"doi":"10.1371/journal.pone.0165147","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"For lower-expertise reviewer pairs, the two reviewers' scientific-merit scores differed on average by 0.57 points on the 1.0-5.0 scale for top-15% (fundable) applications.","vf":"unverified"},{"key":"9EBX8PWV","au":"Gallo, Stephen A","y":2016,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"average scoring differences (absolute differences) in SM score between the two assigned reviewers","estd":"AD index","v":0.95,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1.0-5.0, 1 = highest merit, 5 = lowest","field":"biomedical","wr":"mean absolute score difference between two reviewers, unfundable apps","conf":"high","self":false,"doi":"10.1371/journal.pone.0165147","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For lower-expertise reviewer pairs, the two reviewers' scores differed on average by 0.95 points on the 1.0-5.0 scale for bottom-85% (unfundable) applications.","vf":"unverified"},{"key":"N99QJWN2","au":"Gallo, Stephen A.","y":2022,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC derived from the variance estimate of the final mixed ordinal regression model; form unspecified","estd":"ICC (single/unspec)","v":0.15,"n":"4","k":"","samp":"special","blind":"single","agg":"unspecified","scale":"1-9 whole numbers, 1 = exceptional, 9 = poor","field":"biomedical","wr":"NIH reviewers scoring mock proposal impact statements","conf":"med","self":false,"doi":"10.1371/journal.pone.0273813","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":true,"ms":"Six hundred and five experienced NIH grant reviewers each scored constructed mock impact statements on the 1 to 9 NIH scale; an intra-class coefficient of 0.15, derived from a mixed regression model, indicates low agreement between reviewers. The rated set was four distinct OISs (one control plus three manipulated), each reviewer rating two.","vf":"unverified"},{"key":"N99QJWN2","au":"Gallo, Stephen A.","y":2022,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"directional agreement: proportion of reviewers scoring the manipulated OIS worse than the control OIS","estd":"percent agreement","v":0.79,"n":"2","k":"204","samp":"special","blind":"single","agg":"unspecified","scale":"1-9 whole numbers, 1 = exceptional, 9 = poor","field":"biomedical","wr":"reviewers judging mock OIS worse than control (approach risk)","conf":"med","self":false,"doi":"10.1371/journal.pone.0273813","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The 204 reviewers in the approach-risk condition each independently scored the control and the manipulated impact statement; 79 per cent scored the manipulated statement worse than the control, showing majority agreement on the direction of the risk penalty. This is a directional agreement proportion, not point agreement.","vf":"unverified"},{"key":"N99QJWN2","au":"Gallo, Stephen A.","y":2022,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"directional agreement: proportion of reviewers scoring the manipulated OIS worse than the control OIS","estd":"percent agreement","v":0.88,"n":"2","k":"202","samp":"special","blind":"single","agg":"unspecified","scale":"1-9 whole numbers, 1 = exceptional, 9 = poor","field":"biomedical","wr":"reviewers judging mock OIS worse than control (PI-approach risk)","conf":"med","self":false,"doi":"10.1371/journal.pone.0273813","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The 202 reviewers in the combined PI-and-approach-risk condition each independently scored the control and the manipulated impact statement; 88 per cent scored the manipulated statement worse than the control, the strongest directional agreement across the three conditions. This is a directional agreement proportion, not point agreement.","vf":"unverified"},{"key":"N99QJWN2","au":"Gallo, Stephen A.","y":2022,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"directional agreement: proportion of reviewers scoring the manipulated OIS worse than the control OIS","estd":"percent agreement","v":0.65,"n":"2","k":"199","samp":"special","blind":"single","agg":"unspecified","scale":"1-9 whole numbers, 1 = exceptional, 9 = poor","field":"biomedical","wr":"reviewers judging mock OIS worse than control (PI risk)","conf":"med","self":false,"doi":"10.1371/journal.pone.0273813","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The 199 reviewers in the PI-risk condition each independently scored the control and the manipulated impact statement; 65 per cent scored the manipulated statement worse than the control, showing majority but not unanimous agreement on the direction of the risk penalty. This is a directional agreement proportion, not point agreement.","vf":"unverified"},{"key":"WJIKJ2M5","au":"Garfunkel, Joseph M.","y":1990,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"overall proportion of agreement (sum of proportions of manuscripts on which both reviewers agreed); exact agreement on dichotomised recommendation","estd":"percent agreement","v":0.84,"n":"25","k":"2","samp":"funded-only","blind":"unclear","agg":"single-rater","scale":"accept vs reject or submit to another journal","field":"biomedical (paediatrics)","wr":"two new referees' publication recommendations on manuscripts","conf":"high","self":false,"doi":"10.1001/jama.1990.03440100077011","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":true,"he":true,"ms":"Twenty-five manuscripts already accepted after revision were each sent to two new referees who did not know the manuscripts were accepted. With each recommendation reduced to accept versus reject or send elsewhere, the two referees gave the same verdict for 0.84 of the manuscripts, an uncorrected proportion of agreement.","vf":"unverified"},{"key":"WJIKJ2M5","au":"Garfunkel, Joseph M.","y":1990,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"overall proportion of agreement (sum of proportions of manuscripts on which both reviewers agreed); exact agreement on dichotomised recommendation","estd":"percent agreement","v":0.84,"n":"25","k":"2","samp":"funded-only","blind":"unclear","agg":"single-rater","scale":"accept without revision vs with revision","field":"biomedical (paediatrics)","wr":"two new referees' revision recommendations on manuscripts","conf":"high","self":false,"doi":"10.1001/jama.1990.03440100077011","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"For the same 25 already-accepted manuscripts rereviewed by two new referees each, recommendations were reduced to publish without revision versus with revision. The two referees agreed on this distinction for 0.84 of the manuscripts, again an uncorrected proportion of agreement.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"Cronbach-alpha","form":"Cronbach's alpha","estd":"Cronbach alpha","v":0.85,"n":"380","k":"1","samp":"funded-only","blind":"single","agg":"unspecified","scale":"7-point, no impact to great impact","field":"psychology","wr":"internal consistency of 2-item impact scale","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"The two-item overall impact scale had a Cronbach's alpha of 0.85 across 380 rated articles, indicating high internal consistency of the impact instrument.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass coefficient (averaged judgments)","estd":"ICC (average)","v":0.52,"n":"","k":"2","samp":"funded-only","blind":"single","agg":"average-of-k","scale":"7-point, no impact to great impact","field":"psychology","wr":"averaged expert judgment of article impact","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"The reliability of impact judgments averaged across two experts was 0.52, higher than the single-judge coefficient because averaging pools two independent ratings.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass coefficient","estd":"ICC (single/unspec)","v":0.35,"n":"122","k":"2","samp":"funded-only","blind":"single","agg":"single-rater","scale":"7-point, no impact to great impact","field":"psychology","wr":"experts on published-article impact","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Agreement across two expert judges on the overall impact scale was an intraclass coefficient of 0.35, indicating modest interjudge reliability of a single expert's impact rating.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"Cronbach-alpha","form":"Cronbach's alpha","estd":"Cronbach alpha","v":0.92,"n":"382","k":"1","samp":"funded-only","blind":"single","agg":"unspecified","scale":"7-point, clearly inferior to clearly superior","field":"psychology","wr":"internal consistency of 3-item quality scale","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"The three-item overall quality scale had a Cronbach's alpha of 0.92 across 382 rated articles, indicating high internal consistency of the quality instrument.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass coefficient (averaged judgments)","estd":"ICC (average)","v":0.58,"n":"","k":"2","samp":"funded-only","blind":"single","agg":"average-of-k","scale":"7-point, clearly inferior to clearly superior","field":"psychology","wr":"averaged expert judgment of article quality","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"The reliability of quality judgments averaged across two experts was 0.58, higher than the single-judge coefficient because averaging pools two independent ratings.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass coefficient","estd":"ICC (single/unspec)","v":0.41,"n":"121","k":"2","samp":"funded-only","blind":"single","agg":"single-rater","scale":"7-point, clearly inferior to clearly superior","field":"psychology","wr":"experts on published-article quality","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":true,"he":true,"ms":"For articles with at least two expert judges, agreement across judges on the overall quality scale was an intraclass coefficient of 0.41, indicating modest interjudge reliability of a single expert's quality rating.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"Cronbach-alpha","form":"Cronbach's alpha","estd":"Cronbach alpha","v":0.78,"n":"331","k":"1","samp":"funded-only","blind":"single","agg":"unspecified","scale":"6-point, strongly disagree to strongly agree","field":"psychology","wr":"internal consistency of 5-item Don't's scale","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"The five-item Don't's evaluative scale had a Cronbach's alpha of 0.78 across 331 articles, an acceptable internal consistency.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass coefficient","estd":"ICC (single/unspec)","v":0.16,"n":"92","k":"2","samp":"funded-only","blind":"single","agg":"single-rater","scale":"6-point, strongly disagree to strongly agree","field":"psychology","wr":"experts on Don't's evaluative scale","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Agreement across two experts on the Don't's evaluative scale was an intraclass coefficient of 0.16, low interjudge reliability attributed to lack of variance in this subsample.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"Cronbach-alpha","form":"Cronbach's alpha","estd":"Cronbach alpha","v":0.58,"n":"335","k":"1","samp":"funded-only","blind":"single","agg":"unspecified","scale":"6-point, strongly disagree to strongly agree","field":"psychology","wr":"internal consistency of 4-item Substantive do's scale","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"The four-item Substantive do's evaluative scale had a Cronbach's alpha of 0.58 across 335 articles, a modest internal consistency.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass coefficient","estd":"ICC (single/unspec)","v":0.5,"n":"95","k":"2","samp":"funded-only","blind":"single","agg":"single-rater","scale":"6-point, strongly disagree to strongly agree","field":"psychology","wr":"experts on Substantive do's evaluative scale","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Agreement across two experts on the Substantive do's evaluative scale was an intraclass coefficient of 0.50, the highest interjudge reliability among the individual scales.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"Cronbach-alpha","form":"Cronbach's alpha","estd":"Cronbach alpha","v":0.74,"n":"335","k":"1","samp":"funded-only","blind":"single","agg":"unspecified","scale":"6-point, strongly disagree to strongly agree","field":"psychology","wr":"internal consistency of 5-item Stylistic do's scale","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"The five-item Stylistic/compositional do's evaluative scale had a Cronbach's alpha of 0.74 across 335 articles, an acceptable internal consistency.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass coefficient","estd":"ICC (single/unspec)","v":0.2,"n":"99","k":"2","samp":"funded-only","blind":"single","agg":"single-rater","scale":"6-point, strongly disagree to strongly agree","field":"psychology","wr":"experts on Stylistic/compositional do's scale","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Agreement across two experts on the Stylistic/compositional do's evaluative scale was an intraclass coefficient of 0.20, low interjudge reliability attributed to lack of variance.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"Cronbach-alpha","form":"Cronbach's alpha","estd":"Cronbach alpha","v":0.86,"n":"340","k":"1","samp":"funded-only","blind":"single","agg":"unspecified","scale":"6-point, strongly disagree to strongly agree","field":"psychology","wr":"internal consistency of 4-item Originality scale","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"The four-item Originality/heurism evaluative scale had a Cronbach's alpha of 0.86 across 340 articles, a high internal consistency.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass coefficient","estd":"ICC (single/unspec)","v":0.37,"n":"97","k":"2","samp":"funded-only","blind":"single","agg":"single-rater","scale":"6-point, strongly disagree to strongly agree","field":"psychology","wr":"experts on Originality/heurism evaluative scale","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Agreement across two experts on the Originality/heurism evaluative scale was an intraclass coefficient of 0.37, a moderate interjudge reliability.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"Cronbach-alpha","form":"Cronbach's alpha","estd":"Cronbach alpha","v":0.89,"n":"339","k":"1","samp":"funded-only","blind":"single","agg":"unspecified","scale":"6-point, strongly disagree to strongly agree","field":"psychology","wr":"internal consistency of 4-item Trivia scale","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"The four-item Trivia evaluative scale had a Cronbach's alpha of 0.89 across 339 articles, a high internal consistency.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass coefficient","estd":"ICC (single/unspec)","v":0.4,"n":"97","k":"2","samp":"funded-only","blind":"single","agg":"single-rater","scale":"6-point, strongly disagree to strongly agree","field":"psychology","wr":"experts on Trivia evaluative scale","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Agreement across two experts on the Trivia evaluative scale was an intraclass coefficient of 0.40, a moderate interjudge reliability.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"Cronbach-alpha","form":"Cronbach's alpha","estd":"Cronbach alpha","v":0.64,"n":"321","k":"1","samp":"funded-only","blind":"single","agg":"unspecified","scale":"6-point, strongly disagree to strongly agree","field":"psychology","wr":"internal consistency of 4-item Where do we go scale","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"The four-item Where do we go from here? evaluative scale had a Cronbach's alpha of 0.64 across 321 articles, a moderate internal consistency.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass coefficient","estd":"ICC (single/unspec)","v":0.45,"n":"83","k":"2","samp":"funded-only","blind":"single","agg":"single-rater","scale":"6-point, strongly disagree to strongly agree","field":"psychology","wr":"experts on Where do we go from here? scale","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Agreement across two experts on the Where do we go from here? evaluative scale was an intraclass coefficient of 0.45, a moderate interjudge reliability.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"Cronbach-alpha","form":"Cronbach's alpha","estd":"Cronbach alpha","v":0.7,"n":"339","k":"1","samp":"funded-only","blind":"single","agg":"unspecified","scale":"6-point, strongly disagree to strongly agree","field":"psychology","wr":"internal consistency of 3-item Data grinders scale","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"The three-item Data grinders evaluative scale had a Cronbach's alpha of 0.70 across 339 articles, an acceptable internal consistency.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass coefficient","estd":"ICC (single/unspec)","v":0.49,"n":"95","k":"2","samp":"funded-only","blind":"single","agg":"single-rater","scale":"6-point, strongly disagree to strongly agree","field":"psychology","wr":"experts on Data grinders evaluative scale","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Agreement across two experts on the Data grinders evaluative scale was an intraclass coefficient of 0.49, among the higher interjudge reliabilities of the individual scales.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"Cronbach-alpha","form":"Cronbach's alpha","estd":"Cronbach alpha","v":0.1,"n":"320","k":"1","samp":"funded-only","blind":"single","agg":"unspecified","scale":"6-point, strongly disagree to strongly agree","field":"psychology","wr":"internal consistency of 3-item Ho-hum research scale","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"The three-item Ho-hum research evaluative scale had a Cronbach's alpha of 0.10 across 320 articles, a very low internal consistency the paper attributes to the scale's unreliability.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass coefficient","estd":"ICC (single/unspec)","v":0.22,"n":"83","k":"2","samp":"funded-only","blind":"single","agg":"single-rater","scale":"6-point, strongly disagree to strongly agree","field":"psychology","wr":"experts on Ho-hum research evaluative scale","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Agreement across two experts on the Ho-hum research evaluative scale was an intraclass coefficient of 0.22, low interjudge reliability consistent with this scale's unreliability.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"Cronbach-alpha","form":"Cronbach's alpha","estd":"Cronbach alpha","v":0.13,"n":"319","k":"1","samp":"funded-only","blind":"single","agg":"unspecified","scale":"6-point, strongly disagree to strongly agree","field":"psychology","wr":"internal consistency of 4-item Magnitude scale","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"The four-item Magnitude of problem/interest evaluative scale had a Cronbach's alpha of 0.13 across 319 articles, a very low internal consistency the paper attributes to the scale's unreliability.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass coefficient","estd":"ICC (single/unspec)","v":0.19,"n":"91","k":"2","samp":"funded-only","blind":"single","agg":"single-rater","scale":"6-point, strongly disagree to strongly agree","field":"psychology","wr":"experts on Magnitude of problem/interest scale","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Agreement across two experts on the Magnitude of problem/interest evaluative scale was an intraclass coefficient of 0.19, low interjudge reliability consistent with this scale's unreliability.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"Cronbach-alpha","form":"Cronbach's alpha","estd":"Cronbach alpha","v":0.89,"n":"340","k":"1","samp":"funded-only","blind":"single","agg":"unspecified","scale":"6-point, strongly disagree to strongly agree","field":"psychology","wr":"internal consistency of seven reliable combined scales","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"The composite of the seven reliable evaluative scales (20 to 29 items) had a Cronbach's alpha of 0.89 across 340 articles, a high internal consistency.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass coefficient","estd":"ICC (single/unspec)","v":0.49,"n":"96","k":"2","samp":"funded-only","blind":"single","agg":"single-rater","scale":"6-point, strongly disagree to strongly agree","field":"psychology","wr":"experts on seven reliable combined evaluative scales","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Agreement across two experts on the composite of the seven reliable evaluative scales was an intraclass coefficient of 0.49, slightly higher than for all nine scales.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"Cronbach-alpha","form":"Cronbach's alpha","estd":"Cronbach alpha","v":0.86,"n":"340","k":"1","samp":"funded-only","blind":"single","agg":"unspecified","scale":"6-point, strongly disagree to strongly agree","field":"psychology","wr":"internal consistency of all nine combined scales","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"The composite of all nine evaluative scales (25 to 36 items) had a Cronbach's alpha of 0.86 across 340 articles, a high internal consistency.","vf":"unverified"},{"key":"HH53MPTF","au":"Gottfredson, Stephen D.","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass coefficient","estd":"ICC (single/unspec)","v":0.46,"n":"96","k":"2","samp":"funded-only","blind":"single","agg":"single-rater","scale":"6-point, strongly disagree to strongly agree","field":"psychology","wr":"experts on all nine combined evaluative scales","conf":"med","self":false,"doi":"10.1037/0003-066x.33.10.920","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Agreement across two experts on the composite of all nine evaluative scales was an intraclass coefficient of 0.46, a moderate interjudge reliability.","vf":"unverified"},{"key":"AKUVGJ36","au":"Graham, Chris L. B.","y":2022,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"median correlation between all pairs of reviews within a project (Spearman per Methods, 'Pearson' per Results/Discussion)","estd":"correlation","v":0.28,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-5","field":"multi-field","wr":"community reviewers on grant/project proposal scores","conf":"med","self":false,"doi":"10.12688/f1000research.125886.1","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Community reviewers, a mix of applicants and non-applicants with a median of four per project, scored micro-grant project proposals on 1 to 5 criteria via online forms. The median pairwise correlation between reviews of the same project was 0.28, weak agreement the authors describe as in line with traditional funding schemes; the pairwise design gives k = 2.","vf":"unverified"},{"key":"AKUVGJ36","au":"Graham, Chris L. B.","y":2022,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Spearman correlation between full project ranking and ranking with one review per project (bootstrap, 50 iterations)","estd":"correlation","v":0.75,"n":"","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"projects ranked by decreasing average review score","field":"multi-field","wr":"stability of project ranking to review removal","conf":"med","self":false,"doi":"10.12688/f1000research.125886.1","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":true,"he":false,"ms":"A bootstrap analysis of the paper's own empirical review data repeatedly removed reviews and recomputed each project's mean score and rank over 50 iterations. The Spearman correlation of 0.75 between the full ranking and the ranking based on only one review per project shows the aggregate ranking is largely preserved; the Discussion reports this as an average correlation of 0.7.","vf":"unverified"},{"key":"LJCKBTRE","au":"Gupta, Piyush","y":2013,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"kappa for inter-rater agreement","estd":"kappa","v":0.35,"n":"431","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"‘accept’, ‘resubmit with revision’ and ‘reject’","field":"paediatrics","wr":"reviewers on manuscript accept/revise/reject recommendations","conf":"med","self":false,"doi":"10.1007/s13312-013-0001-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":true,"ms":"External peer reviewers independently recommended acceptance, revision or rejection for the 431 manuscripts that had at least two reviewers. A pairwise kappa of 0.35 indicates fair chance-corrected agreement between any two sets of reviewers.","vf":"unverified"},{"key":"LJCKBTRE","au":"Gupta, Piyush","y":2013,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"kappa for inter-rater agreement","estd":"kappa","v":0.21,"n":"203","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"‘accept’, ‘resubmit with revision’ and ‘reject’","field":"paediatrics","wr":"reviewers on manuscript accept/revise/reject recommendations","conf":"med","self":false,"doi":"10.1007/s13312-013-0001-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For the 203 manuscripts sent to three or more reviewers, agreement between any two sets of reviewers on accept, revise or reject was a pairwise kappa of 0.21, indicating slight chance-corrected agreement.","vf":"unverified"},{"key":"LJCKBTRE","au":"Gupta, Piyush","y":2013,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"kappa for inter-rater agreement","estd":"kappa","v":0.17,"n":"228","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"‘accept’, ‘resubmit with revision’ and ‘reject’","field":"paediatrics","wr":"reviewers on manuscript accept/revise/reject recommendations","conf":"med","self":false,"doi":"10.1007/s13312-013-0001-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Two external reviewers independently recommended acceptance, revision or rejection for each of the 228 manuscripts sent to exactly two reviewers. Their pairwise kappa was 0.17, indicating slight chance-corrected agreement.","vf":"unverified"},{"key":"BCKFAVIX","au":"Hargens, Lowell L.","y":1988,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"gamma coefficient (Goodman-Kruskal ordinal association)","estd":"other","v":0.89,"n":"342","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"referee accept/split/reject vs editor accept/revise/reject","field":"sociology","wr":"referee recommendations vs editor initial dispositions on manuscripts","conf":"high","self":false,"doi":"10.2307/2095739","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":true,"he":false,"ms":"A Goodman-Kruskal gamma of 0.89 shows strong ordinal association between the initial referees' recommendations and the editor's initial disposition of the same 342 ASR manuscripts, indicating editors largely followed the referees' advice.","vf":"unverified"},{"key":"BCKFAVIX","au":"Hargens, Lowell L.","y":1988,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"Cohen's kappa, a PRE measure of rater agreement (Cohen 1960)","estd":"kappa","v":0.15,"n":"342","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"accept or accept conditional on minor revisions; otherwise reject","field":"sociology","wr":"referees' accept/reject recommendations on ASR manuscripts","conf":"high","self":false,"doi":"10.2307/2095739","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Two initial referees independently recommended acceptance or rejection (recommendations dichotomised) for 342 manuscripts submitted to the American Sociological Review; a Cohen's kappa of 0.15 indicates only slight chance-corrected agreement between referees, close to statistical independence.","vf":"unverified"},{"key":"LJ3DDLY7","au":"Hargens, Lowell L.","y":1990,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient","estd":"ICC (single/unspec)","v":0.59,"n":"71","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"accept / acceptable with minor revisions / revise and resubmit / reject","field":"multi-field (sociology, zoology, law, psychology)","wr":"referees on submitted manuscript recommendations","conf":"med","self":false,"doi":"10.1016/0049-089x(90)90012-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":true,"ms":"Pairs of referee recommendations were analysed for 71 American Psychologist manuscripts; the intraclass correlation using equal-interval scores is 0.59, an unusually high single-referee reliability for a behavioural-science journal.","vf":"unverified"},{"key":"LJ3DDLY7","au":"Hargens, Lowell L.","y":1990,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient","estd":"ICC (single/unspec)","v":0.64,"n":"71","k":"2","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"accept / acceptable with minor revisions / revise and resubmit / reject","field":"multi-field (sociology, zoology, law, psychology)","wr":"referees on submitted manuscript recommendations","conf":"med","self":false,"doi":"10.1016/0049-089x(90)90012-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":true,"ms":"For the same 71 American Psychologist manuscripts, RC association model scale values raise the intraclass correlation to 0.64; the authors attribute the larger value partly to categories departing from their ostensible order.","vf":"unverified"},{"key":"LJ3DDLY7","au":"Hargens, Lowell L.","y":1990,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient","estd":"ICC (single/unspec)","v":0.28,"n":"322","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"accept minor changes / accept after substantial revision / revise and resubmit / reject","field":"multi-field (sociology, zoology, law, psychology)","wr":"referees on submitted manuscript recommendations","conf":"med","self":false,"doi":"10.1016/0049-089x(90)90012-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Two initial referees independently gave recommendations on 322 manuscripts submitted to American Sociological Review; the intraclass correlation of 0.28 using equal-interval category scores indicates low single-referee reliability.","vf":"unverified"},{"key":"LJ3DDLY7","au":"Hargens, Lowell L.","y":1990,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient","estd":"ICC (single/unspec)","v":0.29,"n":"322","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"accept minor changes / accept after substantial revision / revise and resubmit / reject","field":"multi-field (sociology, zoology, law, psychology)","wr":"referees on submitted manuscript recommendations","conf":"med","self":false,"doi":"10.1016/0049-089x(90)90012-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the same 322 American Sociological Review manuscripts, scale values derived from the RC association model give an intraclass correlation of 0.29, almost identical to the equal-interval estimate.","vf":"unverified"},{"key":"LJ3DDLY7","au":"Hargens, Lowell L.","y":1990,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient","estd":"ICC (single/unspec)","v":0.17,"n":"251","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"publish as is / publish with revisions / perhaps publish / do not publish","field":"multi-field (sociology, zoology, law, psychology)","wr":"referees on submitted manuscript recommendations","conf":"med","self":false,"doi":"10.1016/0049-089x(90)90012-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two referees judged 251 manuscripts submitted to Law and Society Review; the intraclass correlation using equal-interval scores is 0.17, the lowest single-referee reliability among the five journals.","vf":"unverified"},{"key":"LJ3DDLY7","au":"Hargens, Lowell L.","y":1990,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient","estd":"ICC (single/unspec)","v":0.23,"n":"251","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"publish as is / publish with revisions / perhaps publish / do not publish","field":"multi-field (sociology, zoology, law, psychology)","wr":"referees on submitted manuscript recommendations","conf":"med","self":false,"doi":"10.1016/0049-089x(90)90012-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the same 251 Law and Society Review manuscripts, RC association model scale values raise the intraclass correlation to 0.23, which the authors note is 35% larger than the equal-interval estimate.","vf":"unverified"},{"key":"LJ3DDLY7","au":"Hargens, Lowell L.","y":1990,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient","estd":"ICC (single/unspec)","v":0.27,"n":"177","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"combined accept / combined revise-resubmit and reject-revise / definitely reject","field":"multi-field (sociology, zoology, law, psychology)","wr":"referees on submitted manuscript recommendations","conf":"med","self":false,"doi":"10.1016/0049-089x(90)90012-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Collapsing the P&SPB categories to three statistically indistinguishable groups for the same 177 manuscripts gives an equal-interval intraclass correlation of 0.27.","vf":"unverified"},{"key":"LJ3DDLY7","au":"Hargens, Lowell L.","y":1990,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient","estd":"ICC (single/unspec)","v":0.28,"n":"177","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"combined accept / combined revise-resubmit and reject-revise / definitely reject","field":"multi-field (sociology, zoology, law, psychology)","wr":"referees on submitted manuscript recommendations","conf":"med","self":false,"doi":"10.1016/0049-089x(90)90012-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the same collapsed three-category P&SPB data, RC association model scale values give an intraclass correlation of 0.28, almost identical to the equal-interval estimate.","vf":"unverified"},{"key":"LJ3DDLY7","au":"Hargens, Lowell L.","y":1990,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient","estd":"ICC (single/unspec)","v":0.23,"n":"177","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"definitely accept / probably accept / revise-resubmit / reject but maybe revise / definitely reject","field":"multi-field (sociology, zoology, law, psychology)","wr":"referees on submitted manuscript recommendations","conf":"med","self":false,"doi":"10.1016/0049-089x(90)90012-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Pairs of referee recommendations were analysed for 177 Personality and Social Psychology Bulletin manuscripts using its five categories; the equal-interval intraclass correlation is 0.23, a low single-referee reliability.","vf":"unverified"},{"key":"LJ3DDLY7","au":"Hargens, Lowell L.","y":1990,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient","estd":"ICC (single/unspec)","v":0.3,"n":"177","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"definitely accept / probably accept / revise-resubmit / reject but maybe revise / definitely reject","field":"multi-field (sociology, zoology, law, psychology)","wr":"referees on submitted manuscript recommendations","conf":"med","self":false,"doi":"10.1016/0049-089x(90)90012-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the same 177 P&SPB manuscripts with five categories, RC association model scale values raise the intraclass correlation to 0.30, larger because the categories depart from their ostensible order.","vf":"unverified"},{"key":"LJ3DDLY7","au":"Hargens, Lowell L.","y":1990,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient","estd":"ICC (single/unspec)","v":0.28,"n":"209","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"acceptable / acceptable with suggestions / acceptable only if adequately revised / unacceptable","field":"multi-field (sociology, zoology, law, psychology)","wr":"referees on submitted manuscript recommendations","conf":"med","self":false,"doi":"10.1016/0049-089x(90)90012-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two initial referees independently gave recommendations on 209 manuscripts submitted to Physiological Zoology; the intraclass correlation of 0.28 using equal-interval category scores indicates low single-referee reliability.","vf":"unverified"},{"key":"LJ3DDLY7","au":"Hargens, Lowell L.","y":1990,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient","estd":"ICC (single/unspec)","v":0.31,"n":"209","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"acceptable / acceptable with suggestions / acceptable only if adequately revised / unacceptable","field":"multi-field (sociology, zoology, law, psychology)","wr":"referees on submitted manuscript recommendations","conf":"med","self":false,"doi":"10.1016/0049-089x(90)90012-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the same 209 Physiological Zoology manuscripts, RC association model scale values give an intraclass correlation of 0.31, close to the equal-interval estimate.","vf":"unverified"},{"key":"BHZQYW9A","au":"Hasbahçeci, Mustafa","y":2016,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"Kappa analysis","estd":"kappa","v":0.354,"n":"50","k":"2","samp":"funded-only","blind":"double","agg":"single-rater","scale":"each parameter assessed as 0 or 1; total 0-13","field":"surgery","wr":"two reviewers on congress abstract reporting quality","conf":"med","self":false,"doi":"10.5152/UCD.2016.3195","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":true,"he":true,"ms":"Two general surgeons independently scored 50 accepted observational congress abstracts using the proposed National Evaluation System. A kappa of 0.354 indicates acceptable chance-corrected agreement for the paper's focal instrument.","vf":"unverified"},{"key":"BHZQYW9A","au":"Hasbahçeci, Mustafa","y":2016,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"Kappa analysis","estd":"kappa","v":0.393,"n":"50","k":"2","samp":"funded-only","blind":"double","agg":"single-rater","scale":"each parameter assessed as 0 or 1; total 0-11","field":"surgery","wr":"two reviewers on congress abstract reporting quality","conf":"med","self":false,"doi":"10.5152/UCD.2016.3195","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Two general surgeons independently scored 50 accepted observational congress abstracts using the Turkish STROBE-based checklist. A kappa of 0.393 indicates acceptable chance-corrected agreement between the two reviewers.","vf":"unverified"},{"key":"BHZQYW9A","au":"Hasbahçeci, Mustafa","y":2016,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"Kappa analysis","estd":"kappa","v":0.523,"n":"50","k":"2","samp":"funded-only","blind":"double","agg":"single-rater","scale":"each parameter assessed as 0 or 1; total 0-15","field":"surgery","wr":"two reviewers on congress abstract reporting quality","conf":"med","self":false,"doi":"10.5152/UCD.2016.3195","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Two general surgeons independently scored 50 accepted observational congress abstracts using the Timmer quality score. A kappa of 0.523 indicates moderate chance-corrected agreement, the highest of the three systems.","vf":"unverified"},{"key":"SEPHFSGV","au":"Herbert, Danielle L.","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"raw percent agreement on binary funding decision, exact agreement, not chance-corrected","estd":"percent agreement","v":0.64,"n":"36","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"panel-consensus","scale":"funded: yes / no","field":"biomedical (basic science and public health)","wr":"two processes' fund decisions on same grant proposals","conf":"high","self":false,"doi":"10.1136/bmjopen-2015-008380","ciLow":0.47,"ciHigh":0.78,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For the 36 Basic Science proposals, the two-reviewer journal-style process and the official panel agreed on the binary funding decision 64% of the time (raw agreement).","vf":"unverified"},{"key":"SEPHFSGV","au":"Herbert, Danielle L.","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"raw percent agreement on binary funding decision, exact agreement, not chance-corrected","estd":"percent agreement","v":0.74,"n":"72","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"panel-consensus","scale":"funded: yes / no","field":"biomedical (basic science and public health)","wr":"two processes' fund decisions on same grant proposals","conf":"high","self":false,"doi":"10.1136/bmjopen-2015-008380","ciLow":0.62,"ciHigh":0.83,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"A journal-style process using two independent reviewers per proposal agreed with the official NHMRC panel on the binary funding decision for 74% of the 72 proposals (raw agreement, not chance-corrected).","vf":"unverified"},{"key":"SEPHFSGV","au":"Herbert, Danielle L.","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"raw percent agreement on binary funding decision, exact agreement, not chance-corrected","estd":"percent agreement","v":0.83,"n":"36","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"panel-consensus","scale":"funded: yes / no","field":"biomedical (basic science and public health)","wr":"two processes' fund decisions on same grant proposals","conf":"high","self":false,"doi":"10.1136/bmjopen-2015-008380","ciLow":0.69,"ciHigh":0.94,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For the 36 Public Health proposals, the two-reviewer journal-style process and the official panel agreed on the binary funding decision 83% of the time (raw agreement).","vf":"unverified"},{"key":"SEPHFSGV","au":"Herbert, Danielle L.","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"raw percent agreement on binary funding decision, exact agreement, not chance-corrected","estd":"percent agreement","v":0.78,"n":"36","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"panel-consensus","scale":"funded: yes / no","field":"biomedical (basic science and public health)","wr":"two processes' fund decisions on same grant proposals","conf":"high","self":false,"doi":"10.1136/bmjopen-2015-008380","ciLow":0.64,"ciHigh":0.92,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For the 36 Basic Science proposals, the simplified panel and the journal-style process agreed on the binary funding decision 78% of the time (raw agreement).","vf":"unverified"},{"key":"SEPHFSGV","au":"Herbert, Danielle L.","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"raw percent agreement on binary funding decision, exact agreement, not chance-corrected","estd":"percent agreement","v":0.79,"n":"72","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"panel-consensus","scale":"funded: yes / no","field":"biomedical (basic science and public health)","wr":"two processes' fund decisions on same grant proposals","conf":"high","self":false,"doi":"10.1136/bmjopen-2015-008380","ciLow":0.68,"ciHigh":0.89,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"The simplified seven-member panel and the two-reviewer journal-style process agreed on the binary funding decision for 79% of the 72 proposals (raw agreement); the text loosely calls these 'the two simplified panels' but Table 4 defines the comparison as simplified versus journal.","vf":"unverified"},{"key":"SEPHFSGV","au":"Herbert, Danielle L.","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"raw percent agreement on binary funding decision, exact agreement, not chance-corrected","estd":"percent agreement","v":0.81,"n":"36","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"panel-consensus","scale":"funded: yes / no","field":"biomedical (basic science and public health)","wr":"two processes' fund decisions on same grant proposals","conf":"high","self":false,"doi":"10.1136/bmjopen-2015-008380","ciLow":0.67,"ciHigh":0.92,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For the 36 Public Health proposals, the simplified panel and the journal-style process agreed on the binary funding decision 81% of the time (raw agreement).","vf":"unverified"},{"key":"SEPHFSGV","au":"Herbert, Danielle L.","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"raw percent agreement on binary funding decision, exact agreement, not chance-corrected","estd":"percent agreement","v":0.69,"n":"36","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"panel-consensus","scale":"funded: yes / no","field":"biomedical (basic science and public health)","wr":"two processes' fund decisions on same grant proposals","conf":"high","self":false,"doi":"10.1136/bmjopen-2015-008380","ciLow":0.56,"ciHigh":0.83,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For the 36 Basic Science proposals, the simplified seven-member panel and the official panel agreed on the binary funding decision 69% of the time (raw agreement).","vf":"unverified"},{"key":"SEPHFSGV","au":"Herbert, Danielle L.","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"raw percent agreement on binary funding decision, exact agreement, not chance-corrected","estd":"percent agreement","v":0.72,"n":"72","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"panel-consensus","scale":"funded: yes / no","field":"biomedical (basic science and public health)","wr":"two processes' fund decisions on same grant proposals","conf":"high","self":false,"doi":"10.1136/bmjopen-2015-008380","ciLow":0.61,"ciHigh":0.82,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":true,"he":false,"ms":"In parallel with the official NHMRC review, 72 Project Grant proposals were assessed by a simplified seven-member panel; the simplified and official processes agreed on the binary fund/not-fund decision for 72% of proposals (raw agreement, not chance-corrected).","vf":"unverified"},{"key":"SEPHFSGV","au":"Herbert, Danielle L.","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"raw percent agreement on binary funding decision, exact agreement, not chance-corrected","estd":"percent agreement","v":0.75,"n":"36","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"panel-consensus","scale":"funded: yes / no","field":"biomedical (basic science and public health)","wr":"two processes' fund decisions on same grant proposals","conf":"high","self":false,"doi":"10.1136/bmjopen-2015-008380","ciLow":0.61,"ciHigh":0.89,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For the 36 Public Health proposals, the simplified seven-member panel and the official panel agreed on the binary funding decision 75% of the time (raw agreement).","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, average measure","estd":"ICC (average)","v":0.404,"n":"86","k":"3","samp":"full-pool","blind":"double","agg":"average-of-k","scale":"1 (very low/very poor) to 5 (very high/excellent)","field":"anthrozoology","wr":"reviewers on conference abstract clarity scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 86 programs/applications abstracts submitted to the IAHAIO 2004 conference on overall clarity using a five-point scale. The average measure intraclass correlation of 0.404 gives the reliability of the mean rating of the three reviewers for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, average measure","estd":"ICC (average)","v":0.487,"n":"119","k":"3","samp":"full-pool","blind":"double","agg":"average-of-k","scale":"1 (very low/very poor) to 5 (very high/excellent)","field":"anthrozoology","wr":"reviewers on conference abstract clarity scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 119 research presentations abstracts submitted to the IAHAIO 2004 conference on overall clarity using a five-point scale. The average measure intraclass correlation of 0.487 gives the reliability of the mean rating of the three reviewers for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, average measure","estd":"ICC (average)","v":0.814,"n":"15","k":"3","samp":"full-pool","blind":"double","agg":"average-of-k","scale":"1 (very low/very poor) to 5 (very high/excellent)","field":"anthrozoology","wr":"reviewers on conference abstract clarity scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 15 critical reviews abstracts submitted to the IAHAIO 2004 conference on overall clarity using a five-point scale. The average measure intraclass correlation of 0.814 gives the reliability of the mean rating of the three reviewers for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, average measure","estd":"ICC (average)","v":0.164,"n":"86","k":"3","samp":"full-pool","blind":"double","agg":"average-of-k","scale":"1 (very low/very poor) to 5 (very high/excellent)","field":"anthrozoology","wr":"reviewers on abstract importance scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 86 programs/applications abstracts submitted to the IAHAIO 2004 conference on implications and importance for the field using a five-point scale. The average measure intraclass correlation of 0.164 gives the reliability of the mean rating of the three reviewers for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, average measure","estd":"ICC (average)","v":0.414,"n":"119","k":"3","samp":"full-pool","blind":"double","agg":"average-of-k","scale":"1 (very low/very poor) to 5 (very high/excellent)","field":"anthrozoology","wr":"reviewers on abstract importance scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 119 research presentations abstracts submitted to the IAHAIO 2004 conference on implications and importance for the field using a five-point scale. The average measure intraclass correlation of 0.414 gives the reliability of the mean rating of the three reviewers for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, average measure","estd":"ICC (average)","v":0.695,"n":"15","k":"3","samp":"full-pool","blind":"double","agg":"average-of-k","scale":"1 (very low/very poor) to 5 (very high/excellent)","field":"anthrozoology","wr":"reviewers on abstract importance scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 15 critical reviews abstracts submitted to the IAHAIO 2004 conference on implications and importance for the field using a five-point scale. The average measure intraclass correlation of 0.695 gives the reliability of the mean rating of the three reviewers for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, average measure","estd":"ICC (average)","v":0.279,"n":"86","k":"3","samp":"full-pool","blind":"double","agg":"average-of-k","scale":"1 (very low/very poor) to 5 (very high/excellent)","field":"anthrozoology","wr":"reviewers on abstract results/interpretation scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 86 programs/applications abstracts submitted to the IAHAIO 2004 conference on quality of results and interpretation using a five-point scale. The average measure intraclass correlation of 0.279 gives the reliability of the mean rating of the three reviewers for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, average measure","estd":"ICC (average)","v":0.643,"n":"119","k":"3","samp":"full-pool","blind":"double","agg":"average-of-k","scale":"1 (very low/very poor) to 5 (very high/excellent)","field":"anthrozoology","wr":"reviewers on abstract results/interpretation scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 119 research presentations abstracts submitted to the IAHAIO 2004 conference on quality of results and interpretation using a five-point scale. The average measure intraclass correlation of 0.643 gives the reliability of the mean rating of the three reviewers for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, average measure","estd":"ICC (average)","v":0.651,"n":"15","k":"3","samp":"full-pool","blind":"double","agg":"average-of-k","scale":"1 (very low/very poor) to 5 (very high/excellent)","field":"anthrozoology","wr":"reviewers on abstract results/interpretation scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 15 critical reviews abstracts submitted to the IAHAIO 2004 conference on quality of results and interpretation using a five-point scale. The average measure intraclass correlation of 0.651 gives the reliability of the mean rating of the three reviewers for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, average measure","estd":"ICC (average)","v":0.432,"n":"86","k":"3","samp":"full-pool","blind":"double","agg":"average-of-k","scale":"1 (very low/very poor) to 5 (very high/excellent)","field":"anthrozoology","wr":"reviewers on abstract methodological rigor scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 86 programs/applications abstracts submitted to the IAHAIO 2004 conference on methodological, procedural and theoretical rigour using a five-point scale. The average measure intraclass correlation of 0.432 gives the reliability of the mean rating of the three reviewers for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, average measure","estd":"ICC (average)","v":0.569,"n":"119","k":"3","samp":"full-pool","blind":"double","agg":"average-of-k","scale":"1 (very low/very poor) to 5 (very high/excellent)","field":"anthrozoology","wr":"reviewers on abstract methodological rigor scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 119 research presentations abstracts submitted to the IAHAIO 2004 conference on methodological, procedural and theoretical rigour using a five-point scale. The average measure intraclass correlation of 0.569 gives the reliability of the mean rating of the three reviewers for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, average measure","estd":"ICC (average)","v":0.668,"n":"15","k":"3","samp":"full-pool","blind":"double","agg":"average-of-k","scale":"1 (very low/very poor) to 5 (very high/excellent)","field":"anthrozoology","wr":"reviewers on abstract methodological rigor scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 15 critical reviews abstracts submitted to the IAHAIO 2004 conference on methodological, procedural and theoretical rigour using a five-point scale. The average measure intraclass correlation of 0.668 gives the reliability of the mean rating of the three reviewers for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, average measure","estd":"ICC (average)","v":0.327,"n":"86","k":"3","samp":"full-pool","blind":"double","agg":"average-of-k","scale":"total score 4-20, sum of four 1-5 sub-scores","field":"anthrozoology","wr":"reviewers on abstract total proposal scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 86 programs/applications abstracts submitted to the IAHAIO 2004 conference on total proposal score using summed total score, which could range from 4 to 20. The average measure intraclass correlation of 0.327 gives the reliability of the mean rating of the three reviewers for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, average measure","estd":"ICC (average)","v":0.594,"n":"119","k":"3","samp":"full-pool","blind":"double","agg":"average-of-k","scale":"total score 4-20, sum of four 1-5 sub-scores","field":"anthrozoology","wr":"reviewers on abstract total proposal scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 119 research presentations abstracts submitted to the IAHAIO 2004 conference on total proposal score using summed total score, which could range from 4 to 20. The average measure intraclass correlation of 0.594 gives the reliability of the mean rating of the three reviewers for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, average measure","estd":"ICC (average)","v":0.765,"n":"15","k":"3","samp":"full-pool","blind":"double","agg":"average-of-k","scale":"total score 4-20, sum of four 1-5 sub-scores","field":"anthrozoology","wr":"reviewers on abstract total proposal scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 15 critical reviews abstracts submitted to the IAHAIO 2004 conference on total proposal score using summed total score, which could range from 4 to 20. The average measure intraclass correlation of 0.765 gives the reliability of the mean rating of the three reviewers for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, single measure","estd":"ICC (single/unspec)","v":0.185,"n":"86","k":"3","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 (very low/very poor) to 5 (very high/excellent)","field":"anthrozoology","wr":"reviewers on conference abstract clarity scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 86 programs/applications abstracts submitted to the IAHAIO 2004 conference on overall clarity using a five-point scale. The single measure intraclass correlation of 0.185 gives the reliability of a single reviewer's rating for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, single measure","estd":"ICC (single/unspec)","v":0.241,"n":"119","k":"3","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 (very low/very poor) to 5 (very high/excellent)","field":"anthrozoology","wr":"reviewers on conference abstract clarity scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 119 research presentations abstracts submitted to the IAHAIO 2004 conference on overall clarity using a five-point scale. The single measure intraclass correlation of 0.241 gives the reliability of a single reviewer's rating for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, single measure","estd":"ICC (single/unspec)","v":0.593,"n":"15","k":"3","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 (very low/very poor) to 5 (very high/excellent)","field":"anthrozoology","wr":"reviewers on conference abstract clarity scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 15 critical reviews abstracts submitted to the IAHAIO 2004 conference on overall clarity using a five-point scale. The single measure intraclass correlation of 0.593 gives the reliability of a single reviewer's rating for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, single measure","estd":"ICC (single/unspec)","v":0.061,"n":"86","k":"3","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 (very low/very poor) to 5 (very high/excellent)","field":"anthrozoology","wr":"reviewers on abstract importance scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 86 programs/applications abstracts submitted to the IAHAIO 2004 conference on implications and importance for the field using a five-point scale. The single measure intraclass correlation of 0.061 gives the reliability of a single reviewer's rating for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, single measure","estd":"ICC (single/unspec)","v":0.19,"n":"119","k":"3","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 (very low/very poor) to 5 (very high/excellent)","field":"anthrozoology","wr":"reviewers on abstract importance scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 119 research presentations abstracts submitted to the IAHAIO 2004 conference on implications and importance for the field using a five-point scale. The single measure intraclass correlation of 0.190 gives the reliability of a single reviewer's rating for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, single measure","estd":"ICC (single/unspec)","v":0.431,"n":"15","k":"3","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 (very low/very poor) to 5 (very high/excellent)","field":"anthrozoology","wr":"reviewers on abstract importance scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 15 critical reviews abstracts submitted to the IAHAIO 2004 conference on implications and importance for the field using a five-point scale. The single measure intraclass correlation of 0.431 gives the reliability of a single reviewer's rating for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, single measure","estd":"ICC (single/unspec)","v":0.114,"n":"86","k":"3","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 (very low/very poor) to 5 (very high/excellent)","field":"anthrozoology","wr":"reviewers on abstract results/interpretation scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 86 programs/applications abstracts submitted to the IAHAIO 2004 conference on quality of results and interpretation using a five-point scale. The single measure intraclass correlation of 0.114 gives the reliability of a single reviewer's rating for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, single measure","estd":"ICC (single/unspec)","v":0.376,"n":"119","k":"3","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 (very low/very poor) to 5 (very high/excellent)","field":"anthrozoology","wr":"reviewers on abstract results/interpretation scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 119 research presentations abstracts submitted to the IAHAIO 2004 conference on quality of results and interpretation using a five-point scale. The single measure intraclass correlation of 0.376 gives the reliability of a single reviewer's rating for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, single measure","estd":"ICC (single/unspec)","v":0.383,"n":"15","k":"3","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 (very low/very poor) to 5 (very high/excellent)","field":"anthrozoology","wr":"reviewers on abstract results/interpretation scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 15 critical reviews abstracts submitted to the IAHAIO 2004 conference on quality of results and interpretation using a five-point scale. The single measure intraclass correlation of 0.383 gives the reliability of a single reviewer's rating for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, single measure","estd":"ICC (single/unspec)","v":0.203,"n":"86","k":"3","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 (very low/very poor) to 5 (very high/excellent)","field":"anthrozoology","wr":"reviewers on abstract methodological rigor scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 86 programs/applications abstracts submitted to the IAHAIO 2004 conference on methodological, procedural and theoretical rigour using a five-point scale. The single measure intraclass correlation of 0.203 gives the reliability of a single reviewer's rating for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, single measure","estd":"ICC (single/unspec)","v":0.301,"n":"119","k":"3","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 (very low/very poor) to 5 (very high/excellent)","field":"anthrozoology","wr":"reviewers on abstract methodological rigor scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 119 research presentations abstracts submitted to the IAHAIO 2004 conference on methodological, procedural and theoretical rigour using a five-point scale. The single measure intraclass correlation of 0.301 gives the reliability of a single reviewer's rating for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, single measure","estd":"ICC (single/unspec)","v":0.402,"n":"15","k":"3","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 (very low/very poor) to 5 (very high/excellent)","field":"anthrozoology","wr":"reviewers on abstract methodological rigor scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 15 critical reviews abstracts submitted to the IAHAIO 2004 conference on methodological, procedural and theoretical rigour using a five-point scale. The single measure intraclass correlation of 0.402 gives the reliability of a single reviewer's rating for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, single measure","estd":"ICC (single/unspec)","v":0.14,"n":"86","k":"3","samp":"full-pool","blind":"double","agg":"single-rater","scale":"total score 4-20, sum of four 1-5 sub-scores","field":"anthrozoology","wr":"reviewers on abstract total proposal scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 86 programs/applications abstracts submitted to the IAHAIO 2004 conference on total proposal score using summed total score, which could range from 4 to 20. The single measure intraclass correlation of 0.140 gives the reliability of a single reviewer's rating for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, single measure","estd":"ICC (single/unspec)","v":0.328,"n":"119","k":"3","samp":"full-pool","blind":"double","agg":"single-rater","scale":"total score 4-20, sum of four 1-5 sub-scores","field":"anthrozoology","wr":"reviewers on abstract total proposal scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 119 research presentations abstracts submitted to the IAHAIO 2004 conference on total proposal score using summed total score, which could range from 4 to 20. The single measure intraclass correlation of 0.328 gives the reliability of a single reviewer's rating for this category.","vf":"unverified"},{"key":"ZPZW8PR3","au":"Herzog, Harold A.","y":2005,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"ICC Model I, SPSS one-way random model, single measure","estd":"ICC (single/unspec)","v":0.52,"n":"15","k":"3","samp":"full-pool","blind":"double","agg":"single-rater","scale":"total score 4-20, sum of four 1-5 sub-scores","field":"anthrozoology","wr":"reviewers on abstract total proposal scores","conf":"high","self":false,"doi":"10.2752/089279305785594180","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three experts, blind to author identity, independently rated each of the 15 critical reviews abstracts submitted to the IAHAIO 2004 conference on total proposal score using summed total score, which could range from 4 to 20. The single measure intraclass correlation of 0.520 gives the reliability of a single reviewer's rating for this category.","vf":"unverified"},{"key":"LVKA3ZWD","au":"Hesselberg, Jan-Ole","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"Gwet-AC","form":"Gwet's agreement coefficient 1 (AC1)","estd":"Gwet AC","v":0.829,"n":"511","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"eligible: yes/no","field":"health","wr":"reviewers on grant proposal eligibility","conf":"high","self":true,"doi":"10.1186/s41073-021-00115-5","ciLow":0.794,"ciHigh":0.864,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Committee reviewers independently judged whether grant proposals were eligible for the programme. For the general feedback group at baseline in 2017, a Gwet's AC1 of 0.83 indicates high chance-corrected agreement on eligibility between paired reviewers.","vf":"unverified"},{"key":"LVKA3ZWD","au":"Hesselberg, Jan-Ole","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"Gwet-AC","form":"Gwet's agreement coefficient 1 (AC1)","estd":"Gwet AC","v":0.927,"n":"450","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"eligible: yes/no","field":"health","wr":"reviewers on grant proposal eligibility","conf":"high","self":true,"doi":"10.1186/s41073-021-00115-5","ciLow":0.905,"ciHigh":0.95,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"In 2018, after the general (control) feedback report, the general feedback group's paired reviewers showed a Gwet's AC1 of 0.93 on the eligibility judgment, the highest agreement observed in the study.","vf":"unverified"},{"key":"LVKA3ZWD","au":"Hesselberg, Jan-Ole","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"Gwet-AC","form":"Gwet's agreement coefficient 1 (AC1)","estd":"Gwet AC","v":0.786,"n":"409","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"eligible: yes/no","field":"health","wr":"reviewers on grant proposal eligibility","conf":"high","self":true,"doi":"10.1186/s41073-021-00115-5","ciLow":0.762,"ciHigh":0.852,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"For the individual feedback group at baseline in 2017, Gwet's AC1 for the binary eligibility judgment between paired reviewers was 0.79. The reported confidence interval is identical to the 2018 individual-group value, which looks like a typographical error.","vf":"unverified"},{"key":"LVKA3ZWD","au":"Hesselberg, Jan-Ole","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"Gwet-AC","form":"Gwet's agreement coefficient 1 (AC1)","estd":"Gwet AC","v":0.807,"n":"376","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"eligible: yes/no","field":"health","wr":"reviewers on grant proposal eligibility","conf":"high","self":true,"doi":"10.1186/s41073-021-00115-5","ciLow":0.762,"ciHigh":0.852,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"In 2018, after two individual feedback reports, the individual feedback group's paired reviewers showed a Gwet's AC1 of 0.81 on the eligibility judgment, little changed from baseline.","vf":"unverified"},{"key":"LVKA3ZWD","au":"Hesselberg, Jan-Ole","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"average absolute difference between paired reviews (= 2 x AD index)","estd":"AD index","v":2.1,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-10, 10 = best","field":"health","wr":"reviewers on grant proposal quality scores","conf":"high","self":true,"doi":"10.1186/s41073-021-00115-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Pooled across both study arms at baseline in 2017, the average absolute difference between paired reviewers' quality scores was 2.1 points (SD of differences 1.56), over 1035 both-eligible paired reviews.","vf":"unverified"},{"key":"LVKA3ZWD","au":"Hesselberg, Jan-Ole","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"average absolute difference between paired reviews (= 2 x AD index)","estd":"AD index","v":2,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-10, 10 = best","field":"health","wr":"reviewers on grant proposal quality scores","conf":"high","self":true,"doi":"10.1186/s41073-021-00115-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Average absolute difference between paired reviewers' 1-10 quality scores (twice the AD index) in the general feedback group in 2017 was 2.0 points (SD of differences 1.54); larger values mean less agreement.","vf":"unverified"},{"key":"LVKA3ZWD","au":"Hesselberg, Jan-Ole","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"average absolute difference between paired reviews (= 2 x AD index)","estd":"AD index","v":1.8,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-10, 10 = best","field":"health","wr":"reviewers on grant proposal quality scores","conf":"high","self":true,"doi":"10.1186/s41073-021-00115-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Average absolute difference between paired reviewers' quality scores in the general feedback group in 2018 was 1.8 points (SD of differences 1.47), a decrease from 2.0 at baseline.","vf":"unverified"},{"key":"LVKA3ZWD","au":"Hesselberg, Jan-Ole","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"average absolute difference between paired reviews (= 2 x AD index)","estd":"AD index","v":2.2,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-10, 10 = best","field":"health","wr":"reviewers on grant proposal quality scores","conf":"high","self":true,"doi":"10.1186/s41073-021-00115-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Average absolute difference between paired reviewers' quality scores in the individual feedback group in 2017 was 2.2 points (SD of differences 1.59), the largest disagreement observed.","vf":"unverified"},{"key":"LVKA3ZWD","au":"Hesselberg, Jan-Ole","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"average absolute difference between paired reviews (= 2 x AD index)","estd":"AD index","v":1.9,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-10, 10 = best","field":"health","wr":"reviewers on grant proposal quality scores","conf":"high","self":true,"doi":"10.1186/s41073-021-00115-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Average absolute difference between paired reviewers' quality scores in the individual feedback group in 2018 was 1.9 points (SD of differences 1.48), a decrease from 2.2 at baseline.","vf":"unverified"},{"key":"LVKA3ZWD","au":"Hesselberg, Jan-Ole","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"one-way random effects model, ICC (1, 2)","estd":"ICC (average)","v":0.276,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1-10, 10 = best","field":"health","wr":"reviewers on grant proposal quality scores","conf":"high","self":true,"doi":"10.1186/s41073-021-00115-5","ciLow":0.151,"ciHigh":0.383,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Reliability of the average of two reviewers' 1-10 quality scores, one-way random ICC(1,2), for the general feedback group at baseline in 2017 was 0.28, indicating low agreement. Computed over 601 both-eligible paired reviews.","vf":"unverified"},{"key":"LVKA3ZWD","au":"Hesselberg, Jan-Ole","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"one-way random effects model, ICC (1, 2)","estd":"ICC (average)","v":0.303,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1-10, 10 = best","field":"health","wr":"reviewers on grant proposal quality scores","conf":"high","self":true,"doi":"10.1186/s41073-021-00115-5","ciLow":0.181,"ciHigh":0.406,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"One-way random ICC(1,2) for the general feedback group in 2018, after the control feedback report, was 0.30 for the average of two reviewers' quality scores, computed over 594 both-eligible paired reviews.","vf":"unverified"},{"key":"LVKA3ZWD","au":"Hesselberg, Jan-Ole","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"one-way random effects model, ICC (1, 2)","estd":"ICC (average)","v":0.323,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1-10, 10 = best","field":"health","wr":"reviewers on grant proposal quality scores","conf":"high","self":true,"doi":"10.1186/s41073-021-00115-5","ciLow":0.182,"ciHigh":0.439,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"One-way random ICC(1,2) for the individual feedback group at baseline in 2017 was 0.32 for the average of two reviewers' 1-10 quality scores, computed over 434 both-eligible paired reviews.","vf":"unverified"},{"key":"LVKA3ZWD","au":"Hesselberg, Jan-Ole","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"one-way random effects model, ICC (1, 2)","estd":"ICC (average)","v":0.401,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1-10, 10 = best","field":"health","wr":"reviewers on grant proposal quality scores","conf":"high","self":true,"doi":"10.1186/s41073-021-00115-5","ciLow":0.272,"ciHigh":0.506,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"One-way random ICC(1,2) for the individual feedback group in 2018, after two individual feedback reports, was 0.40 for the average of two reviewers' quality scores, computed over 409 both-eligible paired reviews.","vf":"unverified"},{"key":"LVKA3ZWD","au":"Hesselberg, Jan-Ole","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"one-way random effects model, ICC (1, 3)","estd":"ICC (average)","v":0.334,"n":"1197","k":"3","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1-10, 10 = best","field":"health","wr":"reviewers on grant proposal quality scores","conf":"high","self":true,"doi":"10.1186/s41073-021-00115-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Across all 1197 proposals submitted in 2017 (whole calls, including non-study reviews), the one-way random ICC(1,3) for the average of three reviewers' 1-10 quality scores was 0.33. This is the study's overall estimate of the ordinary review process before any feedback was delivered.","vf":"unverified"},{"key":"LVKA3ZWD","au":"Hesselberg, Jan-Ole","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"one-way random effects model, ICC (1, 3)","estd":"ICC (average)","v":0.428,"n":"1043","k":"3","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1-10, 10 = best","field":"health","wr":"reviewers on grant proposal quality scores","conf":"high","self":true,"doi":"10.1186/s41073-021-00115-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Across all 1043 proposals submitted in 2018, after feedback reports had been delivered to both study arms, the one-way random ICC(1,3) for the average of three reviewers' quality scores was 0.43, up from 0.33 in 2017. Presented as descriptive real-funder data, regardless of the interventions.","vf":"unverified"},{"key":"LVKA3ZWD","au":"Hesselberg, Jan-Ole","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement on eligibility (agree/disagree)","estd":"percent agreement","v":0.86,"n":"511","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"eligible: yes/no","field":"health","wr":"reviewers on grant proposal eligibility","conf":"high","self":true,"doi":"10.1186/s41073-021-00115-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Raw proportion of paired reviewers agreeing on proposal eligibility in the general feedback group in 2017 was 0.86 (612 of 715 pairs), before any chance correction.","vf":"unverified"},{"key":"LVKA3ZWD","au":"Hesselberg, Jan-Ole","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement on eligibility (agree/disagree)","estd":"percent agreement","v":0.93,"n":"450","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"eligible: yes/no","field":"health","wr":"reviewers on grant proposal eligibility","conf":"high","self":true,"doi":"10.1186/s41073-021-00115-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Raw proportion of paired reviewers agreeing on proposal eligibility in the general feedback group in 2018 was 0.93 (599 of 642 pairs).","vf":"unverified"},{"key":"LVKA3ZWD","au":"Hesselberg, Jan-Ole","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement on eligibility (agree/disagree)","estd":"percent agreement","v":0.83,"n":"409","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"eligible: yes/no","field":"health","wr":"reviewers on grant proposal eligibility","conf":"high","self":true,"doi":"10.1186/s41073-021-00115-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Raw proportion of paired reviewers agreeing on proposal eligibility in the individual feedback group in 2017 was 0.83 (453 of 545 pairs).","vf":"unverified"},{"key":"LVKA3ZWD","au":"Hesselberg, Jan-Ole","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement on eligibility (agree/disagree)","estd":"percent agreement","v":0.84,"n":"376","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"eligible: yes/no","field":"health","wr":"reviewers on grant proposal eligibility","conf":"high","self":true,"doi":"10.1186/s41073-021-00115-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Raw proportion of paired reviewers agreeing on proposal eligibility in the individual feedback group in 2018 was 0.84 (417 of 496 pairs).","vf":"unverified"},{"key":"LVKA3ZWD","au":"Hesselberg, Jan-Ole","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"agreement within 1 point (score difference of 0 or 1)","estd":"percent agreement","v":0.497,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-10, 10 = best","field":"health","wr":"reviewers on grant proposal quality scores","conf":"high","self":true,"doi":"10.1186/s41073-021-00115-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Preregistered dichotomous score-agreement measure, reported as a sensitivity analysis of the 2018 group comparison: the proportion of paired reviewers scoring within one point on the 1-10 quality scale was 0.50 in the general feedback group.","vf":"unverified"},{"key":"LVKA3ZWD","au":"Hesselberg, Jan-Ole","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"agreement within 1 point (score difference of 0 or 1)","estd":"percent agreement","v":0.496,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-10, 10 = best","field":"health","wr":"reviewers on grant proposal quality scores","conf":"high","self":true,"doi":"10.1186/s41073-021-00115-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Preregistered dichotomous score-agreement measure, reported as a sensitivity analysis of the 2018 group comparison: the proportion of paired reviewers scoring within one point on the 1-10 quality scale was 0.50 in the individual feedback group.","vf":"unverified"},{"key":"A58BN9MR","au":"Heyard, Rachel","y":2022,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC (Shrout & Fleiss 1979), computed with function ICC in R package psych","estd":"ICC (single/unspec)","v":0.82,"n":"87","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"1, poor, to 6, outstanding","field":"multi-field","wr":"panel members on project grant proposal scores","conf":"high","self":false,"doi":"10.1080/2330443X.2022.2086190","ciLow":0.78,"ciHigh":0.86,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Nine panel members each scored all 87 project grant proposals in sub-panel one unless they had a conflict of interest. An ICC of 0.82 indicates good inter-rater reliability across the full score range.","vf":"unverified"},{"key":"A58BN9MR","au":"Heyard, Rachel","y":2022,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC (Shrout & Fleiss 1979), computed with function ICC in R package psych","estd":"ICC (single/unspec)","v":0.85,"n":"92","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"1, poor, to 6, outstanding","field":"multi-field","wr":"panel members on project grant proposal scores","conf":"high","self":false,"doi":"10.1080/2330443X.2022.2086190","ciLow":0.82,"ciHigh":0.89,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Nine panel members each scored all 92 project grant proposals in sub-panel two unless they had a conflict of interest. An ICC of 0.85 indicates good inter-rater reliability across the full score range.","vf":"unverified"},{"key":"A58BN9MR","au":"Heyard, Rachel","y":2022,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC (Shrout & Fleiss 1979), computed with function ICC in R package psych","estd":"ICC (single/unspec)","v":0.85,"n":"86","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"1, poor, to 6, outstanding","field":"multi-field","wr":"panel members on project grant proposal scores","conf":"high","self":false,"doi":"10.1080/2330443X.2022.2086190","ciLow":0.81,"ciHigh":0.88,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Nine panel members each scored all 86 project grant proposals in sub-panel three unless they had a conflict of interest. An ICC of 0.85 indicates good inter-rater reliability across the full score range.","vf":"unverified"},{"key":"A58BN9MR","au":"Heyard, Rachel","y":2022,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC (Shrout & Fleiss 1979), computed with function ICC in R package psych","estd":"ICC (single/unspec)","v":0.82,"n":"88","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"1, poor, to 6, outstanding","field":"multi-field","wr":"panel members on project grant proposal scores","conf":"high","self":false,"doi":"10.1080/2330443X.2022.2086190","ciLow":0.78,"ciHigh":0.86,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Nine panel members each scored all 88 project grant proposals in sub-panel four unless they had a conflict of interest. An ICC of 0.82 indicates good inter-rater reliability across the full score range.","vf":"unverified"},{"key":"A58BN9MR","au":"Heyard, Rachel","y":2022,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC (Shrout & Fleiss 1979), computed with function ICC in R package psych","estd":"ICC (single/unspec)","v":0.5,"n":"18","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"1, poor, to 6, outstanding","field":"multi-field","wr":"panel members on junior fellowship proposal scores","conf":"high","self":false,"doi":"10.1080/2330443X.2022.2086190","ciLow":0.37,"ciHigh":0.67,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Twenty-six Biology panel members independently scored the 18 discussed junior fellowship proposals, each voting unless conflicted or absent. An ICC of 0.50 indicates fair inter-rater reliability.","vf":"unverified"},{"key":"A58BN9MR","au":"Heyard, Rachel","y":2022,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC (Shrout & Fleiss 1979), computed with function ICC in R package psych","estd":"ICC (single/unspec)","v":0.33,"n":"11","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"1, poor, to 6, outstanding","field":"multi-field","wr":"panel members on junior fellowship proposal scores","conf":"high","self":false,"doi":"10.1080/2330443X.2022.2086190","ciLow":0.18,"ciHigh":0.58,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Fourteen Humanities panel members independently scored the 11 discussed junior fellowship proposals in a remote SNSF meeting, each voting unless conflicted or absent. An ICC of 0.33 indicates poor inter-rater reliability.","vf":"unverified"},{"key":"A58BN9MR","au":"Heyard, Rachel","y":2022,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC (Shrout & Fleiss 1979), computed with function ICC in R package psych","estd":"ICC (single/unspec)","v":0.43,"n":"14","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"1, poor, to 6, outstanding","field":"multi-field","wr":"panel members on junior fellowship proposal scores","conf":"high","self":false,"doi":"10.1080/2330443X.2022.2086190","ciLow":0.28,"ciHigh":0.63,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Seventeen Medicine panel members independently scored the 14 discussed junior fellowship proposals, each voting unless conflicted or absent. An ICC of 0.43 indicates fair inter-rater reliability.","vf":"unverified"},{"key":"A58BN9MR","au":"Heyard, Rachel","y":2022,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC (Shrout & Fleiss 1979), computed with function ICC in R package psych","estd":"ICC (single/unspec)","v":0.4,"n":"18","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"1, poor, to 6, outstanding","field":"multi-field","wr":"panel members on junior fellowship proposal scores","conf":"high","self":false,"doi":"10.1080/2330443X.2022.2086190","ciLow":0.27,"ciHigh":0.58,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":true,"ms":"Fourteen Social Sciences panel members independently scored the 18 discussed junior fellowship proposals, each voting unless conflicted or absent. An ICC of 0.40 indicates fair inter-rater reliability. Marked primary as the representative near-funding-line result; no overall ICC is reported.","vf":"unverified"},{"key":"A58BN9MR","au":"Heyard, Rachel","y":2022,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC (Shrout & Fleiss 1979), computed with function ICC in R package psych","estd":"ICC (single/unspec)","v":0.31,"n":"18","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"1, poor, to 6, outstanding","field":"multi-field","wr":"panel members on junior fellowship proposal scores","conf":"high","self":false,"doi":"10.1080/2330443X.2022.2086190","ciLow":0.2,"ciHigh":0.47,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Twenty-nine STEM panel members independently scored the 18 discussed junior fellowship proposals, each voting unless conflicted or absent. An ICC of 0.31 indicates poor inter-rater reliability.","vf":"unverified"},{"key":"YXXSYUJR","au":"Heyard, Rachel","y":2025,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"share of consensus reports falling within the 95% credible interval predicted by the Bayesian hierarchical model from the individual evaluation reports","estd":"other","v":0.564,"n":"381","k":"3","samp":"full-pool","blind":"unclear","agg":"panel-consensus","scale":"0-5 per criterion, converted to 0-100 total score","field":"multi-field","wr":"model prediction versus consensus score on grant proposals","conf":"med","self":false,"doi":"10.1371/journal.pone.0317772","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":true,"he":false,"ms":"For the 381 proposals of the 2015 Life Sciences panel, the largest sample in the study, 56.4 per cent of the consensus meeting scores fell inside the 95 per cent credible interval predicted from the experts' individual evaluation reports. Marked primary because the paper reports no pooled overall value for its central comparison.","vf":"unverified"},{"key":"YXXSYUJR","au":"Heyard, Rachel","y":2025,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"share of consensus reports falling within the 95% credible interval predicted by the Bayesian hierarchical model from the individual evaluation reports","estd":"other","v":0.504,"n":"355","k":"3","samp":"full-pool","blind":"unclear","agg":"panel-consensus","scale":"0-5 per criterion, converted to 0-100 total score","field":"multi-field","wr":"model prediction versus consensus score on grant proposals","conf":"med","self":false,"doi":"10.1371/journal.pone.0317772","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the 355 proposals of the 2019 Life Sciences panel, 50.4 per cent of the consensus meeting scores fell inside the 95 per cent credible interval predicted from the individual evaluation reports.","vf":"unverified"},{"key":"YXXSYUJR","au":"Heyard, Rachel","y":2025,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"share of consensus reports falling within the 95% credible interval predicted by the Bayesian hierarchical model from the individual evaluation reports","estd":"other","v":0.722,"n":"18","k":"3","samp":"full-pool","blind":"unclear","agg":"panel-consensus","scale":"0-5 per criterion, converted to 0-100 total score","field":"multi-field","wr":"model prediction versus consensus score on grant proposals","conf":"med","self":false,"doi":"10.1371/journal.pone.0317772","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the 18 proposals of the 2015 Mathematics panel, 72.2 per cent of the consensus meeting scores fell inside the 95 per cent credible interval predicted from the experts' individual evaluation reports. It shows how far the panel's agreed score matched what the individual scores implied.","vf":"unverified"},{"key":"YXXSYUJR","au":"Heyard, Rachel","y":2025,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"share of consensus reports falling within the 95% credible interval predicted by the Bayesian hierarchical model from the individual evaluation reports","estd":"other","v":0.688,"n":"16","k":"3","samp":"full-pool","blind":"unclear","agg":"panel-consensus","scale":"0-5 per criterion, converted to 0-100 total score","field":"multi-field","wr":"model prediction versus consensus score on grant proposals","conf":"med","self":false,"doi":"10.1371/journal.pone.0317772","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the 16 proposals of the 2019 Mathematics panel, 68.8 per cent of the consensus meeting scores fell inside the 95 per cent credible interval predicted from the individual evaluation reports.","vf":"unverified"},{"key":"YXXSYUJR","au":"Heyard, Rachel","y":2025,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"share of consensus reports falling within the 95% credible interval predicted by the Bayesian hierarchical model from the individual evaluation reports","estd":"other","v":0.272,"n":"114","k":"3","samp":"full-pool","blind":"unclear","agg":"panel-consensus","scale":"0-5 per criterion, converted to 0-100 total score","field":"multi-field","wr":"model prediction versus consensus score on grant proposals","conf":"med","self":false,"doi":"10.1371/journal.pone.0317772","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the 114 proposals of the 2015 Social Sciences and Humanities panel, only 27.2 per cent of the consensus meeting scores fell inside the 95 per cent credible interval predicted from the individual evaluation reports.","vf":"unverified"},{"key":"YXXSYUJR","au":"Heyard, Rachel","y":2025,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"share of consensus reports falling within the 95% credible interval predicted by the Bayesian hierarchical model from the individual evaluation reports","estd":"other","v":0.385,"n":"122","k":"3","samp":"full-pool","blind":"unclear","agg":"panel-consensus","scale":"0-5 per criterion, converted to 0-100 total score","field":"multi-field","wr":"model prediction versus consensus score on grant proposals","conf":"med","self":false,"doi":"10.1371/journal.pone.0317772","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the 122 proposals of the 2019 Social Sciences and Humanities panel, 38.5 per cent of the consensus meeting scores fell inside the 95 per cent credible interval predicted from the individual evaluation reports.","vf":"unverified"},{"key":"YXXSYUJR","au":"Heyard, Rachel","y":2025,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"share of agreement between Bayesian ranking recommendation and MSCA funding ranking group, exact agreement on the three-category grouping","estd":"percent agreement","v":0.882,"n":"381","k":"3","samp":"full-pool","blind":"unclear","agg":"panel-consensus","scale":"Accepted/Lottery/Rejected versus Main/Reserve/Rejected","field":"multi-field","wr":"Bayesian funding recommendation versus panel funding decision","conf":"med","self":false,"doi":"10.1371/journal.pone.0317772","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the 381 proposals of the 2015 Life Sciences panel, the funding group recommended by the Bayesian ranking matched the group assigned after the MSCA consensus meeting for 88.2 per cent of proposals.","vf":"unverified"},{"key":"YXXSYUJR","au":"Heyard, Rachel","y":2025,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"share of agreement between Bayesian ranking recommendation and MSCA funding ranking group, exact agreement on the three-category grouping","estd":"percent agreement","v":0.803,"n":"355","k":"3","samp":"full-pool","blind":"unclear","agg":"panel-consensus","scale":"Accepted/Lottery/Rejected versus Main/Reserve/Rejected","field":"multi-field","wr":"Bayesian funding recommendation versus panel funding decision","conf":"med","self":false,"doi":"10.1371/journal.pone.0317772","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the 355 proposals of the 2019 Life Sciences panel, the funding group recommended by the Bayesian ranking matched the group assigned after the MSCA consensus meeting for 80.3 per cent of proposals.","vf":"unverified"},{"key":"YXXSYUJR","au":"Heyard, Rachel","y":2025,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"share of agreement between Bayesian ranking recommendation and MSCA funding ranking group, exact agreement on the three-category grouping","estd":"percent agreement","v":0.889,"n":"18","k":"3","samp":"full-pool","blind":"unclear","agg":"panel-consensus","scale":"Accepted/Lottery/Rejected versus Main/Reserve/Rejected","field":"multi-field","wr":"Bayesian funding recommendation versus panel funding decision","conf":"med","self":false,"doi":"10.1371/journal.pone.0317772","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the 18 proposals of the 2015 Mathematics panel, the funding group recommended by the Bayesian ranking matched the group assigned after the MSCA consensus meeting for 88.9 per cent of proposals.","vf":"unverified"},{"key":"YXXSYUJR","au":"Heyard, Rachel","y":2025,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"share of agreement between Bayesian ranking recommendation and MSCA funding ranking group, exact agreement on the three-category grouping","estd":"percent agreement","v":0.938,"n":"16","k":"3","samp":"full-pool","blind":"unclear","agg":"panel-consensus","scale":"Accepted/Lottery/Rejected versus Main/Reserve/Rejected","field":"multi-field","wr":"Bayesian funding recommendation versus panel funding decision","conf":"med","self":false,"doi":"10.1371/journal.pone.0317772","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the 16 proposals of the 2019 Mathematics panel, the funding group recommended by the Bayesian ranking matched the group assigned after the MSCA consensus meeting for 93.8 per cent of proposals.","vf":"unverified"},{"key":"YXXSYUJR","au":"Heyard, Rachel","y":2025,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"share of agreement between Bayesian ranking recommendation and MSCA funding ranking group, exact agreement on the three-category grouping","estd":"percent agreement","v":0.825,"n":"114","k":"3","samp":"full-pool","blind":"unclear","agg":"panel-consensus","scale":"Accepted/Lottery/Rejected versus Main/Reserve/Rejected","field":"multi-field","wr":"Bayesian funding recommendation versus panel funding decision","conf":"med","self":false,"doi":"10.1371/journal.pone.0317772","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the 114 proposals of the 2015 Social Sciences and Humanities panel, the funding group recommended by the Bayesian ranking matched the group assigned after the MSCA consensus meeting for 82.5 per cent of proposals.","vf":"unverified"},{"key":"YXXSYUJR","au":"Heyard, Rachel","y":2025,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"share of agreement between Bayesian ranking recommendation and MSCA funding ranking group, exact agreement on the three-category grouping","estd":"percent agreement","v":0.861,"n":"122","k":"3","samp":"full-pool","blind":"unclear","agg":"panel-consensus","scale":"Accepted/Lottery/Rejected versus Main/Reserve/Rejected","field":"multi-field","wr":"Bayesian funding recommendation versus panel funding decision","conf":"med","self":false,"doi":"10.1371/journal.pone.0317772","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the 122 proposals of the 2019 Social Sciences and Humanities panel, the funding group recommended by the Bayesian ranking matched the group assigned after the MSCA consensus meeting for 86.1 per cent of proposals.","vf":"unverified"},{"key":"YXXSYUJR","au":"Heyard, Rachel","y":2025,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"share of official final ranks falling within the 95% credible interval of the rank expected from the Bayesian ranking","estd":"other","v":0.643,"n":"381","k":"3","samp":"full-pool","blind":"unclear","agg":"panel-consensus","scale":"unique final rank within panel, derived from 0-100 consensus scores","field":"multi-field","wr":"Bayesian expected rank versus official final rank of proposals","conf":"med","self":false,"doi":"10.1371/journal.pone.0317772","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the 381 proposals of the 2015 Life Sciences panel, 64.3 per cent of the official final ranks fell inside the 95 per cent credible interval of the expected Bayesian rank.","vf":"unverified"},{"key":"YXXSYUJR","au":"Heyard, Rachel","y":2025,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"share of official final ranks falling within the 95% credible interval of the rank expected from the Bayesian ranking","estd":"other","v":0.583,"n":"355","k":"3","samp":"full-pool","blind":"unclear","agg":"panel-consensus","scale":"unique final rank within panel, derived from 0-100 consensus scores","field":"multi-field","wr":"Bayesian expected rank versus official final rank of proposals","conf":"med","self":false,"doi":"10.1371/journal.pone.0317772","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the 355 proposals of the 2019 Life Sciences panel, 58.3 per cent of the official final ranks fell inside the 95 per cent credible interval of the expected Bayesian rank.","vf":"unverified"},{"key":"YXXSYUJR","au":"Heyard, Rachel","y":2025,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"share of official final ranks falling within the 95% credible interval of the rank expected from the Bayesian ranking","estd":"other","v":0.5,"n":"18","k":"3","samp":"full-pool","blind":"unclear","agg":"panel-consensus","scale":"unique final rank within panel, derived from 0-100 consensus scores","field":"multi-field","wr":"Bayesian expected rank versus official final rank of proposals","conf":"med","self":false,"doi":"10.1371/journal.pone.0317772","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the 18 proposals of the 2015 Mathematics panel, half of the official final ranks fell inside the 95 per cent credible interval of the rank expected from the individual evaluation reports.","vf":"unverified"},{"key":"YXXSYUJR","au":"Heyard, Rachel","y":2025,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"share of official final ranks falling within the 95% credible interval of the rank expected from the Bayesian ranking","estd":"other","v":0.812,"n":"16","k":"3","samp":"full-pool","blind":"unclear","agg":"panel-consensus","scale":"unique final rank within panel, derived from 0-100 consensus scores","field":"multi-field","wr":"Bayesian expected rank versus official final rank of proposals","conf":"med","self":false,"doi":"10.1371/journal.pone.0317772","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the 16 proposals of the 2019 Mathematics panel, 81.2 per cent of the official final ranks fell inside the 95 per cent credible interval of the expected Bayesian rank.","vf":"unverified"},{"key":"YXXSYUJR","au":"Heyard, Rachel","y":2025,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"share of official final ranks falling within the 95% credible interval of the rank expected from the Bayesian ranking","estd":"other","v":0.316,"n":"114","k":"3","samp":"full-pool","blind":"unclear","agg":"panel-consensus","scale":"unique final rank within panel, derived from 0-100 consensus scores","field":"multi-field","wr":"Bayesian expected rank versus official final rank of proposals","conf":"med","self":false,"doi":"10.1371/journal.pone.0317772","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the 114 proposals of the 2015 Social Sciences and Humanities panel, 31.6 per cent of the official final ranks fell inside the 95 per cent credible interval of the expected Bayesian rank.","vf":"unverified"},{"key":"YXXSYUJR","au":"Heyard, Rachel","y":2025,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"share of official final ranks falling within the 95% credible interval of the rank expected from the Bayesian ranking","estd":"other","v":0.525,"n":"122","k":"3","samp":"full-pool","blind":"unclear","agg":"panel-consensus","scale":"unique final rank within panel, derived from 0-100 consensus scores","field":"multi-field","wr":"Bayesian expected rank versus official final rank of proposals","conf":"med","self":false,"doi":"10.1371/journal.pone.0317772","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For the 122 proposals of the 2019 Social Sciences and Humanities panel, 52.5 per cent of the official final ranks fell inside the 95 per cent credible interval of the expected Bayesian rank.","vf":"unverified"},{"key":"YXXSYUJR","au":"Heyard, Rachel","y":2025,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"share of agreement on membership of the best ranked 10% of proposals, exact agreement","estd":"percent agreement","v":0.95,"n":"","k":"3","samp":"full-pool","blind":"unclear","agg":"panel-consensus","scale":"percentile groups of the final ranking","field":"multi-field","wr":"best 10% by Bayesian ranking versus official ranking","conf":"med","self":false,"doi":"10.1371/journal.pone.0317772","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Across the three panels and two calls, the proposals placed in the best ranked 10 per cent by the Bayesian ranking and by the official MSCA ranking overlapped by 84 to 95 per cent. This row records the upper bound of that reported range.","vf":"unverified"},{"key":"YXXSYUJR","au":"Heyard, Rachel","y":2025,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"share of agreement on membership of the best ranked 10% of proposals, exact agreement","estd":"percent agreement","v":0.84,"n":"","k":"3","samp":"full-pool","blind":"unclear","agg":"panel-consensus","scale":"percentile groups of the final ranking","field":"multi-field","wr":"best 10% by Bayesian ranking versus official ranking","conf":"med","self":false,"doi":"10.1371/journal.pone.0317772","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Across the three panels and two calls, the proposals placed in the best ranked 10 per cent by the Bayesian ranking and by the official MSCA ranking overlapped by 84 to 95 per cent. This row records the lower bound of that reported range.","vf":"unverified"},{"key":"S5ZNPKAX","au":"Hodgson, C","y":1997,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation coefficient","estd":"correlation","v":0.592,"n":"248","k":"2","samp":"special","blind":"single","agg":"panel-consensus","scale":"0-5 scientific merit, decimal committee-averaged score","field":"biomedical (cardiovascular)","wr":"two agencies' merit scores on same proposals","conf":"high","self":false,"doi":"10.1016/S0895-4356(97)00167-4","ciLow":0.511,"ciHigh":0.673,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two Canadian funding agencies independently scored the same 248 grant proposals for scientific merit on a 0-5 scale through their ordinary committee processes; a Pearson correlation of 0.592 indicates a moderate linear relationship between the two agencies' final scores.","vf":"unverified"},{"key":"S5ZNPKAX","au":"Hodgson, C","y":1997,"cx":"Grant","ob":"grant-proposal","fam":"kappa","form":"Cohen's kappa (unweighted)","estd":"kappa","v":0.444,"n":"248","k":"2","samp":"special","blind":"single","agg":"panel-consensus","scale":"clearly fundable (>=3.0) / not clearly fundable (<3.0)","field":"biomedical (cardiovascular)","wr":"two agencies' fund/not-fund dichotomy at 3.0","conf":"high","self":false,"doi":"10.1016/S0895-4356(97)00167-4","ciLow":0.328,"ciHigh":0.552,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"When proposals were dichotomised as clearly fundable or not at a score of 3.0, a Cohen kappa of 0.444 indicates moderate chance-corrected agreement between the two agencies on fundability; this is the agreement figure the paper's conclusions repeatedly cite as its headline result.","vf":"unverified"},{"key":"S5ZNPKAX","au":"Hodgson, C","y":1997,"cx":"Grant","ob":"grant-proposal","fam":"kappa","form":"Cohen's kappa (unweighted)","estd":"kappa","v":0.29,"n":"248","k":"2","samp":"special","blind":"single","agg":"panel-consensus","scale":"whole-digit ranges 1.0-1.9 to 4.0-4.9 (Poor to Very good)","field":"biomedical (cardiovascular)","wr":"two agencies' whole-digit merit categories","conf":"high","self":false,"doi":"10.1016/S0895-4356(97)00167-4","ciLow":0.198,"ciHigh":0.382,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Comparing the two agencies' scores collapsed into the four whole-digit merit ranges of the 0-5 scale, a Cohen kappa of 0.29 indicates only fair chance-corrected agreement between the two funding systems.","vf":"unverified"},{"key":"S5ZNPKAX","au":"Hodgson, C","y":1997,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"raw exact agreement on fundability dichotomy","estd":"percent agreement","v":0.725,"n":"248","k":"2","samp":"special","blind":"single","agg":"panel-consensus","scale":"clearly fundable (>=3.0) / not clearly fundable (<3.0)","field":"biomedical (cardiovascular)","wr":"two agencies agreeing on fund/not-fund","conf":"high","self":false,"doi":"10.1016/S0895-4356(97)00167-4","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The two agencies agreed on whether a proposal was clearly fundable (dichotomised at 3.0) for 72.5 percent of the 248 proposals, an uncorrected raw agreement figure; a later passage rounds the same result to 72.6 percent.","vf":"unverified"},{"key":"S5ZNPKAX","au":"Hodgson, C","y":1997,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"raw exact agreement on fundability dichotomy","estd":"percent agreement","v":0.742,"n":"128","k":"2","samp":"special","blind":"single","agg":"panel-consensus","scale":"clearly fundable (>=3.0) / not clearly fundable (<3.0)","field":"biomedical (cardiovascular)","wr":"two agencies agreeing on fundability, Ontario subsample","conf":"high","self":false,"doi":"10.1016/S0895-4356(97)00167-4","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"In the 128 Ontario-based proposal pairs accessible to the author, the two agencies agreed on fundability for 74.2 percent of cases, close to the full-sample raw agreement.","vf":"unverified"},{"key":"S5ZNPKAX","au":"Hodgson, C","y":1997,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"raw exact agreement within whole-digit ranges","estd":"percent agreement","v":0.528,"n":"248","k":"2","samp":"special","blind":"single","agg":"panel-consensus","scale":"whole-digit ranges 1.0-1.9 to 4.0-4.9 (Poor to Very good)","field":"biomedical (cardiovascular)","wr":"two agencies agreeing on whole-digit merit category","conf":"high","self":false,"doi":"10.1016/S0895-4356(97)00167-4","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The two agencies placed a proposal in the same whole-digit merit range for 52.8 percent of the 248 proposals, an uncorrected raw agreement figure.","vf":"unverified"},{"key":"S5ZNPKAX","au":"Hodgson, C","y":1997,"cx":"Grant","ob":"grant-proposal","fam":"weighted-kappa","form":"weighted kappa (including partial agreement)","estd":"weighted kappa","v":0.537,"n":"248","k":"2","samp":"special","blind":"single","agg":"panel-consensus","scale":"whole-digit ranges 1.0-1.9 to 4.0-4.9 (Poor to Very good)","field":"biomedical (cardiovascular)","wr":"two agencies' whole-digit merit categories (weighted)","conf":"high","self":false,"doi":"10.1016/S0895-4356(97)00167-4","ciLow":0.449,"ciHigh":0.616,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"A weighted kappa crediting partial agreement between the two agencies' whole-digit merit categories was 0.537 in the text, indicating moderate chance-corrected agreement; Table 2 reports 0.527 for the same coefficient with identical SE and CI, an internal inconsistency the paper does not resolve.","vf":"unverified"},{"key":"XTSCEH5T","au":"Hogan, Thomas P.","y":2021,"cx":"General","ob":"other","fam":"correlation","form":"Pearson correlation between the average ratings of the two independent reviews of each test","estd":"correlation","v":0.344,"n":"50","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1-5, 5 = entirely or almost entirely positive conclusion","field":"psychometrics","wr":"MMY reviewers on overall test quality","conf":"high","self":false,"doi":"10.1080/08957347.2021.1890742","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Each of 50 randomly sampled tests in the 19th Mental Measurements Yearbook had two independent published reviews, and trained raters scored every review summary from 1 to 5 for how positive its conclusion was. The correlation of 0.34 between the paired reviews' ratings shows weak agreement between two independent reviewers about a test's quality.","vf":"unverified"},{"key":"XTSCEH5T","au":"Hogan, Thomas P.","y":2021,"cx":"General","ob":"other","fam":"percent-agreement","form":"difference of 1.0 point or less between the two reviews' average ratings (within one point, not exact agreement)","estd":"percent agreement","v":0.58,"n":"50","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1-5, 5 = entirely or almost entirely positive conclusion","field":"psychometrics","wr":"MMY reviewers on overall test quality","conf":"high","self":false,"doi":"10.1080/08957347.2021.1890742","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the same 50 tests, the quality ratings of the two independent reviews differed by one point or less for 58 per cent of the tests. The paper additionally reports differences of two or more points for 22 per cent of tests and of three or more points for 10 per cent.","vf":"unverified"},{"key":"FPDF4D3R","au":"Howard, Louise M.","y":1998,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"kappa (SPSS), one, two or three v. four or five","estd":"kappa","v":0.19,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"ranks 1-3 vs ranks 4-5 (dichotomised)","field":"psychiatry","wr":"assessors ranking journal manuscripts","conf":"high","self":false,"doi":"10.1192/bjp.173.2.110","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Agreement between two external assessors on whether a manuscript was ranked one to three versus four or five; kappa 0.19 indicates poor agreement.","vf":"unverified"},{"key":"FPDF4D3R","au":"Howard, Louise M.","y":1998,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"kappa (SPSS), ranks one or two v. three, four or five","estd":"kappa","v":0.16,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"ranks 1-2 vs ranks 3-5 (dichotomised)","field":"psychiatry","wr":"assessors ranking journal manuscripts","conf":"high","self":false,"doi":"10.1192/bjp.173.2.110","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Agreement between two external assessors on whether a manuscript was ranked one or two versus three, four or five; kappa 0.16 indicates poor agreement.","vf":"unverified"},{"key":"FPDF4D3R","au":"Howard, Louise M.","y":1998,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"kappa (SPSS), 1-4 v. five (reject vs not)","estd":"kappa","v":0.27,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"ranks 1-4 vs rank 5 (reject; dichotomised)","field":"psychiatry","wr":"assessors ranking journal manuscripts","conf":"high","self":false,"doi":"10.1192/bjp.173.2.110","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Agreement between two external assessors on whether a manuscript should be rejected (rank five) versus not; kappa 0.27 indicates fair agreement, the best of the collapsed cut-points.","vf":"unverified"},{"key":"FPDF4D3R","au":"Howard, Louise M.","y":1998,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"kappa (SPSS), rank one v. others","estd":"kappa","v":0.11,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"rank 1 vs ranks 2-6 (dichotomised)","field":"psychiatry","wr":"assessors ranking journal manuscripts","conf":"high","self":false,"doi":"10.1192/bjp.173.2.110","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Agreement between two external assessors on whether a manuscript was ranked one (publish unamended) versus any other rank; kappa 0.11 indicates poor agreement.","vf":"unverified"},{"key":"FPDF4D3R","au":"Howard, Louise M.","y":1998,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"kappa (SPSS), whether accepted or rejected using categories 1-6","estd":"kappa","v":0.1,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1-6 ranking, 1=publish unamended, 5=reject, 6=other","field":"psychiatry","wr":"assessors ranking journal manuscripts","conf":"high","self":false,"doi":"10.1192/bjp.173.2.110","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":false,"ms":"Two external assessors ranked manuscripts submitted to the British Journal of Psychiatry on a 1-6 scale; an overall kappa of 0.10 indicates poor chance-corrected agreement between them.","vf":"unverified"},{"key":"QUNM7567","au":"Ingelfinger, Franz J.","y":1974,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement (both reviewers gave identical ratings, i.e., A-A, B-B, etc.)","estd":"percent agreement","v":0.505,"n":"95","k":"2","samp":"funded-only","blind":"single","agg":"single-rater","scale":"A, B, C or D","field":"biomedical","wr":"two reviewers on accepted manuscript ratings","conf":"high","self":false,"doi":"10.1016/0002-9343(74)90635-4","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Among the 95 subsequently accepted manuscripts, the two reviewers gave identical A to D ratings for 50.5 per cent of papers.","vf":"unverified"},{"key":"QUNM7567","au":"Ingelfinger, Franz J.","y":1974,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement (both reviewers gave identical ratings, i.e., A-A, B-B, etc.)","estd":"percent agreement","v":0.418,"n":"496","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"A, B, C or D","field":"biomedical","wr":"two reviewers on submitted manuscript quality ratings","conf":"high","self":false,"doi":"10.1016/0002-9343(74)90635-4","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Pairs of reviewers independently rated 496 consecutive manuscripts submitted to the NEJM on an A to D quality scale; the two reviewers gave identical ratings for 41.8 per cent of papers, against roughly 30 per cent expected by chance.","vf":"unverified"},{"key":"QUNM7567","au":"Ingelfinger, Franz J.","y":1974,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement (both reviewers gave identical ratings, i.e., A-A, B-B, etc.)","estd":"percent agreement","v":0.397,"n":"401","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"A, B, C or D","field":"biomedical","wr":"two reviewers on rejected manuscript ratings","conf":"high","self":false,"doi":"10.1016/0002-9343(74)90635-4","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Among the 401 subsequently rejected manuscripts, the two reviewers gave identical A to D ratings for 39.7 per cent of papers.","vf":"unverified"},{"key":"LEGPWSWB","au":"Jackson, Jeffrey L","y":2011,"cx":"Journal","ob":"journal-manuscript","fam":"Cronbach-alpha","form":"Cronbach alpha","estd":"Cronbach alpha","v":0.79,"n":"379","k":"1","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"five domains 5-point; study design 3-point","field":"general medicine","wr":"six manuscript quality domains within each review","conf":"med","self":false,"doi":"10.1371/journal.pone.0022475","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Each external reviewer scored manuscripts on six quality domains (five 5-point, one 3-point); a Cronbach alpha of 0.79 indicates good internal consistency across the domains within a review. The item set is the six domains scored by a single reviewer.","vf":"unverified"},{"key":"LEGPWSWB","au":"Jackson, Jeffrey L","y":2011,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficients","estd":"ICC (single/unspec)","v":0.13,"n":"379","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"mean of six domains: five 5-point, one 3-point","field":"general medicine","wr":"reviewer average manuscript quality ratings","conf":"med","self":false,"doi":"10.1371/journal.pone.0022475","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"External reviewers independently rated the same manuscripts on an average of six quality domains; the reported intraclass correlations ranged from 0.09 to 0.13, and this row carries the range maximum, so agreement between individual reviewers was low across the whole range.","vf":"unverified"},{"key":"LEGPWSWB","au":"Jackson, Jeffrey L","y":2011,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficients","estd":"ICC (single/unspec)","v":0.09,"n":"379","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"mean of six domains: five 5-point, one 3-point","field":"general medicine","wr":"reviewer average manuscript quality ratings","conf":"med","self":false,"doi":"10.1371/journal.pone.0022475","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":true,"ms":"External reviewers independently rated the same manuscripts on an average of six quality domains; the reported intraclass correlations ranged from 0.09 to 0.13, and this row carries the range minimum. No single overall value was reported, so this headline range endpoint is marked primary. Reviewers per manuscript varied from one to four, mostly three.","vf":"unverified"},{"key":"LEGPWSWB","au":"Jackson, Jeffrey L","y":2011,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"quadratic kappas","estd":"weighted kappa","v":0.15,"n":"379","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"reject / major revision / minor revision / conditional accept / accept as is","field":"general medicine","wr":"reviewer manuscript publication recommendations","conf":"med","self":false,"doi":"10.1371/journal.pone.0022475","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Pairs of external reviewers gave publication recommendations for the same manuscripts on a five-category scale; the quadratic weighted kappas between reviewers ranged from 0.11 to 0.15, and this row carries the range maximum, so agreement was low across the whole range.","vf":"unverified"},{"key":"LEGPWSWB","au":"Jackson, Jeffrey L","y":2011,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"quadratic kappas","estd":"weighted kappa","v":0.11,"n":"379","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"reject / major revision / minor revision / conditional accept / accept as is","field":"general medicine","wr":"reviewer manuscript publication recommendations","conf":"med","self":false,"doi":"10.1371/journal.pone.0022475","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Pairs of external reviewers gave publication recommendations for the same manuscripts on a five-category scale; the quadratic weighted kappas between reviewers ranged from 0.11 to 0.15, and this row carries the range minimum, indicating low chance-corrected agreement.","vf":"unverified"},{"key":"XZ8J64TH","au":"Jayasinghe, U. W.","y":2001,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability stepped up with the Spearman-Brown equation to an average of 4.3 external reviewers","estd":"ICC (average)","v":0.44,"n":"2331","k":"4.3","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"scale of 1 to 100, quality of the proposal","field":"multi-field (all disciplines except medical research)","wr":"external reviewers on grant proposal quality","conf":"high","self":false,"doi":"10.3102/01623737023004343","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Stepping the single-rater project reliability up to the mean of 4.3 external reviewers gives 0.44. This is the reliability of the averaged project rating that each proposal actually carried into the panel stage.","vf":"unverified"},{"key":"XZ8J64TH","au":"Jayasinghe, U. W.","y":2001,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability stepped up with the Spearman-Brown equation to an average of 4.3 external reviewers","estd":"ICC (average)","v":0.53,"n":"2331","k":"4.3","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"scale of 1 to 100, research track record of the team","field":"multi-field (all disciplines except medical research)","wr":"external reviewers on research team track record","conf":"high","self":false,"doi":"10.3102/01623737023004343","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":true,"ms":"Averaging the researcher (track record) ratings of the 4.3 external reviewers who typically assessed each proposal yields a reliability of 0.53. The abstract names this value as the study's headline result on the reliability of ARC peer review.","vf":"unverified"},{"key":"XZ8J64TH","au":"Jayasinghe, U. W.","y":2001,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"correlation (r) between committee rating and one reviewer's rating; coefficient type not stated","estd":"correlation","v":0.595,"n":"346","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"scale of 1 to 100 global reviewer and panel ratings","field":"multi-field (all disciplines except medical research)","wr":"Australian reviewer against ARC committee rating","conf":"high","self":false,"doi":"10.3102/01623737023004343","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"In the same 346 proposals, the randomly chosen Australian reviewer's rating correlated 0.595 with the committee rating, which the authors read as North American reviews being somewhat less reliable.","vf":"unverified"},{"key":"XZ8J64TH","au":"Jayasinghe, U. W.","y":2001,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"correlation (r) between final panel rating and one reviewer's average rating; coefficient type not stated","estd":"correlation","v":0.61,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 to 100 weighted average of project and researcher ratings","field":"multi-field (all disciplines except medical research)","wr":"first-round reviewer against ARC panel rating","conf":"high","self":false,"doi":"10.3102/01623737023004343","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Reviewers invited in the first round, assumed to be the closest experts, produced ratings correlating 0.61 with the final panel rating of the same proposals.","vf":"unverified"},{"key":"XZ8J64TH","au":"Jayasinghe, U. W.","y":2001,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"correlation (r) between committee rating and one reviewer's rating; coefficient type not stated","estd":"correlation","v":0.533,"n":"346","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"scale of 1 to 100 global reviewer and panel ratings","field":"multi-field (all disciplines except medical research)","wr":"North American reviewer against ARC committee rating","conf":"high","self":false,"doi":"10.3102/01623737023004343","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"For 346 proposals that had both a North American and an Australian reviewer, one reviewer of each nationality was drawn at random. The randomly chosen North American reviewer's rating correlated 0.533 with the committee rating.","vf":"unverified"},{"key":"XZ8J64TH","au":"Jayasinghe, U. W.","y":2001,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"correlation (r) between final panel rating and reviewer rating; coefficient type not stated","estd":"correlation","v":0.62,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"scale of 1 to 100 global reviewer and panel ratings","field":"multi-field (all disciplines except medical research)","wr":"panel-nominated reviewer against ARC panel rating","conf":"high","self":false,"doi":"10.3102/01623737023004343","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"The first panel-nominated reviewer's ratings correlated 0.62 with the ARC panel's final rating of the same proposals, higher than the corresponding figure for applicant-nominated reviewers.","vf":"unverified"},{"key":"XZ8J64TH","au":"Jayasinghe, U. W.","y":2001,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"correlation (r) between final panel rating and reviewer rating; coefficient type not stated","estd":"correlation","v":0.61,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"scale of 1 to 100 global reviewer and panel ratings","field":"multi-field (all disciplines except medical research)","wr":"panel-nominated reviewer against ARC panel rating","conf":"high","self":false,"doi":"10.3102/01623737023004343","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"The second panel-nominated reviewer's ratings correlated 0.61 with the ARC panel's final rating of the same proposals.","vf":"unverified"},{"key":"XZ8J64TH","au":"Jayasinghe, U. W.","y":2001,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"correlation (r) between final panel rating and reviewer rating; coefficient type not stated","estd":"correlation","v":0.59,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"scale of 1 to 100 global reviewer and panel ratings","field":"multi-field (all disciplines except medical research)","wr":"panel-nominated reviewer against ARC panel rating","conf":"high","self":false,"doi":"10.3102/01623737023004343","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"The third panel-nominated reviewer's ratings correlated 0.59 with the ARC panel's final rating of the same proposals.","vf":"unverified"},{"key":"XZ8J64TH","au":"Jayasinghe, U. W.","y":2001,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"correlation (r) between final panel rating and reviewer rating; coefficient type not stated","estd":"correlation","v":0.52,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"scale of 1 to 100 global reviewer and panel ratings","field":"multi-field (all disciplines except medical research)","wr":"applicant-nominated reviewer against ARC panel rating","conf":"high","self":false,"doi":"10.3102/01623737023004343","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Ratings by the reviewer nominated by the applicant correlated 0.52 with the ARC panel's final rating of the same proposals. The panel had already seen those reviews, so the two judgments are not independent.","vf":"unverified"},{"key":"XZ8J64TH","au":"Jayasinghe, U. W.","y":2001,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"correlation (r) between final panel rating and one reviewer's average rating; coefficient type not stated","estd":"correlation","v":0.56,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 to 100 weighted average of project and researcher ratings","field":"multi-field (all disciplines except medical research)","wr":"replacement reviewer against ARC panel rating","conf":"high","self":false,"doi":"10.3102/01623737023004343","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Reviewers approached only after the first round, typically as replacements, produced ratings correlating 0.56 with the final panel rating of the same proposals.","vf":"unverified"},{"key":"XZ8J64TH","au":"Jayasinghe, U. W.","y":2001,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"correlation (r); coefficient type not stated","estd":"correlation","v":0.95,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 to 100 weighted average (.6 project, .4 researcher)","field":"multi-field (all disciplines except medical research)","wr":"panel score against averaged external reviewer score","conf":"high","self":false,"doi":"10.3102/01623737023004343","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"The weighted average of the external reviewer scores for each proposal correlated 0.95 with the final score the ARC panel assigned afterwards. The panel could revise the reviewer average but in practice almost never changed it, so the two are not independent judgments.","vf":"unverified"},{"key":"XZ8J64TH","au":"Jayasinghe, U. W.","y":2001,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"correlation (r) among the three panel-nominated reviewers; pairing and rating type not stated","estd":"correlation","v":0.22,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"scale of 1 to 100 global reviewer ratings","field":"multi-field (all disciplines except medical research)","wr":"pairs of panel-nominated reviewers on the same proposals","conf":"high","self":false,"doi":"10.3102/01623737023004343","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The first of three reported correlations between pairs of panel-nominated external reviewers of the same proposals was 0.22. The paper does not say which reviewer pair each value refers to.","vf":"unverified"},{"key":"XZ8J64TH","au":"Jayasinghe, U. W.","y":2001,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"correlation (r) among the three panel-nominated reviewers; pairing and rating type not stated","estd":"correlation","v":0.21,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"scale of 1 to 100 global reviewer ratings","field":"multi-field (all disciplines except medical research)","wr":"pairs of panel-nominated reviewers on the same proposals","conf":"high","self":false,"doi":"10.3102/01623737023004343","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The second of three reported correlations between pairs of panel-nominated external reviewers of the same proposals was 0.21.","vf":"unverified"},{"key":"XZ8J64TH","au":"Jayasinghe, U. W.","y":2001,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"correlation (r) among the three panel-nominated reviewers; pairing and rating type not stated","estd":"correlation","v":0.22,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"scale of 1 to 100 global reviewer ratings","field":"multi-field (all disciplines except medical research)","wr":"pairs of panel-nominated reviewers on the same proposals","conf":"high","self":false,"doi":"10.3102/01623737023004343","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The third of three reported correlations between pairs of panel-nominated external reviewers of the same proposals was 0.22.","vf":"unverified"},{"key":"XZ8J64TH","au":"Jayasinghe, U. W.","y":2001,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"correlation (r) between two reviewers' ratings; coefficient type and rating type not stated","estd":"correlation","v":0.2,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"scale of 1 to 100 global reviewer ratings","field":"multi-field (all disciplines except medical research)","wr":"pairs of external reviewers on the same proposals","conf":"high","self":false,"doi":"10.3102/01623737023004343","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Ratings by the reviewer the applicant nominated were correlated with those of the first panel-selected reviewer of the same proposal, giving 0.20. The paper reports this as evidence that researcher-nominated reviews are less reliable.","vf":"unverified"},{"key":"XZ8J64TH","au":"Jayasinghe, U. W.","y":2001,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"correlation (r) between two reviewers' ratings; coefficient type and rating type not stated","estd":"correlation","v":0.15,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"scale of 1 to 100 global reviewer ratings","field":"multi-field (all disciplines except medical research)","wr":"pairs of external reviewers on the same proposals","conf":"high","self":false,"doi":"10.3102/01623737023004343","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Ratings by the applicant-nominated reviewer correlated 0.15 with those of the second panel-selected reviewer of the same proposal.","vf":"unverified"},{"key":"XZ8J64TH","au":"Jayasinghe, U. W.","y":2001,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"correlation (r) between two reviewers' ratings; coefficient type and rating type not stated","estd":"correlation","v":0.16,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"scale of 1 to 100 global reviewer ratings","field":"multi-field (all disciplines except medical research)","wr":"pairs of external reviewers on the same proposals","conf":"high","self":false,"doi":"10.3102/01623737023004343","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Ratings by the applicant-nominated reviewer correlated 0.16 with those of the third panel-selected reviewer of the same proposal.","vf":"unverified"},{"key":"XZ8J64TH","au":"Jayasinghe, U. W.","y":2001,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability of overall external reviewer rating from ANOVA, citing Shrout & Fleiss (1979)","estd":"ICC (single/unspec)","v":0.223,"n":"","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1 to 100 weighted average of project and researcher ratings","field":"multi-field (all disciplines except medical research)","wr":"first-round reviewers on overall proposal rating","conf":"high","self":false,"doi":"10.3102/01623737023004343","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For reviewers invited in the first round, the reliability of a single reviewer's overall rating was 0.223, marginally above that of later-invited reviewers.","vf":"unverified"},{"key":"XZ8J64TH","au":"Jayasinghe, U. W.","y":2001,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability of overall external reviewer rating from ANOVA, citing Shrout & Fleiss (1979)","estd":"ICC (single/unspec)","v":0.206,"n":"","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1 to 100 weighted average of project and researcher ratings","field":"multi-field (all disciplines except medical research)","wr":"replacement reviewers on overall proposal rating","conf":"high","self":false,"doi":"10.3102/01623737023004343","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For reviewers approached only after the first round, the reliability of a single reviewer's overall rating was 0.206.","vf":"unverified"},{"key":"XZ8J64TH","au":"Jayasinghe, U. W.","y":2001,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability (intra-proposal correlation) derived from ANOVA, citing Shrout & Fleiss (1979)","estd":"ICC (single/unspec)","v":0.154,"n":"2331","k":"4.3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"scale of 1 to 100, quality of the proposal","field":"multi-field (all disciplines except medical research)","wr":"external reviewers on grant proposal quality","conf":"high","self":false,"doi":"10.3102/01623737023004343","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"External reviewers, on average 4.3 per proposal, rated the quality of 2,331 Australian Research Council proposals on a 1 to 100 scale. The reliability of one reviewer's project rating was 0.15, meaning a single reviewer's judgment reproduces very little of the between-proposal differences.","vf":"unverified"},{"key":"XZ8J64TH","au":"Jayasinghe, U. W.","y":2001,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability (r11) from ANOVA, citing Shrout & Fleiss (1979)","estd":"ICC (single/unspec)","v":0.159,"n":"","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"scale of 1 to 100, quality of the proposal","field":"multi-field (all disciplines except medical research)","wr":"experienced reviewers on grant proposal quality","conf":"high","self":false,"doi":"10.3102/01623737023004343","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Among external reviewers who assessed three or more proposals, the reliability of one reviewer's project rating was 0.159, slightly higher than for reviewers who assessed fewer proposals.","vf":"unverified"},{"key":"XZ8J64TH","au":"Jayasinghe, U. W.","y":2001,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability (r11) from ANOVA, citing Shrout & Fleiss (1979)","estd":"ICC (single/unspec)","v":0.147,"n":"","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"scale of 1 to 100, quality of the proposal","field":"multi-field (all disciplines except medical research)","wr":"occasional reviewers on grant proposal quality","conf":"high","self":false,"doi":"10.3102/01623737023004343","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Among external reviewers who assessed fewer than three proposals, the reliability of one reviewer's project rating was 0.147.","vf":"unverified"},{"key":"XZ8J64TH","au":"Jayasinghe, U. W.","y":2001,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability (intra-proposal correlation) derived from ANOVA, citing Shrout & Fleiss (1979)","estd":"ICC (single/unspec)","v":0.21,"n":"2331","k":"4.3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"scale of 1 to 100, research track record of the team","field":"multi-field (all disciplines except medical research)","wr":"external reviewers on research team track record","conf":"high","self":false,"doi":"10.3102/01623737023004343","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"External reviewers rated the research track record of the applicant team on a 1 to 100 scale for 2,331 proposals. The reliability of a single reviewer's researcher rating was 0.21, slightly higher than for the proposal itself but still very low.","vf":"unverified"},{"key":"XZ8J64TH","au":"Jayasinghe, U. W.","y":2001,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability (r11) from ANOVA, citing Shrout & Fleiss (1979)","estd":"ICC (single/unspec)","v":0.267,"n":"","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"scale of 1 to 100, research track record of the team","field":"multi-field (all disciplines except medical research)","wr":"experienced reviewers on research team track record","conf":"high","self":false,"doi":"10.3102/01623737023004343","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Among external reviewers who assessed three or more proposals, the reliability of one reviewer's researcher rating was 0.267, the highest single-rater value reported in the study.","vf":"unverified"},{"key":"XZ8J64TH","au":"Jayasinghe, U. W.","y":2001,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability (r11) from ANOVA, citing Shrout & Fleiss (1979)","estd":"ICC (single/unspec)","v":0.192,"n":"","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"scale of 1 to 100, research track record of the team","field":"multi-field (all disciplines except medical research)","wr":"occasional reviewers on research team track record","conf":"high","self":false,"doi":"10.3102/01623737023004343","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Among external reviewers who assessed fewer than three proposals, the reliability of one reviewer's researcher rating was 0.192.","vf":"unverified"},{"key":"KWATA3RR","au":"Jayasinghe, Upali W.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"correlation between final panel score and average assessor rating","estd":"correlation","v":0.95,"n":"","k":"","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"1-100","field":"multi-field","wr":"panel final score vs assessor mean on same proposals","conf":"high","self":false,"doi":"10.1007/s11192-006-0171-4","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"restricted-other","pr":false,"he":false,"ms":"In the traditional ARC process the panel could revise scores after reviewing assessor ratings and applicant rejoinders, but the final panel score correlated 0.95 with the average assessor rating, showing the panel rarely changed the assessor-based scores.","vf":"unverified"},{"key":"KWATA3RR","au":"Jayasinghe, Upali W.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"reliability of mean of 4.3 raters via Spearman-Brown equation","estd":"ICC (average)","v":0.643,"n":"196","k":"2","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"-4 to +4, nine-point rank-derived","field":"multi-field","wr":"mean of 4.3 readers on project quality scores","conf":"high","self":false,"doi":"10.1007/s11192-006-0171-4","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Spearman-Brown projection of the reader-system project-rating reliability to the mean of 4.3 readers gives 0.643, matching the average assessor count of the traditional process; observed data had two readers per proposal. This is the abstract's headline reader-system figure.","vf":"unverified"},{"key":"KWATA3RR","au":"Jayasinghe, Upali W.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability (intra-proposal correlation) from multilevel cross-classified model","estd":"ICC (single/unspec)","v":0.296,"n":"196","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"-4 to +4, nine-point rank-derived","field":"multi-field","wr":"readers on grant proposal (project quality) scores","conf":"high","self":false,"doi":"10.1007/s11192-006-0171-4","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"In the 1997 ARC reader-system trial, 19 expert readers each read all proposals in their subdiscipline (two per subdiscipline, three in one); the single-reader reliability of project ratings across 196 proposals was 0.296. The 860 ratings in the data file cover both project and researcher response variables.","vf":"unverified"},{"key":"KWATA3RR","au":"Jayasinghe, Upali W.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"reliability of mean of 4.3 raters via Spearman-Brown equation","estd":"ICC (average)","v":0.881,"n":"196","k":"2","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"-4 to +4, nine-point rank-derived","field":"multi-field","wr":"mean of 4.3 readers on researcher track-record scores","conf":"high","self":false,"doi":"10.1007/s11192-006-0171-4","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Spearman-Brown projection of the reader-system researcher-rating reliability to the mean of 4.3 readers gives 0.881, the highest reliability reported in the study; observed data had two readers per proposal.","vf":"unverified"},{"key":"KWATA3RR","au":"Jayasinghe, Upali W.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability (intra-proposal correlation) from multilevel cross-classified model","estd":"ICC (single/unspec)","v":0.634,"n":"196","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"-4 to +4, nine-point rank-derived","field":"multi-field","wr":"readers on researcher/team track-record scores","conf":"high","self":false,"doi":"10.1007/s11192-006-0171-4","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"In the ARC reader-system trial, the single-reader reliability of researcher (track record) ratings across 196 proposals rated by 19 readers was 0.634, higher than the corresponding project rating reliability.","vf":"unverified"},{"key":"KWATA3RR","au":"Jayasinghe, Upali W.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"reliability of mean of 4.3 raters via Spearman-Brown equation","estd":"ICC (average)","v":0.475,"n":"2331","k":"4.3","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1-100","field":"multi-field","wr":"mean of 4.3 assessors on project quality scores (all disciplines)","conf":"high","self":false,"doi":"10.1007/s11192-006-0171-4","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Reliability of the mean of 4.3 assessors (Spearman-Brown) for project ratings under the traditional all-disciplines process was 0.475, the figure the abstract calls unacceptably low.","vf":"unverified"},{"key":"KWATA3RR","au":"Jayasinghe, Upali W.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability (intra-proposal correlation) from multilevel cross-classified model","estd":"ICC (single/unspec)","v":0.174,"n":"2331","k":"4.3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-100","field":"multi-field","wr":"external assessors on project quality scores (all disciplines)","conf":"high","self":false,"doi":"10.1007/s11192-006-0171-4","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For 2,331 externally reviewed proposals across all nine ARC discipline panels in 1996 (average 4.3 assessors each, from a pool of 6,233), the single-assessor reliability of project ratings was 0.174.","vf":"unverified"},{"key":"KWATA3RR","au":"Jayasinghe, Upali W.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"reliability of mean of 4.3 raters via Spearman-Brown equation","estd":"ICC (average)","v":0.572,"n":"2331","k":"4.3","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1-100","field":"multi-field","wr":"mean of 4.3 assessors on researcher track-record scores (all disciplines)","conf":"high","self":false,"doi":"10.1007/s11192-006-0171-4","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Reliability of the mean of 4.3 assessors (Spearman-Brown) for researcher ratings under the traditional all-disciplines process was 0.572.","vf":"unverified"},{"key":"KWATA3RR","au":"Jayasinghe, Upali W.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability (intra-proposal correlation) from multilevel cross-classified model","estd":"ICC (single/unspec)","v":0.237,"n":"2331","k":"4.3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-100","field":"multi-field","wr":"external assessors on researcher track-record scores (all disciplines)","conf":"high","self":false,"doi":"10.1007/s11192-006-0171-4","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For 2,331 proposals across all ARC discipline panels under the traditional 1996 process, the single-assessor reliability of researcher (track record) ratings was 0.237.","vf":"unverified"},{"key":"KWATA3RR","au":"Jayasinghe, Upali W.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"reliability of mean of 4.3 raters via Spearman-Brown equation","estd":"ICC (average)","v":0.449,"n":"385","k":"4.3","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1-100","field":"multi-field","wr":"mean of 4.3 assessors on project quality scores (social sciences)","conf":"high","self":false,"doi":"10.1007/s11192-006-0171-4","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Spearman-Brown projection of the traditional social-science project-rating reliability to the mean of 4.3 assessors gives 0.449.","vf":"unverified"},{"key":"KWATA3RR","au":"Jayasinghe, Upali W.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability (intra-proposal correlation) from multilevel cross-classified model","estd":"ICC (single/unspec)","v":0.159,"n":"385","k":"4.3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-100","field":"multi-field","wr":"external assessors on project quality scores (social sciences)","conf":"high","self":false,"doi":"10.1007/s11192-006-0171-4","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"A pool of 1,110 external assessors gave 1,643 project ratings to 385 social-science proposals (average 4.3 assessors per proposal) in the 1996 traditional ARC round; the single-assessor reliability of project ratings on the 1-100 scale was 0.159.","vf":"unverified"},{"key":"KWATA3RR","au":"Jayasinghe, Upali W.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"reliability of mean of 4.3 raters via Spearman-Brown equation","estd":"ICC (average)","v":0.591,"n":"385","k":"4.3","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1-100","field":"multi-field","wr":"mean of 4.3 assessors on researcher track-record scores (social sciences)","conf":"high","self":false,"doi":"10.1007/s11192-006-0171-4","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Spearman-Brown projection of the traditional social-science researcher-rating reliability to the mean of 4.3 assessors gives 0.591.","vf":"unverified"},{"key":"KWATA3RR","au":"Jayasinghe, Upali W.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability (intra-proposal correlation) from multilevel cross-classified model","estd":"ICC (single/unspec)","v":0.252,"n":"385","k":"4.3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-100","field":"multi-field","wr":"external assessors on researcher track-record scores (social sciences)","conf":"high","self":false,"doi":"10.1007/s11192-006-0171-4","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For the 385 social-science proposals in the 1996 traditional ARC round, the single-assessor reliability of researcher (track record) ratings was 0.252, higher than the project rating reliability but still low.","vf":"unverified"},{"key":"YYCAX8PN","au":"Jerrim, John","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach's alpha","estd":"Cronbach alpha","v":0.48,"n":"","k":"5","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1 = poor to 6 = outstanding","field":"social science","wr":"five reviewers' overall grades of the same proposal","conf":"med","self":false,"doi":"10.1080/03623319.2020.1728506","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"For the rare proposals carrying five peer reviews, Cronbach's alpha across the five reviewers' overall grades was 0.48. Even with five reviews the authors judge internal consistency to be low.","vf":"unverified"},{"key":"YYCAX8PN","au":"Jerrim, John","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach's alpha","estd":"Cronbach alpha","v":0.53,"n":"","k":"5","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1 = poor to 6 = outstanding","field":"social science","wr":"five independent reviewers' grades of the same proposal","conf":"med","self":false,"doi":"10.1080/03623319.2020.1728506","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Excluding applicant-nominated reviewers, Cronbach's alpha across five reviewers' overall grades of the same proposal rose to 0.53. The figure is reported only in a footnote.","vf":"unverified"},{"key":"YYCAX8PN","au":"Jerrim, John","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach's alpha","estd":"Cronbach alpha","v":0.44,"n":"","k":"4","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1 = poor to 6 = outstanding","field":"social science","wr":"four reviewers' overall grades of the same proposal","conf":"med","self":false,"doi":"10.1080/03623319.2020.1728506","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"For proposals carrying four peer reviews, Cronbach's alpha across the four reviewers' overall grades was 0.44, estimating the reliability of their combined judgment. The authors read this as low internal consistency, on the boundary of the poor and unacceptable classifications.","vf":"unverified"},{"key":"YYCAX8PN","au":"Jerrim, John","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach's alpha","estd":"Cronbach alpha","v":0.48,"n":"","k":"4","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1 = poor to 6 = outstanding","field":"social science","wr":"four independent reviewers' grades of the same proposal","conf":"med","self":false,"doi":"10.1080/03623319.2020.1728506","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Excluding applicant-nominated reviewers, Cronbach's alpha across four reviewers' overall grades of the same proposal rose to 0.48. The figure is reported only in a footnote.","vf":"unverified"},{"key":"YYCAX8PN","au":"Jerrim, John","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"polychoric correlation between each pair of reviewers, averaged across all pairs, weighted by sample size","estd":"correlation","v":0.07,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1 = poor to 6 = outstanding","field":"social science","wr":"independent versus applicant-nominated reviewers on proposals","conf":"med","self":false,"doi":"10.1080/03623319.2020.1728506","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The Results text states that the correlation between the scores of independent and applicant-nominated reviewers of the same proposal falls to 0.07, indicating they are barely associated at all. This conflicts with the 0.11 tabulated for the same comparison in Table 3; both reported values are retained.","vf":"unverified"},{"key":"YYCAX8PN","au":"Jerrim, John","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"weighted-kappa","form":"Kappa statistic","estd":"weighted kappa","v":0.03,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1 = poor to 6 = outstanding","field":"social science","wr":"independent versus applicant-nominated reviewers on proposals","conf":"med","self":false,"doi":"10.1080/03623319.2020.1728506","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The Results text states a kappa of 0.03 between independent and applicant-nominated reviewer scores of the same proposal, which the paper calls no better than chance. This conflicts with the 0.05 tabulated in Table 3; both reported values are retained. The Methods describe only weighted kappa.","vf":"unverified"},{"key":"YYCAX8PN","au":"Jerrim, John","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intra-cluster correlation coefficient from a multi-level (random-effects) model, reviews nested within proposals","estd":"ICC (single/unspec)","v":0.17,"n":"4144","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1 = poor to 6 = outstanding","field":"social science","wr":"academic reviewers on ESRC grant proposal grades","conf":"med","self":false,"doi":"10.1080/03623319.2020.1728506","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"A multi-level random-effects model with all 15,047 reviews nested within 4,144 proposals gave an intra-cluster correlation of 0.17, so only 17 per cent of the variation in review scores lies between proposals. Proposals carried between two and six or more reviews, and the average number per proposal is not stated.","vf":"unverified"},{"key":"YYCAX8PN","au":"Jerrim, John","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intra-cluster correlation coefficient from a multi-level (random-effects) model, reviews nested within proposals","estd":"ICC (single/unspec)","v":0.18,"n":"","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1 = poor to 6 = outstanding","field":"social science","wr":"independent reviewers on ESRC grant proposal grades","conf":"med","self":false,"doi":"10.1080/03623319.2020.1728506","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The multi-level model of reviews nested within proposals gave an intra-cluster correlation of 0.18 when restricted to independent (non-nominated) reviews. The number of proposals and reviews behind this column is not stated.","vf":"unverified"},{"key":"YYCAX8PN","au":"Jerrim, John","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"polychoric correlation between each pair of reviewers, averaged across all pairs, weighted by sample size","estd":"correlation","v":0.17,"n":"4144","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1 = poor to 6 = outstanding","field":"social science","wr":"academic reviewers on ESRC grant proposal grades","conf":"med","self":false,"doi":"10.1080/03623319.2020.1728506","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Academic peer reviewers graded 4,144 ESRC grant proposals (15,047 reviews) on a six-point scale, and the polychoric correlation between any two reviewers of the same proposal, averaged over all pairs, was 0.17. This is the paper's headline result, described in the abstract as a correlation of only 0.2.","vf":"unverified"},{"key":"YYCAX8PN","au":"Jerrim, John","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"polychoric correlation between each pair of reviewers, averaged across all pairs, weighted by sample size","estd":"correlation","v":0.11,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1 = poor to 6 = outstanding","field":"social science","wr":"independent versus applicant-nominated reviewers on proposals","conf":"med","self":false,"doi":"10.1080/03623319.2020.1728506","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For pairs made up of one independent and one applicant-nominated reviewer of the same proposal, Table 3 gives an average polychoric correlation of 0.11. The Results text states 0.07 for the same comparison, so table and text disagree.","vf":"unverified"},{"key":"YYCAX8PN","au":"Jerrim, John","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"polychoric correlation between each pair of reviewers, averaged across all pairs, weighted by sample size","estd":"correlation","v":0.19,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1 = poor to 6 = outstanding","field":"social science","wr":"independent reviewers on ESRC grant proposal grades","conf":"med","self":false,"doi":"10.1080/03623319.2020.1728506","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Restricting the pairs to two independent (non-nominated) reviewers of the same proposal, the average polychoric correlation between their six-point grades was 0.19. This is essentially the same as the value for all reviewer pairs.","vf":"unverified"},{"key":"YYCAX8PN","au":"Jerrim, John","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"weighted-kappa","form":"weighted Kappa calculated for each pair of reviews and then averaged, weighted by sample size","estd":"weighted kappa","v":0.1,"n":"4144","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1 = poor to 6 = outstanding","field":"social science","wr":"academic reviewers on ESRC grant proposal grades","conf":"med","self":false,"doi":"10.1080/03623319.2020.1728506","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"A weighted kappa was computed for every pair of reviews of the same ESRC proposal and then averaged, giving 0.10 across all reviewer pairs. On the Landis and Koch rules of thumb the paper cites, this is only slight chance-corrected agreement.","vf":"unverified"},{"key":"YYCAX8PN","au":"Jerrim, John","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"weighted-kappa","form":"weighted Kappa calculated for each pair of reviews and then averaged, weighted by sample size","estd":"weighted kappa","v":0.05,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1 = poor to 6 = outstanding","field":"social science","wr":"independent versus applicant-nominated reviewers on proposals","conf":"med","self":false,"doi":"10.1080/03623319.2020.1728506","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The average weighted kappa between one independent and one applicant-nominated reviewer of the same proposal is 0.05 in Table 3. The Results text states 0.03 for the same comparison and calls it no better than chance.","vf":"unverified"},{"key":"YYCAX8PN","au":"Jerrim, John","y":2020,"cx":"Grant","ob":"grant-proposal","fam":"weighted-kappa","form":"weighted Kappa calculated for each pair of reviews and then averaged, weighted by sample size","estd":"weighted kappa","v":0.12,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1 = poor to 6 = outstanding","field":"social science","wr":"independent reviewers on ESRC grant proposal grades","conf":"med","self":false,"doi":"10.1080/03623319.2020.1728506","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The average weighted kappa between pairs of independent (non-nominated) reviewers grading the same proposal was 0.12. The paper reads all its kappa values as showing only slight agreement.","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"Gwet-AC","form":"AC1 coefficient as calculated by Baethge et al. (2013) for more than two raters","estd":"Gwet AC","v":0.15,"n":"145","k":"3.06","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"weak or strong acceptance versus all categories below weak acceptance","field":"multi-field","wr":"conference reviewers on accept versus non-accept recommendation","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Using the variant of AC1 for more than two raters that Baethge and colleagues applied, agreement on the dichotomised accept or not-accept recommendation was .15 across the same 443 reviews of 145 papers. The authors report it to allow direct comparison with that earlier journal study.","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"Gwet-AC","form":"AC1 coefficient for multiple raters (Gwet 2008, 2014; see Fleiss 1971)","estd":"Gwet AC","v":0.14,"n":"145","k":"3.06","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"weak or strong acceptance versus all categories below weak acceptance","field":"multi-field","wr":"conference reviewers on accept versus non-accept recommendation","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"The authors collapsed the five-point overall evaluation into acceptance versus non-acceptance and computed Gwet's AC1 for multiple raters over the all-reviewer sample (443 reviews of 145 papers by 130 reviewers, on average 3.06 reviews per paper). The value of .14 shows very little chance-corrected agreement on whether a paper should be accepted.","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"Cohen’s Kappa as calculated by Baethge et al. (2013) for more than two raters","estd":"kappa","v":0.15,"n":"145","k":"3.06","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"weak or strong acceptance versus all categories below weak acceptance","field":"multi-field","wr":"conference reviewers on accept versus non-accept recommendation","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"A multi-rater version of Cohen's kappa, computed as in Baethge and colleagues, gave .15 for agreement on the dichotomised accept or not-accept recommendation across 443 reviews of 145 conference papers. The authors treat this as confirming the poor agreement seen in the G-coefficients.","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"G-theory","form":"mean single-rater reliability G(qk, k = 1) across the five rating dimensions, estimated by the procedure of Bornmann et al. (2010, p. 3)","estd":"G-theory","v":0.22,"n":"145","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"five dimensions: 5-point overall evaluation plus four 3-point scales","field":"multi-field","wr":"conference reviewers on submitted papers, five rating dimensions","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":false,"ms":"This is the study's headline figure: the mean of the five single-rater G-coefficients based on all reviewers of 145 submissions to an interdisciplinary conference. At .22 it says a single reviewer's rating carries little reliable information, and the authors note it falls below the mean of .34 reported in Bornmann and colleagues' meta-analysis. Per-dimension sample sizes vary slightly (142 to 145 papers, Table 2) because cannot judge responses were dropped.","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"G-theory","form":"mean single-rater reliability G(qk, k = 1) across the five rating dimensions, estimated by the procedure of Bornmann et al. (2010, p. 3)","estd":"G-theory","v":0.35,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"five dimensions: 5-point overall evaluation plus four 3-point scales","field":"multi-field","wr":"different-discipline reviewers on papers, five rating dimensions","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Averaged across the five rating dimensions, single-rater reliability among reviewers from a discipline other than the paper's was .35, higher than the .16 for same-discipline reviewers, though the difference was not statistically significant. Sample sizes differ by dimension (108 to 119 papers, Table 2) and are not reported for this average.","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"G-theory","form":"mean single-rater reliability G(qk, k = 1) across the five rating dimensions, estimated by the procedure of Bornmann et al. (2010, p. 3)","estd":"G-theory","v":0.16,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"five dimensions: 5-point overall evaluation plus four 3-point scales","field":"multi-field","wr":"same-discipline reviewers on papers, five rating dimensions","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Averaged across the five rating dimensions, single-rater reliability among reviewers from the same discipline as the paper was .16, which the authors call poor and not even distinguishable from chance. Sample sizes differ by dimension (102 to 119 papers, Table 2) and are not reported for this average.","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"monotrait-heteromethod correlation, estimated by MLR with full information maximum likelihood","estd":"correlation","v":0.25,"n":"145","k":"","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"recoded 3-point: not novel, some novelty, highly novel","field":"multi-field","wr":"two reviewer groups on the same conference papers","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Same-discipline and different-discipline mean novelty scores for the same papers correlated at .25 across 145 papers (73 with both scores present). This was the strongest of the five between-group correlations and remained significant after adjustment.","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"monotrait-heteromethod correlation, estimated by MLR with full information maximum likelihood","estd":"correlation","v":0.23,"n":"145","k":"","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"-2 (strong reject) to 2 (strong accept), recoded 1 to 5","field":"multi-field","wr":"two reviewer groups on the same conference papers","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"For each paper the authors averaged the overall evaluations of same-discipline reviewers and, separately, of different-discipline reviewers, then correlated the two sets of paper scores across 145 papers (93 with both scores present). The correlation of .23 was significant after adjustment but shows the two reviewer groups largely disagreed on paper quality.","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"monotrait-heteromethod correlation, estimated by MLR with full information maximum likelihood","estd":"correlation","v":0.18,"n":"145","k":"","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1 = not relevant, 2 = some relevance, 3 = highly relevant","field":"multi-field","wr":"two reviewer groups on the same conference papers","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Same-discipline and different-discipline mean relevance scores for the same papers correlated at .18 across 145 papers (85 with both scores present), which was not significant after adjustment for multiple testing.","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"monotrait-heteromethod correlation, estimated by MLR with full information maximum likelihood","estd":"correlation","v":0.07,"n":"145","k":"","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"recoded 3-point: not significant, some aspects significant, highly significant","field":"multi-field","wr":"two reviewer groups on the same conference papers","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Same-discipline and different-discipline mean significance scores for the same papers correlated at only .07 across 145 papers (75 with both scores present), a non-significant relationship indicating the two reviewer groups did not agree.","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"monotrait-heteromethod correlation, estimated by MLR with full information maximum likelihood","estd":"correlation","v":0.05,"n":"145","k":"","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"recoded 3-point: unacceptable major flaws, good but some flaws, excellent","field":"multi-field","wr":"two reviewer groups on the same conference papers","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Same-discipline and different-discipline mean soundness scores for the same papers correlated at only .05 across 145 papers (69 with both scores present), the weakest of the five between-group correlations and not significant.","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"G-theory","form":"single-rater reliability G(qk, k = 1) per Putka et al. (2008), REML variance components","estd":"G-theory","v":0.28,"n":"143","k":"2.11","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"recoded 3-point: not novel, some novelty, highly novel","field":"multi-field","wr":"conference reviewers on submitted papers, novelty","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":0.13,"ciHigh":0.42,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Reviewers rated the novelty of 143 conference submissions, 344 usable ratings from 120 reviewers. The single-rater G-coefficient of .28 was the highest of the five rating dimensions but still indicates poor agreement between individual reviewers.","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"G-theory","form":"single-rater reliability G(qk, k = 1) per Putka et al. (2008), REML variance components","estd":"G-theory","v":0.35,"n":"109","k":"1.34","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"recoded 3-point: not novel, some novelty, highly novel","field":"multi-field","wr":"different-discipline reviewers on papers, novelty","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":0.01,"ciHigh":0.59,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Among reviewers from a discipline other than the paper's, 176 novelty ratings by 89 reviewers of 109 papers, the single-rater G-coefficient was .35. The unadjusted interval excluded zero but the Holm-adjusted interval did not.","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"G-theory","form":"single-rater reliability G(qk, k = 1) per Putka et al. (2008), REML variance components","estd":"G-theory","v":0.2,"n":"107","k":"1.35","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"recoded 3-point: not novel, some novelty, highly novel","field":"multi-field","wr":"same-discipline reviewers on papers, novelty","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":0,"ciHigh":0.48,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Among reviewers from the same discipline as the paper, 168 novelty ratings by 86 reviewers of 107 papers, the single-rater G-coefficient was .20, with a confidence interval that included zero.","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"G-theory","form":"single-rater reliability G(qk, k = 1) per Putka et al. (2008), REML variance components","estd":"G-theory","v":0.21,"n":"145","k":"3.0","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"-2 (strong reject) to 2 (strong accept), recoded 1 to 5","field":"multi-field","wr":"conference reviewers on submitted papers, overall evaluation","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":0.1,"ciHigh":0.31,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Reviewers at an interdisciplinary conference gave overall evaluations to 145 submissions, 443 reviews from 130 reviewers, on average three reviewers per paper. The single-rater G-coefficient of .21 means one reviewer's overall score is a poor indicator of the paper's mean score; the Holm-adjusted interval was [.07, .35].","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"G-theory","form":"single-rater reliability G(qk, k = 1) per Putka et al. (2008), REML variance components","estd":"G-theory","v":0.18,"n":"119","k":"1.58","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"-2 (strong reject) to 2 (strong accept), recoded 1 to 5","field":"multi-field","wr":"different-discipline reviewers on papers, overall evaluation","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":0,"ciHigh":0.4,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Restricting the analysis to reviewers from a discipline other than the paper's, 227 reviews by 103 reviewers of 119 papers, the single-rater G-coefficient for overall evaluation was .18, with a confidence interval that included zero.","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"G-theory","form":"single-rater reliability G(qk, k = 1) per Putka et al. (2008), REML variance components","estd":"G-theory","v":0.23,"n":"119","k":"1.48","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"-2 (strong reject) to 2 (strong accept), recoded 1 to 5","field":"multi-field","wr":"same-discipline reviewers on papers, overall evaluation","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":0,"ciHigh":0.46,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Restricting the analysis to reviewers from the same discipline as the paper, 216 reviews by 99 reviewers of 119 papers, the single-rater G-coefficient for overall evaluation was .23. Its confidence interval included zero, so agreement was not distinguishable from chance.","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"G-theory","form":"single-rater reliability G(qk, k = 1) per Putka et al. (2008), REML variance components","estd":"G-theory","v":0.21,"n":"144","k":"2.63","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = not relevant, 2 = some relevance, 3 = highly relevant","field":"multi-field","wr":"conference reviewers on submitted papers, relevance","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":0.09,"ciHigh":0.32,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Reviewers rated the relevance of 144 conference submissions, 402 usable ratings from 128 reviewers, with a harmonic mean of 2.63 reviewers per paper. A single-rater G-coefficient of .21 indicates poor agreement between individual reviewers on relevance.","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"G-theory","form":"single-rater reliability G(qk, k = 1) per Putka et al. (2008), REML variance components","estd":"G-theory","v":0.37,"n":"114","k":"1.45","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = not relevant, 2 = some relevance, 3 = highly relevant","field":"multi-field","wr":"different-discipline reviewers on papers, relevance","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":0.1,"ciHigh":0.58,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Among reviewers from a discipline other than the paper's, 199 relevance ratings by 99 reviewers of 114 papers, the single-rater G-coefficient was .37, which the authors describe as fair agreement and significantly above chance.","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"G-theory","form":"single-rater reliability G(qk, k = 1) per Putka et al. (2008), REML variance components","estd":"G-theory","v":0.13,"n":"115","k":"1.45","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = not relevant, 2 = some relevance, 3 = highly relevant","field":"multi-field","wr":"same-discipline reviewers on papers, relevance","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":0,"ciHigh":0.39,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Among reviewers from the same discipline as the paper, 203 relevance ratings by 94 reviewers of 115 papers, the single-rater G-coefficient was .13, with a confidence interval that included zero.","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"G-theory","form":"single-rater reliability G(qk, k = 1) per Putka et al. (2008), REML variance components","estd":"G-theory","v":0.17,"n":"143","k":"2.33","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"recoded 3-point: not significant, some aspects significant, highly significant","field":"multi-field","wr":"conference reviewers on submitted papers, significance","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":0.04,"ciHigh":0.3,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Reviewers rated the significance of 143 conference submissions, 369 usable ratings from 124 reviewers. The single-rater G-coefficient of .17 was the lowest of the five dimensions and shows very poor agreement between individual reviewers.","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"G-theory","form":"single-rater reliability G(qk, k = 1) per Putka et al. (2008), REML variance components","estd":"G-theory","v":0.46,"n":"108","k":"1.41","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"recoded 3-point: not significant, some aspects significant, highly significant","field":"multi-field","wr":"different-discipline reviewers on papers, significance","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":0.19,"ciHigh":0.65,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Among reviewers from a discipline other than the paper's, 185 significance ratings by 96 reviewers of 108 papers, the single-rater G-coefficient was .46, the highest value in the study and significantly above chance.","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"G-theory","form":"single-rater reliability G(qk, k = 1) per Putka et al. (2008), REML variance components","estd":"G-theory","v":0.14,"n":"110","k":"1.39","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"recoded 3-point: not significant, some aspects significant, highly significant","field":"multi-field","wr":"same-discipline reviewers on papers, significance","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":0,"ciHigh":0.42,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Among reviewers from the same discipline as the paper, 184 significance ratings by 89 reviewers of 110 papers, the single-rater G-coefficient was .14, with a confidence interval that included zero.","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"G-theory","form":"single-rater reliability G(qk, k = 1) per Putka et al. (2008), REML variance components","estd":"G-theory","v":0.21,"n":"142","k":"1.92","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"recoded 3-point: unacceptable major flaws, good but some flaws, excellent","field":"multi-field","wr":"conference reviewers on submitted papers, soundness","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":0.04,"ciHigh":0.37,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Reviewers rated the soundness of 142 conference submissions, 325 usable ratings from 123 reviewers. The single-rater G-coefficient of .21 indicates poor agreement between individual reviewers on methodological soundness.","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"G-theory","form":"single-rater reliability G(qk, k = 1) per Putka et al. (2008), REML variance components","estd":"G-theory","v":0.39,"n":"109","k":"1.31","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"recoded 3-point: unacceptable major flaws, good but some flaws, excellent","field":"multi-field","wr":"different-discipline reviewers on papers, soundness","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":0.05,"ciHigh":0.63,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Among reviewers from a discipline other than the paper's, 172 soundness ratings by 92 reviewers of 109 papers, the single-rater G-coefficient was .39. The unadjusted interval excluded zero but the Holm-adjusted interval did not.","vf":"unverified"},{"key":"WC3AX6UH","au":"Jirschitzka, Jens","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"G-theory","form":"single-rater reliability G(qk, k = 1) per Putka et al. (2008), REML variance components","estd":"G-theory","v":0.13,"n":"102","k":"1.27","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"recoded 3-point: unacceptable major flaws, good but some flaws, excellent","field":"multi-field","wr":"same-discipline reviewers on papers, soundness","conf":"high","self":false,"doi":"10.1007/s11192-017-2516-6","ciLow":0,"ciHigh":0.48,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Among reviewers from the same discipline as the paper, 153 soundness ratings by 83 reviewers of 102 papers, the single-rater G-coefficient was .13, with a confidence interval that included zero.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"average CA-ICC, two-way random-effects, consistency of agreement, mean across 12 pools","estd":"ICC (average)","v":0.92,"n":"159","k":"5","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"A Plus (5), A High (4), A Standard (3), B (2), C (1)","field":"multi-field","wr":"committee assessors on Academic Merit domain scores","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":0.89,"ciHigh":0.95,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Averaged across 12 pools, the average consistency ICC for committee Academic Merit scores was 0.92, the highest of the three domains and indicating excellent agreement.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"average CA-ICC, two-way random-effects, consistency of agreement, mean across 12 pools","estd":"ICC (average)","v":0.78,"n":"159","k":"5","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"A Plus (5), A High (4), A Standard (3), B (2), C (1)","field":"multi-field","wr":"committee assessors on Character/Leadership domain scores","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":0.71,"ciHigh":0.85,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Averaged across 12 pools, the average consistency ICC for committee Character/Leadership scores was 0.78, indicating good agreement, lower than for Academic Merit.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"average CA-ICC, two-way random-effects, consistency of agreement","estd":"ICC (average)","v":0.98,"n":"12","k":"5","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"A Plus (5), A High (4), A Standard (3), B (2), C (1)","field":"multi-field","wr":"committee assessors on Academic Merit scores, Shirtcliffe 2011","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":0.96,"ciHigh":0.99,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The highest of the 36 component-by-round strata: for Academic Merit in the 2011 Shirtcliffe round the average consistency ICC was 0.98, indicating near-perfect agreement.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"average CA-ICC, two-way random-effects, consistency of agreement","estd":"ICC (average)","v":0.47,"n":"16","k":"5","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"A Plus (5), A High (4), A Standard (3), B (2), C (1)","field":"multi-field","wr":"committee assessors on Quality of Study Plans scores, GW 2010","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":-0.12,"ciHigh":0.8,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The lowest of the 36 component-by-round strata: for Quality of Study Plans in the 2010 Gordon Watson round the average consistency ICC was 0.47, with a wide interval spanning zero.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"average CA-ICC, two-way random-effects, consistency of agreement, mean across 12 pools","estd":"ICC (average)","v":0.75,"n":"159","k":"5","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"A Plus (5), A High (4), A Standard (3), B (2), C (1)","field":"multi-field","wr":"committee assessors on Quality of Study Plans domain scores","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":0.67,"ciHigh":0.82,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Averaged across 12 pools, the average consistency ICC for committee Quality of Study Plans scores was 0.75, indicating good agreement, lower than for Academic Merit.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"average CA-ICC, two-way random-effects, consistency of agreement","estd":"ICC (average)","v":0.96,"n":"13","k":"5","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"standardised assessor total score, 0-1 range","field":"multi-field","wr":"committee assessors on Gordon Watson applicants' total scores","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":0.91,"ciHigh":0.99,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 2009 Gordon Watson round, five assessors scored 13 applicants; the average consistency ICC was 0.96, indicating excellent agreement on the committee's mean rating.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"average CA-ICC, two-way random-effects, consistency of agreement","estd":"ICC (average)","v":0.75,"n":"16","k":"5","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"standardised assessor total score, 0-1 range","field":"multi-field","wr":"committee assessors on Gordon Watson applicants' total scores","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":0.48,"ciHigh":0.9,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 2010 Gordon Watson round, five assessors scored 16 applicants; the average consistency ICC was 0.75, indicating good agreement on the committee's mean rating.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"average CA-ICC, two-way random-effects, consistency of agreement","estd":"ICC (average)","v":0.94,"n":"10","k":"5","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"standardised assessor total score, 0-1 range","field":"multi-field","wr":"committee assessors on Gordon Watson applicants' total scores","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":0.86,"ciHigh":0.98,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 2011 Gordon Watson round, five assessors scored 10 applicants; the average consistency ICC was 0.94, indicating excellent agreement on the committee's mean rating.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"average CA-ICC, two-way random-effects, consistency of agreement","estd":"ICC (average)","v":0.74,"n":"11","k":"5","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"standardised assessor total score, 0-1 range","field":"multi-field","wr":"committee assessors on Gordon Watson applicants' total scores","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":0.37,"ciHigh":0.92,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 2012 Gordon Watson round, five assessors scored 11 applicants; the average consistency ICC was 0.74, the lowest round observed, indicating good agreement on the committee's mean rating.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"average CA-ICC, two-way random-effects, consistency of agreement","estd":"ICC (average)","v":0.79,"n":"10","k":"5","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"standardised assessor total score, 0-1 range","field":"multi-field","wr":"committee assessors on Gordon Watson applicants' total scores","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":0.48,"ciHigh":0.94,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 2013 Gordon Watson round, five assessors scored 10 applicants; the average consistency ICC was 0.79, indicating good agreement on the committee's mean rating.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"average CA-ICC, two-way random-effects, consistency of agreement","estd":"ICC (average)","v":0.85,"n":"11","k":"5","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"standardised assessor total score, 0-1 range","field":"multi-field","wr":"committee assessors on Gordon Watson applicants' total scores","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":0.65,"ciHigh":0.95,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 2014 Gordon Watson round, five assessors scored 11 applicants; the average consistency ICC was 0.85, indicating good agreement on the committee's mean rating.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"individual CA-ICC, two-way random-effects, consistency of agreement (equivalent to Pearson correlation), mean across 12 pools","estd":"ICC (single/unspec)","v":0.76,"n":"159","k":"3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"A Plus (5), A High (4), A Standard (3), B (2), C (1)","field":"multi-field","wr":"single academic assessor on Academic Merit scores","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Averaged across 12 pools, the individual consistency ICC for a single academic assessor's Academic Merit score was 0.76, significantly higher and less variable (SD 0.11) than for non-academics.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"individual CA-ICC, two-way random-effects, consistency of agreement (equivalent to Pearson correlation), mean across 12 pools","estd":"ICC (single/unspec)","v":0.55,"n":"159","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"A Plus (5), A High (4), A Standard (3), B (2), C (1)","field":"multi-field","wr":"single non-academic assessor on Academic Merit scores","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Averaged across 12 pools, the individual consistency ICC for a single non-academic assessor's Academic Merit score was 0.55, significantly lower and more variable (SD 0.26) than for academics.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"individual CA-ICC, two-way random-effects, consistency of agreement (equivalent to Pearson correlation), mean across 12 pools","estd":"ICC (single/unspec)","v":0.39,"n":"159","k":"3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"A Plus (5), A High (4), A Standard (3), B (2), C (1)","field":"multi-field","wr":"single academic assessor on Character/Leadership scores","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Averaged across 12 pools, the individual consistency ICC for a single academic assessor's Character/Leadership score was 0.39, not significantly different from non-academics.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"individual CA-ICC, two-way random-effects, consistency of agreement (equivalent to Pearson correlation), mean across 12 pools","estd":"ICC (single/unspec)","v":0.52,"n":"159","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"A Plus (5), A High (4), A Standard (3), B (2), C (1)","field":"multi-field","wr":"single non-academic assessor on Character/Leadership scores","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Averaged across 12 pools, the individual consistency ICC for a single non-academic assessor's Character/Leadership score was 0.52, not significantly different from academics.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"individual CA-ICC, two-way random-effects, consistency of agreement (equivalent to Pearson correlation), mean across 12 pools","estd":"ICC (single/unspec)","v":0.42,"n":"159","k":"3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"A Plus (5), A High (4), A Standard (3), B (2), C (1)","field":"multi-field","wr":"single academic assessor on Quality of Study Plans scores","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Averaged across 12 pools, the individual consistency ICC for a single academic assessor's Quality of Study Plans score was 0.42, not significantly different from non-academics.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"individual CA-ICC, two-way random-effects, consistency of agreement (equivalent to Pearson correlation), mean across 12 pools","estd":"ICC (single/unspec)","v":0.37,"n":"159","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"A Plus (5), A High (4), A Standard (3), B (2), C (1)","field":"multi-field","wr":"single non-academic assessor on Quality of Study Plans scores","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Averaged across 12 pools, the individual consistency ICC for a single non-academic assessor's Quality of Study Plans score was 0.37, not significantly different from academics.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"individual CA-ICC, two-way random-effects, consistency of agreement (equivalent to Pearson correlation), mean across 12 pools","estd":"ICC (single/unspec)","v":0.64,"n":"159","k":"3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"standardised assessor total score, 0-1 range","field":"multi-field","wr":"single academic assessor on applicants' total scores","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Averaged across 12 pools, the individual consistency ICC for a single academic assessor's total score was 0.64, not significantly different from the non-academic value.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"individual CA-ICC, two-way random-effects, consistency of agreement (equivalent to Pearson correlation), mean across 12 pools","estd":"ICC (single/unspec)","v":0.59,"n":"159","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"standardised assessor total score, 0-1 range","field":"multi-field","wr":"single non-academic assessor on applicants' total scores","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Averaged across 12 pools, the individual consistency ICC for a single non-academic assessor's total score was 0.59, not significantly different from the academic value.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"average CA-ICC, two-way random-effects, consistency of agreement (equivalent to Cronbach's alpha)","estd":"ICC (average)","v":0.87,"n":"159","k":"5","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"standardised assessor total score, 0-1 range","field":"multi-field","wr":"committee assessors on scholarship applicants' total scores","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":0.82,"ciHigh":0.91,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Five committee assessors independently scored 159 doctoral scholarship applicants over 12 rounds; the mean average consistency ICC of 0.87 indicates good to excellent reliability of the committee's mean rating.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"average CA-ICC, two-way random-effects, consistency of agreement","estd":"ICC (average)","v":0.9,"n":"6","k":"5","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"standardised assessor total score, 0-1 range","field":"multi-field","wr":"committee assessors on Shirtcliffe applicants' total scores","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":0.67,"ciHigh":0.98,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 2009 Shirtcliffe round, five assessors scored 6 applicants; the average consistency ICC was 0.90, indicating excellent agreement on the committee's mean rating.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"average CA-ICC, two-way random-effects, consistency of agreement","estd":"ICC (average)","v":0.92,"n":"13","k":"5","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"standardised assessor total score, 0-1 range","field":"multi-field","wr":"committee assessors on Shirtcliffe applicants' total scores","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":0.81,"ciHigh":0.97,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 2010 Shirtcliffe round, five assessors scored 13 applicants; the average consistency ICC was 0.92, indicating excellent agreement on the committee's mean rating.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"average CA-ICC, two-way random-effects, consistency of agreement","estd":"ICC (average)","v":0.96,"n":"12","k":"5","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"standardised assessor total score, 0-1 range","field":"multi-field","wr":"committee assessors on Shirtcliffe applicants' total scores","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":0.91,"ciHigh":0.99,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 2011 Shirtcliffe round, five assessors scored 12 applicants; the average consistency ICC was 0.96, the highest round observed, indicating excellent agreement on the committee's mean rating.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"average CA-ICC, two-way random-effects, consistency of agreement","estd":"ICC (average)","v":0.93,"n":"20","k":"5","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"standardised assessor total score, 0-1 range","field":"multi-field","wr":"committee assessors on Shirtcliffe applicants' total scores","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":0.87,"ciHigh":0.97,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 2012 Shirtcliffe round, five assessors scored 20 applicants; the average consistency ICC was 0.93, indicating excellent agreement on the committee's mean rating.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"average CA-ICC, two-way random-effects, consistency of agreement","estd":"ICC (average)","v":0.79,"n":"19","k":"5","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"standardised assessor total score, 0-1 range","field":"multi-field","wr":"committee assessors on Shirtcliffe applicants' total scores","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":0.59,"ciHigh":0.91,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 2013 Shirtcliffe round, five assessors scored 19 applicants; the average consistency ICC was 0.79, indicating good agreement on the committee's mean rating.","vf":"unverified"},{"key":"CM2NSUHH","au":"Johnston, Lucy","y":2015,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"average CA-ICC, two-way random-effects, consistency of agreement","estd":"ICC (average)","v":0.94,"n":"18","k":"5","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"standardised assessor total score, 0-1 range","field":"multi-field","wr":"committee assessors on Shirtcliffe applicants' total scores","conf":"high","self":false,"doi":"10.1080/03075079.2015.1124849","ciLow":0.89,"ciHigh":0.98,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 2014 Shirtcliffe round, five assessors scored 18 applicants; the average consistency ICC was 0.94, indicating excellent agreement on the committee's mean rating.","vf":"unverified"},{"key":"GEWF5FVV","au":"Justice, Amy C.","y":1994,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"weighted observed agreement, not exact: perfect agreement 9, one away 8, decreasing to perfect disagreement 0","estd":"percent agreement","v":0.69,"n":"113","k":"2","samp":"funded-only","blind":"unclear","agg":"single-rater","scale":"10-point: poor, fair, acceptable, good, superb","field":"biomedical","wr":"readers and experts on manuscript quality grade","conf":"high","self":false,"doi":"10.1001/jama.1994.03520020043011","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Distance-weighted observed agreement between readers and experts on the 10-point manuscript grade was 69%, equal to the agreement expected by chance.","vf":"unverified"},{"key":"GEWF5FVV","au":"Justice, Amy C.","y":1994,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"kappa index, weighted to give credit for near misses after Kramer and Feinstein","estd":"weighted kappa","v":-0.01,"n":"113","k":"2","samp":"funded-only","blind":"unclear","agg":"single-rater","scale":"10-point: poor, fair, acceptable, good, superb","field":"biomedical","wr":"readers and experts on manuscript quality grade","conf":"high","self":false,"doi":"10.1001/jama.1994.03520020043011","ciLow":-0.2,"ciHigh":0.1,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Journal readers and clinical-research-methods experts independently graded accepted manuscripts on a 10-point quality scale; across 371 pairs the weighted kappa of -0.01 indicates no agreement beyond chance.","vf":"unverified"},{"key":"GEWF5FVV","au":"Justice, Amy C.","y":1994,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"weighted observed agreement, not exact: perfect agreement 9, one away 8, decreasing to perfect disagreement 0","estd":"percent agreement","v":0.74,"n":"113","k":"2","samp":"funded-only","blind":"unclear","agg":"single-rater","scale":"10-point: poor, fair, acceptable, good, superb","field":"biomedical","wr":"readers and peer reviewers on manuscript quality grade","conf":"high","self":false,"doi":"10.1001/jama.1994.03520020043011","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Distance-weighted observed agreement between readers and peer reviewers on the 10-point manuscript grade was 74%, equal to the agreement expected by chance.","vf":"unverified"},{"key":"GEWF5FVV","au":"Justice, Amy C.","y":1994,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"kappa index, weighted to give credit for near misses after Kramer and Feinstein","estd":"weighted kappa","v":-0.02,"n":"113","k":"2","samp":"funded-only","blind":"unclear","agg":"single-rater","scale":"10-point: poor, fair, acceptable, good, superb","field":"biomedical","wr":"readers and peer reviewers on manuscript quality grade","conf":"high","self":false,"doi":"10.1001/jama.1994.03520020043011","ciLow":-0.2,"ciHigh":0.2,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":true,"he":true,"ms":"Journal readers and peer reviewers independently graded accepted manuscripts on a 10-point quality scale; across 352 pairs the weighted kappa of -0.02 shows no agreement beyond chance, the study's headline reader-versus-reviewer result.","vf":"unverified"},{"key":"GEWF5FVV","au":"Justice, Amy C.","y":1994,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"weighted observed agreement, not exact: perfect agreement 9, one away 8, decreasing to perfect disagreement 0","estd":"percent agreement","v":0.77,"n":"113","k":"2","samp":"funded-only","blind":"unclear","agg":"single-rater","scale":"10-point: poor, fair, acceptable, good, superb","field":"biomedical","wr":"pairs of readers on manuscript quality grade","conf":"high","self":false,"doi":"10.1001/jama.1994.03520020043011","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Distance-weighted observed agreement between pairs of readers on the 10-point manuscript grade was 77%, barely above the 76% expected by chance.","vf":"unverified"},{"key":"GEWF5FVV","au":"Justice, Amy C.","y":1994,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"kappa index, weighted to give credit for near misses after Kramer and Feinstein","estd":"weighted kappa","v":0.05,"n":"113","k":"2","samp":"funded-only","blind":"unclear","agg":"single-rater","scale":"10-point: poor, fair, acceptable, good, superb","field":"biomedical","wr":"pairs of readers on manuscript quality grade","conf":"high","self":false,"doi":"10.1001/jama.1994.03520020043011","ciLow":-0.2,"ciHigh":0.3,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Pairs of journal readers independently graded accepted manuscripts on a 10-point quality scale; across 159 pairs the weighted kappa of 0.05 indicates agreement barely above chance.","vf":"unverified"},{"key":"SUR92HZC","au":"Kaatz, Anna","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraclass correlation coefficient (ICC), model unspecified","estd":"ICC (single/unspec)","v":0.183,"n":"","k":"","samp":"special","blind":"single","agg":"single-rater","scale":"percentage of words in the category, 0 to 100","field":"biomedical","wr":"reviewers' ability-word percentage in R01 critiques","conf":"med","self":false,"doi":"10.1097/acm.0000000000000442","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"restricted-other","pr":false,"he":false,"ms":"Two to five reviewers critiqued each NIH R01 application; an ICC of 0.18 (18% between-applicant variance) indicates weak consistency among reviewers of the same applicant's applications in their use of ability words.","vf":"unverified"},{"key":"SUR92HZC","au":"Kaatz, Anna","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraclass correlation coefficient (ICC), model unspecified","estd":"ICC (single/unspec)","v":0.14,"n":"","k":"","samp":"special","blind":"single","agg":"single-rater","scale":"percentage of words in the category, 0 to 100","field":"biomedical","wr":"reviewers' achievement-word percentage in R01 critiques","conf":"med","self":false,"doi":"10.1097/acm.0000000000000442","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"restricted-other","pr":false,"he":false,"ms":"Two to five reviewers critiqued each NIH R01 application; an ICC of 0.14 (14% between-applicant variance) indicates weak consistency among reviewers of the same applicant's applications in their use of achievement words.","vf":"unverified"},{"key":"SUR92HZC","au":"Kaatz, Anna","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraclass correlation coefficient (ICC), model unspecified","estd":"ICC (single/unspec)","v":0.223,"n":"","k":"","samp":"special","blind":"single","agg":"single-rater","scale":"percentage of words in the category, 0 to 100","field":"biomedical","wr":"reviewers' agentic-word percentage in R01 critiques","conf":"med","self":false,"doi":"10.1097/acm.0000000000000442","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"restricted-other","pr":false,"he":false,"ms":"Two to five reviewers critiqued each NIH R01 application; an ICC of 0.22 (22% between-applicant variance) indicates modest consistency among reviewers of the same applicant's applications in their use of agentic words.","vf":"unverified"},{"key":"SUR92HZC","au":"Kaatz, Anna","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraclass correlation coefficient (ICC), model unspecified","estd":"ICC (single/unspec)","v":0.228,"n":"","k":"","samp":"special","blind":"single","agg":"single-rater","scale":"percentage of words in the category, 0 to 100","field":"biomedical","wr":"reviewers' negative-evaluation-word percentage in R01 critiques","conf":"med","self":false,"doi":"10.1097/acm.0000000000000442","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"restricted-other","pr":true,"he":false,"ms":"Two to five reviewers critiqued each NIH R01 application; an ICC of 0.23 (23% between-applicant variance) indicates modest consistency among reviewers of the same applicant's applications in their use of negative evaluation words. Flagged primary only because no overall reliability result exists.","vf":"unverified"},{"key":"SUR92HZC","au":"Kaatz, Anna","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraclass correlation coefficient (ICC), model unspecified","estd":"ICC (single/unspec)","v":0.11,"n":"","k":"","samp":"special","blind":"single","agg":"single-rater","scale":"percentage of words in the category, 0 to 100","field":"biomedical","wr":"reviewers' positive-evaluation-word percentage in R01 critiques","conf":"med","self":false,"doi":"10.1097/acm.0000000000000442","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"restricted-other","pr":false,"he":false,"ms":"Two to five reviewers critiqued each NIH R01 application; an ICC of 0.11 (11% between-applicant variance) indicates weak consistency among reviewers of the same applicant's applications in their use of positive evaluation words.","vf":"unverified"},{"key":"SUR92HZC","au":"Kaatz, Anna","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraclass correlation coefficient (ICC), model unspecified","estd":"ICC (single/unspec)","v":0.37,"n":"","k":"","samp":"special","blind":"single","agg":"single-rater","scale":"percentage of words in the category, 0 to 100","field":"biomedical","wr":"reviewers' research-word percentage in R01 critiques","conf":"med","self":false,"doi":"10.1097/acm.0000000000000442","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"restricted-other","pr":false,"he":false,"ms":"Two to five reviewers critiqued each NIH R01 application; an ICC of 0.37 (37% between-applicant variance) indicates moderate consistency among reviewers of the same applicant's applications in their use of research words.","vf":"unverified"},{"key":"SUR92HZC","au":"Kaatz, Anna","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraclass correlation coefficient (ICC), model unspecified","estd":"ICC (single/unspec)","v":0.41,"n":"","k":"","samp":"special","blind":"single","agg":"single-rater","scale":"percentage of words in the category, 0 to 100","field":"biomedical","wr":"reviewers' standout-adjective percentage in R01 critiques","conf":"med","self":false,"doi":"10.1097/acm.0000000000000442","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"restricted-other","pr":false,"he":false,"ms":"Two to five reviewers critiqued each NIH R01 application; an ICC of 0.41 (41% between-applicant variance) indicates moderate consistency among reviewers of the same applicant's applications in their use of standout adjectives.","vf":"unverified"},{"key":"SUR92HZC","au":"Kaatz, Anna","y":2014,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraclass correlation coefficient (ICC), model unspecified","estd":"ICC (single/unspec)","v":0.113,"n":"","k":"","samp":"special","blind":"single","agg":"single-rater","scale":"percentage of words in the category, 0 to 100","field":"biomedical","wr":"reviewers' total word count in R01 critiques","conf":"med","self":false,"doi":"10.1097/acm.0000000000000442","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"restricted-other","pr":false,"he":false,"ms":"Two to five reviewers critiqued each NIH R01 application; an ICC of 0.11 means only 11% of the variation in critique word count lay between applicants, indicating weak consistency among reviewers of the same applicant's applications.","vf":"unverified"},{"key":"KULA287G","au":"Kemp, Simon","y":2005,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Spearman-Brown 'effective' reliability (Rosenthal, 1991, Table 3.10)","estd":"other","v":0.81,"n":"","k":"8","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"four-point scale from 1 (accept) to 4 (reject)","field":"economic psychology","wr":"projected mean recommendation reliability","conf":"high","self":false,"doi":"10.1016/j.joep.2005.05.004","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"restricted-other","pr":false,"he":false,"ms":"Using the empirical recommendation reliability, the paper projected an effective reliability of 0.81 for the mean recommendation from eight reviewers.","vf":"unverified"},{"key":"KULA287G","au":"Kemp, Simon","y":2005,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Spearman-Brown 'effective' reliability (Rosenthal, 1991, Table 3.10)","estd":"other","v":0.9,"n":"","k":"16","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"four-point scale from 1 (accept) to 4 (reject)","field":"economic psychology","wr":"projected mean recommendation reliability","conf":"high","self":false,"doi":"10.1016/j.joep.2005.05.004","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"restricted-other","pr":false,"he":false,"ms":"Using the empirical recommendation reliability, the paper projected an effective reliability of 0.90 for the mean recommendation from sixteen reviewers.","vf":"unverified"},{"key":"KULA287G","au":"Kemp, Simon","y":2005,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Spearman-Brown 'effective' reliability (Rosenthal, 1991, Table 3.10)","estd":"other","v":0.62,"n":"","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"four-point scale from 1 (accept) to 4 (reject)","field":"economic psychology","wr":"projected mean recommendation reliability","conf":"high","self":false,"doi":"10.1016/j.joep.2005.05.004","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"restricted-other","pr":false,"he":false,"ms":"Using the empirical recommendation reliability, the paper projected an effective reliability of 0.62 for the mean recommendation from three reviewers.","vf":"unverified"},{"key":"KULA287G","au":"Kemp, Simon","y":2005,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Spearman-Brown 'effective' reliability (Rosenthal, 1991, Table 3.10)","estd":"other","v":0.52,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"four-point scale from 1 (accept) to 4 (reject)","field":"economic psychology","wr":"effective reliability of mean recommendations","conf":"high","self":false,"doi":"10.1016/j.joep.2005.05.004","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"restricted-other","pr":false,"he":false,"ms":"The paper presents 0.52 as the Spearman-Brown effective reliability when two reviewers are employed, the same value as the empirical recommendation ICC; a projection fed by the paper's own data.","vf":"unverified"},{"key":"KULA287G","au":"Kemp, Simon","y":2005,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intra-class correlation (R1) between two reviewers; different papers reviewed by different referees (cf. Shrout & Fleiss 1979)","estd":"ICC (single/unspec)","v":0.52,"n":"263","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"four-point, 1 (accept) to 4 (reject); recoded","field":"economic psychology","wr":"two reviewers on manuscript accept/reject recommendation","conf":"high","self":false,"doi":"10.1016/j.joep.2005.05.004","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":true,"ms":"For 263 manuscripts each given an overall recommendation (recoded accept-to-reject 4-point) by two reviewers, the intraclass correlation was 0.52, the paper's headline agreement figure.","vf":"unverified"},{"key":"KULA287G","au":"Kemp, Simon","y":2005,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intra-class correlation (R1) between two reviewers; different papers reviewed by different referees (cf. Shrout & Fleiss 1979)","estd":"ICC (single/unspec)","v":0.51,"n":"271","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"five-point, 1 (excellent) to 5 (very poor)","field":"economic psychology","wr":"two reviewers on research quality of manuscripts","conf":"high","self":false,"doi":"10.1016/j.joep.2005.05.004","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For 271 manuscripts each rated by two reviewers on research quality, the intraclass correlation between reviewers was 0.51, indicating moderate agreement.","vf":"unverified"},{"key":"KULA287G","au":"Kemp, Simon","y":2005,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intra-class correlation (R1) between two reviewers; different papers reviewed by different referees (cf. Shrout & Fleiss 1979)","estd":"ICC (single/unspec)","v":0.63,"n":"271","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"five-point, 1 (excellent) to 5 (very poor)","field":"economic psychology","wr":"two reviewers on subject-interest of manuscripts","conf":"high","self":false,"doi":"10.1016/j.joep.2005.05.004","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For 271 manuscripts each rated by two reviewers on subject interest, the intraclass correlation between reviewers was 0.63, indicating moderate agreement.","vf":"unverified"},{"key":"KULA287G","au":"Kemp, Simon","y":2005,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intra-class correlation (R1) between two reviewers; different papers reviewed by different referees (cf. Shrout & Fleiss 1979)","estd":"ICC (single/unspec)","v":0.57,"n":"271","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"five-point, 1 (excellent) to 5 (very poor)","field":"economic psychology","wr":"two reviewers on writing quality of manuscripts","conf":"high","self":false,"doi":"10.1016/j.joep.2005.05.004","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For 271 manuscripts each rated by two reviewers on writing quality, the intraclass correlation between reviewers was 0.57, indicating moderate agreement.","vf":"unverified"},{"key":"KULA287G","au":"Kemp, Simon","y":2005,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"percentage of occasions on which the two reviewers made identical judgements (exact agreement), original six categories","estd":"percent agreement","v":0.62,"n":"279","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"six categories (publishable as is ... reject)","field":"economic psychology","wr":"two reviewers exact-matching six-category recommendation","conf":"high","self":false,"doi":"10.1016/j.joep.2005.05.004","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The two reviewers made identical six-category final recommendations for 62.0% of manuscripts; the exact denominator is not stated (279 pairs available, but Table 2 sums to 276).","vf":"unverified"},{"key":"KULA287G","au":"Kemp, Simon","y":2005,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"overall agreement (exact) on final recommendations reduced to reject vs not reject","estd":"percent agreement","v":0.724,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"reject / not reject","field":"economic psychology","wr":"two reviewers agreeing on reject vs not-reject","conf":"high","self":false,"doi":"10.1016/j.joep.2005.05.004","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Reducing the two reviewers' final recommendations to reject vs not-reject (Table 2), overall exact agreement was 72.4%; the denominator is not stated (Table 2 cells sum to 276).","vf":"unverified"},{"key":"KULA287G","au":"Kemp, Simon","y":2005,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"percentage of occasions on which the two reviewers made identical judgements (exact agreement)","estd":"percent agreement","v":0.539,"n":"271","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"five-point, 1 (excellent) to 5 (very poor)","field":"economic psychology","wr":"two reviewers exact-matching research-quality ratings","conf":"high","self":false,"doi":"10.1016/j.joep.2005.05.004","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The two reviewers gave identical research-quality ratings for 53.9% of manuscripts.","vf":"unverified"},{"key":"KULA287G","au":"Kemp, Simon","y":2005,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"percentage of occasions on which the two reviewers made identical judgements (exact agreement)","estd":"percent agreement","v":0.642,"n":"271","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"five-point, 1 (excellent) to 5 (very poor)","field":"economic psychology","wr":"two reviewers exact-matching subject-interest ratings","conf":"high","self":false,"doi":"10.1016/j.joep.2005.05.004","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The two reviewers gave identical subject-interest ratings for 64.2% of manuscripts, a figure the author notes is inflated by no reviewer using the lowest category.","vf":"unverified"},{"key":"KULA287G","au":"Kemp, Simon","y":2005,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"percentage of occasions on which the two reviewers made identical judgements (exact agreement)","estd":"percent agreement","v":0.579,"n":"271","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"five-point, 1 (excellent) to 5 (very poor)","field":"economic psychology","wr":"two reviewers exact-matching writing-quality ratings","conf":"high","self":false,"doi":"10.1016/j.joep.2005.05.004","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The two reviewers gave identical writing-quality ratings for 57.9% of manuscripts.","vf":"unverified"},{"key":"625FMAPA","au":"Kemper, Kathi J.","y":1996,"cx":"Journal","ob":"conference-abstract","fam":"weighted-kappa","form":"weighted kappa calculated for each pair of raters and averaged into an overall weighted k score","estd":"weighted kappa","v":0.19,"n":"246","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 to 5; 1 unsuitable; 5 a 'must'","field":"paediatrics","wr":"reviewers on submitted abstract scores","conf":"high","self":false,"doi":"10.1001/archpedi.1996.02170290046007","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"All 11 board members scored each of the 246 abstracts submitted in 1990 on a 1 to 5 scale; weighted kappa was computed for every rater pair and averaged, giving 0.19, indicating poor chance-corrected agreement.","vf":"unverified"},{"key":"625FMAPA","au":"Kemper, Kathi J.","y":1996,"cx":"Journal","ob":"conference-abstract","fam":"percent-agreement","form":"weighted percentage agreement averaged over all rater pairs; weights: perfect=1, 1pt=0.75, 2pt=0.5, 3pt=0.25, 4pt=0","estd":"percent agreement","v":0.79,"n":"246","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 to 5; 1 unsuitable; 5 a 'must'","field":"paediatrics","wr":"reviewers on submitted abstract scores","conf":"high","self":false,"doi":"10.1001/archpedi.1996.02170290046007","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Across all pairs of the 11 board members scoring the 246 abstracts submitted in 1990, weighted percentage agreement averaged 79 percent, which the authors classed as only fair.","vf":"unverified"},{"key":"625FMAPA","au":"Kemper, Kathi J.","y":1996,"cx":"Journal","ob":"conference-abstract","fam":"weighted-kappa","form":"weighted kappa calculated for each pair of raters and averaged into an overall weighted k score","estd":"weighted kappa","v":0.24,"n":"43","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 to 5; 1 unsuitable; 5 a 'must'","field":"paediatrics","wr":"reviewers on submitted abstract scores","conf":"high","self":false,"doi":"10.1001/archpedi.1996.02170290046007","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"In 1995, five raters (the behavioral-pediatrics SIG chairperson, two regional chairpersons, and two board members) scored 43 behavioral-pediatrics abstracts; weighted kappa averaged across rater pairs was 0.24, indicating poor chance-corrected agreement.","vf":"unverified"},{"key":"625FMAPA","au":"Kemper, Kathi J.","y":1996,"cx":"Journal","ob":"conference-abstract","fam":"percent-agreement","form":"weighted percentage agreement averaged over all rater pairs; weights: perfect=1, 1pt=0.75, 2pt=0.5, 3pt=0.25, 4pt=0","estd":"percent agreement","v":0.79,"n":"43","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 to 5; 1 unsuitable; 5 a 'must'","field":"paediatrics","wr":"reviewers on submitted abstract scores","conf":"high","self":false,"doi":"10.1001/archpedi.1996.02170290046007","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Across pairs of the five raters scoring the 43 behavioral-pediatrics abstracts in 1995, weighted percentage agreement averaged 79 percent.","vf":"unverified"},{"key":"625FMAPA","au":"Kemper, Kathi J.","y":1996,"cx":"Journal","ob":"conference-abstract","fam":"weighted-kappa","form":"weighted kappa calculated for each pair of raters and averaged into an overall weighted k score","estd":"weighted kappa","v":0.2,"n":"118","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 to 5; 1 unsuitable; 5 a 'must'","field":"paediatrics","wr":"reviewers on submitted abstract scores","conf":"high","self":false,"doi":"10.1001/archpedi.1996.02170290046007","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"In 1995, four raters (the emergency-medicine SIG chairperson, two regional chairpersons, and one board member) scored 118 emergency-medicine abstracts; weighted kappa averaged across rater pairs was 0.20, indicating poor chance-corrected agreement.","vf":"unverified"},{"key":"625FMAPA","au":"Kemper, Kathi J.","y":1996,"cx":"Journal","ob":"conference-abstract","fam":"percent-agreement","form":"weighted percentage agreement averaged over all rater pairs; weights: perfect=1, 1pt=0.75, 2pt=0.5, 3pt=0.25, 4pt=0","estd":"percent agreement","v":0.79,"n":"118","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 to 5; 1 unsuitable; 5 a 'must'","field":"paediatrics","wr":"reviewers on submitted abstract scores","conf":"high","self":false,"doi":"10.1001/archpedi.1996.02170290046007","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Across pairs of the four raters scoring the 118 emergency-medicine abstracts in 1995, weighted percentage agreement averaged 79 percent.","vf":"unverified"},{"key":"625FMAPA","au":"Kemper, Kathi J.","y":1996,"cx":"Journal","ob":"conference-abstract","fam":"weighted-kappa","form":"weighted kappa calculated for each pair of raters and averaged into an overall weighted k score","estd":"weighted kappa","v":0.15,"n":"246","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 to 5; 1 unsuitable; 5 a 'must'","field":"paediatrics","wr":"reviewers on submitted abstract scores","conf":"high","self":false,"doi":"10.1001/archpedi.1996.02170290046007","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"In 1995, an 11-member committee (five board members and six regional chairpersons) scored 246 general-pediatrics abstracts, each abstract by at least five raters; weighted kappa averaged across rater pairs was 0.15, indicating poor chance-corrected agreement.","vf":"unverified"},{"key":"625FMAPA","au":"Kemper, Kathi J.","y":1996,"cx":"Journal","ob":"conference-abstract","fam":"percent-agreement","form":"weighted percentage agreement averaged over all rater pairs; weights: perfect=1, 1pt=0.75, 2pt=0.5, 3pt=0.25, 4pt=0","estd":"percent agreement","v":0.77,"n":"246","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 to 5; 1 unsuitable; 5 a 'must'","field":"paediatrics","wr":"reviewers on submitted abstract scores","conf":"high","self":false,"doi":"10.1001/archpedi.1996.02170290046007","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 246 general-pediatrics abstracts scored by the 11-member committee in 1995, weighted percentage agreement averaged 77 percent across rater pairs.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC for the average combined ratings, agreement not consistency","estd":"ICC (single/unspec)","v":0.14,"n":"","k":"2","samp":"full-pool","blind":"double","agg":"unspecified","scale":"mean of eight items, each 0-5","field":"social work","wr":"reviewers on eight-item mean manuscript rating","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":-0.28,"ciHigh":0.5,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"With item 7 (substantial missing data) excluded, the agreement-form ICC between the two reviewers' eight-item composite scores was 0.14, again indicating low overall reliability.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC for the average combined ratings, agreement not consistency","estd":"ICC (single/unspec)","v":0.13,"n":"54","k":"2","samp":"full-pool","blind":"double","agg":"unspecified","scale":"mean of nine items, each 0-5","field":"social work","wr":"reviewers on nine-item mean manuscript rating","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":-0.31,"ciHigh":0.51,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Each reviewer's ratings were averaged across all nine items into one composite score for 54 manuscripts; the agreement-form ICC between the two reviewers' composites was 0.13, indicating low overall reliability.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Correlations","estd":"correlation","v":0.33,"n":"54","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"mean of nine items, each 0-5","field":"social work","wr":"reviewers on nine-item mean manuscript rating","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Raters 1 and 2 each produced a nine-item composite rating for 54 manuscripts; the correlation between their composite scores was 0.33.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Correlations","estd":"correlation","v":0.36,"n":"32","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"mean of nine items, each 0-5","field":"social work","wr":"reviewers on nine-item mean manuscript rating","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Raters 1 and 3 both rated 32 manuscripts; the correlation between their nine-item composite scores was 0.36.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Correlations","estd":"correlation","v":0.57,"n":"32","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"mean of nine items, each 0-5","field":"social work","wr":"reviewers on nine-item mean manuscript rating","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Raters 2 and 3 both rated 32 manuscripts; the correlation between their nine-item composite scores was 0.57.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement (two reviewers gave identical rating)","estd":"percent agreement","v":0.315,"n":"54","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript importance to social work","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers each rated 54 manuscripts on item 1 (importance to social work) on a 0-5 scale; they gave exactly the same rating 31.5% of the time.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement (two reviewers gave identical rating)","estd":"percent agreement","v":0.326,"n":"43","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript prior studies cited","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers each rated 43 manuscripts on item 2 (prior studies cited) on a 0-5 scale; they gave exactly the same rating 32.6% of the time.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement (two reviewers gave identical rating)","estd":"percent agreement","v":0.279,"n":"43","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript novelty of information","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers each rated 43 manuscripts on item 3 (novelty of information) on a 0-5 scale; they gave exactly the same rating 27.9% of the time.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement (two reviewers gave identical rating)","estd":"percent agreement","v":0.262,"n":"42","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript research design suitability","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers each rated 42 manuscripts on item 4 (research design suitability) on a 0-5 scale; they gave exactly the same rating 26.2% of the time.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement (two reviewers gave identical rating)","estd":"percent agreement","v":0.277,"n":"47","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript methods described adequately","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers each rated 47 manuscripts on item 5 (methods described adequately) on a 0-5 scale; they gave exactly the same rating 27.7% of the time.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement (two reviewers gave identical rating)","estd":"percent agreement","v":0.233,"n":"43","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript statistics described adequately","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers each rated 43 manuscripts on item 6 (statistics described adequately) on a 0-5 scale; they gave exactly the same rating 23.3% of the time.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement (two reviewers gave identical rating)","estd":"percent agreement","v":0.147,"n":"34","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript statistics chosen/carried out correctly","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers each rated 34 manuscripts on item 7 (statistics chosen/carried out correctly) on a 0-5 scale; they gave exactly the same rating 14.7% of the time.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement (two reviewers gave identical rating)","estd":"percent agreement","v":0.159,"n":"44","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript data support conclusions","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers each rated 44 manuscripts on item 8 (data support conclusions) on a 0-5 scale; they gave exactly the same rating 15.9% of the time.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement (two reviewers gave identical rating)","estd":"percent agreement","v":0.185,"n":"54","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript clear, organized writing style","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers each rated 54 manuscripts on item 9 (clear, organized writing style) on a 0-5 scale; they gave exactly the same rating 18.5% of the time.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC designed to measure agreement rather than consistency","estd":"ICC (single/unspec)","v":0.175,"n":"54","k":"2","samp":"full-pool","blind":"double","agg":"unspecified","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript importance to social work","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":-0.092,"ciHigh":0.423,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated 54 manuscripts on item 1 (importance to social work); an absolute-agreement ICC of 0.175 indicates low chance-corrected agreement between the two raters.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC designed to measure agreement rather than consistency","estd":"ICC (single/unspec)","v":0.345,"n":"43","k":"2","samp":"full-pool","blind":"double","agg":"unspecified","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript prior studies cited","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":0.05,"ciHigh":0.584,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated 43 manuscripts on item 2 (prior studies cited); an absolute-agreement ICC of 0.345 indicates low chance-corrected agreement between the two raters.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC designed to measure agreement rather than consistency","estd":"ICC (single/unspec)","v":0.283,"n":"43","k":"2","samp":"full-pool","blind":"double","agg":"unspecified","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript novelty of information","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":-0.017,"ciHigh":0.535,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated 43 manuscripts on item 3 (novelty of information); an absolute-agreement ICC of 0.283 indicates low chance-corrected agreement between the two raters.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC designed to measure agreement rather than consistency","estd":"ICC (single/unspec)","v":0.404,"n":"42","k":"2","samp":"full-pool","blind":"double","agg":"unspecified","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript research design suitability","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":0.125,"ciHigh":0.626,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated 42 manuscripts on item 4 (research design suitability); an absolute-agreement ICC of 0.404 indicates low chance-corrected agreement between the two raters.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC designed to measure agreement rather than consistency","estd":"ICC (single/unspec)","v":0.128,"n":"47","k":"2","samp":"full-pool","blind":"double","agg":"unspecified","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript methods described adequately","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":-0.166,"ciHigh":0.4,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated 47 manuscripts on item 5 (methods described adequately); an absolute-agreement ICC of 0.128 indicates low chance-corrected agreement between the two raters.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC designed to measure agreement rather than consistency","estd":"ICC (single/unspec)","v":0.127,"n":"43","k":"2","samp":"full-pool","blind":"double","agg":"unspecified","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript statistics described adequately","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":-0.181,"ciHigh":0.411,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated 43 manuscripts on item 6 (statistics described adequately); an absolute-agreement ICC of 0.127 indicates low chance-corrected agreement between the two raters.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC designed to measure agreement rather than consistency","estd":"ICC (single/unspec)","v":-0.015,"n":"34","k":"2","samp":"full-pool","blind":"double","agg":"unspecified","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript statistics chosen/carried out correctly","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":-0.36,"ciHigh":0.327,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated 34 manuscripts on item 7 (statistics chosen/carried out correctly); an absolute-agreement ICC of -0.015 indicates low chance-corrected agreement between the two raters.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC designed to measure agreement rather than consistency","estd":"ICC (single/unspec)","v":0.175,"n":"44","k":"2","samp":"full-pool","blind":"double","agg":"unspecified","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript data support conclusions","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":-0.123,"ciHigh":0.444,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated 44 manuscripts on item 8 (data support conclusions); an absolute-agreement ICC of 0.175 indicates low chance-corrected agreement between the two raters.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC designed to measure agreement rather than consistency","estd":"ICC (single/unspec)","v":0.279,"n":"54","k":"2","samp":"full-pool","blind":"double","agg":"unspecified","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript clear, organized writing style","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":0.011,"ciHigh":0.509,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated 54 manuscripts on item 9 (clear, organized writing style); an absolute-agreement ICC of 0.279 indicates low chance-corrected agreement between the two raters.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"within one point (one-step agreement)","estd":"percent agreement","v":0.648,"n":"54","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript importance to social work","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated 54 manuscripts on item 1 (importance to social work); their ratings fell within one scale point of each other 64.8% of the time (one-step agreement).","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"within one point (one-step agreement)","estd":"percent agreement","v":0.674,"n":"43","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript prior studies cited","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated 43 manuscripts on item 2 (prior studies cited); their ratings fell within one scale point of each other 67.4% of the time (one-step agreement).","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"within one point (one-step agreement)","estd":"percent agreement","v":0.698,"n":"43","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript novelty of information","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated 43 manuscripts on item 3 (novelty of information); their ratings fell within one scale point of each other 69.8% of the time (one-step agreement).","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"within one point (one-step agreement)","estd":"percent agreement","v":0.619,"n":"42","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript research design suitability","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated 42 manuscripts on item 4 (research design suitability); their ratings fell within one scale point of each other 61.9% of the time (one-step agreement).","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"within one point (one-step agreement)","estd":"percent agreement","v":0.574,"n":"47","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript methods described adequately","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated 47 manuscripts on item 5 (methods described adequately); their ratings fell within one scale point of each other 57.4% of the time (one-step agreement).","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"within one point (one-step agreement)","estd":"percent agreement","v":0.581,"n":"43","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript statistics described adequately","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated 43 manuscripts on item 6 (statistics described adequately); their ratings fell within one scale point of each other 58.1% of the time (one-step agreement).","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"within one point (one-step agreement)","estd":"percent agreement","v":0.559,"n":"34","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript statistics chosen/carried out correctly","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated 34 manuscripts on item 7 (statistics chosen/carried out correctly); their ratings fell within one scale point of each other 55.9% of the time (one-step agreement).","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"within one point (one-step agreement)","estd":"percent agreement","v":0.591,"n":"44","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript data support conclusions","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated 44 manuscripts on item 8 (data support conclusions); their ratings fell within one scale point of each other 59.1% of the time (one-step agreement).","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"within one point (one-step agreement)","estd":"percent agreement","v":0.63,"n":"54","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript clear, organized writing style","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated 54 manuscripts on item 9 (clear, organized writing style); their ratings fell within one scale point of each other 63.0% of the time (one-step agreement).","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement, averaged over the three reviewer pairs","estd":"percent agreement","v":0.3233,"n":"","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript importance to social work","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For manuscripts with three reviewers, exact agreement on item 1 (importance to social work) averaged across the three reviewer pairs (A-B, B-C, A-C) was 32.33%; the per-item number of manuscripts is not reported in Table 2.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement, averaged over the three reviewer pairs","estd":"percent agreement","v":0.236,"n":"","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript prior studies cited","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For manuscripts with three reviewers, exact agreement on item 2 (prior studies cited) averaged across the three reviewer pairs (A-B, B-C, A-C) was 23.6%; the per-item number of manuscripts is not reported in Table 2.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement, averaged over the three reviewer pairs","estd":"percent agreement","v":0.244,"n":"","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript novelty of information","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For manuscripts with three reviewers, exact agreement on item 3 (novelty of information) averaged across the three reviewer pairs (A-B, B-C, A-C) was 24.4%; the per-item number of manuscripts is not reported in Table 2.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement, averaged over the three reviewer pairs","estd":"percent agreement","v":0.1973,"n":"","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript research design suitability","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For manuscripts with three reviewers, exact agreement on item 4 (research design suitability) averaged across the three reviewer pairs (A-B, B-C, A-C) was 19.73%; the per-item number of manuscripts is not reported in Table 2.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement, averaged over the three reviewer pairs","estd":"percent agreement","v":0.274,"n":"","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript methods described adequately","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For manuscripts with three reviewers, exact agreement on item 5 (methods described adequately) averaged across the three reviewer pairs (A-B, B-C, A-C) was 27.4%; the per-item number of manuscripts is not reported in Table 2.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement, averaged over the three reviewer pairs","estd":"percent agreement","v":0.2213,"n":"","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript statistics described adequately","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For manuscripts with three reviewers, exact agreement on item 6 (statistics described adequately) averaged across the three reviewer pairs (A-B, B-C, A-C) was 22.13%; the per-item number of manuscripts is not reported in Table 2.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement, averaged over the three reviewer pairs","estd":"percent agreement","v":0.1757,"n":"","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript statistics chosen/carried out correctly","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For manuscripts with three reviewers, exact agreement on item 7 (statistics chosen/carried out correctly) averaged across the three reviewer pairs (A-B, B-C, A-C) was 17.57%; the per-item number of manuscripts is not reported in Table 2.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement, averaged over the three reviewer pairs","estd":"percent agreement","v":0.1977,"n":"","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript data support conclusions","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For manuscripts with three reviewers, exact agreement on item 8 (data support conclusions) averaged across the three reviewer pairs (A-B, B-C, A-C) was 19.77%; the per-item number of manuscripts is not reported in Table 2.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement, averaged over the three reviewer pairs","estd":"percent agreement","v":0.3117,"n":"","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript clear, organized writing style","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For manuscripts with three reviewers, exact agreement on item 9 (clear, organized writing style) averaged across the three reviewer pairs (A-B, B-C, A-C) was 31.17%; the per-item number of manuscripts is not reported in Table 2.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"within one point (one-step agreement), averaged over the three reviewer pairs","estd":"percent agreement","v":0.6067,"n":"","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript importance to social work","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For manuscripts with three reviewers, one-step (within one point) agreement on item 1 (importance to social work), averaged across the three reviewer pairs, was 60.67%; the per-item number of manuscripts is not reported in Table 2.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"within one point (one-step agreement), averaged over the three reviewer pairs","estd":"percent agreement","v":0.526,"n":"","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript prior studies cited","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For manuscripts with three reviewers, one-step (within one point) agreement on item 2 (prior studies cited), averaged across the three reviewer pairs, was 52.6%; the per-item number of manuscripts is not reported in Table 2.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"within one point (one-step agreement), averaged over the three reviewer pairs","estd":"percent agreement","v":0.6037,"n":"","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript novelty of information","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For manuscripts with three reviewers, one-step (within one point) agreement on item 3 (novelty of information), averaged across the three reviewer pairs, was 60.37%; the per-item number of manuscripts is not reported in Table 2.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"within one point (one-step agreement), averaged over the three reviewer pairs","estd":"percent agreement","v":0.5257,"n":"","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript research design suitability","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For manuscripts with three reviewers, one-step (within one point) agreement on item 4 (research design suitability), averaged across the three reviewer pairs, was 52.57%; the per-item number of manuscripts is not reported in Table 2.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"within one point (one-step agreement), averaged over the three reviewer pairs","estd":"percent agreement","v":0.5413,"n":"","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript methods described adequately","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For manuscripts with three reviewers, one-step (within one point) agreement on item 5 (methods described adequately), averaged across the three reviewer pairs, was 54.13%; the per-item number of manuscripts is not reported in Table 2.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"within one point (one-step agreement), averaged over the three reviewer pairs","estd":"percent agreement","v":0.586,"n":"","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript statistics described adequately","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For manuscripts with three reviewers, one-step (within one point) agreement on item 6 (statistics described adequately), averaged across the three reviewer pairs, was 58.6%; the per-item number of manuscripts is not reported in Table 2.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"within one point (one-step agreement), averaged over the three reviewer pairs","estd":"percent agreement","v":0.546,"n":"","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript statistics chosen/carried out correctly","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For manuscripts with three reviewers, one-step (within one point) agreement on item 7 (statistics chosen/carried out correctly), averaged across the three reviewer pairs, was 54.6%; the per-item number of manuscripts is not reported in Table 2.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"within one point (one-step agreement), averaged over the three reviewer pairs","estd":"percent agreement","v":0.6173,"n":"","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript data support conclusions","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For manuscripts with three reviewers, one-step (within one point) agreement on item 8 (data support conclusions), averaged across the three reviewer pairs, was 61.73%; the per-item number of manuscripts is not reported in Table 2.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"within one point (one-step agreement), averaged over the three reviewer pairs","estd":"percent agreement","v":0.679,"n":"","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript clear, organized writing style","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For manuscripts with three reviewers, one-step (within one point) agreement on item 9 (clear, organized writing style), averaged across the three reviewer pairs, was 67.9%; the per-item number of manuscripts is not reported in Table 2.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC across all three raters, agreement not consistency","estd":"ICC (single/unspec)","v":0.116,"n":"29","k":"3","samp":"full-pool","blind":"double","agg":"unspecified","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript importance to social work","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":-0.092,"ciHigh":0.374,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Across all three reviewers of 29 manuscripts, an absolute-agreement ICC of 0.116 on item 1 (importance to social work) indicates low chance-corrected agreement among the three raters.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC across all three raters, agreement not consistency","estd":"ICC (single/unspec)","v":0.239,"n":"24","k":"3","samp":"full-pool","blind":"double","agg":"unspecified","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript prior studies cited","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":-0.01,"ciHigh":0.517,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Across all three reviewers of 24 manuscripts, an absolute-agreement ICC of 0.239 on item 2 (prior studies cited) indicates low chance-corrected agreement among the three raters.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC across all three raters, agreement not consistency","estd":"ICC (single/unspec)","v":0.082,"n":"20","k":"3","samp":"full-pool","blind":"double","agg":"unspecified","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript novelty of information","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":-0.163,"ciHigh":0.403,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Across all three reviewers of 20 manuscripts, an absolute-agreement ICC of 0.082 on item 3 (novelty of information) indicates low chance-corrected agreement among the three raters.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC across all three raters, agreement not consistency","estd":"ICC (single/unspec)","v":0.066,"n":"22","k":"3","samp":"full-pool","blind":"double","agg":"unspecified","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript research design suitability","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":-0.162,"ciHigh":0.368,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Across all three reviewers of 22 manuscripts, an absolute-agreement ICC of 0.066 on item 4 (research design suitability) indicates low chance-corrected agreement among the three raters.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC across all three raters, agreement not consistency","estd":"ICC (single/unspec)","v":0.227,"n":"27","k":"3","samp":"full-pool","blind":"double","agg":"unspecified","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript methods described adequately","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":-0.003,"ciHigh":0.488,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Across all three reviewers of 27 manuscripts, an absolute-agreement ICC of 0.227 on item 5 (methods described adequately) indicates low chance-corrected agreement among the three raters.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC across all three raters, agreement not consistency","estd":"ICC (single/unspec)","v":0.302,"n":"25","k":"3","samp":"full-pool","blind":"double","agg":"unspecified","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript statistics described adequately","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":0.06,"ciHigh":0.561,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Across all three reviewers of 25 manuscripts, an absolute-agreement ICC of 0.302 on item 6 (statistics described adequately) indicates low chance-corrected agreement among the three raters.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC across all three raters, agreement not consistency","estd":"ICC (single/unspec)","v":0.014,"n":"16","k":"3","samp":"full-pool","blind":"double","agg":"unspecified","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript statistics chosen/carried out correctly","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":-0.243,"ciHigh":0.383,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Across all three reviewers of 16 manuscripts, an absolute-agreement ICC of 0.014 on item 7 (statistics chosen/carried out correctly) indicates low chance-corrected agreement among the three raters.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC across all three raters, agreement not consistency","estd":"ICC (single/unspec)","v":0.296,"n":"21","k":"3","samp":"full-pool","blind":"double","agg":"unspecified","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript data support conclusions","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":0.044,"ciHigh":0.575,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Across all three reviewers of 21 manuscripts, an absolute-agreement ICC of 0.296 on item 8 (data support conclusions) indicates low chance-corrected agreement among the three raters.","vf":"unverified"},{"key":"STUQWYY7","au":"Kirk, S. A.","y":1997,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC across all three raters, agreement not consistency","estd":"ICC (single/unspec)","v":0.403,"n":"32","k":"3","samp":"full-pool","blind":"double","agg":"unspecified","scale":"0 = fails by a large amount to 5 = succeeds","field":"social work","wr":"reviewers on manuscript clear, organized writing style","conf":"high","self":false,"doi":"10.1093/swr/21.2.121","ciLow":0.185,"ciHigh":0.616,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Across all three reviewers of 32 manuscripts, an absolute-agreement ICC of 0.403 on item 9 (clear, organized writing style) indicates low chance-corrected agreement among the three raters.","vf":"unverified"},{"key":"RJR9QVGF","au":"Koch, Thomas","y":2019,"cx":"Journal","ob":"conference-abstract","fam":"other","form":"Brennan and Prediger's kappa, ordinal (Gwet 2014 programming)","estd":"other","v":0.463,"n":"","k":"2.57","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5, higher = better (harmonised)","field":"communication science","wr":"reviewers on submission fit with conference theme","conf":"high","self":false,"doi":"10.5771/2192-4007-2019-2-203","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For fit with conference theme, Brennan and Prediger's ordinal kappa was 0.463, the lowest of the criteria on this measure. Sample is 904 rating pairs/triples.","vf":"unverified"},{"key":"RJR9QVGF","au":"Koch, Thomas","y":2019,"cx":"Journal","ob":"conference-abstract","fam":"other","form":"Brennan and Prediger's kappa, ordinal (Gwet 2014 programming)","estd":"other","v":0.757,"n":"","k":"2.57","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5, higher = better (harmonised)","field":"communication science","wr":"reviewers on each submission's mean score across criteria","conf":"high","self":false,"doi":"10.5771/2192-4007-2019-2-203","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"When reviewers' scores were averaged across the six criteria, Brennan and Prediger's ordinal kappa reached 0.757, the highest value and the only one meeting the 0.667 minimum threshold. Sample is 1311 rating pairs/triples.","vf":"unverified"},{"key":"RJR9QVGF","au":"Koch, Thomas","y":2019,"cx":"Journal","ob":"conference-abstract","fam":"other","form":"Brennan and Prediger's kappa, ordinal (Gwet 2014 programming)","estd":"other","v":0.506,"n":"","k":"2.57","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5, higher = better (harmonised)","field":"communication science","wr":"reviewers on submission method","conf":"high","self":false,"doi":"10.5771/2192-4007-2019-2-203","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For method, Brennan and Prediger's ordinal kappa was 0.506, within the mid-range of criteria on this measure. Sample is 1306 rating pairs/triples.","vf":"unverified"},{"key":"RJR9QVGF","au":"Koch, Thomas","y":2019,"cx":"Journal","ob":"conference-abstract","fam":"other","form":"Brennan and Prediger's kappa, ordinal (Gwet 2014 programming)","estd":"other","v":0.524,"n":"","k":"2.57","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5, higher = better (harmonised)","field":"communication science","wr":"reviewers on submission originality","conf":"high","self":false,"doi":"10.5771/2192-4007-2019-2-203","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For originality, Brennan and Prediger's ordinal kappa was 0.524, within the mid-range of criteria on this measure. Sample is 815 rating pairs/triples.","vf":"unverified"},{"key":"RJR9QVGF","au":"Koch, Thomas","y":2019,"cx":"Journal","ob":"conference-abstract","fam":"other","form":"Brennan and Prediger's kappa, ordinal (Gwet 2014 programming)","estd":"other","v":0.649,"n":"1409","k":"2.57","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5, higher = better (harmonised)","field":"communication science","wr":"reviewers on conference submission scores","conf":"high","self":false,"doi":"10.5771/2192-4007-2019-2-203","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Pooling all criteria, Brennan and Prediger's ordinal kappa (uniform-chance model, Gwet programming) was 0.649, more favourable than Krippendorff's alpha but still below the recommended threshold. Fallzahl 8486 is rating pairs/triples.","vf":"unverified"},{"key":"RJR9QVGF","au":"Koch, Thomas","y":2019,"cx":"Journal","ob":"conference-abstract","fam":"other","form":"Brennan and Prediger's kappa, ordinal (Gwet 2014 programming)","estd":"other","v":0.529,"n":"","k":"2.57","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5, higher = better (harmonised)","field":"communication science","wr":"reviewers on a separate overall submission rating","conf":"high","self":false,"doi":"10.5771/2192-4007-2019-2-203","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For a separate overall rating, Brennan and Prediger's ordinal kappa was 0.529, comparable to the single criteria; requested at few conferences, so interpret with caution. Sample is 278 rating pairs/triples.","vf":"unverified"},{"key":"RJR9QVGF","au":"Koch, Thomas","y":2019,"cx":"Journal","ob":"conference-abstract","fam":"other","form":"Brennan and Prediger's kappa, ordinal (Gwet 2014 programming)","estd":"other","v":0.522,"n":"","k":"2.57","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5, higher = better (harmonised)","field":"communication science","wr":"reviewers on submission clarity of presentation","conf":"high","self":false,"doi":"10.5771/2192-4007-2019-2-203","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For clarity of presentation, Brennan and Prediger's ordinal kappa was 0.522, within the mid-range of criteria on this measure. Sample is 1265 rating pairs/triples.","vf":"unverified"},{"key":"RJR9QVGF","au":"Koch, Thomas","y":2019,"cx":"Journal","ob":"conference-abstract","fam":"other","form":"Brennan and Prediger's kappa, ordinal (Gwet 2014 programming)","estd":"other","v":0.581,"n":"","k":"2.57","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5, higher = better (harmonised)","field":"communication science","wr":"reviewers on submission relevance","conf":"high","self":false,"doi":"10.5771/2192-4007-2019-2-203","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For relevance, Brennan and Prediger's ordinal kappa was 0.581, the highest among the single criteria on this measure. Sample is 1342 rating pairs/triples.","vf":"unverified"},{"key":"RJR9QVGF","au":"Koch, Thomas","y":2019,"cx":"Journal","ob":"conference-abstract","fam":"other","form":"Brennan and Prediger's kappa, ordinal (Gwet 2014 programming)","estd":"other","v":0.525,"n":"","k":"2.57","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5, higher = better (harmonised)","field":"communication science","wr":"reviewers on submission theoretical grounding","conf":"high","self":false,"doi":"10.5771/2192-4007-2019-2-203","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For theoretical grounding, Brennan and Prediger's ordinal kappa was 0.525, within the mid-range of criteria on this measure. Sample is 1265 rating pairs/triples.","vf":"unverified"},{"key":"RJR9QVGF","au":"Koch, Thomas","y":2019,"cx":"Journal","ob":"conference-abstract","fam":"Krippendorff","form":"Krippendorff's alpha, ordinal variant","estd":"Krippendorff","v":0.207,"n":"","k":"2.57","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5, higher = better (harmonised)","field":"communication science","wr":"reviewers on submission fit with conference theme","conf":"high","self":false,"doi":"10.5771/2192-4007-2019-2-203","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the criterion fit with the conference theme, ordinal Krippendorff's alpha was 0.207 across the reviewers' scores, indicating low agreement. Sample is 904 rating pairs/triples.","vf":"unverified"},{"key":"RJR9QVGF","au":"Koch, Thomas","y":2019,"cx":"Journal","ob":"conference-abstract","fam":"Krippendorff","form":"Krippendorff's alpha, ordinal variant","estd":"Krippendorff","v":0.293,"n":"","k":"2.57","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5, higher = better (harmonised)","field":"communication science","wr":"reviewers on each submission's mean score across criteria","conf":"high","self":false,"doi":"10.5771/2192-4007-2019-2-203","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"When reviewers' scores were averaged across the six criteria, ordinal Krippendorff's alpha rose to 0.293, the highest alpha, though still low agreement. Sample is 1311 rating pairs/triples.","vf":"unverified"},{"key":"RJR9QVGF","au":"Koch, Thomas","y":2019,"cx":"Journal","ob":"conference-abstract","fam":"Krippendorff","form":"Krippendorff's alpha, ordinal variant","estd":"Krippendorff","v":0.229,"n":"","k":"2.57","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5, higher = better (harmonised)","field":"communication science","wr":"reviewers on submission method","conf":"high","self":false,"doi":"10.5771/2192-4007-2019-2-203","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the method criterion, ordinal Krippendorff's alpha was 0.229 across reviewers' scores, indicating low agreement. Sample is 1306 rating pairs/triples.","vf":"unverified"},{"key":"RJR9QVGF","au":"Koch, Thomas","y":2019,"cx":"Journal","ob":"conference-abstract","fam":"Krippendorff","form":"Krippendorff's alpha, ordinal variant","estd":"Krippendorff","v":0.165,"n":"","k":"2.57","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5, higher = better (harmonised)","field":"communication science","wr":"reviewers on submission originality","conf":"high","self":false,"doi":"10.5771/2192-4007-2019-2-203","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the originality criterion, ordinal Krippendorff's alpha was 0.165, the weakest of the criteria, indicating very low agreement between reviewers. Sample is 815 rating pairs/triples.","vf":"unverified"},{"key":"RJR9QVGF","au":"Koch, Thomas","y":2019,"cx":"Journal","ob":"conference-abstract","fam":"Krippendorff","form":"Krippendorff's alpha, ordinal variant","estd":"Krippendorff","v":0.276,"n":"1409","k":"2.57","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5, higher = better (harmonised)","field":"communication science","wr":"reviewers on conference submission scores","conf":"high","self":false,"doi":"10.5771/2192-4007-2019-2-203","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Reviewers (usually two or three per paper) independently scored 1409 communication-science conference submissions; pooling all criteria, an ordinal Krippendorff's alpha of 0.276 indicates low inter-rater agreement. Fallzahl 8486 is rating pairs/triples.","vf":"unverified"},{"key":"RJR9QVGF","au":"Koch, Thomas","y":2019,"cx":"Journal","ob":"conference-abstract","fam":"Krippendorff","form":"Krippendorff's alpha, ordinal variant","estd":"Krippendorff","v":0.237,"n":"","k":"2.57","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5, higher = better (harmonised)","field":"communication science","wr":"reviewers on a separate overall submission rating","conf":"high","self":false,"doi":"10.5771/2192-4007-2019-2-203","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For a separate overall rating (distinct from the computed mean), ordinal Krippendorff's alpha was 0.237, still low; requested at few conferences, so interpret with caution. Sample is 278 rating pairs/triples.","vf":"unverified"},{"key":"RJR9QVGF","au":"Koch, Thomas","y":2019,"cx":"Journal","ob":"conference-abstract","fam":"Krippendorff","form":"Krippendorff's alpha, ordinal variant","estd":"Krippendorff","v":0.246,"n":"","k":"2.57","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5, higher = better (harmonised)","field":"communication science","wr":"reviewers on submission clarity of presentation","conf":"high","self":false,"doi":"10.5771/2192-4007-2019-2-203","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the clarity of presentation criterion, ordinal Krippendorff's alpha was 0.246 across reviewers' scores, the best of the single criteria but still low. Sample is 1265 rating pairs/triples.","vf":"unverified"},{"key":"RJR9QVGF","au":"Koch, Thomas","y":2019,"cx":"Journal","ob":"conference-abstract","fam":"Krippendorff","form":"Krippendorff's alpha, ordinal variant","estd":"Krippendorff","v":0.173,"n":"","k":"2.57","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5, higher = better (harmonised)","field":"communication science","wr":"reviewers on submission relevance","conf":"high","self":false,"doi":"10.5771/2192-4007-2019-2-203","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the relevance criterion, ordinal Krippendorff's alpha was 0.173, among the weakest criteria, indicating very low agreement between reviewers. Sample is 1342 rating pairs/triples.","vf":"unverified"},{"key":"RJR9QVGF","au":"Koch, Thomas","y":2019,"cx":"Journal","ob":"conference-abstract","fam":"Krippendorff","form":"Krippendorff's alpha, ordinal variant","estd":"Krippendorff","v":0.227,"n":"","k":"2.57","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5, higher = better (harmonised)","field":"communication science","wr":"reviewers on submission theoretical grounding","conf":"high","self":false,"doi":"10.5771/2192-4007-2019-2-203","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the theoretical grounding criterion, ordinal Krippendorff's alpha was 0.227 across reviewers' scores, indicating low agreement. Sample is 1265 rating pairs/triples.","vf":"unverified"},{"key":"RJR9QVGF","au":"Koch, Thomas","y":2019,"cx":"Journal","ob":"conference-abstract","fam":"percent-agreement","form":"exact agreement: score range 0 across the 2-3 reviewers (identical scores)","estd":"percent agreement","v":0.27,"n":"1409","k":"2.57","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5, higher = better (harmonised)","field":"communication science","wr":"reviewers on conference submission scores (exact match)","conf":"high","self":false,"doi":"10.5771/2192-4007-2019-2-203","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Across all pooled ratings, 27 per cent of paper-criterion rating sets had a range of zero, meaning the usually two or three reviewers gave exactly identical scores; uniform-chance agreement would be about 20 per cent. Sample is 8,458 rating pairs/triples on 1,409 papers.","vf":"unverified"},{"key":"EJFLDVJU","au":"Kowalczuk, Maria","y":2015,"cx":"Journal","ob":"review-report","fam":"weighted-kappa","form":"weighted kappa with quadratic weights","estd":"weighted kappa","v":0.4,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1-5 per RQI question, e.g. not at all to discussed extensively","field":"biomedical","wr":"two staff rating quality of reviewer reports","conf":"low","self":false,"doi":"10.1136/bmjopen-2015-008707","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":true,"he":false,"ms":"Two senior editorial staff independently rated the quality of reviewer reports on eight RQI questions for two BMC journals; weighted kappa values, reported as generally around 0.4 or higher, indicate moderate inter-rater agreement. The main text gives only this approximate characterisation; exact per-question coefficients sit in supplementary tables not present here.","vf":"unverified"},{"key":"MM5CTRWM","au":"Kravitz, Richard L","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"Deputy Editor ICC","estd":"ICC (single/unspec)","v":0.03,"n":"2264","k":"1","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"reject or not reject","field":"biomedical, general internal medicine","wr":"deputy editors on reject vs not-reject decision","conf":"high","self":false,"doi":"10.1371/journal.pone.0010072","ciLow":0.02,"ciHigh":0.08,"mt":"other","tgt":"editor-decisions","rr":"restricted-other","pr":false,"he":false,"ms":"After adjusting for reviewer agreement, manuscript year and article type, the deputy-editor ICC rose slightly to 0.03, still indicating little deputy-editor style effect on rejection decisions.","vf":"unverified"},{"key":"MM5CTRWM","au":"Kravitz, Richard L","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"Deputy Editor ICC","estd":"ICC (single/unspec)","v":0.02,"n":"2264","k":"1","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"reject or not reject","field":"biomedical, general internal medicine","wr":"deputy editors on reject vs not-reject decision","conf":"high","self":false,"doi":"10.1371/journal.pone.0010072","ciLow":0.01,"ciHigh":0.06,"mt":"other","tgt":"editor-decisions","rr":"restricted-other","pr":true,"he":false,"ms":"A random-effects model with deputy editor as random effect gave a deputy-editor ICC of 0.02 for the initial reject decision, an editor style effect rather than agreement on shared manuscripts. Fifty-seven deputy editors handled the 2264 manuscripts, one decision each.","vf":"unverified"},{"key":"MM5CTRWM","au":"Kravitz, Richard L","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"rho coefficient (ICC) for manuscript identity, mixed effects logistic regression","estd":"ICC (single/unspec)","v":0.17,"n":"2264","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"reject vs. accept/revise","field":"biomedical, general internal medicine","wr":"reviewers on manuscript reject vs accept/revise","conf":"high","self":false,"doi":"10.1371/journal.pone.0010072","ciLow":0.13,"ciHigh":0.22,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"A mixed-effects logistic model with manuscript as random effect, adjusting for study year and manuscript type, gave a manuscript-level ICC of 0.17, indicating modest agreement among reviewers rating the same manuscript. The abstract reports CI 0.12 to 0.22.","vf":"unverified"},{"key":"MM5CTRWM","au":"Kravitz, Richard L","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"reviewer ICC, mixed effects model, reviewer identity as random effect","estd":"ICC (single/unspec)","v":0.23,"n":"2264","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"reject vs. accept/revise","field":"biomedical, general internal medicine","wr":"reviewer style across manuscripts","conf":"high","self":false,"doi":"10.1371/journal.pone.0010072","ciLow":0.18,"ciHigh":0.29,"mt":"other","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"A mixed-effects logistic model with reviewer as random effect gave a reviewer-level ICC of 0.23, reflecting a reviewer style effect: variance in reject recommendations attributable to individual reviewers rather than inter-rater agreement. The abstract reports CI 0.19 to 0.29.","vf":"unverified"},{"key":"MM5CTRWM","au":"Kravitz, Richard L","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"kappa statistic","estd":"kappa","v":0.08,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"reject vs. accept/revise","field":"biomedical, general internal medicine","wr":"reviewers on manuscript reject vs accept/revise","conf":"high","self":false,"doi":"10.1371/journal.pone.0010072","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For manuscripts that received two independent reviews, chance-corrected agreement among reviewers on reject vs accept/revise was 0.08.","vf":"unverified"},{"key":"MM5CTRWM","au":"Kravitz, Richard L","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"kappa statistic","estd":"kappa","v":0.12,"n":"","k":"3","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"reject vs. accept/revise","field":"biomedical, general internal medicine","wr":"reviewers on manuscript reject vs accept/revise","conf":"high","self":false,"doi":"10.1371/journal.pone.0010072","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For manuscripts that received three independent reviews, chance-corrected agreement among reviewers on reject vs accept/revise was 0.12.","vf":"unverified"},{"key":"MM5CTRWM","au":"Kravitz, Richard L","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"kappa statistic","estd":"kappa","v":0.14,"n":"","k":"4","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"reject vs. accept/revise","field":"biomedical, general internal medicine","wr":"reviewers on manuscript reject vs accept/revise","conf":"high","self":false,"doi":"10.1371/journal.pone.0010072","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For manuscripts that received four independent reviews, chance-corrected agreement among reviewers on reject vs accept/revise was 0.14.","vf":"unverified"},{"key":"MM5CTRWM","au":"Kravitz, Richard L","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"kappa statistic","estd":"kappa","v":0.11,"n":"2264","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"reject vs. accept/revise","field":"biomedical, general internal medicine","wr":"reviewers on manuscript reject vs accept/revise","conf":"high","self":false,"doi":"10.1371/journal.pone.0010072","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":true,"ms":"External reviewers independently recommended rejecting or accepting/revising 2264 manuscripts sent for external review; a kappa of 0.11 shows chance-corrected agreement barely above chance. Manuscripts received 1 to 4 reviews each.","vf":"unverified"},{"key":"MM5CTRWM","au":"Kravitz, Richard L","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"Cronbach-alpha","form":"alpha","estd":"Cronbach alpha","v":0.8,"n":"","k":"18","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"reject vs. accept/revise","field":"biomedical, general internal medicine","wr":"projected reliability of mean of 18 reviewer recommendations","conf":"high","self":false,"doi":"10.1371/journal.pone.0010072","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"other","rr":"restricted-other","pr":false,"he":false,"ms":"Scaling the observed manuscript-level ICC of 0.17, the authors projected that eighteen reviewers per manuscript would yield an alpha reliability of 0.8. A projection from the paper's own empirical data, not a new measurement.","vf":"unverified"},{"key":"MM5CTRWM","au":"Kravitz, Richard L","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"Cronbach-alpha","form":"Cronbach’s alpha reliability","estd":"Cronbach alpha","v":0.6,"n":"","k":"7","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"reject vs. accept/revise","field":"biomedical, general internal medicine","wr":"projected reliability of mean of 7 reviewer recommendations","conf":"high","self":false,"doi":"10.1371/journal.pone.0010072","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"other","rr":"restricted-other","pr":false,"he":false,"ms":"Scaling the observed manuscript-level ICC of 0.17, the authors projected that seven reviewers per manuscript would yield a Cronbach's alpha reliability of 0.6. A projection from the paper's own empirical data, not a new measurement.","vf":"unverified"},{"key":"DRPY6FB5","au":"Linden, Wolfgang","y":1992,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Pearson Product-Moment correlations","estd":"correlation","v":0.31,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"Excellent (4), Good (3), Adequate (2), or Poor (1), with intermediates","field":"psychology","wr":"adequacy of manuscript analysis","conf":"med","self":false,"doi":"10.1037/h0078757","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two external reviewers independently rated each manuscript's adequacy of analysis on a 1-4 scale used as seven points; the Pearson correlation between reviewers was 0.31.","vf":"unverified"},{"key":"DRPY6FB5","au":"Linden, Wolfgang","y":1992,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Pearson Product-Moment correlations","estd":"correlation","v":0.21,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"Excellent (4), Good (3), Adequate (2), or Poor (1), with intermediates","field":"psychology","wr":"appropriateness for CJBS","conf":"med","self":false,"doi":"10.1037/h0078757","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two external reviewers independently rated each manuscript's appropriateness for the journal on a 1-4 scale used as seven points; the Pearson correlation of 0.21 was non-significant.","vf":"unverified"},{"key":"DRPY6FB5","au":"Linden, Wolfgang","y":1992,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Pearson Product-Moment correlations","estd":"correlation","v":0.27,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"Excellent (4), Good (3), Adequate (2), or Poor (1), with intermediates","field":"psychology","wr":"manuscript design quality","conf":"med","self":false,"doi":"10.1037/h0078757","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two external reviewers independently rated each manuscript's quality of design on a 1-4 scale used as seven points; the Pearson correlation between reviewers was 0.27.","vf":"unverified"},{"key":"DRPY6FB5","au":"Linden, Wolfgang","y":1992,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Pearson Product-Moment correlations","estd":"correlation","v":0.16,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"Excellent (4), Good (3), Adequate (2), or Poor (1), with intermediates","field":"psychology","wr":"importance of manuscript research","conf":"med","self":false,"doi":"10.1037/h0078757","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two external reviewers independently rated each manuscript's importance on a 1-4 scale used as seven points; the Pearson correlation of 0.16 was the lowest inter-rater agreement among the criteria and non-significant.","vf":"unverified"},{"key":"DRPY6FB5","au":"Linden, Wolfgang","y":1992,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Pearson Product-Moment correlations","estd":"correlation","v":0.34,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"Excellent (4), Good (3), Adequate (2), or Poor (1), with intermediates","field":"psychology","wr":"interpretation of manuscript results","conf":"med","self":false,"doi":"10.1037/h0078757","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two external reviewers independently rated each manuscript's interpretation of results on a 1-4 scale used as seven points; the Pearson correlation between reviewers was 0.34.","vf":"unverified"},{"key":"DRPY6FB5","au":"Linden, Wolfgang","y":1992,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Pearson Product-Moment correlations","estd":"correlation","v":0.44,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"Excellent (4), Good (3), Adequate (2), or Poor (1), with intermediates","field":"psychology","wr":"adequacy of literature review","conf":"med","self":false,"doi":"10.1037/h0078757","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two external reviewers independently rated each manuscript's adequacy of literature review on a 1-4 scale used as seven points; the Pearson correlation between reviewers was 0.44.","vf":"unverified"},{"key":"DRPY6FB5","au":"Linden, Wolfgang","y":1992,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Pearson Product-Moment correlations","estd":"correlation","v":0.23,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"Excellent (4), Good (3), Adequate (2), or Poor (1), with intermediates","field":"psychology","wr":"manuscript presentation quality","conf":"med","self":false,"doi":"10.1037/h0078757","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two external reviewers independently rated each manuscript's clarity of presentation on a 1-4 scale used as seven points; the Pearson correlation between reviewers was 0.23, indicating low agreement.","vf":"unverified"},{"key":"DRPY6FB5","au":"Linden, Wolfgang","y":1992,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Pearson Product-Moment correlations","estd":"correlation","v":0.42,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1. Accept outright ... 5. Reject","field":"psychology","wr":"reviewers' overall publication recommendation","conf":"med","self":false,"doi":"10.1037/h0078757","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Two external reviewers each gave an overall publication recommendation on a five-category scale; the Pearson correlation of 0.42 was the highest inter-rater agreement in the study.","vf":"unverified"},{"key":"DRPY6FB5","au":"Linden, Wolfgang","y":1992,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Pearson Product-Moment correlations","estd":"correlation","v":0.61,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"1. Accept outright ... 5. Reject","field":"psychology","wr":"reviewer 1 recommendation versus editorial decision","conf":"med","self":false,"doi":"10.1037/h0078757","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"Reviewer 1 recommended a disposition and the Editorial Committee made the final decision using the same five categories; their Pearson correlation was 0.61, with the decision informed by the reviews.","vf":"unverified"},{"key":"DRPY6FB5","au":"Linden, Wolfgang","y":1992,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Pearson Product-Moment correlations","estd":"correlation","v":0.65,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"1. Accept outright ... 5. Reject","field":"psychology","wr":"reviewer 2 recommendation versus editorial decision","conf":"med","self":false,"doi":"10.1037/h0078757","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":true,"he":false,"ms":"Reviewer 2 recommended a disposition and the Editorial Committee made the final decision using the same five categories; their Pearson correlation was 0.65, with the decision informed by the reviews.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.23,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"ChatGPT and human reviewers on applicability scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"ChatGPT and the two randomly assigned human reviewers of each abstract (three raters per abstract) scored 160 abstracts on the applicability criterion; the ICC(1,k) of 0.23 indicates fair human-LLM agreement.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.54,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"ChatGPT and human reviewers on clarity scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"ChatGPT and the two randomly assigned human reviewers of each abstract (three raters per abstract) scored 160 abstracts on the clarity criterion; the ICC(1,k) of 0.54 indicates moderate human-LLM agreement.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.5,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"mean of seven 1-5 criterion scores","field":"biomedical (pediatric)","wr":"ChatGPT and human reviewers on composite scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"ChatGPT and the two randomly assigned human reviewers of each abstract (three raters per abstract) scored 160 abstracts; the ICC(1,k) of 0.50 for the composite score indicates moderate human-LLM agreement.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.33,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"ChatGPT and human reviewers on engagement scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"ChatGPT and the two randomly assigned human reviewers of each abstract (three raters per abstract) scored 160 abstracts on the engagement criterion; the ICC(1,k) of 0.33 indicates fair human-LLM agreement.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.27,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"ChatGPT and human reviewers on impact scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"ChatGPT and the two randomly assigned human reviewers of each abstract (three raters per abstract) scored 160 abstracts on the impact criterion; the ICC(1,k) of 0.27 indicates fair human-LLM agreement.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.49,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"ChatGPT and human reviewers on impression scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"ChatGPT and the two randomly assigned human reviewers of each abstract (three raters per abstract) scored 160 abstracts on the impression criterion; the ICC(1,k) of 0.49 indicates moderate human-LLM agreement.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.45,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"ChatGPT and human reviewers on objective scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"ChatGPT and the two randomly assigned human reviewers of each abstract (three raters per abstract) scored 160 abstracts on the objective criterion; the ICC(1,k) of 0.45 indicates moderate human-LLM agreement.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.62,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"ChatGPT and human reviewers on results scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"ChatGPT and the two randomly assigned human reviewers of each abstract (three raters per abstract) scored 160 abstracts on the results criterion; the ICC(1,k) of 0.62 indicates good human-LLM agreement.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.31,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"Claude and human reviewers on applicability scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Claude and the two randomly assigned human reviewers of each abstract (three raters per abstract) scored 160 abstracts on the applicability criterion; the ICC(1,k) of 0.31 indicates fair human-LLM agreement.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.54,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"Claude and human reviewers on clarity scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Claude and the two randomly assigned human reviewers of each abstract (three raters per abstract) scored 160 abstracts on the clarity criterion; the ICC(1,k) of 0.54 indicates moderate human-LLM agreement.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.55,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"mean of seven 1-5 criterion scores","field":"biomedical (pediatric)","wr":"Claude and human reviewers on composite scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Claude and the two randomly assigned human reviewers of each abstract (three raters per abstract) scored 160 abstracts; the ICC(1,k) of 0.55 for the composite score indicates moderate human-LLM agreement.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.38,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"Claude and human reviewers on engagement scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Claude and the two randomly assigned human reviewers of each abstract (three raters per abstract) scored 160 abstracts on the engagement criterion; the ICC(1,k) of 0.38 indicates fair human-LLM agreement.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.31,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"Claude and human reviewers on impact scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Claude and the two randomly assigned human reviewers of each abstract (three raters per abstract) scored 160 abstracts on the impact criterion; the ICC(1,k) of 0.31 indicates fair human-LLM agreement.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.51,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"Claude and human reviewers on impression scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Claude and the two randomly assigned human reviewers of each abstract (three raters per abstract) scored 160 abstracts on the impression criterion; the ICC(1,k) of 0.51 indicates moderate human-LLM agreement.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.51,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"Claude and human reviewers on objective scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Claude and the two randomly assigned human reviewers of each abstract (three raters per abstract) scored 160 abstracts on the objective criterion; the ICC(1,k) of 0.51 indicates moderate human-LLM agreement.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.56,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"Claude and human reviewers on results scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Claude and the two randomly assigned human reviewers of each abstract (three raters per abstract) scored 160 abstracts on the results criterion; the ICC(1,k) of 0.56 indicates moderate human-LLM agreement.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.06,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"Gemini and human reviewers on applicability scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Gemini and the two randomly assigned human reviewers of each abstract (three raters per abstract) scored 160 abstracts on the applicability criterion; the ICC(1,k) of 0.06 indicates poor human-LLM agreement.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.47,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"Gemini and human reviewers on clarity scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Gemini and the two randomly assigned human reviewers of each abstract (three raters per abstract) scored 160 abstracts on the clarity criterion; the ICC(1,k) of 0.47 indicates moderate human-LLM agreement.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.38,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"mean of seven 1-5 criterion scores","field":"biomedical (pediatric)","wr":"Gemini and human reviewers on composite scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Gemini and the two randomly assigned human reviewers of each abstract (three raters per abstract) scored 160 abstracts; the ICC(1,k) of 0.38 for the composite score indicates fair human-LLM agreement.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.29,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"Gemini and human reviewers on engagement scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Gemini and the two randomly assigned human reviewers of each abstract (three raters per abstract) scored 160 abstracts on the engagement criterion; the ICC(1,k) of 0.29 indicates fair human-LLM agreement.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.04,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"Gemini and human reviewers on impact scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Gemini and the two randomly assigned human reviewers of each abstract (three raters per abstract) scored 160 abstracts on the impact criterion; the ICC(1,k) of 0.04 indicates poor human-LLM agreement.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.34,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"Gemini and human reviewers on impression scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Gemini and the two randomly assigned human reviewers of each abstract (three raters per abstract) scored 160 abstracts on the impression criterion; the ICC(1,k) of 0.34 indicates fair human-LLM agreement.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.31,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"Gemini and human reviewers on objective scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Gemini and the two randomly assigned human reviewers of each abstract (three raters per abstract) scored 160 abstracts on the objective criterion; the ICC(1,k) of 0.31 indicates fair human-LLM agreement.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.57,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"Gemini and human reviewers on results scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Gemini and the two randomly assigned human reviewers of each abstract (three raters per abstract) scored 160 abstracts on the results criterion; the ICC(1,k) of 0.57 indicates moderate human-LLM agreement.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.553,"n":"160","k":"2","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"two human reviewers on abstract scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Fourteen human reviewers, two randomly assigned per abstract, independently scored the 160 abstracts; across the eight criteria the ICC(1,k) values ranged from 0.151 to 0.553, and this row records the reported maximum. Per-criterion values sit in Supplementary Table 3, which is not available.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"one-way random effects model, specified as ICC (1, k)","estd":"ICC (average)","v":0.151,"n":"160","k":"2","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"two human reviewers on abstract scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Fourteen human reviewers, two randomly assigned per abstract, independently scored the 160 abstracts; across the eight criteria the ICC(1,k) values ranged from 0.151 to 0.553, and this row records the reported minimum. Per-criterion values sit in Supplementary Table 3, which is not available.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"two-way random effects model with absolute agreement definition and rater average unit, specified as ICC (2, k)","estd":"ICC (average)","v":0.69,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"three LLMs on abstract applicability scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three large language models each independently scored all 160 conference abstracts on the applicability criterion; the ICC(2,k) of 0.69 indicates good agreement among the models.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"two-way random effects model with absolute agreement definition and rater average unit, specified as ICC (2, k)","estd":"ICC (average)","v":0.65,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"three LLMs on abstract clarity scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three large language models each independently scored all 160 conference abstracts on the clarity criterion; the ICC(2,k) of 0.65 indicates good agreement among the models.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"two-way random effects model with absolute agreement definition and rater average unit, specified as ICC (2, k)","estd":"ICC (average)","v":0.8,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"mean of seven 1-5 criterion scores","field":"biomedical (pediatric)","wr":"three LLMs on abstract composite scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Three large language models each independently scored all 160 conference abstracts; the ICC(2,k) of 0.80 for the composite score indicates excellent agreement among the models. The 480 ratings are derived from the stated complete crossing (3 models by 160 abstracts).","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"two-way random effects model with absolute agreement definition and rater average unit, specified as ICC (2, k)","estd":"ICC (average)","v":0.74,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"three LLMs on abstract engagement scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three large language models each independently scored all 160 conference abstracts on the engagement criterion; the ICC(2,k) of 0.74 indicates good agreement among the models.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"two-way random effects model with absolute agreement definition and rater average unit, specified as ICC (2, k)","estd":"ICC (average)","v":0.76,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"three LLMs on abstract impact scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three large language models each independently scored all 160 conference abstracts on the impact criterion; the ICC(2,k) of 0.76 indicates good agreement among the models.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"two-way random effects model with absolute agreement definition and rater average unit, specified as ICC (2, k)","estd":"ICC (average)","v":0.79,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"three LLMs on abstract impression scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three large language models each independently scored all 160 conference abstracts on the impression criterion; the ICC(2,k) of 0.79 indicates good agreement among the models.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"two-way random effects model with absolute agreement definition and rater average unit, specified as ICC (2, k)","estd":"ICC (average)","v":0.59,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"three LLMs on abstract objective scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three large language models each independently scored all 160 conference abstracts on the objective criterion; the ICC(2,k) of 0.59 indicates moderate agreement among the models.","vf":"unverified"},{"key":"9QWVUD3K","au":"Liu, Yinuo","y":2026,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"two-way random effects model with absolute agreement definition and rater average unit, specified as ICC (2, k)","estd":"ICC (average)","v":0.87,"n":"160","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical (pediatric)","wr":"three LLMs on abstract results scores","conf":"high","self":false,"doi":"10.3389/frma.2026.1807672","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three large language models each independently scored all 160 conference abstracts on the results criterion; the ICC(2,k) of 0.87 indicates excellent agreement among the models.","vf":"unverified"},{"key":"LB5K9AI7","au":"Lobb, Rebecca","y":2013,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC, Shrout and Fleiss (1979) method","estd":"ICC (single/unspec)","v":-0.04,"n":"25","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"criterion scores 0-15 to 0-40 (+2 optional 0-10), summed","field":"public health","wr":"cross-discipline reviewers on same proposals (n=25)","conf":"med","self":false,"doi":"10.1097/PHH.0b013e31823991c2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"In the 25-proposal subset scored by the full panel, agreement across adjudicators from different disciplines gave an intraclass correlation of -0.04, effectively no cross-discipline agreement.","vf":"unverified"},{"key":"LB5K9AI7","au":"Lobb, Rebecca","y":2013,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC, Shrout and Fleiss (1979) method","estd":"ICC (single/unspec)","v":0.12,"n":"7","k":"3","samp":"funded-only","blind":"single","agg":"unspecified","scale":"criterion scores 0-15 to 0-40 (+2 optional 0-10), summed","field":"public health","wr":"cross-discipline panel on funded proposal scores","conf":"med","self":false,"doi":"10.1097/PHH.0b013e31823991c2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"For the seven highly ranked proposals selected for funding, cross-discipline agreement in adjudicators' rankings gave an intraclass correlation of 0.12.","vf":"unverified"},{"key":"LB5K9AI7","au":"Lobb, Rebecca","y":2013,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC, Shrout and Fleiss (1979) method","estd":"ICC (single/unspec)","v":0.01,"n":"34","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"criterion scores 0-15 to 0-40 (+2 optional 0-10), summed","field":"public health","wr":"cross-discipline panel on not-funded proposal scores","conf":"med","self":false,"doi":"10.1097/PHH.0b013e31823991c2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"For the 34 proposals not selected for funding, cross-discipline agreement in adjudicators' rankings gave an intraclass correlation of 0.01, effectively no agreement.","vf":"unverified"},{"key":"LB5K9AI7","au":"Lobb, Rebecca","y":2013,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC, Shrout and Fleiss (1979) method","estd":"ICC (single/unspec)","v":-0.08,"n":"41","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"criterion scores 0-15 to 0-40 (+2 optional 0-10), summed","field":"public health","wr":"policy vs practice reviewers on proposal scores","conf":"med","self":false,"doi":"10.1097/PHH.0b013e31823991c2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Agreement in proposal rankings between the policy and practice expert reviewers of 41 proposals gave an intraclass correlation of -0.08; the negative value reflects within-proposal rater variance exceeding between-proposal variance.","vf":"unverified"},{"key":"LB5K9AI7","au":"Lobb, Rebecca","y":2013,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC, Shrout and Fleiss (1979) method","estd":"ICC (single/unspec)","v":0.21,"n":"41","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"criterion scores 0-15 to 0-40 (+2 optional 0-10), summed","field":"public health","wr":"research vs policy reviewers on proposal scores","conf":"med","self":false,"doi":"10.1097/PHH.0b013e31823991c2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Agreement in proposal rankings between the research and policy expert reviewers of 41 proposals gave an intraclass correlation of 0.21, higher than the other cross-discipline pairings but still low.","vf":"unverified"},{"key":"LB5K9AI7","au":"Lobb, Rebecca","y":2013,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC, Shrout and Fleiss (1979) method","estd":"ICC (single/unspec)","v":0.11,"n":"41","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"criterion scores 0-15 to 0-40 (+2 optional 0-10), summed","field":"public health","wr":"research vs practice reviewers on proposal scores","conf":"med","self":false,"doi":"10.1097/PHH.0b013e31823991c2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Agreement in proposal rankings between the research and practice expert reviewers of 41 proposals gave an intraclass correlation of 0.11, indicating low cross-discipline agreement.","vf":"unverified"},{"key":"LB5K9AI7","au":"Lobb, Rebecca","y":2013,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC, Shrout and Fleiss (1979) method","estd":"ICC (single/unspec)","v":0.71,"n":"25","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"criterion scores 0-15 to 0-40 (+2 optional 0-10), summed","field":"public health","wr":"policy experts rating same proposals (within-discipline)","conf":"med","self":false,"doi":"10.1097/PHH.0b013e31823991c2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Among policy expert reviewers rating the 25-proposal subset scored by the full panel, within-discipline agreement gave an intraclass correlation of 0.71, substantially higher than cross-discipline agreement.","vf":"unverified"},{"key":"LB5K9AI7","au":"Lobb, Rebecca","y":2013,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC, Shrout and Fleiss (1979) method","estd":"ICC (single/unspec)","v":0.56,"n":"25","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"criterion scores 0-15 to 0-40 (+2 optional 0-10), summed","field":"public health","wr":"practice experts rating same proposals (within-discipline)","conf":"med","self":false,"doi":"10.1097/PHH.0b013e31823991c2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Among practice expert reviewers rating the 25-proposal subset scored by the full panel, within-discipline agreement gave an intraclass correlation of 0.56, substantially higher than cross-discipline agreement.","vf":"unverified"},{"key":"LB5K9AI7","au":"Lobb, Rebecca","y":2013,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC, Shrout and Fleiss (1979) method","estd":"ICC (single/unspec)","v":0.66,"n":"25","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"criterion scores 0-15 to 0-40 (+2 optional 0-10), summed","field":"public health","wr":"researchers rating same proposals (within-discipline)","conf":"med","self":false,"doi":"10.1097/PHH.0b013e31823991c2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Among research expert reviewers rating the 25-proposal subset scored by the full panel, within-discipline agreement gave an intraclass correlation of 0.66, substantially higher than cross-discipline agreement.","vf":"unverified"},{"key":"LB5K9AI7","au":"Lobb, Rebecca","y":2013,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC, Shrout and Fleiss (1979) method","estd":"ICC (single/unspec)","v":0.12,"n":"41","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"criterion scores 0-15 to 0-40 (+2 optional 0-10), summed","field":"public health","wr":"transdisciplinary panel on grant proposal scores","conf":"med","self":false,"doi":"10.1097/PHH.0b013e31823991c2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":false,"ms":"A transdisciplinary adjudication panel (at least one research, practice, and policy expert per proposal) scored 41 grant proposals; an intraclass correlation of 0.12 indicates low agreement in rankings across the three reviewer disciplines.","vf":"unverified"},{"key":"DK7Z8PE6","au":"Loonen, Martijn P. J.","y":2005,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"kappa statistic","estd":"kappa","v":0.8,"n":"40","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / reject","field":"biomedical (plastic and reconstructive surgery)","wr":"reviewer vs editor accept/reject on manuscripts","conf":"high","self":false,"doi":"10.1097/01.prs.0000178796.82273.7c","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":true,"he":false,"ms":"For the 40 manuscripts where the reviewer received the editor's actual decision, the reviewer's accept/reject advice and the editor's decision showed a Cohen kappa of 0.80, indicating high chance-corrected agreement.","vf":"unverified"},{"key":"DK7Z8PE6","au":"Loonen, Martijn P. J.","y":2005,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"actual proportion of exact agreement on accept/reject","estd":"percent agreement","v":0.9,"n":"40","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / reject","field":"biomedical (plastic and reconstructive surgery)","wr":"reviewer vs editor accept/reject on manuscripts","conf":"high","self":false,"doi":"10.1097/01.prs.0000178796.82273.7c","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For the 40 manuscripts with actual editor feedback, the reviewer's advice and the editor's decision matched on 36, an exact agreement proportion of 0.90.","vf":"unverified"},{"key":"DK7Z8PE6","au":"Loonen, Martijn P. J.","y":2005,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"kappa statistic","estd":"kappa","v":0.56,"n":"92","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / reject","field":"biomedical (plastic and reconstructive surgery)","wr":"reviewer vs editor apparent accept/reject on manuscripts","conf":"high","self":false,"doi":"10.1097/01.prs.0000178796.82273.7c","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"Assuming unpublished manuscripts were rejected, the reviewer's advice and the editor's apparent decision across 92 manuscripts gave a Cohen kappa of 0.56, indicating moderate chance-corrected agreement.","vf":"unverified"},{"key":"DK7Z8PE6","au":"Loonen, Martijn P. J.","y":2005,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"actual proportion of exact agreement on accept/reject","estd":"percent agreement","v":0.77,"n":"92","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / reject","field":"biomedical (plastic and reconstructive surgery)","wr":"reviewer vs editor apparent accept/reject on manuscripts","conf":"high","self":false,"doi":"10.1097/01.prs.0000178796.82273.7c","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"Across 92 manuscripts, with non-publication treated as rejection, the reviewer's advice and the editor's apparent decision matched on 71, an exact agreement proportion of 0.77.","vf":"unverified"},{"key":"DK7Z8PE6","au":"Loonen, Martijn P. J.","y":2005,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"kappa statistic","estd":"kappa","v":0.42,"n":"29","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / reject","field":"biomedical (plastic and reconstructive surgery)","wr":"second reviewer vs editor accept/reject on manuscripts","conf":"high","self":false,"doi":"10.1097/01.prs.0000178796.82273.7c","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"For the 29 manuscripts with a second reviewer, the second reviewer's advice and the editor's decision showed a Cohen kappa of 0.42, indicating fair chance-corrected agreement.","vf":"unverified"},{"key":"DK7Z8PE6","au":"Loonen, Martijn P. J.","y":2005,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"actual proportion of exact agreement on accept/reject","estd":"percent agreement","v":0.72,"n":"29","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / reject","field":"biomedical (plastic and reconstructive surgery)","wr":"second reviewer vs editor accept/reject on manuscripts","conf":"high","self":false,"doi":"10.1097/01.prs.0000178796.82273.7c","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"Across the 29 manuscripts, the second reviewer's advice and the editor's decision matched on 21, an exact agreement proportion of 0.72.","vf":"unverified"},{"key":"DK7Z8PE6","au":"Loonen, Martijn P. J.","y":2005,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"kappa statistic","estd":"kappa","v":0.29,"n":"29","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / reject","field":"biomedical (plastic and reconstructive surgery)","wr":"two reviewers' accept/reject recommendations on manuscripts","conf":"high","self":false,"doi":"10.1097/01.prs.0000178796.82273.7c","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"For the 29 manuscripts where a second reviewer's opinion was available, the two reviewers' accept/reject recommendations showed a Cohen kappa of 0.29, indicating poor chance-corrected agreement.","vf":"unverified"},{"key":"DK7Z8PE6","au":"Loonen, Martijn P. J.","y":2005,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"actual proportion of exact agreement on accept/reject","estd":"percent agreement","v":0.72,"n":"29","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / reject","field":"biomedical (plastic and reconstructive surgery)","wr":"two reviewers' accept/reject recommendations on manuscripts","conf":"high","self":false,"doi":"10.1097/01.prs.0000178796.82273.7c","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Across the 29 manuscripts with a second reviewer, the two reviewers' accept/reject recommendations matched on 21, an exact agreement proportion of 0.72.","vf":"unverified"},{"key":"GIY7LUIX","au":"Love, Jeffrey N.","y":2024,"cx":"Journal","ob":"review-report","fam":"percent-agreement","form":"exact agreement on accept, reject, or revise","estd":"percent agreement","v":0.235,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"reject, revise, or accept","field":"medical education (emergency medicine)","wr":"reviewer recommendations and final manuscript decisions","conf":"high","self":false,"doi":"10.5811/westjem.18432","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"Among 17 comparisons whose review had a summary HESR rating of 1.00, reviewer recommendations exactly matched final editorial decisions in 23.5%.","vf":"unverified"},{"key":"GIY7LUIX","au":"Love, Jeffrey N.","y":2024,"cx":"Journal","ob":"review-report","fam":"percent-agreement","form":"exact agreement on accept, reject, or revise","estd":"percent agreement","v":0.286,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"reject, revise, or accept","field":"medical education (emergency medicine)","wr":"reviewer recommendations and final manuscript decisions","conf":"high","self":false,"doi":"10.5811/westjem.18432","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"Among 21 comparisons whose review had a summary HESR rating of 2.00, reviewer recommendations exactly matched final editorial decisions in 28.6%.","vf":"unverified"},{"key":"GIY7LUIX","au":"Love, Jeffrey N.","y":2024,"cx":"Journal","ob":"review-report","fam":"percent-agreement","form":"exact agreement on accept, reject, or revise","estd":"percent agreement","v":0.607,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"reject, revise, or accept","field":"medical education (emergency medicine)","wr":"reviewer recommendations and final manuscript decisions","conf":"high","self":false,"doi":"10.5811/westjem.18432","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"Among 28 comparisons whose review had a summary HESR rating of 3.00, reviewer recommendations exactly matched final editorial decisions in 60.7%.","vf":"unverified"},{"key":"GIY7LUIX","au":"Love, Jeffrey N.","y":2024,"cx":"Journal","ob":"review-report","fam":"percent-agreement","form":"exact agreement on accept, reject, or revise","estd":"percent agreement","v":0.588,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"reject, revise, or accept","field":"medical education (emergency medicine)","wr":"reviewer recommendations and final manuscript decisions","conf":"high","self":false,"doi":"10.5811/westjem.18432","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"Among 17 comparisons whose review had a summary HESR rating of 4.00, reviewer recommendations exactly matched final editorial decisions in 58.8%.","vf":"unverified"},{"key":"GIY7LUIX","au":"Love, Jeffrey N.","y":2024,"cx":"Journal","ob":"review-report","fam":"percent-agreement","form":"exact agreement on accept, reject, or revise","estd":"percent agreement","v":1,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"reject, revise, or accept","field":"medical education (emergency medicine)","wr":"reviewer recommendations and final manuscript decisions","conf":"high","self":false,"doi":"10.5811/westjem.18432","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"Among seven comparisons whose review had a summary HESR rating of 5.00, reviewer recommendations exactly matched final editorial decisions in all cases.","vf":"unverified"},{"key":"GIY7LUIX","au":"Love, Jeffrey N.","y":2024,"cx":"Journal","ob":"review-report","fam":"percent-agreement","form":"exact agreement on accept, reject, or revise","estd":"percent agreement","v":0.489,"n":"47","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"reject, revise, or accept","field":"medical education (emergency medicine)","wr":"reviewer recommendations and final manuscript decisions","conf":"high","self":false,"doi":"10.5811/westjem.18432","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":true,"he":false,"ms":"Reviewer recommendations were compared with final editorial decisions for 47 manuscripts across 90 recommendation-decision comparisons. Exact decision agreement occurred in 48.9% of comparisons.","vf":"unverified"},{"key":"GIY7LUIX","au":"Love, Jeffrey N.","y":2024,"cx":"Journal","ob":"review-report","fam":"ICC","form":"one-way random effects model reflecting absolute agreement and the unit of analysis related to average measures","estd":"ICC (average)","v":0.795,"n":"90","k":"2","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"5-Exceptional to 1-Unacceptable","field":"medical education (emergency medicine)","wr":"editors rating peer-review quality with HESR","conf":"high","self":false,"doi":"10.5811/westjem.18432","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":true,"he":false,"ms":"The senior editor and one of three associate editors independently rated the quality of 90 peer reviews on a five-point rubric; an average-measures ICC of 0.795 indicates high relative consistency between editors.","vf":"unverified"},{"key":"GIY7LUIX","au":"Love, Jeffrey N.","y":2024,"cx":"Journal","ob":"review-report","fam":"percent-agreement","form":"percent absolute (exact) agreement between editors","estd":"percent agreement","v":0.378,"n":"90","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"5-Exceptional to 1-Unacceptable","field":"medical education (emergency medicine)","wr":"editors rating peer-review quality with HESR","conf":"high","self":false,"doi":"10.5811/westjem.18432","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"The two editors assigned identical five-point rubric scores to 37.8% of the 90 peer reviews, an exact (absolute) agreement rate that is low to moderate.","vf":"unverified"},{"key":"GIY7LUIX","au":"Love, Jeffrey N.","y":2024,"cx":"Journal","ob":"review-report","fam":"percent-agreement","form":"percent of ratings differing by exactly one point (within-one-point disagreement indicator)","estd":"percent agreement","v":0.478,"n":"90","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"5-Exceptional to 1-Unacceptable","field":"medical education (emergency medicine)","wr":"editors rating peer-review quality with HESR","conf":"high","self":false,"doi":"10.5811/westjem.18432","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"A further 47.8% of the 90 peer reviews received editor scores that differed by exactly one point, the paper's within-one-point consistency indicator.","vf":"unverified"},{"key":"GIY7LUIX","au":"Love, Jeffrey N.","y":2024,"cx":"Journal","ob":"review-report","fam":"correlation","form":"Spearman rho correlation","estd":"correlation","v":0.703,"n":"90","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"5-Exceptional to 1-Unacceptable","field":"medical education (emergency medicine)","wr":"editors rating peer-review quality with HESR","conf":"high","self":false,"doi":"10.5811/westjem.18432","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"The senior and associate editors' quality ratings of the same 90 peer reviews correlated at a Spearman rho of 0.703, a moderate-to-strong rank association between the two editor types.","vf":"unverified"},{"key":"2H7BLLRW","au":"Lyons‐Warren, Ariel M.","y":2024,"cx":"Journal","ob":"review-report","fam":"ICC","form":"intraclass correlation coefficient (ICC), model unspecified","estd":"ICC (single/unspec)","v":0.861,"n":"","k":"3","samp":"special","blind":"unclear","agg":"unspecified","scale":"modified RQI total, 10-54","field":"biomedical (neurology)","wr":"three scorers on review-report quality (post)","conf":"high","self":false,"doi":"10.1186/s41073-024-00143-x","ciLow":0.765,"ciHigh":0.925,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Three independent scorers applied the modified Review Quality Index to mentees' post-programme manuscript reviews; an ICC of 0.861 indicates strong agreement between the scorers on total review-quality scores.","vf":"unverified"},{"key":"2H7BLLRW","au":"Lyons‐Warren, Ariel M.","y":2024,"cx":"Journal","ob":"review-report","fam":"ICC","form":"intraclass correlation coefficient (ICC), model unspecified","estd":"ICC (single/unspec)","v":0.786,"n":"","k":"3","samp":"special","blind":"unclear","agg":"unspecified","scale":"modified RQI total, 10-54","field":"biomedical (neurology)","wr":"three scorers on review-report quality (pre)","conf":"high","self":false,"doi":"10.1186/s41073-024-00143-x","ciLow":0.649,"ciHigh":0.88,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":true,"he":false,"ms":"Three independent scorers applied the modified Review Quality Index to mentees' pre-programme manuscript reviews; an ICC of 0.786 indicates strong agreement between the scorers on total review-quality scores.","vf":"unverified"},{"key":"2H7BLLRW","au":"Lyons‐Warren, Ariel M.","y":2024,"cx":"Journal","ob":"review-report","fam":"correlation","form":"R2 (coefficient of determination) between total mRQI score and question 14 mean score","estd":"correlation","v":0.902,"n":"","k":"3","samp":"special","blind":"unclear","agg":"unspecified","scale":"mRQI total (10-54) vs question 14 (1-5)","field":"biomedical (neurology)","wr":"total mRQI vs overall-quality item (Q14)","conf":"high","self":false,"doi":"10.1186/s41073-024-00143-x","ciLow":null,"ciHigh":null,"mt":"other","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Across all scored reviews, the total modified RQI score correlated strongly (R2 = 0.902) with the single overall-quality item, which the authors label within-rater reliability and interpret as the summed score capturing overall quality judgment.","vf":"unverified"},{"key":"WH32N544","au":"Mahoney, MJ","y":1977,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (a)","estd":"ICC (single/unspec)","v":0.3,"n":"4","k":"","samp":"special","blind":"double","agg":"unspecified","scale":"4-point: poor, marginal, adequate, good; scored 0, 2, 4, 6","field":"psychology (behaviour modification)","wr":"referees rating a fabricated manuscript's data presentation","conf":"med","self":false,"doi":"10.1007/BF01173636","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Only the four groups whose manuscript contained a results section rated data presentation, which the paper's obtained-value list labels groups 1, 2, 4 and 5. The intraclass correlation of 0.30 indicates modest agreement between referees; the quoted sentence begins 'Ratings of data presenta-' at the foot of page 170.","vf":"unverified"},{"key":"WH32N544","au":"Mahoney, MJ","y":1977,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"unspecified","estd":"correlation","v":0.6,"n":"","k":"1","samp":"special","blind":"double","agg":"unspecified","scale":"4-point criterion scales scored 0, 2, 4, 6","field":"psychology (behaviour modification)","wr":"one referee's data and methodology ratings","conf":"med","self":false,"doi":"10.1007/BF01173636","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Each referee's rating of the data section correlated 0.60 with that referee's own methodology rating, even though the methodology text was identical across versions. This is consistency between two criteria of the review form rather than agreement between referees.","vf":"unverified"},{"key":"WH32N544","au":"Mahoney, MJ","y":1977,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"unspecified","estd":"correlation","v":0.56,"n":"","k":"1","samp":"special","blind":"double","agg":"unspecified","scale":"4-point criterion scales scored 0, 2, 4, 6","field":"psychology (behaviour modification)","wr":"one referee's data and recommendation ratings","conf":"med","self":false,"doi":"10.1007/BF01173636","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Each referee's rating of the data section correlated 0.56 with that referee's own publication recommendation. This is consistency between two criteria of the review form rather than agreement between referees, and the paper does not state which groups contributed.","vf":"unverified"},{"key":"WH32N544","au":"Mahoney, MJ","y":1977,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (a)","estd":"ICC (single/unspec)","v":0.01,"n":"2","k":"","samp":"special","blind":"double","agg":"unspecified","scale":"4-point: poor, marginal, adequate, good; scored 0, 2, 4, 6","field":"psychology (behaviour modification)","wr":"referees rating a fabricated manuscript's discussion section","conf":"med","self":false,"doi":"10.1007/BF01173636","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Only the two mixed-results groups received a discussion section, one arguing the results supported the reviewers' perspective and one arguing the opposite, and rated it on the four-point scale. The intraclass correlation is near zero, and note that the paper's obtained-value list prints this same figure as -.01 rather than .01.","vf":"unverified"},{"key":"WH32N544","au":"Mahoney, MJ","y":1977,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (a)","estd":"ICC (single/unspec)","v":0.03,"n":"5","k":"","samp":"special","blind":"double","agg":"unspecified","scale":"4-point: poor, marginal, adequate, good; scored 0, 2, 4, 6","field":"psychology (behaviour modification)","wr":"referees rating a fabricated manuscript's methodology","conf":"med","self":false,"doi":"10.1007/BF01173636","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The same referees rated the methodology section, which was identical across all five manuscript versions, on the four-point scale. The intraclass correlation of 0.03 shows almost no agreement between referees about methodological quality.","vf":"unverified"},{"key":"WH32N544","au":"Mahoney, MJ","y":1977,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"unspecified","estd":"correlation","v":0.94,"n":"","k":"1","samp":"special","blind":"double","agg":"unspecified","scale":"4-point criterion scales scored 0, 2, 4, 6","field":"psychology (behaviour modification)","wr":"one referee's methodology and recommendation ratings","conf":"med","self":false,"doi":"10.1007/BF01173636","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Among referees who read a clear positive or negative results manuscript, each referee's own methodology rating and publication recommendation correlated at 0.94. This is consistency between two criteria of the review form rather than agreement between referees, and the author reads it as a halo effect.","vf":"unverified"},{"key":"WH32N544","au":"Mahoney, MJ","y":1977,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (a)","estd":"ICC (single/unspec)","v":0.3,"n":"5","k":"","samp":"special","blind":"double","agg":"unspecified","scale":"4-point: poor, marginal, adequate, good; scored 0, 2, 4, 6","field":"psychology (behaviour modification)","wr":"referees rating a fabricated manuscript's scientific contribution","conf":"med","self":false,"doi":"10.1007/BF01173636","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Referees rated the manuscript's overall scientific contribution on the four-point scale, and several failed to rate this factor so its sample was slightly reduced. The intraclass correlation of 0.30 indicates modest agreement between referees.","vf":"unverified"},{"key":"WH32N544","au":"Mahoney, MJ","y":1977,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (a)","estd":"ICC (single/unspec)","v":0.3,"n":"5","k":"","samp":"special","blind":"double","agg":"unspecified","scale":"accept, accept with minor revisions, accept with major revisions, reject","field":"psychology (behaviour modification)","wr":"referees' publication recommendations on a fabricated manuscript","conf":"med","self":false,"doi":"10.1007/BF01173636","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":true,"ms":"Referees gave a summary publication recommendation on a four-point accept-to-reject scale, which the paper calls the critical dependent variable because it decides publication. The intraclass correlation of 0.30 indicates only modest agreement between referees on whether to publish; it is coded primary because the paper reports no single overall agreement figure.","vf":"unverified"},{"key":"WH32N544","au":"Mahoney, MJ","y":1977,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (a)","estd":"ICC (single/unspec)","v":-0.07,"n":"5","k":"","samp":"special","blind":"double","agg":"unspecified","scale":"4-point: poor, marginal, adequate, good; scored 0, 2, 4, 6","field":"psychology (behaviour modification)","wr":"referees rating a fabricated manuscript's topic relevance","conf":"med","self":false,"doi":"10.1007/BF01173636","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Journal referees each refereed one of five versions of a fabricated manuscript and rated its topic relevance on a four-point scale. The intraclass correlation of -0.07 pooled across the five groups indicates essentially no agreement between referees about topic relevance.","vf":"unverified"},{"key":"T5FYTZ8E","au":"Malički, Mario","y":2023,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"absolute agreement: exact match of recommendation choice","estd":"percent agreement","v":0.31,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"final recommendation choice (categories not enumerated)","field":"multi-field","wr":"reviewers' recommendations, 2019-2021 baseline","conf":"med","self":false,"doi":"10.15291/pubmet.4268","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"At the same 23 journals before structured peer review was introduced, reviewers' final recommendations matched exactly 31% of the time across 2019 to 2021; the pilot's 41% was significantly higher (P=0.0275).","vf":"unverified"},{"key":"T5FYTZ8E","au":"Malički, Mario","y":2023,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"absolute agreement: exact match of recommendation choice","estd":"percent agreement","v":0.41,"n":"107","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"final recommendation choice (categories not enumerated)","field":"multi-field","wr":"reviewers' final recommendations on manuscripts","conf":"med","self":false,"doi":"10.15291/pubmet.4268","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":false,"ms":"Two reviewers each assessed 107 manuscripts across 23 Elsevier journals piloting structured peer review; their final recommendation choices matched exactly for 41% of manuscripts, the study's headline agreement result.","vf":"unverified"},{"key":"T5FYTZ8E","au":"Malički, Mario","y":2023,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"partial agreement: same or similar answer","estd":"percent agreement","v":0.72,"n":"107","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"yes, no, NA, etc., with open-ended fields","field":"multi-field","wr":"reviewers on manuscript flow and structure","conf":"med","self":false,"doi":"10.15291/pubmet.4268","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Two reviewers answered structured questions on each of 107 manuscripts; their answers on manuscript flow and structure were the same or similar 72% of the time, the highest agreement across the questions.","vf":"unverified"},{"key":"T5FYTZ8E","au":"Malički, Mario","y":2023,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"partial agreement: same or similar answer","estd":"percent agreement","v":0.53,"n":"107","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"yes, no, NA, etc., with open-ended fields","field":"multi-field","wr":"reviewers on interpretation of results supported by data","conf":"med","self":false,"doi":"10.15291/pubmet.4268","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Two reviewers answered structured questions on each of 107 manuscripts; their answers on whether the interpretation of results was supported by the data were the same or similar 53% of the time, the lowest agreement.","vf":"unverified"},{"key":"T5FYTZ8E","au":"Malički, Mario","y":2023,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"partial agreement: same or similar answer","estd":"percent agreement","v":0.53,"n":"107","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"yes, no, NA, etc., with open-ended fields","field":"multi-field","wr":"reviewers on statistical analyses adequacy and reporting","conf":"med","self":false,"doi":"10.15291/pubmet.4268","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Two reviewers answered structured questions on each of 107 manuscripts; their answers on whether statistical analyses were appropriate and sufficiently reported were the same or similar 53% of the time, tied for the lowest agreement.","vf":"unverified"},{"key":"93G8L2PU","au":"Malički, Mario","y":2024,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"absolute agreement, exact match of final recommendation choice","estd":"percent agreement","v":0.41,"n":"107","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"accept / revise / reject (options in Appendix Table 2)","field":"multi-field","wr":"reviewers on final publication recommendation","conf":"high","self":false,"doi":"10.7717/peerj.17514","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":false,"ms":"Two reviewers each gave a final publication recommendation (accept, revise, reject) for each of 107 manuscripts across 23 Elsevier journals under piloted structured peer review; the two reviewers made exactly the same recommendation for 41% of manuscripts.","vf":"unverified"},{"key":"93G8L2PU","au":"Malički, Mario","y":2024,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"absolute agreement (exact match of recommendation choice)","estd":"percent agreement","v":0.31,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"e.g., accept, revise, reject","field":"multi-field","wr":"pre-pilot reviewer final recommendations","conf":"high","self":false,"doi":"10.7717/peerj.17514","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Using Elsevier Peer Review Workbench data for the same 23 journals from 2019 to 2021, before structured peer review, reviewers' final recommendations matched exactly for 31% of manuscripts; the number of manuscripts and reviewers per manuscript was not stated.","vf":"unverified"},{"key":"93G8L2PU","au":"Malički, Mario","y":2024,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"(partial) agreement, combining full agreement (category A) and partial agreement (category C)","estd":"percent agreement","v":0.58,"n":"107","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"coded: yes / no / N/A / yes+comments / no+comments","field":"multi-field","wr":"reviewers on study objectives/rationale question","conf":"high","self":false,"doi":"10.7717/peerj.17514","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Two reviewers each answered whether each of 107 manuscripts clearly stated its objectives and rationale; their coded answers fully or partially agreed for 58% of manuscripts.","vf":"unverified"},{"key":"93G8L2PU","au":"Malički, Mario","y":2024,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"(partial) agreement, combining full agreement (category A) and partial agreement (category C)","estd":"percent agreement","v":0.6,"n":"107","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"coded: yes / no / N/A / yes+comments / no+comments","field":"multi-field","wr":"reviewers on replicability/reproducibility question","conf":"high","self":false,"doi":"10.7717/peerj.17514","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Two reviewers each judged whether each of 107 manuscripts reported methods in sufficient detail for replicability; their coded answers fully or partially agreed for 60% of manuscripts.","vf":"unverified"},{"key":"93G8L2PU","au":"Malički, Mario","y":2024,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"(partial) agreement, combining full agreement (category A) and partial agreement (category C)","estd":"percent agreement","v":0.52,"n":"107","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"coded: yes / no / N/A / yes+comments / no+comments","field":"multi-field","wr":"reviewers on statistical analyses question","conf":"high","self":false,"doi":"10.7717/peerj.17514","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Two reviewers each judged whether statistical analyses were appropriate and well described in each of 107 manuscripts; their coded answers fully or partially agreed for 52% of manuscripts, the lowest agreement of any question.","vf":"unverified"},{"key":"93G8L2PU","au":"Malički, Mario","y":2024,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"(partial) agreement, combining full agreement (category A) and partial agreement (category C)","estd":"percent agreement","v":0.64,"n":"107","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"coded: yes / no / N/A / yes+comments / no+comments","field":"multi-field","wr":"reviewers on tables/figures question","conf":"high","self":false,"doi":"10.7717/peerj.17514","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Two reviewers each judged whether each of 107 manuscripts would benefit from additional or improved tables or figures; their coded answers fully or partially agreed for 64% of manuscripts.","vf":"unverified"},{"key":"93G8L2PU","au":"Malički, Mario","y":2024,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"(partial) agreement, combining full agreement (category A) and partial agreement (category C)","estd":"percent agreement","v":0.53,"n":"107","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"coded: yes / no / N/A / yes+comments / no+comments","field":"multi-field","wr":"reviewers on interpretation-of-results question","conf":"high","self":false,"doi":"10.7717/peerj.17514","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Two reviewers each judged whether the interpretation of results was supported by the data in each of 107 manuscripts; their coded answers fully or partially agreed for 53% of manuscripts.","vf":"unverified"},{"key":"93G8L2PU","au":"Malički, Mario","y":2024,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"(partial) agreement, combining full agreement (category A) and partial agreement (category C)","estd":"percent agreement","v":0.64,"n":"107","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"coded: yes / no / N/A / yes+comments / no+comments","field":"multi-field","wr":"reviewers on study-strengths question","conf":"high","self":false,"doi":"10.7717/peerj.17514","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Two reviewers each judged whether the authors clearly emphasised their study's strengths in each of 107 manuscripts; their coded answers fully or partially agreed for 64% of manuscripts.","vf":"unverified"},{"key":"93G8L2PU","au":"Malički, Mario","y":2024,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"(partial) agreement, combining full agreement (category A) and partial agreement (category C)","estd":"percent agreement","v":0.67,"n":"107","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"coded: yes / no / N/A / yes+comments / no+comments","field":"multi-field","wr":"reviewers on study-limitations question","conf":"high","self":false,"doi":"10.7717/peerj.17514","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Two reviewers each judged whether the authors clearly stated their study's limitations in each of 107 manuscripts; their coded answers fully or partially agreed for 67% of manuscripts.","vf":"unverified"},{"key":"93G8L2PU","au":"Malički, Mario","y":2024,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"(partial) agreement, combining full agreement (category A) and partial agreement (category C)","estd":"percent agreement","v":0.72,"n":"107","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"coded: yes / no / N/A / yes+comments / no+comments","field":"multi-field","wr":"reviewers on manuscript structure/flow question","conf":"high","self":false,"doi":"10.7717/peerj.17514","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Two reviewers each judged whether each of 107 manuscripts needed structural, flow or writing improvement; their coded answers fully or partially agreed for 72% of manuscripts, the highest agreement of any question.","vf":"unverified"},{"key":"93G8L2PU","au":"Malički, Mario","y":2024,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"absolute (exact) agreement on the yes/no answer","estd":"percent agreement","v":0.58,"n":"107","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"yes / no (drop-down menu)","field":"multi-field","wr":"reviewers on need-for-language-editing question","conf":"high","self":false,"doi":"10.7717/peerj.17514","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Two reviewers each answered yes or no on a drop-down whether each of 107 manuscripts needed language editing; they gave the same answer for 58% of manuscripts.","vf":"unverified"},{"key":"F68GYWKS","au":"Marsh, Herbert W","y":2008,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability (correlation between two independent readers' ratings), reader trial system","estd":"ICC (single/unspec)","v":0.3,"n":"","k":"","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"unspecified","field":"multi-field","wr":"expert readers on proposal quality (reader system)","conf":"med","self":false,"doi":"10.1037/0003-066X.63.3.160","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"In the reader trial, small teams of expert readers (typically three or four) each reviewed all 16 to 25 proposals in their subdiscipline; the single-rater reliability for proposal quality rose to 0.30.","vf":"unverified"},{"key":"F68GYWKS","au":"Marsh, Herbert W","y":2008,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"reliability of researcher ratings based on an average of 4.3 readers, reader system","estd":"ICC (average)","v":0.88,"n":"","k":"4.3","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"unspecified","field":"multi-field","wr":"mean of 4.3 readers on researcher track record","conf":"med","self":false,"doi":"10.1037/0003-066X.63.3.160","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Standardised to an average of 4.3 readers per proposal, the reader system reached an acceptable researcher-track-record reliability of 0.88, versus 0.53 for the traditional approach.","vf":"unverified"},{"key":"F68GYWKS","au":"Marsh, Herbert W","y":2008,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability (correlation between two independent readers' ratings), reader trial system","estd":"ICC (single/unspec)","v":0.63,"n":"","k":"","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"unspecified","field":"multi-field","wr":"expert readers on researcher track record (reader system)","conf":"med","self":false,"doi":"10.1037/0003-066X.63.3.160","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"In the reader trial, the single-rater reliability of researcher-track-record ratings rose to 0.63, far above the traditional ARC approach.","vf":"unverified"},{"key":"F68GYWKS","au":"Marsh, Herbert W","y":2008,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"reliability of the mean rating of 4.3 assessors via Spearman-Brown equation","estd":"ICC (average)","v":0.44,"n":"2331","k":"4.3","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"0-100 scale","field":"multi-field","wr":"mean of 4.3 assessors on proposal quality","conf":"med","self":false,"doi":"10.1037/0003-066X.63.3.160","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Applying the Spearman-Brown equation to the observed single-rater reliability, the mean rating of 4.3 external assessors per proposal has an estimated proposal-quality reliability of 0.44, still below acceptable levels.","vf":"unverified"},{"key":"F68GYWKS","au":"Marsh, Herbert W","y":2008,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"projected reliability of the mean rating of 6 assessors via Spearman-Brown equation","estd":"ICC (average)","v":0.71,"n":"2331","k":"6","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"0-100 scale","field":"multi-field","wr":"projected mean of 6 assessors on proposal quality","conf":"med","self":false,"doi":"10.1037/0003-066X.63.3.160","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"A Spearman-Brown projection from the same ARC database estimates that six assessors per proposal would be needed to reach a proposal-quality reliability of 0.71.","vf":"unverified"},{"key":"F68GYWKS","au":"Marsh, Herbert W","y":2008,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability (correlation between two independent assessors' ratings), derived from multilevel cross-classified models","estd":"ICC (single/unspec)","v":0.15,"n":"2331","k":"4.3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"0-100 scale","field":"multi-field","wr":"external assessors on grant proposal quality","conf":"med","self":false,"doi":"10.1037/0003-066X.63.3.160","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":true,"ms":"External assessors (on average 4.3 per proposal) independently rated the quality of 2,331 ARC grant proposals; a single-rater reliability of 0.15 indicates very poor agreement between two assessors of the same proposal.","vf":"unverified"},{"key":"F68GYWKS","au":"Marsh, Herbert W","y":2008,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability (correlation between two independent assessors' ratings), traditional ARC approach","estd":"ICC (single/unspec)","v":0.17,"n":"","k":"","samp":"special","blind":"single","agg":"single-rater","scale":"0-100 scale","field":"multi-field","wr":"assessors on proposal quality, traditional approach","conf":"med","self":false,"doi":"10.1037/0003-066X.63.3.160","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"In the reader-trial comparison, the traditional ARC approach yielded a single-rater proposal-quality reliability of 0.17.","vf":"unverified"},{"key":"F68GYWKS","au":"Marsh, Herbert W","y":2008,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability (correlation between two independent assessors' ratings), derived from multilevel cross-classified models","estd":"ICC (single/unspec)","v":0.17,"n":"","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"0-100 scale","field":"multi-field","wr":"assessors on proposal quality, science panels","conf":"med","self":false,"doi":"10.1037/0003-066X.63.3.160","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For science-panel proposals in the ARC database, the single-rater reliability of proposal-quality ratings was 0.17, no higher than in the social sciences and humanities.","vf":"unverified"},{"key":"F68GYWKS","au":"Marsh, Herbert W","y":2008,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability (correlation between two independent assessors' ratings), derived from multilevel cross-classified models","estd":"ICC (single/unspec)","v":0.18,"n":"","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"0-100 scale","field":"multi-field","wr":"assessors on proposal quality, social science/humanities","conf":"med","self":false,"doi":"10.1037/0003-066X.63.3.160","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For social science and humanities proposals in the ARC database, the single-rater reliability of proposal-quality ratings was 0.18, marginally above the science panels.","vf":"unverified"},{"key":"F68GYWKS","au":"Marsh, Herbert W","y":2008,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"reliability of the mean rating of 4.3 assessors via Spearman-Brown equation","estd":"ICC (average)","v":0.53,"n":"2331","k":"4.3","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"0-100 scale","field":"multi-field","wr":"mean of 4.3 assessors on researcher track record","conf":"med","self":false,"doi":"10.1037/0003-066X.63.3.160","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Applying the Spearman-Brown equation, the mean rating of 4.3 external assessors per proposal has an estimated researcher-track-record reliability of 0.53.","vf":"unverified"},{"key":"F68GYWKS","au":"Marsh, Herbert W","y":2008,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"projected reliability of the mean rating of 6 assessors via Spearman-Brown equation","estd":"ICC (average)","v":0.82,"n":"2331","k":"6","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"0-100 scale","field":"multi-field","wr":"projected mean of 6 assessors on researcher track record","conf":"med","self":false,"doi":"10.1037/0003-066X.63.3.160","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"A Spearman-Brown projection from the same ARC database estimates that six assessors per proposal would be needed to reach a researcher-track-record reliability of 0.82.","vf":"unverified"},{"key":"F68GYWKS","au":"Marsh, Herbert W","y":2008,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability (correlation between two independent assessors' ratings), derived from multilevel cross-classified models","estd":"ICC (single/unspec)","v":0.21,"n":"2331","k":"4.3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"0-100 scale","field":"multi-field","wr":"external assessors on researcher track-record quality","conf":"med","self":false,"doi":"10.1037/0003-066X.63.3.160","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"External assessors independently rated the research track record of applicants across 2,331 ARC proposals; a single-rater reliability of 0.21 indicates poor agreement between two assessors of the same application.","vf":"unverified"},{"key":"F68GYWKS","au":"Marsh, Herbert W","y":2008,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability (correlation between two independent assessors' ratings), traditional ARC approach","estd":"ICC (single/unspec)","v":0.24,"n":"","k":"","samp":"special","blind":"single","agg":"single-rater","scale":"0-100 scale","field":"multi-field","wr":"assessors on researcher track record, traditional approach","conf":"med","self":false,"doi":"10.1037/0003-066X.63.3.160","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"In the reader-trial comparison, the traditional ARC approach yielded a single-rater researcher-track-record reliability of 0.24.","vf":"unverified"},{"key":"F68GYWKS","au":"Marsh, Herbert W","y":2008,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability (correlation between two independent assessors' ratings), derived from multilevel cross-classified models","estd":"ICC (single/unspec)","v":0.23,"n":"","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"0-100 scale","field":"multi-field","wr":"assessors on researcher track record, science panels","conf":"med","self":false,"doi":"10.1037/0003-066X.63.3.160","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For science-panel proposals in the ARC database, the single-rater reliability of researcher-track-record ratings was 0.23.","vf":"unverified"},{"key":"F68GYWKS","au":"Marsh, Herbert W","y":2008,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-rater reliability (correlation between two independent assessors' ratings), derived from multilevel cross-classified models","estd":"ICC (single/unspec)","v":0.26,"n":"","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"0-100 scale","field":"multi-field","wr":"assessors on researcher track record, social science/humanities","conf":"med","self":false,"doi":"10.1037/0003-066X.63.3.160","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For social science and humanities proposals in the ARC database, the single-rater reliability of researcher-track-record ratings was 0.26.","vf":"unverified"},{"key":"TIAHXM4B","au":"Marsh, Herbert W.","y":1981,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"single-rater reliabilities (intraclass correlations)","estd":"ICC (single/unspec)","v":0.21,"n":"325","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"high, medium, low; expanded to 5 points","field":"educational psychology","wr":"reviewers on journal appropriateness","conf":"med","self":false,"doi":"10.1037/0022-0663.73.6.872","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Pairs of external reviewers independently rated the appropriateness of 325 manuscripts for the journal; the convergent coefficient of .21 is the single-rater reliability for this subscale.","vf":"unverified"},{"key":"TIAHXM4B","au":"Marsh, Herbert W.","y":1981,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"reliability of the average overall recommendation from two reviewers","estd":"ICC (average)","v":0.51,"n":"325","k":"2","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"5 categories analysed 1-9; 1 = strongly recommend acceptance, 9 = reject","field":"educational psychology","wr":"two reviewers' averaged overall recommendation","conf":"med","self":false,"doi":"10.1037/0022-0663.73.6.872","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":true,"ms":"The reliability of the average overall recommendation from two reviewers was .51 across the same 325 manuscripts; the paper presents this as its headline estimate while arguing it is artificially low.","vf":"unverified"},{"key":"TIAHXM4B","au":"Marsh, Herbert W.","y":1981,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"reliability of the average rating by the two reviewers (rxx), corrected for reviewer response bias","estd":"ICC (average)","v":0.5,"n":"325","k":"2","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"5 categories analysed 1-9; 1 = strongly recommend acceptance, 9 = reject","field":"educational psychology","wr":"two reviewers' averaged overall recommendation","conf":"med","self":false,"doi":"10.1037/0022-0663.73.6.872","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"After the same response-bias correction, the reliability of the two-reviewer average overall recommendation was .50, essentially unchanged from the uncorrected .51.","vf":"unverified"},{"key":"TIAHXM4B","au":"Marsh, Herbert W.","y":1981,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"single-rater reliability (rii), ratings corrected for reviewer response bias","estd":"ICC (single/unspec)","v":0.34,"n":"325","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"5 categories analysed 1-9; 1 = strongly recommend acceptance, 9 = reject","field":"educational psychology","wr":"one reviewer on manuscript overall recommendation","conf":"med","self":false,"doi":"10.1037/0022-0663.73.6.872","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"After correcting the ratings of 31 frequent reviewers for systematic response bias, the single-reviewer intraclass correlation for the overall recommendation was still .34, so the correction did not improve agreement.","vf":"unverified"},{"key":"TIAHXM4B","au":"Marsh, Herbert W.","y":1981,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"single-reviewer reliability (intraclass correlation)","estd":"ICC (single/unspec)","v":0.34,"n":"325","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"5 categories analysed 1-9; 1 = strongly recommend acceptance, 9 = reject","field":"educational psychology","wr":"one reviewer on manuscript overall recommendation","conf":"med","self":false,"doi":"10.1037/0022-0663.73.6.872","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Two external reviewers independently rated each of 325 Journal of Educational Psychology manuscripts on an overall publication recommendation; the single-reviewer intraclass correlation was .34, indicating low agreement between individual reviewers.","vf":"unverified"},{"key":"TIAHXM4B","au":"Marsh, Herbert W.","y":1981,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"single-rater reliabilities (intraclass correlations)","estd":"ICC (single/unspec)","v":0.27,"n":"325","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"high, medium, low; expanded to 5 points","field":"educational psychology","wr":"reviewers on manuscript research quality","conf":"med","self":false,"doi":"10.1037/0022-0663.73.6.872","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Pairs of external reviewers independently rated the research quality of 325 manuscripts; the convergent coefficient of .27 is the single-rater reliability for this subscale.","vf":"unverified"},{"key":"TIAHXM4B","au":"Marsh, Herbert W.","y":1981,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"single-rater reliabilities (intraclass correlations)","estd":"ICC (single/unspec)","v":0.2,"n":"325","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"high, medium, low; expanded to 5 points","field":"educational psychology","wr":"reviewers on manuscript significance","conf":"med","self":false,"doi":"10.1037/0022-0663.73.6.872","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Pairs of external reviewers independently rated the significance of 325 manuscripts on a three-point scale expanded to five; the convergent coefficient of .20 is the single-rater reliability for this subscale.","vf":"unverified"},{"key":"TIAHXM4B","au":"Marsh, Herbert W.","y":1981,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"reliability of the average rating by the two reviewers (rxx)","estd":"ICC (average)","v":0.53,"n":"325","k":"2","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"first principal component of the five ratings","field":"educational psychology","wr":"two reviewers' averaged composite score","conf":"med","self":false,"doi":"10.1037/0022-0663.73.6.872","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The reliability of the two-reviewer average of the composite total score was .53 across the 325 manuscripts.","vf":"unverified"},{"key":"TIAHXM4B","au":"Marsh, Herbert W.","y":1981,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"reliability of the average rating (rxx) of the total score, corrected for reviewer response bias","estd":"ICC (average)","v":0.52,"n":"325","k":"2","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"first principal component of the five ratings","field":"educational psychology","wr":"two reviewers' averaged composite score","conf":"med","self":false,"doi":"10.1037/0022-0663.73.6.872","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"After correcting for reviewer response bias, the two-reviewer average reliability of the composite total score was .52, essentially unchanged from the uncorrected .53.","vf":"unverified"},{"key":"TIAHXM4B","au":"Marsh, Herbert W.","y":1981,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"reliability of the average of three reviewers (using the Spearman-Brown equation)","estd":"ICC (average)","v":0.65,"n":"325","k":"2","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"first principal component of the five ratings","field":"educational psychology","wr":"three reviewers' averaged composite score (projected)","conf":"med","self":false,"doi":"10.1037/0022-0663.73.6.872","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Extrapolating the composite total score's single-reviewer reliability with the Spearman-Brown formula, the reliability of an average of three reviewers would be .65; only two reviewers per manuscript were in the actual data.","vf":"unverified"},{"key":"TIAHXM4B","au":"Marsh, Herbert W.","y":1981,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"reliability of the average of four reviewers (using the Spearman-Brown equation)","estd":"ICC (average)","v":0.69,"n":"325","k":"2","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"first principal component of the five ratings","field":"educational psychology","wr":"four reviewers' averaged composite score (projected)","conf":"med","self":false,"doi":"10.1037/0022-0663.73.6.872","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Extrapolating with the Spearman-Brown formula, the reliability of an average of four reviewers' composite total score would be .69; only two reviewers per manuscript were in the actual data.","vf":"unverified"},{"key":"TIAHXM4B","au":"Marsh, Herbert W.","y":1981,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"single-rater reliability for the total score (first principal component)","estd":"ICC (single/unspec)","v":0.36,"n":"325","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"first principal component of the five ratings","field":"educational psychology","wr":"one reviewer on manuscript composite score","conf":"med","self":false,"doi":"10.1037/0022-0663.73.6.872","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Using a weighted composite (first principal component) of the five rating items, the single-reviewer intraclass correlation was .36, only marginally higher than the .34 for the overall recommendation alone.","vf":"unverified"},{"key":"TIAHXM4B","au":"Marsh, Herbert W.","y":1981,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"single-rater reliability (rii) of the total score, corrected for reviewer response bias","estd":"ICC (single/unspec)","v":0.35,"n":"325","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"first principal component of the five ratings","field":"educational psychology","wr":"one reviewer on manuscript composite score","conf":"med","self":false,"doi":"10.1037/0022-0663.73.6.872","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"After correcting for reviewer response bias, the single-reviewer reliability of the composite total score was .35, essentially unchanged from the uncorrected .36.","vf":"unverified"},{"key":"TIAHXM4B","au":"Marsh, Herbert W.","y":1981,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"single-rater reliabilities (intraclass correlations)","estd":"ICC (single/unspec)","v":0.24,"n":"325","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"high, medium, low; expanded to 5 points","field":"educational psychology","wr":"reviewers on manuscript writing quality","conf":"med","self":false,"doi":"10.1037/0022-0663.73.6.872","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Pairs of external reviewers independently rated the writing quality of 325 manuscripts; the convergent coefficient of .24 is the single-rater reliability for this subscale.","vf":"unverified"},{"key":"XHLJRX3Y","au":"Marsh, Herbert W.","y":1989,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"canonical correlation relating four manuscript review factors for the first reviewers to those for the second reviewers","estd":"correlation","v":0.31,"n":"278","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"optimally weighted composite of four factor scores","field":"educational psychology","wr":"external reviewers on weighted composites of four factors","conf":"med","self":false,"doi":"10.1080/00220973.1989.10806503","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Canonical correlation was used to find the optimally weighted combination of the four review factors that best matched between the two reviewers of the same manuscript. The first canonical correlation of .31 was about the same size as the single-reviewer reliability of the overall recommendation.","vf":"unverified"},{"key":"XHLJRX3Y","au":"Marsh, Herbert W.","y":1989,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"canonical correlation for the first set of canonical variates, overall recommendation added to four factor scores","estd":"correlation","v":0.32,"n":"278","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"weighted composite of four factor scores plus overall recommendation","field":"educational psychology","wr":"external reviewers on weighted composite including overall recommendation","conf":"med","self":false,"doi":"10.1080/00220973.1989.10806503","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"When the overall recommendation was added to the four factor scores, the best weighted combination matching between the two reviewers of the same manuscript correlated only .32, showing that optimal weighting did not improve on the overall recommendation alone.","vf":"unverified"},{"key":"XHLJRX3Y","au":"Marsh, Herbert W.","y":1989,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"canonical correlation relating four manuscript review factors for the first reviewers to those for the second reviewers","estd":"correlation","v":0.25,"n":"278","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"optimally weighted composite of four factor scores","field":"educational psychology","wr":"external reviewers on weighted composites of four factors","conf":"med","self":false,"doi":"10.1080/00220973.1989.10806503","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The second pair of canonical variates, which reflected mainly writing style and presentation clarity, correlated .25 between the two reviewers of the same manuscript.","vf":"unverified"},{"key":"XHLJRX3Y","au":"Marsh, Herbert W.","y":1989,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"canonical correlation relating four manuscript review factors for the first reviewers to those for the second reviewers","estd":"correlation","v":0.17,"n":"278","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"optimally weighted composite of four factor scores","field":"educational psychology","wr":"external reviewers on weighted composites of four factors","conf":"med","self":false,"doi":"10.1080/00220973.1989.10806503","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The third pair of canonical variates, which reflected mainly relevance to readers, correlated .17 between the two reviewers of the same manuscript.","vf":"unverified"},{"key":"XHLJRX3Y","au":"Marsh, Herbert W.","y":1989,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Pearson product-moment correlation between independent reviewers' responses to the same item (single-reviewer reliability)","estd":"correlation","v":0.12,"n":"278","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"3-point high, medium, low, expanded to 5-point response scale","field":"educational psychology","wr":"external reviewers on manuscript significance item","conf":"med","self":false,"doi":"10.1080/00220973.1989.10806503","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Table 2 reports single-reviewer reliabilities for all 21 rating items, ranging from .12 to .30; because that stratification exceeds the 15-row cap, only the overall result plus the lowest and highest items are coded. Item 1, significance of the paper, was the least reliable at .12.","vf":"unverified"},{"key":"XHLJRX3Y","au":"Marsh, Herbert W.","y":1989,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Pearson product-moment correlation between independent reviewers' responses to the same item (single-reviewer reliability)","estd":"correlation","v":0.3,"n":"278","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"9-point: very poor (1) to very good (9)","field":"educational psychology","wr":"external reviewers on population and sampling item","conf":"med","self":false,"doi":"10.1080/00220973.1989.10806503","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Item 12, appropriateness of population, sampling and data gathering techniques, was the most reliable of the 20 items other than the overall recommendation, at .30. It is coded as the maximum stratum of the 21 item-level reliabilities in Table 2.","vf":"unverified"},{"key":"XHLJRX3Y","au":"Marsh, Herbert W.","y":1989,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Pearson product-moment correlation (single-reviewer reliability), after correction for idiosyncratic response bias","estd":"correlation","v":0.29,"n":"278","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"9-point: very poor (1) to very good (9)","field":"educational psychology","wr":"external reviewers on experimental-form overall evaluation","conf":"med","self":false,"doi":"10.1080/00220973.1989.10806503","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"After correction for each reviewer's leniency or harshness, agreement between the two reviewers on the experimental form's overall evaluation rose from .21 to .29.","vf":"unverified"},{"key":"XHLJRX3Y","au":"Marsh, Herbert W.","y":1989,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Pearson product-moment correlation (single-reviewer reliability), before correction for response bias","estd":"correlation","v":0.21,"n":"278","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"9-point: very poor (1) to very good (9)","field":"educational psychology","wr":"external reviewers on experimental-form overall evaluation","conf":"med","self":false,"doi":"10.1080/00220973.1989.10806503","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Item 20 asked reviewers for an overall evaluation of the manuscript in its present form on the experimental nine point scale. The two reviewers of the same manuscript agreed at .21, well below the standard form's overall recommendation.","vf":"unverified"},{"key":"XHLJRX3Y","au":"Marsh, Herbert W.","y":1989,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Pearson product-moment correlation (single-reviewer reliability), after correction for idiosyncratic response bias","estd":"correlation","v":0.31,"n":"278","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"9-point: very poor (1) to very good (9)","field":"educational psychology","wr":"external reviewers on likely evaluation after revision","conf":"med","self":false,"doi":"10.1080/00220973.1989.10806503","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"After correction for each reviewer's leniency or harshness, agreement between the two reviewers on the likely evaluation after revision rose from .23 to .31.","vf":"unverified"},{"key":"XHLJRX3Y","au":"Marsh, Herbert W.","y":1989,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Pearson product-moment correlation (single-reviewer reliability), before correction for response bias","estd":"correlation","v":0.23,"n":"278","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"9-point: very poor (1) to very good (9)","field":"educational psychology","wr":"external reviewers on likely evaluation after revision","conf":"med","self":false,"doi":"10.1080/00220973.1989.10806503","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Item 21 asked reviewers how they would likely evaluate the manuscript once the authors had made feasible modifications. The two reviewers of the same manuscript agreed at .23 before any correction for response bias.","vf":"unverified"},{"key":"XHLJRX3Y","au":"Marsh, Herbert W.","y":1989,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Pearson product-moment correlation (single-reviewer reliability), after correction for idiosyncratic response bias","estd":"correlation","v":0.33,"n":"278","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"strongly recommend acceptance (1) to reject (5), expanded to nine categories","field":"educational psychology","wr":"external reviewers on manuscript overall publication recommendation","conf":"med","self":false,"doi":"10.1080/00220973.1989.10806503","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"After correcting each reviewer's scores for their estimated leniency or harshness, the single-reviewer reliability of the overall recommendation rose from .30 to .33, the smallest improvement among the six scores in Table 3.","vf":"unverified"},{"key":"XHLJRX3Y","au":"Marsh, Herbert W.","y":1989,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Pearson product-moment correlation between matching scores for first and second reviewers (single-reviewer reliability)","estd":"correlation","v":0.3,"n":"278","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"strongly recommend acceptance (1) to reject (5), expanded to nine categories","field":"educational psychology","wr":"external reviewers on manuscript overall publication recommendation","conf":"med","self":false,"doi":"10.1080/00220973.1989.10806503","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":true,"ms":"Two external reviewers independently rated each of 278 manuscripts submitted to the Journal of Educational Psychology on a five point overall publication recommendation that was expanded to nine categories. The correlation of .30 between the two reviewers is the reliability of a single reviewer's recommendation and is the study's headline result.","vf":"unverified"},{"key":"XHLJRX3Y","au":"Marsh, Herbert W.","y":1989,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Pearson product-moment correlation (single-reviewer reliability), after correction for idiosyncratic response bias","estd":"correlation","v":0.33,"n":"278","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"factor-score total of all 21 items (5- and 9-point)","field":"educational psychology","wr":"external reviewers on total of all 21 items","conf":"med","self":false,"doi":"10.1080/00220973.1989.10806503","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"After correction for each reviewer's leniency or harshness, the two reviewers' total scores across all 21 items correlated .33.","vf":"unverified"},{"key":"XHLJRX3Y","au":"Marsh, Herbert W.","y":1989,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Pearson product-moment correlation (single-reviewer reliability) for total score, before correction for response bias","estd":"correlation","v":0.27,"n":"278","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"factor-score total of all 21 items (5- and 9-point)","field":"educational psychology","wr":"external reviewers on total of all 21 items","conf":"med","self":false,"doi":"10.1080/00220973.1989.10806503","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"All 21 items from both rating forms were combined into one total score. The two reviewers of the same manuscript agreed on that total at .27, no better than on the overall recommendation alone.","vf":"unverified"},{"key":"XHLJRX3Y","au":"Marsh, Herbert W.","y":1989,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Pearson product-moment correlation (single-reviewer reliability), after correction for idiosyncratic response bias","estd":"correlation","v":0.31,"n":"278","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"factor-score total of four 5-point standard-form items","field":"educational psychology","wr":"external reviewers on total of four standard items","conf":"med","self":false,"doi":"10.1080/00220973.1989.10806503","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Each reviewer's responses were adjusted for their estimated leniency or harshness across the manuscripts they reviewed. After that adjustment the two reviewers' total scores on standard-form items 1 to 4 correlated .31.","vf":"unverified"},{"key":"XHLJRX3Y","au":"Marsh, Herbert W.","y":1989,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Pearson product-moment correlation (single-reviewer reliability) for total score, before correction for response bias","estd":"correlation","v":0.24,"n":"278","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"factor-score total of four 5-point standard-form items","field":"educational psychology","wr":"external reviewers on total of four standard items","conf":"med","self":false,"doi":"10.1080/00220973.1989.10806503","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The four specific standard-form items were combined into a total score. The two reviewers of the same manuscript agreed on that total at .24, lower than for the overall recommendation on its own.","vf":"unverified"},{"key":"XHLJRX3Y","au":"Marsh, Herbert W.","y":1989,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Pearson product-moment correlation (single-reviewer reliability), after correction for idiosyncratic response bias","estd":"correlation","v":0.33,"n":"278","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"factor-score total of sixteen 9-point experimental items","field":"educational psychology","wr":"external reviewers on total of 16 experimental items","conf":"med","self":false,"doi":"10.1080/00220973.1989.10806503","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"After each reviewer's scores were adjusted for their estimated leniency or harshness, the two reviewers' total scores across the 16 experimental items correlated .33.","vf":"unverified"},{"key":"XHLJRX3Y","au":"Marsh, Herbert W.","y":1989,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Pearson product-moment correlation (single-reviewer reliability) for total score, before correction for response bias","estd":"correlation","v":0.27,"n":"278","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"factor-score total of sixteen 9-point experimental items","field":"educational psychology","wr":"external reviewers on total of 16 experimental items","conf":"med","self":false,"doi":"10.1080/00220973.1989.10806503","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The 16 items on the experimental rating form were combined into a total score. The two reviewers of the same manuscript agreed on that total at .27 before any correction for reviewer response bias.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass R (R_i), Model III, based on average number of reviews per proposal","estd":"ICC (single/unspec)","v":0.18,"n":"50","k":"4.24","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"10-50 (poor to excellent)","field":"multi-field","wr":"reviewers on NSF/COSPUP grant proposal scores","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Mail reviewers independently rated NSF grant proposals in chemical dynamics; the intraclass R of 0.18 (average of about 4.24 reviews per proposal) indicates low chance-corrected agreement among reviewers.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass R (R_i), Model III, based on average number of reviews per proposal","estd":"ICC (single/unspec)","v":0.37,"n":"49","k":"4.04","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"10-50 (poor to excellent)","field":"multi-field","wr":"reviewers on NSF/COSPUP grant proposal scores","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Mail reviewers independently rated NSF grant proposals in economics; the intraclass R of 0.37 (average of about 4.04 reviews per proposal) indicates low chance-corrected agreement among reviewers.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass R (R_i), Model III, based on average number of reviews per proposal","estd":"ICC (single/unspec)","v":0.33,"n":"50","k":"4.06","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"10-50 (poor to excellent)","field":"multi-field","wr":"reviewers on NSF/COSPUP grant proposal scores","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Mail reviewers independently rated NSF grant proposals in solid state physics; the intraclass R of 0.33 (average of about 4.06 reviews per proposal) indicates low chance-corrected agreement among reviewers.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass R (R_i), Model III, based on average number of reviews per proposal","estd":"ICC (single/unspec)","v":0.32,"n":"50","k":"4.26","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"10-50 (poor to excellent)","field":"multi-field","wr":"reviewers on NSF/COSPUP grant proposal scores","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Mail reviewers independently rated NSF grant proposals in chemical dynamics; the intraclass R of 0.32 (average of about 4.26 reviews per proposal) indicates low chance-corrected agreement among reviewers.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass R (R_i), Model III, based on average number of reviews per proposal","estd":"ICC (single/unspec)","v":0.36,"n":"49","k":"3.69","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"10-50 (poor to excellent)","field":"multi-field","wr":"reviewers on NSF/COSPUP grant proposal scores","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Mail reviewers independently rated NSF grant proposals in economics; the intraclass R of 0.36 (average of about 3.69 reviews per proposal) indicates low chance-corrected agreement among reviewers.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass R (R_i), Model III, based on average number of reviews per proposal","estd":"ICC (single/unspec)","v":0.34,"n":"49","k":"3.86","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"10-50 (poor to excellent)","field":"multi-field","wr":"reviewers on NSF/COSPUP grant proposal scores","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Mail reviewers independently rated NSF grant proposals in solid state physics; the intraclass R of 0.34 (average of about 3.86 reviews per proposal) indicates low chance-corrected agreement among reviewers.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass R (R_i), Model III, based on average number of reviews per proposal","estd":"ICC (single/unspec)","v":0.25,"n":"50","k":"4.84","samp":"unclear","blind":"single","agg":"single-rater","scale":"10-50 (poor to excellent)","field":"multi-field","wr":"reviewers on NSF/COSPUP grant proposal scores","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Mail reviewers independently rated NSF grant proposals in chemical dynamics; the intraclass R of 0.25 (average of about 4.84 reviews per proposal) indicates low chance-corrected agreement among reviewers.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass R (R_i), Model III, based on average number of reviews per proposal","estd":"ICC (single/unspec)","v":0.37,"n":"42","k":"3.69","samp":"unclear","blind":"single","agg":"single-rater","scale":"10-50 (poor to excellent)","field":"multi-field","wr":"reviewers on NSF/COSPUP grant proposal scores","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Mail reviewers independently rated NSF grant proposals in economics; the intraclass R of 0.37 (average of about 3.69 reviews per proposal) indicates low chance-corrected agreement among reviewers.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass R (R_i), Model III, based on average number of reviews per proposal","estd":"ICC (single/unspec)","v":0.32,"n":"50","k":"3.84","samp":"unclear","blind":"single","agg":"single-rater","scale":"10-50 (poor to excellent)","field":"multi-field","wr":"reviewers on NSF/COSPUP grant proposal scores","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Mail reviewers independently rated NSF grant proposals in solid state physics; the intraclass R of 0.32 (average of about 3.84 reviews per proposal) indicates low chance-corrected agreement among reviewers.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"observed percent agreement (exact) on acceptance","estd":"percent agreement","v":0.66,"n":"62","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / reject","field":"multi-field","wr":"two reviewers agreeing on accept of manuscripts","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Among manuscripts submitted to \"American Psychologist\", two independent reviewers showed 66% observed agreement on accept decisions.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"overall observed percent agreement (exact) on accept vs reject","estd":"percent agreement","v":0.74,"n":"159","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / reject","field":"multi-field","wr":"two reviewers agreeing on combined of manuscripts","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Among manuscripts submitted to \"American Psychologist\", two independent reviewers showed 74% observed agreement on combined decisions.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"observed percent agreement (exact) on rejection","estd":"percent agreement","v":0.78,"n":"97","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / reject","field":"multi-field","wr":"two reviewers agreeing on reject of manuscripts","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Among manuscripts submitted to \"American Psychologist\", two independent reviewers showed 78% observed agreement on reject decisions.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"observed percent agreement (exact) on acceptance","estd":"percent agreement","v":0.52,"n":"25","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / reject","field":"multi-field","wr":"two reviewers agreeing on accept of manuscripts","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Among manuscripts submitted to \"Developmental Review\", two independent reviewers showed 52% observed agreement on accept decisions.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"overall observed percent agreement (exact) on accept vs reject","estd":"percent agreement","v":0.67,"n":"72","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / reject","field":"multi-field","wr":"two reviewers agreeing on combined of manuscripts","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Among manuscripts submitted to \"Developmental Review\", two independent reviewers showed 67% observed agreement on combined decisions.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"observed percent agreement (exact) on rejection","estd":"percent agreement","v":0.74,"n":"47","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / reject","field":"multi-field","wr":"two reviewers agreeing on reject of manuscripts","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Among manuscripts submitted to \"Developmental Review\", two independent reviewers showed 74% observed agreement on reject decisions.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass R (R_i), Model I; on dichotomised accept/reject (equals kappa in dichotomous case)","estd":"ICC (single/unspec)","v":0.45,"n":"159","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / reject (dichotomised)","field":"multi-field","wr":"two reviewers on manuscript accept/reject","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers judged manuscripts submitted to \"American Psychologist\" as accept or reject; the intraclass R of 0.45 shows low chance-corrected agreement between them.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass R (R_i), Model I; on dichotomised accept/reject (equals kappa in dichotomous case)","estd":"ICC (single/unspec)","v":0.27,"n":"72","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / reject (dichotomised)","field":"multi-field","wr":"two reviewers on manuscript accept/reject","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers judged manuscripts submitted to \"Developmental Review\" as accept or reject; the intraclass R of 0.27 shows low chance-corrected agreement between them.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass R (R_i), Model I; on dichotomised accept/reject (equals kappa in dichotomous case)","estd":"ICC (single/unspec)","v":0.14,"n":"1319","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / reject (dichotomised)","field":"multi-field","wr":"two reviewers on manuscript accept/reject","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers judged manuscripts submitted to \"Journal of Abnormal Psychology\" as accept or reject; the intraclass R of 0.14 shows low chance-corrected agreement between them.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass R (R_i), Model I; on dichotomised accept/reject (equals kappa in dichotomous case)","estd":"ICC (single/unspec)","v":0.26,"n":"866","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / reject (dichotomised)","field":"multi-field","wr":"two reviewers on manuscript accept/reject","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers judged manuscripts submitted to Untitled Medical Specialty Journal as accept or reject; the intraclass R of 0.26 shows low chance-corrected agreement between them.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"observed percent agreement (exact) on acceptance","estd":"percent agreement","v":0.44,"n":"462","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / reject","field":"multi-field","wr":"two reviewers agreeing on accept of manuscripts","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Among manuscripts submitted to \"Journal of Abnormal Psychology\", two independent reviewers showed 44% observed agreement on accept decisions.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"overall observed percent agreement (exact) on accept vs reject","estd":"percent agreement","v":0.61,"n":"1319","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / reject","field":"multi-field","wr":"two reviewers agreeing on combined of manuscripts","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Among manuscripts submitted to \"Journal of Abnormal Psychology\", two independent reviewers showed 61% observed agreement on combined decisions.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"observed percent agreement (exact) on rejection","estd":"percent agreement","v":0.7,"n":"857","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / reject","field":"multi-field","wr":"two reviewers agreeing on reject of manuscripts","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Among manuscripts submitted to \"Journal of Abnormal Psychology\", two independent reviewers showed 70% observed agreement on reject decisions.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"observed percent agreement (exact) on acceptance","estd":"percent agreement","v":0.5,"n":"289","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / reject","field":"multi-field","wr":"two reviewers agreeing on accept of manuscripts","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Among manuscripts submitted to Untitled Medical Specialty Journal, two independent reviewers showed 50% observed agreement on accept decisions.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"overall observed percent agreement (exact) on accept vs reject","estd":"percent agreement","v":0.67,"n":"866","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / reject","field":"multi-field","wr":"two reviewers agreeing on combined of manuscripts","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Among manuscripts submitted to Untitled Medical Specialty Journal, two independent reviewers showed 67% observed agreement on combined decisions.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"observed percent agreement (exact) on rejection","estd":"percent agreement","v":0.76,"n":"577","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept / reject","field":"multi-field","wr":"two reviewers agreeing on reject of manuscripts","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Among manuscripts submitted to Untitled Medical Specialty Journal, two independent reviewers showed 76% observed agreement on reject decisions.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"observed percent agreement (exact) on all proposals of grant proposals","estd":"percent agreement","v":0.6,"n":"50","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"high (40-50) / low (10-39)","field":"multi-field","wr":"reviewers agreeing on all proposals grant proposals","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For NSF/COSPUP grant proposals in chemical dynamics, reviewers showed 60% observed agreement on all proposals.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"observed percent agreement (exact) on high ratings of grant proposals","estd":"percent agreement","v":0.41,"n":"17","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"high (40-50) / low (10-39)","field":"multi-field","wr":"reviewers agreeing on high ratings grant proposals","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"For NSF/COSPUP grant proposals in chemical dynamics, reviewers showed 41% observed agreement on high ratings.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"observed percent agreement (exact) on low ratings of grant proposals","estd":"percent agreement","v":0.7,"n":"33","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"high (40-50) / low (10-39)","field":"multi-field","wr":"reviewers agreeing on low ratings grant proposals","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For NSF/COSPUP grant proposals in chemical dynamics, reviewers showed 70% observed agreement on low ratings.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"observed percent agreement (exact) on all proposals of grant proposals","estd":"percent agreement","v":0.68,"n":"150","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"high (40-50) / low (10-39)","field":"multi-field","wr":"reviewers agreeing on all proposals grant proposals","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For NSF/COSPUP grant proposals in Combined (3 areas), reviewers showed 68% observed agreement on all proposals.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"observed percent agreement (exact) on high ratings of grant proposals","estd":"percent agreement","v":0.54,"n":"52","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"high (40-50) / low (10-39)","field":"multi-field","wr":"reviewers agreeing on high ratings grant proposals","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"For NSF/COSPUP grant proposals in Combined (3 areas), reviewers showed 54% observed agreement on high ratings.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"observed percent agreement (exact) on low ratings of grant proposals","estd":"percent agreement","v":0.76,"n":"98","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"high (40-50) / low (10-39)","field":"multi-field","wr":"reviewers agreeing on low ratings grant proposals","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For NSF/COSPUP grant proposals in Combined (3 areas), reviewers showed 76% observed agreement on low ratings.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"observed percent agreement (exact) on all proposals of grant proposals","estd":"percent agreement","v":0.76,"n":"50","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"high (40-50) / low (10-39)","field":"multi-field","wr":"reviewers agreeing on all proposals grant proposals","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For NSF/COSPUP grant proposals in economics, reviewers showed 76% observed agreement on all proposals.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"observed percent agreement (exact) on high ratings of grant proposals","estd":"percent agreement","v":0.6,"n":"15","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"high (40-50) / low (10-39)","field":"multi-field","wr":"reviewers agreeing on high ratings grant proposals","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"For NSF/COSPUP grant proposals in economics, reviewers showed 60% observed agreement on high ratings.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"observed percent agreement (exact) on low ratings of grant proposals","estd":"percent agreement","v":0.83,"n":"35","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"high (40-50) / low (10-39)","field":"multi-field","wr":"reviewers agreeing on low ratings grant proposals","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For NSF/COSPUP grant proposals in economics, reviewers showed 83% observed agreement on low ratings.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass R (R_i) on dichotomised high/low grant ratings; not significant","estd":"ICC (single/unspec)","v":0.16,"n":"50","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"high (40-50) / low (10-39)","field":"multi-field","wr":"reviewers on grant proposals (high vs low rated)","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Reviewers rated NSF/COSPUP grant proposals in chemical dynamics as high or low merit; the intraclass R of 0.16 indicates low chance-corrected agreement and was not statistically significant.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass R (R_i) on dichotomised high/low grant ratings","estd":"ICC (single/unspec)","v":0.32,"n":"150","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"high (40-50) / low (10-39)","field":"multi-field","wr":"reviewers on grant proposals (high vs low rated)","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Reviewers rated NSF/COSPUP grant proposals in Combined (3 areas) as high or low merit; the intraclass R of 0.32 indicates low chance-corrected agreement.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass R (R_i) on dichotomised high/low grant ratings","estd":"ICC (single/unspec)","v":0.44,"n":"50","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"high (40-50) / low (10-39)","field":"multi-field","wr":"reviewers on grant proposals (high vs low rated)","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Reviewers rated NSF/COSPUP grant proposals in economics as high or low merit; the intraclass R of 0.44 indicates low chance-corrected agreement.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"ICC","form":"intraclass R (R_i) on dichotomised high/low grant ratings","estd":"ICC (single/unspec)","v":0.34,"n":"50","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"high (40-50) / low (10-39)","field":"multi-field","wr":"reviewers on grant proposals (high vs low rated)","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Reviewers rated NSF/COSPUP grant proposals in solid state physics as high or low merit; the intraclass R of 0.34 indicates low chance-corrected agreement.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"observed percent agreement (exact) on all proposals of grant proposals","estd":"percent agreement","v":0.68,"n":"50","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"high (40-50) / low (10-39)","field":"multi-field","wr":"reviewers agreeing on all proposals grant proposals","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For NSF/COSPUP grant proposals in solid state physics, reviewers showed 68% observed agreement on all proposals.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"observed percent agreement (exact) on high ratings of grant proposals","estd":"percent agreement","v":0.6,"n":"20","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"high (40-50) / low (10-39)","field":"multi-field","wr":"reviewers agreeing on high ratings grant proposals","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"For NSF/COSPUP grant proposals in solid state physics, reviewers showed 60% observed agreement on high ratings.","vf":"unverified"},{"key":"QTGZWS97","au":"Marsh, Herbert W.","y":1991,"cx":"General","ob":"other","fam":"percent-agreement","form":"observed percent agreement (exact) on low ratings of grant proposals","estd":"percent agreement","v":0.73,"n":"30","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"high (40-50) / low (10-39)","field":"multi-field","wr":"reviewers agreeing on low ratings grant proposals","conf":"med","self":false,"doi":"10.1017/s0140525x00065912","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For NSF/COSPUP grant proposals in solid state physics, reviewers showed 73% observed agreement on low ratings.","vf":"unverified"},{"key":"F6AMXHF7","au":"Marsh, Herbert W.","y":2011,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"unspecified","estd":"correlation","v":0.95,"n":"2331","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"0-100; weighted project (.6) + researcher (.4)","field":"multi-field","wr":"external assessors' aggregate vs ARC panel ratings of proposals","conf":"med","self":false,"doi":"10.1016/j.joi.2010.10.004","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":false,"ms":"For 2331 externally reviewed grant proposals, the average rating across each proposal's external assessors correlated .95 with the ARC panel's final rating of the same proposal. The panel rating was itself based substantially on the assessor ratings, so this reflects concordance between two non-independent evaluation stages rather than independent inter-rater agreement.","vf":"unverified"},{"key":"F6AMXHF7","au":"Marsh, Herbert W.","y":2011,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"variance component, multilevel cross-classified model (MCMC), standardised variables","estd":"other","v":0.815,"n":"2331","k":"4.3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"0-100 ratings transformed to normal scores, standardised","field":"multi-field","wr":"external assessors' ratings of grant proposals","conf":"med","self":false,"doi":"10.1016/j.joi.2010.10.004","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"A multilevel cross-classified variance components model of 10,023 external assessor ratings of 2331 ARC grant proposals, which the paper presents as its index of interrater agreement, placed .815 of the standardised rating variance between assessors of the same proposal, indicating substantial disagreement among reviewers.","vf":"unverified"},{"key":"F6AMXHF7","au":"Marsh, Herbert W.","y":2011,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"variance component, multilevel cross-classified model (MCMC), standardised variables","estd":"other","v":0.052,"n":"2331","k":"4.3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"0-100 ratings transformed to normal scores, standardised","field":"multi-field","wr":"between-field-of-study variance in assessor ratings","conf":"med","self":false,"doi":"10.1016/j.joi.2010.10.004","ciLow":null,"ciHigh":null,"mt":"other","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"In the same variance components model, .052 of the standardised rating variance was associated with the 144 fields of study, indicating small but significant systematic differences between fields in assessor ratings.","vf":"unverified"},{"key":"F6AMXHF7","au":"Marsh, Herbert W.","y":2011,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"variance component, multilevel cross-classified model (MCMC), standardised variables","estd":"other","v":0.144,"n":"2331","k":"4.3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"0-100 ratings transformed to normal scores, standardised","field":"multi-field","wr":"systematic between-proposal variance in assessor ratings","conf":"med","self":false,"doi":"10.1016/j.joi.2010.10.004","ciLow":null,"ciHigh":null,"mt":"other","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"In the same variance components model, .144 of the standardised rating variance lay between proposals, the systematic signal that assessors agree on, against .815 of variance between assessors of the same proposal.","vf":"unverified"},{"key":"XFRJ7EYE","au":"Marson, Stephen M.","y":2021,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Kendall's rank order coefficient (Kendall's Tau-b)","estd":"correlation","v":0.719,"n":"","k":"","samp":"special","blind":"double","agg":"unspecified","scale":"1 reject may not resubmit to 5 unconditionally accept","field":"social work","wr":"referee pairs plus imputed editor rejections of manuscripts","conf":"med","self":false,"doi":"10.1177/10497315211052456","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"A hypothetical recalculation in which about 50 manuscripts that the editor or co-editor had rejected as off-mission before referee review, all ranked 1, were added to the referee pairs. Kendall's tau-b rises from 0.68 to 0.72, and the authors state this does not accurately reflect reality because those manuscripts were never assessed by the editorial board.","vf":"unverified"},{"key":"XFRJ7EYE","au":"Marson, Stephen M.","y":2021,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Kendall's rank order coefficient (Kendall's Tau-b)","estd":"correlation","v":0.68,"n":"","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 reject may not resubmit to 5 unconditionally accept","field":"social work","wr":"paired anonymous referees on manuscript decisions","conf":"med","self":false,"doi":"10.1177/10497315211052456","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":true,"ms":"Kendall's tau-b between the first and the second anonymous referee's five point decision for the same manuscripts. The authors prefer this coefficient because the decision form is ordinal and tau-b adjusts for ties, and they treat 0.68 as the study's headline result.","vf":"unverified"},{"key":"XFRJ7EYE","au":"Marson, Stephen M.","y":2021,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Pearson's (rp)","estd":"correlation","v":0.749,"n":"","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 reject may not resubmit to 5 unconditionally accept","field":"social work","wr":"paired anonymous referees on manuscript decisions","conf":"med","self":false,"doi":"10.1177/10497315211052456","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Pearson correlation between the first and the second anonymous referee's five point decision for the same manuscripts. The authors report it only for comparison, judging it less appropriate than Kendall's tau for ordinal decisions.","vf":"unverified"},{"key":"XFRJ7EYE","au":"Marson, Stephen M.","y":2021,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement: paired referees ranked the manuscript in an identical manner","estd":"percent agreement","v":0.66,"n":"","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 reject may not resubmit to 5 unconditionally accept","field":"social work","wr":"paired anonymous referees on manuscript decisions","conf":"med","self":false,"doi":"10.1177/10497315211052456","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Two anonymous referees each rated the same submitted manuscript on a five point reject-to-accept decision scale. In 66 per cent of the referee pairs both referees picked exactly the same category.","vf":"unverified"},{"key":"SAJ2V3PH","au":"Martin, Michael R.","y":2010,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement: no change (0.00) in priority score","estd":"percent agreement","v":0.04,"n":"1395","k":"","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1.0 (highest merit) to 5.0 (lowest merit), 0.1 increments","field":"biomedical","wr":"preliminary and final R01 priority scores","conf":"high","self":false,"doi":"10.1371/journal.pone.0013526","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"The average preliminary score and the post-discussion final panel score were identical for 4% of the 1,395 applications.","vf":"unverified"},{"key":"SAJ2V3PH","au":"Martin, Michael R.","y":2010,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson correlation coefficient","estd":"correlation","v":0.78,"n":"1395","k":"","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1.0 (highest merit) to 5.0 (lowest merit), 0.1 increments","field":"biomedical","wr":"panels on R01 grant proposal priority scores","conf":"high","self":false,"doi":"10.1371/journal.pone.0013526","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":false,"ms":"Across 1,395 NIH R01 applications, the average of the three assigned reviewers' independent preliminary priority scores correlated 0.78 (Pearson) with the full panel's post-discussion final priority score.","vf":"unverified"},{"key":"SAJ2V3PH","au":"Martin, Michael R.","y":2010,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"within 0.05 priority-score points","estd":"percent agreement","v":0.226,"n":"1395","k":"","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1.0 (highest merit) to 5.0 (lowest merit), 0.1 increments","field":"biomedical","wr":"preliminary and final R01 priority scores","conf":"high","self":false,"doi":"10.1371/journal.pone.0013526","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"The average preliminary and post-discussion final panel scores differed by no more than 0.05 points for 22.6% of the 1,395 applications.","vf":"unverified"},{"key":"SAJ2V3PH","au":"Martin, Michael R.","y":2010,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"within preliminary minimum-to-maximum assigned-reviewer score range","estd":"percent agreement","v":0.802,"n":"1395","k":"","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1.0 (highest merit) to 5.0 (lowest merit), 0.1 increments","field":"biomedical","wr":"final scores against assigned-reviewer preliminary ranges","conf":"high","self":false,"doi":"10.1371/journal.pone.0013526","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"For 80.2% of 1,395 applications, the post-discussion final panel score remained within the minimum-to-maximum range of the assigned reviewers' preliminary scores.","vf":"unverified"},{"key":"SAJ2V3PH","au":"Martin, Michael R.","y":2010,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"same predefined priority-score band","estd":"percent agreement","v":0.85,"n":"","k":"","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"priority-score bands 1.00-1.45, 1.46-1.60, 1.61-1.75, 1.76-3.00","field":"biomedical","wr":"preliminary and final priority-score bands","conf":"high","self":false,"doi":"10.1371/journal.pone.0013526","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Among applications with preliminary scores from 1.00 to 1.45, 85% remained in that score band after panel discussion.","vf":"unverified"},{"key":"SAJ2V3PH","au":"Martin, Michael R.","y":2010,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"same predefined priority-score band","estd":"percent agreement","v":0.37,"n":"","k":"","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"priority-score bands 1.00-1.45, 1.46-1.60, 1.61-1.75, 1.76-3.00","field":"biomedical","wr":"preliminary and final priority-score bands","conf":"high","self":false,"doi":"10.1371/journal.pone.0013526","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Among applications with preliminary scores from 1.46 to 1.60, 37% remained in that score band after panel discussion.","vf":"unverified"},{"key":"SAJ2V3PH","au":"Martin, Michael R.","y":2010,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"same predefined priority-score band","estd":"percent agreement","v":0.34,"n":"","k":"","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"priority-score bands 1.00-1.45, 1.46-1.60, 1.61-1.75, 1.76-3.00","field":"biomedical","wr":"preliminary and final priority-score bands","conf":"high","self":false,"doi":"10.1371/journal.pone.0013526","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Among applications with preliminary scores from 1.61 to 1.75, 34% remained in that score band after panel discussion.","vf":"unverified"},{"key":"SAJ2V3PH","au":"Martin, Michael R.","y":2010,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"same predefined priority-score band","estd":"percent agreement","v":0.87,"n":"","k":"","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"priority-score bands 1.00-1.45, 1.46-1.60, 1.61-1.75, 1.76-3.00","field":"biomedical","wr":"preliminary and final priority-score bands","conf":"high","self":false,"doi":"10.1371/journal.pone.0013526","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Among applications in the lowest preliminary score band (defined as 1.76 and above; the table row is printed 1.75-3.00), 87% remained in that band after panel discussion.","vf":"unverified"},{"key":"SZ554632","au":"Martinková, Patrícia","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC(1,1) one-way random, single-rater (IRR = sigma_g^2 / (sigma_g^2 + sigma_e^2))","estd":"ICC (single/unspec)","v":0.4,"n":"25","k":"3","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"unspecified","field":"biomedical","wr":"AIBS proposals from female principal investigators","conf":"high","self":false,"doi":"10.3102/10769986221150517","ciLow":0.22,"ciHigh":0.6,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"Bayesian model averaging over the candidate variance-components models estimated the single-rater reliability (ICC(1,1)) of the three ratings per proposal at 0.40 for the 25 AIBS proposals from female principal investigators.","vf":"unverified"},{"key":"SZ554632","au":"Martinková, Patrícia","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC(1,1) one-way random, single-rater (IRR = sigma_g^2 / (sigma_g^2 + sigma_e^2))","estd":"ICC (single/unspec)","v":0.35,"n":"47","k":"3","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"unspecified","field":"biomedical","wr":"AIBS proposals from male principal investigators","conf":"high","self":false,"doi":"10.3102/10769986221150517","ciLow":0.18,"ciHigh":0.51,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"Bayesian model averaging over the candidate variance-components models estimated the single-rater reliability (ICC(1,1)) of the three ratings per proposal at 0.35 for the 47 AIBS proposals from male principal investigators.","vf":"unverified"},{"key":"SZ554632","au":"Martinková, Patrícia","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC(1,1) one-way random, single-rater (IRR = sigma_g^2 / (sigma_g^2 + sigma_e^2))","estd":"ICC (single/unspec)","v":0.37,"n":"","k":"3","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"unspecified","field":"biomedical","wr":"reviewers on AIBS grant proposal merit scores","conf":"high","self":false,"doi":"10.3102/10769986221150517","ciLow":0.22,"ciHigh":0.52,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":true,"he":false,"ms":"Reviewers scored American Institutes of Biological Sciences grant proposals from 25 female and 47 male principal investigators, with three ratings per proposal; the single-rater inter-rater reliability, an ICC(1,1), was 0.37 under the selected model, which found no difference by applicant gender.","vf":"unverified"},{"key":"SZ554632","au":"Martinková, Patrícia","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC(1,1) one-way random, single-rater (IRR = sigma_g^2 / (sigma_g^2 + sigma_e^2))","estd":"ICC (single/unspec)","v":0.34,"n":"574","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"unspecified","field":"biomedical","wr":"NIH Investigator criterion scores for female PIs","conf":"high","self":false,"doi":"10.3102/10769986221150517","ciLow":0.3,"ciHigh":0.41,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"NIH reviewers' preliminary Investigator criterion scores for 574 proposals from female principal investigators gave a single-rater reliability (ICC(1,1)) of 0.34 under Bayesian model averaging across the candidate variance models.","vf":"unverified"},{"key":"SZ554632","au":"Martinková, Patrícia","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC(1,1) one-way random, single-rater (IRR = sigma_g^2 / (sigma_g^2 + sigma_e^2))","estd":"ICC (single/unspec)","v":0.36,"n":"1310","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"unspecified","field":"biomedical","wr":"NIH Investigator criterion scores for male PIs","conf":"high","self":false,"doi":"10.3102/10769986221150517","ciLow":0.32,"ciHigh":0.39,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"NIH reviewers' preliminary Investigator criterion scores for 1,310 proposals from male principal investigators gave a single-rater reliability (ICC(1,1)) of 0.36 under Bayesian model averaging across the candidate variance models.","vf":"unverified"},{"key":"SZ554632","au":"Martinková, Patrícia","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC(1,1) one-way random, single-rater (IRR = sigma_g^2 / (sigma_g^2 + sigma_e^2))","estd":"ICC (single/unspec)","v":0.33,"n":"574","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"unspecified","field":"biomedical","wr":"reviewers on NIH Investigator criterion scores","conf":"high","self":false,"doi":"10.3102/10769986221150517","ciLow":0.3,"ciHigh":0.36,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"For 574 National Institutes of Health grant proposals with female principal investigators, the single-rater inter-rater reliability (ICC(1,1)) of reviewers' preliminary Investigator criterion scores was 0.33 under the Bayes-factor-selected model; evidence for a gender difference was weak.","vf":"unverified"},{"key":"SZ554632","au":"Martinková, Patrícia","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC(1,1) one-way random, single-rater (IRR = sigma_g^2 / (sigma_g^2 + sigma_e^2))","estd":"ICC (single/unspec)","v":0.36,"n":"1310","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"unspecified","field":"biomedical","wr":"reviewers on NIH Investigator criterion scores","conf":"high","self":false,"doi":"10.3102/10769986221150517","ciLow":0.33,"ciHigh":0.39,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"For 1,310 National Institutes of Health grant proposals with male principal investigators, the single-rater inter-rater reliability (ICC(1,1)) of reviewers' preliminary Investigator criterion scores was 0.36 under the Bayes-factor-selected model.","vf":"unverified"},{"key":"SZ554632","au":"Martinková, Patrícia","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC(1,1) one-way random, single-rater (IRR = sigma_g^2 / (sigma_g^2 + sigma_e^2))","estd":"ICC (single/unspec)","v":0.33,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"unspecified","field":"biomedical","wr":"reviewers on NIH Investigator criterion scores","conf":"high","self":false,"doi":"10.3102/10769986221150517","ciLow":0.27,"ciHigh":0.4,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"Bayesian model-averaged single-rater inter-rater reliability (ICC(1,1)) of NIH reviewers' criterion scores for proposals by experienced female principal investigators was 0.33, with covariates allowed to adjust the mean; the structural between-proposal SD was 0.53.","vf":"unverified"},{"key":"SZ554632","au":"Martinková, Patrícia","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC(1,1) one-way random, single-rater (IRR = sigma_g^2 / (sigma_g^2 + sigma_e^2))","estd":"ICC (single/unspec)","v":0.33,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"unspecified","field":"biomedical","wr":"reviewers on NIH Investigator criterion scores","conf":"high","self":false,"doi":"10.3102/10769986221150517","ciLow":0.27,"ciHigh":0.39,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"Bayesian model-averaged single-rater inter-rater reliability (ICC(1,1)) for NIH proposals by experienced female principal investigators was 0.33 when the mean was held constant (no covariate adjustment); the structural between-proposal SD was 0.53.","vf":"unverified"},{"key":"SZ554632","au":"Martinková, Patrícia","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC(1,1) one-way random, single-rater (IRR = sigma_g^2 / (sigma_g^2 + sigma_e^2))","estd":"ICC (single/unspec)","v":0.29,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"unspecified","field":"biomedical","wr":"reviewers on NIH Investigator criterion scores","conf":"high","self":false,"doi":"10.3102/10769986221150517","ciLow":0.21,"ciHigh":0.37,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"Bayesian model-averaged single-rater inter-rater reliability (ICC(1,1)) of NIH reviewers' criterion scores for proposals by non-experienced female principal investigators was 0.29, with covariates allowed to adjust the mean; the structural between-proposal SD was 0.63.","vf":"unverified"},{"key":"SZ554632","au":"Martinková, Patrícia","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC(1,1) one-way random, single-rater (IRR = sigma_g^2 / (sigma_g^2 + sigma_e^2))","estd":"ICC (single/unspec)","v":0.41,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"unspecified","field":"biomedical","wr":"reviewers on NIH Investigator criterion scores","conf":"high","self":false,"doi":"10.3102/10769986221150517","ciLow":0.36,"ciHigh":0.47,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"Bayesian model-averaged single-rater inter-rater reliability (ICC(1,1)) for NIH proposals by non-experienced female principal investigators was 0.41 when the mean was held constant (no covariate adjustment); the structural between-proposal SD was 0.81.","vf":"unverified"},{"key":"SZ554632","au":"Martinková, Patrícia","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC(1,1) one-way random, single-rater (IRR = sigma_g^2 / (sigma_g^2 + sigma_e^2))","estd":"ICC (single/unspec)","v":0.32,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"unspecified","field":"biomedical","wr":"reviewers on NIH Investigator criterion scores","conf":"high","self":false,"doi":"10.3102/10769986221150517","ciLow":0.26,"ciHigh":0.37,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"Bayesian model-averaged single-rater inter-rater reliability (ICC(1,1)) of NIH reviewers' criterion scores for proposals by experienced male principal investigators was 0.32, with covariates allowed to adjust the mean; the structural between-proposal SD was 0.51.","vf":"unverified"},{"key":"SZ554632","au":"Martinková, Patrícia","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC(1,1) one-way random, single-rater (IRR = sigma_g^2 / (sigma_g^2 + sigma_e^2))","estd":"ICC (single/unspec)","v":0.32,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"unspecified","field":"biomedical","wr":"reviewers on NIH Investigator criterion scores","conf":"high","self":false,"doi":"10.3102/10769986221150517","ciLow":0.27,"ciHigh":0.37,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"Bayesian model-averaged single-rater inter-rater reliability (ICC(1,1)) for NIH proposals by experienced male principal investigators was 0.32 when the mean was held constant (no covariate adjustment); the structural between-proposal SD was 0.51.","vf":"unverified"},{"key":"SZ554632","au":"Martinková, Patrícia","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC(1,1) one-way random, single-rater (IRR = sigma_g^2 / (sigma_g^2 + sigma_e^2))","estd":"ICC (single/unspec)","v":0.28,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"unspecified","field":"biomedical","wr":"reviewers on NIH Investigator criterion scores","conf":"high","self":false,"doi":"10.3102/10769986221150517","ciLow":0.2,"ciHigh":0.34,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"Bayesian model-averaged single-rater inter-rater reliability (ICC(1,1)) of NIH reviewers' criterion scores for proposals by non-experienced male principal investigators was 0.28, with covariates allowed to adjust the mean; the structural between-proposal SD was 0.60.","vf":"unverified"},{"key":"SZ554632","au":"Martinková, Patrícia","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC(1,1) one-way random, single-rater (IRR = sigma_g^2 / (sigma_g^2 + sigma_e^2))","estd":"ICC (single/unspec)","v":0.41,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"unspecified","field":"biomedical","wr":"reviewers on NIH Investigator criterion scores","conf":"high","self":false,"doi":"10.3102/10769986221150517","ciLow":0.35,"ciHigh":0.46,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"Bayesian model-averaged single-rater inter-rater reliability (ICC(1,1)) for NIH proposals by non-experienced male principal investigators was 0.41 when the mean was held constant (no covariate adjustment); the structural between-proposal SD was 0.79.","vf":"unverified"},{"key":"85LYQU4R","au":"Mayo, Nancy E","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"kappa","form":"kappa statistic, adjusting for chance agreement, with 95% CI","estd":"kappa","v":0.36,"n":"32","k":"","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"fund / not fund","field":"biomedical","wr":"funding decisions from RANKING and CLASSIC methods","conf":"high","self":false,"doi":"10.1016/j.jclinepi.2005.12.007","ciLow":0.02,"ciHigh":0.7,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":true,"he":false,"ms":"Fund or not-fund decisions for 32 pilot grant applications were derived in parallel from an all-panel ranking method and the classic two-reviewer method. A kappa of 0.36 indicates poor chance-corrected agreement between the two methods' funding decisions.","vf":"unverified"},{"key":"UFFKJYI4","au":"McReynolds, Paul","y":1971,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"Pearsonian correlation, corrected by Spearman-Brown formula","estd":"correlation","v":0.84,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-10; 10 = brilliant, 1 = extremely poor","field":"clinical psychology","wr":"paired committee judges on conference paper summaries","conf":"high","self":false,"doi":"10.1037/h0037937","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The highest of the nine set-level correlations between the paired judges (Spearman-Brown corrected) was 0.84, for one randomly assigned set of 12 to 14 paper summaries.","vf":"unverified"},{"key":"UFFKJYI4","au":"McReynolds, Paul","y":1971,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"Pearsonian correlation, corrected by Spearman-Brown formula","estd":"correlation","v":0.21,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-10; 10 = brilliant, 1 = extremely poor","field":"clinical psychology","wr":"paired committee judges on conference paper summaries","conf":"high","self":false,"doi":"10.1037/h0037937","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The lowest of the nine set-level correlations between the paired judges (Spearman-Brown corrected) was 0.21, for one randomly assigned set of 12 to 14 paper summaries.","vf":"unverified"},{"key":"UFFKJYI4","au":"McReynolds, Paul","y":1971,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"Pearsonian correlation, corrected by Spearman-Brown formula; nine set correlations averaged by z transformation","estd":"correlation","v":0.62,"n":"118","k":"2","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-10; 10 = brilliant, 1 = extremely poor","field":"clinical psychology","wr":"paired committee judges on conference paper summaries","conf":"high","self":false,"doi":"10.1037/h0037937","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Nine pairs of committee judges (one academic, one applied per pair) independently rated 118 conference paper summaries on a 1-10 scale; the Spearman-Brown-corrected correlation between paired judges, averaged across the nine sets by z transformation, was 0.62, a moderate reliability for the two-judge composite.","vf":"unverified"},{"key":"T6YQ33G6","au":"Montgomery, Alan","y":2002,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"adjusted ICC (corrected for systematic differences and chance)","estd":"ICC (single/unspec)","v":0.45,"n":"52","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"one (poor) to four (excellent)","field":"primary care","wr":"referees scoring abstracts on achievement of aim","conf":"high","self":false,"doi":"10.1186/1472-6963-2-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"After correction for systematic referee differences, the intraclass correlation for achievement of aim was 0.45, the highest adjusted value, moderate agreement.","vf":"unverified"},{"key":"T6YQ33G6","au":"Montgomery, Alan","y":2002,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"crude ICC","estd":"ICC (single/unspec)","v":0.44,"n":"52","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"one (poor) to four (excellent)","field":"primary care","wr":"referees scoring abstracts on achievement of aim","conf":"high","self":false,"doi":"10.1186/1472-6963-2-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two referees independently scored 52 abstracts on achievement of aim (1 to 4). The crude intraclass correlation was 0.44, the highest crude value, moderate agreement.","vf":"unverified"},{"key":"T6YQ33G6","au":"Montgomery, Alan","y":2002,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"adjusted ICC (corrected for systematic differences and chance)","estd":"ICC (single/unspec)","v":0.25,"n":"52","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"one (poor) to four (excellent)","field":"primary care","wr":"referees scoring abstracts on contribution to primary care","conf":"high","self":false,"doi":"10.1186/1472-6963-2-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"After correction for systematic referee differences, the intraclass correlation for contribution to academic primary care was 0.25, fair agreement.","vf":"unverified"},{"key":"T6YQ33G6","au":"Montgomery, Alan","y":2002,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"crude ICC","estd":"ICC (single/unspec)","v":0.2,"n":"52","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"one (poor) to four (excellent)","field":"primary care","wr":"referees scoring abstracts on contribution to primary care","conf":"high","self":false,"doi":"10.1186/1472-6963-2-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two referees independently scored 52 abstracts on contribution to academic primary care (1 to 4). The crude intraclass correlation was 0.20.","vf":"unverified"},{"key":"T6YQ33G6","au":"Montgomery, Alan","y":2002,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"adjusted ICC (corrected for systematic differences and chance)","estd":"ICC (single/unspec)","v":0.41,"n":"52","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"one (poor) to four (excellent)","field":"primary care","wr":"referees scoring abstracts on design appropriateness","conf":"high","self":false,"doi":"10.1186/1472-6963-2-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"After correction for systematic referee differences, the intraclass correlation for appropriateness of the design was 0.41, moderate agreement by the paper's convention.","vf":"unverified"},{"key":"T6YQ33G6","au":"Montgomery, Alan","y":2002,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"crude ICC","estd":"ICC (single/unspec)","v":0.4,"n":"52","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"one (poor) to four (excellent)","field":"primary care","wr":"referees scoring abstracts on design appropriateness","conf":"high","self":false,"doi":"10.1186/1472-6963-2-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two referees independently scored 52 abstracts on appropriateness of the design used (1 to 4). The crude intraclass correlation was 0.40, at the upper bound of fair agreement.","vf":"unverified"},{"key":"T6YQ33G6","au":"Montgomery, Alan","y":2002,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"adjusted ICC (corrected for systematic differences and chance)","estd":"ICC (single/unspec)","v":0.4,"n":"52","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"one (poor) to four (excellent)","field":"primary care","wr":"referees scoring abstracts on study design quality","conf":"high","self":false,"doi":"10.1186/1472-6963-2-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"After correction for systematic referee differences, the intraclass correlation for overall study design quality was 0.40, at the upper bound of fair agreement.","vf":"unverified"},{"key":"T6YQ33G6","au":"Montgomery, Alan","y":2002,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"crude ICC","estd":"ICC (single/unspec)","v":0.3,"n":"52","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"one (poor) to four (excellent)","field":"primary care","wr":"referees scoring abstracts on study design quality","conf":"high","self":false,"doi":"10.1186/1472-6963-2-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two referees independently scored 52 abstracts on overall quality of the study design (1 to 4). The crude intraclass correlation was 0.30, fair agreement.","vf":"unverified"},{"key":"T6YQ33G6","au":"Montgomery, Alan","y":2002,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"adjusted ICC (corrected for systematic differences and chance)","estd":"ICC (single/unspec)","v":0.24,"n":"52","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"one (poor) to four (excellent)","field":"primary care","wr":"referees scoring abstracts on importance of topic","conf":"high","self":false,"doi":"10.1186/1472-6963-2-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"After subtracting the mean systematic difference between the two referees' scores, the intraclass correlation for importance of the topic was 0.24, fair agreement by the paper's convention.","vf":"unverified"},{"key":"T6YQ33G6","au":"Montgomery, Alan","y":2002,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"crude ICC","estd":"ICC (single/unspec)","v":0,"n":"52","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"one (poor) to four (excellent)","field":"primary care","wr":"referees scoring abstracts on importance of topic","conf":"high","self":false,"doi":"10.1186/1472-6963-2-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two referees independently scored 52 anonymised conference abstracts on importance of the topic (1 to 4). The crude intraclass correlation of 0 indicates no chance-corrected agreement.","vf":"unverified"},{"key":"T6YQ33G6","au":"Montgomery, Alan","y":2002,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"adjusted ICC (corrected for systematic differences and chance)","estd":"ICC (single/unspec)","v":0.01,"n":"52","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"one (poor) to four (excellent)","field":"primary care","wr":"referees scoring abstracts on originality","conf":"high","self":false,"doi":"10.1186/1472-6963-2-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"After correction for systematic differences between the two referees, the intraclass correlation for originality was 0.01, essentially no agreement.","vf":"unverified"},{"key":"T6YQ33G6","au":"Montgomery, Alan","y":2002,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"crude ICC","estd":"ICC (single/unspec)","v":0,"n":"52","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"one (poor) to four (excellent)","field":"primary care","wr":"referees scoring abstracts on originality","conf":"high","self":false,"doi":"10.1186/1472-6963-2-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two referees independently scored 52 abstracts on originality (1 to 4). The crude intraclass correlation was 0, indicating no chance-corrected agreement.","vf":"unverified"},{"key":"T6YQ33G6","au":"Montgomery, Alan","y":2002,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"adjusted ICC (corrected for systematic differences and chance)","estd":"ICC (single/unspec)","v":0.41,"n":"52","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"summed total of seven categories, possible 7 to 28","field":"primary care","wr":"referees scoring abstracts: overall total score","conf":"high","self":false,"doi":"10.1186/1472-6963-2-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"After correction for the systematic difference between the two referees, the intraclass correlation of the summed total score across 52 abstracts was 0.41, moderate agreement. The paper's abstract reports its results in adjusted form.","vf":"unverified"},{"key":"T6YQ33G6","au":"Montgomery, Alan","y":2002,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"crude ICC","estd":"ICC (single/unspec)","v":0.31,"n":"52","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"summed total of seven categories, possible 7 to 28","field":"primary care","wr":"referees scoring abstracts: overall total score","conf":"high","self":false,"doi":"10.1186/1472-6963-2-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Each referee's seven category marks were summed to a total score (possible 7 to 28) for each of 52 abstracts. The crude intraclass correlation between the two referees' totals was 0.31, fair agreement.","vf":"unverified"},{"key":"T6YQ33G6","au":"Montgomery, Alan","y":2002,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"adjusted ICC (corrected for systematic differences and chance)","estd":"ICC (single/unspec)","v":0.24,"n":"52","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"one (poor) to four (excellent)","field":"primary care","wr":"referees scoring abstracts on provoking discussion","conf":"high","self":false,"doi":"10.1186/1472-6963-2-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"After correction for systematic referee differences, the intraclass correlation for likelihood of provoking discussion was 0.24, fair agreement.","vf":"unverified"},{"key":"T6YQ33G6","au":"Montgomery, Alan","y":2002,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"crude ICC","estd":"ICC (single/unspec)","v":0.22,"n":"52","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"one (poor) to four (excellent)","field":"primary care","wr":"referees scoring abstracts on provoking discussion","conf":"high","self":false,"doi":"10.1186/1472-6963-2-8","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two referees independently scored 52 abstracts on likelihood of provoking discussion (1 to 4). The crude intraclass correlation was 0.22, fair agreement.","vf":"unverified"},{"key":"LFVVACBY","au":"Munley, Patrick H.","y":1988,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Finn's r (rf; Finn, 1970)","estd":"other","v":0.52,"n":"262","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"poor, marginal, adequate, good, or excellent","field":"counseling psychology","wr":"reviewers on manuscript clarity of presentation","conf":"high","self":false,"doi":"10.1037/0022-0167.35.2.198","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For 262 manuscripts rated on clarity of presentation, Finn's r of 0.52 estimates the proportion of reviewer agreement not attributable to chance.","vf":"unverified"},{"key":"LFVVACBY","au":"Munley, Patrick H.","y":1988,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Finn's r (rf; Finn, 1970)","estd":"other","v":0.37,"n":"207","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"definitely should not / probably should not / probably should / should be published","field":"counseling psychology","wr":"reviewers on manuscript publication recommendation","conf":"high","self":false,"doi":"10.1037/0022-0167.35.2.198","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the same 207 manuscripts, Finn's r of 0.37 estimates the proportion of reviewer agreement on publication recommendation that is not attributable to chance.","vf":"unverified"},{"key":"LFVVACBY","au":"Munley, Patrick H.","y":1988,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Finn's r (rf; Finn, 1970)","estd":"other","v":0.51,"n":"244","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"poor, marginal, adequate, good, or excellent","field":"counseling psychology","wr":"reviewers on manuscript interpretation of results","conf":"high","self":false,"doi":"10.1037/0022-0167.35.2.198","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For 244 manuscripts rated on interpretation of results and conclusions, Finn's r of 0.51 estimates the proportion of reviewer agreement not attributable to chance.","vf":"unverified"},{"key":"LFVVACBY","au":"Munley, Patrick H.","y":1988,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Finn's r (rf; Finn, 1970)","estd":"other","v":0.5,"n":"245","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"poor, marginal, adequate, good, or excellent","field":"counseling psychology","wr":"reviewers on manuscript length","conf":"high","self":false,"doi":"10.1037/0022-0167.35.2.198","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For 245 manuscripts rated on length, Finn's r of 0.50 estimates the proportion of reviewer agreement not attributable to chance.","vf":"unverified"},{"key":"LFVVACBY","au":"Munley, Patrick H.","y":1988,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Finn's r (rf; Finn, 1970)","estd":"other","v":0.52,"n":"236","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"poor, marginal, adequate, good, or excellent","field":"counseling psychology","wr":"reviewers on manuscript quality of methodology","conf":"high","self":false,"doi":"10.1037/0022-0167.35.2.198","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For 236 manuscripts rated on quality of methodology, Finn's r of 0.52 estimates the proportion of reviewer agreement not attributable to chance.","vf":"unverified"},{"key":"LFVVACBY","au":"Munley, Patrick H.","y":1988,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Finn's r (rf; Finn, 1970)","estd":"other","v":0.48,"n":"259","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"poor, marginal, adequate, good, or excellent","field":"counseling psychology","wr":"reviewers on manuscript overall importance","conf":"high","self":false,"doi":"10.1037/0022-0167.35.2.198","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For 259 manuscripts rated on the paper's overall importance, Finn's r of 0.48 estimates the proportion of reviewer agreement not attributable to chance.","vf":"unverified"},{"key":"LFVVACBY","au":"Munley, Patrick H.","y":1988,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Finn's r (rf; Finn, 1970)","estd":"other","v":0.56,"n":"258","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"poor, marginal, adequate, good, or excellent","field":"counseling psychology","wr":"reviewers on manuscript relation to literature","conf":"high","self":false,"doi":"10.1037/0022-0167.35.2.198","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For 258 manuscripts rated on relation to the literature, Finn's r of 0.56 estimates the proportion of reviewer agreement not attributable to chance.","vf":"unverified"},{"key":"LFVVACBY","au":"Munley, Patrick H.","y":1988,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Finn's r (rf; Finn, 1970)","estd":"other","v":0.51,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"poor, marginal, adequate, good, or excellent","field":"counseling psychology","wr":"seven manuscript quality dimensions","conf":"high","self":false,"doi":"10.1037/0022-0167.35.2.198","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers used seven quality scales, with available manuscript pairs varying by scale. The paper reports a mean Finn's r of 0.51 across the seven scales.","vf":"unverified"},{"key":"LFVVACBY","au":"Munley, Patrick H.","y":1988,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Finn's r (rf; Finn, 1970)","estd":"other","v":0.48,"n":"260","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"poor, marginal, adequate, good, or excellent","field":"counseling psychology","wr":"reviewers on manuscript significance of topic","conf":"high","self":false,"doi":"10.1037/0022-0167.35.2.198","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For 260 manuscripts rated on significance of topic, Finn's r of 0.48 estimates the proportion of reviewer agreement not attributable to chance.","vf":"unverified"},{"key":"LFVVACBY","au":"Munley, Patrick H.","y":1988,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (Ri), Whitehurst (1984) procedure","estd":"ICC (single/unspec)","v":0.24,"n":"262","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"poor, marginal, adequate, good, or excellent","field":"counseling psychology","wr":"reviewers on manuscript clarity of presentation","conf":"high","self":false,"doi":"10.1037/0022-0167.35.2.198","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated 262 manuscripts on clarity of presentation; an intraclass correlation of 0.24 reflects the share of rating variance due to true manuscript differences.","vf":"unverified"},{"key":"LFVVACBY","au":"Munley, Patrick H.","y":1988,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (Ri), Whitehurst (1984) procedure","estd":"ICC (single/unspec)","v":0.28,"n":"207","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"definitely should not / probably should not / probably should / should be published","field":"counseling psychology","wr":"reviewers on manuscript publication recommendation","conf":"high","self":false,"doi":"10.1037/0022-0167.35.2.198","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Two independent reviewers rated 207 manuscripts on their overall publication recommendation; an intraclass correlation of 0.28 indicates that about 28 per cent of rating variance reflects true differences between manuscripts.","vf":"unverified"},{"key":"LFVVACBY","au":"Munley, Patrick H.","y":1988,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (Ri), Whitehurst (1984) procedure","estd":"ICC (single/unspec)","v":0.19,"n":"244","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"poor, marginal, adequate, good, or excellent","field":"counseling psychology","wr":"reviewers on manuscript interpretation of results","conf":"high","self":false,"doi":"10.1037/0022-0167.35.2.198","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated 244 manuscripts on interpretation of results and conclusions; an intraclass correlation of 0.19 reflects the share of rating variance due to true manuscript differences.","vf":"unverified"},{"key":"LFVVACBY","au":"Munley, Patrick H.","y":1988,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (Ri), Whitehurst (1984) procedure","estd":"ICC (single/unspec)","v":0.13,"n":"245","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"poor, marginal, adequate, good, or excellent","field":"counseling psychology","wr":"reviewers on manuscript length","conf":"high","self":false,"doi":"10.1037/0022-0167.35.2.198","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated 245 manuscripts on length; an intraclass correlation of 0.13 reflects the share of rating variance due to true manuscript differences.","vf":"unverified"},{"key":"LFVVACBY","au":"Munley, Patrick H.","y":1988,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (Ri), Whitehurst (1984) procedure","estd":"ICC (single/unspec)","v":0.28,"n":"236","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"poor, marginal, adequate, good, or excellent","field":"counseling psychology","wr":"reviewers on manuscript quality of methodology","conf":"high","self":false,"doi":"10.1037/0022-0167.35.2.198","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated 236 manuscripts on quality of methodology; an intraclass correlation of 0.28 reflects the share of rating variance due to true manuscript differences.","vf":"unverified"},{"key":"LFVVACBY","au":"Munley, Patrick H.","y":1988,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (Ri), Whitehurst (1984) procedure","estd":"ICC (single/unspec)","v":0.28,"n":"259","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"poor, marginal, adequate, good, or excellent","field":"counseling psychology","wr":"reviewers on manuscript overall importance","conf":"high","self":false,"doi":"10.1037/0022-0167.35.2.198","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated 259 manuscripts on the paper's overall importance; an intraclass correlation of 0.28 reflects the share of rating variance due to true manuscript differences.","vf":"unverified"},{"key":"LFVVACBY","au":"Munley, Patrick H.","y":1988,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (Ri), Whitehurst (1984) procedure","estd":"ICC (single/unspec)","v":0.35,"n":"258","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"poor, marginal, adequate, good, or excellent","field":"counseling psychology","wr":"reviewers on manuscript relation to literature","conf":"high","self":false,"doi":"10.1037/0022-0167.35.2.198","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated 258 manuscripts on relation to the literature; an intraclass correlation of 0.35 reflects the share of rating variance due to true manuscript differences.","vf":"unverified"},{"key":"LFVVACBY","au":"Munley, Patrick H.","y":1988,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (Ri), Whitehurst (1984) procedure","estd":"ICC (single/unspec)","v":0.24,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"poor, marginal, adequate, good, or excellent","field":"counseling psychology","wr":"seven manuscript quality dimensions","conf":"high","self":false,"doi":"10.1037/0022-0167.35.2.198","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers used seven quality scales, with available manuscript pairs varying by scale. The paper reports a mean intraclass correlation of 0.24 across the seven scales.","vf":"unverified"},{"key":"LFVVACBY","au":"Munley, Patrick H.","y":1988,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (Ri), Whitehurst (1984) procedure","estd":"ICC (single/unspec)","v":0.2,"n":"260","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"poor, marginal, adequate, good, or excellent","field":"counseling psychology","wr":"reviewers on manuscript significance of topic","conf":"high","self":false,"doi":"10.1037/0022-0167.35.2.198","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated 260 manuscripts on significance of topic; an intraclass correlation of 0.20 reflects the share of rating variance due to true manuscript differences.","vf":"unverified"},{"key":"LFVVACBY","au":"Munley, Patrick H.","y":1988,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact (perfect) agreement","estd":"percent agreement","v":0.338,"n":"207","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"definitely should not / probably should not / probably should / should be published","field":"counseling psychology","wr":"reviewers on manuscript publication recommendation","conf":"high","self":false,"doi":"10.1037/0022-0167.35.2.198","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two reviewers gave exactly the same publication recommendation category for 33.8 per cent of the 207 manuscripts.","vf":"unverified"},{"key":"LFVVACBY","au":"Munley, Patrick H.","y":1988,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"disagreement beyond one category ('split reviews')","estd":"percent agreement","v":0.353,"n":"207","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"definitely should not / probably should not / probably should / should be published","field":"counseling psychology","wr":"reviewers on manuscript publication recommendation","conf":"high","self":false,"doi":"10.1037/0022-0167.35.2.198","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two reviewers gave split recommendations, disagreeing beyond one category, for 35.3 per cent of the 207 manuscripts. This is the complement of the within-one-category agreement rate.","vf":"unverified"},{"key":"LFVVACBY","au":"Munley, Patrick H.","y":1988,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"agreement within one category","estd":"percent agreement","v":0.647,"n":"207","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"definitely should not / probably should not / probably should / should be published","field":"counseling psychology","wr":"reviewers on manuscript publication recommendation","conf":"high","self":false,"doi":"10.1037/0022-0167.35.2.198","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two reviewers agreed within one publication recommendation category for 64.7 per cent of the 207 manuscripts.","vf":"unverified"},{"key":"7EKMP546","au":"Mutz, Rüdiger","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"second-order GEE intra-proposal ICC, modified Fisher-z link, base model with no predictors","estd":"ICC (single/unspec)","v":0.22,"n":"8329","k":"2.82","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 to 100 (from poor to excellent)","field":"multi-field","wr":"external reviewers on grant proposal ratings, GEE base model","conf":"high","self":false,"doi":"10.1371/journal.pone.0048509","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The unconditional second-order GEE model estimated a mean single-rater ICC of .22, slightly below the descriptive ICC and flagged by the authors as needing cautious interpretation.","vf":"unverified"},{"key":"7EKMP546","au":"Mutz, Rüdiger","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"second-order GEE intra-proposal ICC, modified Fisher-z link, residual ICC adjusted for ICC, variance and mean covariates","estd":"ICC (single/unspec)","v":0.2,"n":"8329","k":"2.82","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 to 100 (from poor to excellent)","field":"multi-field","wr":"external reviewers on grant proposal ratings, fully adjusted GEE model","conf":"high","self":false,"doi":"10.1371/journal.pone.0048509","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"In the fully adjusted GEE model that includes covariates for the ICC, variance and mean, the residual intercept single-rater ICC shrank to .20.","vf":"unverified"},{"key":"7EKMP546","au":"Mutz, Rüdiger","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC of the mean ratings; ICC(1, k) in the notation by Shrout and Fleiss","estd":"ICC (average)","v":0.383,"n":"1628","k":"2.82","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"0 to 100 (from poor to excellent)","field":"multi-field","wr":"mean of reviewers' ratings on biosciences proposals","conf":"high","self":false,"doi":"10.1371/journal.pone.0048509","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the mean of the reviews per bioscience proposal the ICC was .383, the lowest across research areas but higher than the corresponding single-rater value.","vf":"unverified"},{"key":"7EKMP546","au":"Mutz, Rüdiger","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC of the mean ratings; ICC(1, k) in the notation by Shrout and Fleiss","estd":"ICC (average)","v":0.55,"n":"1413","k":"2.82","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"0 to 100 (from poor to excellent)","field":"multi-field","wr":"mean of reviewers' ratings on humanities proposals","conf":"high","self":false,"doi":"10.1371/journal.pone.0048509","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the mean of the reviews per humanities proposal the ICC was .55, the highest across research areas.","vf":"unverified"},{"key":"7EKMP546","au":"Mutz, Rüdiger","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC of mean ratings; ICC(1, k) in the notation by Shrout and Fleiss","estd":"ICC (average)","v":0.495,"n":"8329","k":"2.82","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"0 to 100 (from poor to excellent)","field":"multi-field","wr":"mean of reviewers' overall ratings per grant proposal","conf":"high","self":false,"doi":"10.1371/journal.pone.0048509","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Scaling the single-rater ICC up to the mean of the 2.82 reviews per proposal via Spearman-Brown gives an ICC of .495 for the reliability of the averaged FWF review score.","vf":"unverified"},{"key":"7EKMP546","au":"Mutz, Rüdiger","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC (single IRR)","estd":"ICC (single/unspec)","v":0.259,"n":"8329","k":"2.82","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 to 100 (from poor to excellent)","field":"multi-field","wr":"external reviewers on grant proposal overall ratings","conf":"high","self":false,"doi":"10.1371/journal.pone.0048509","ciLow":0.249,"ciHigh":0.279,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"About 18,357 external reviewers gave overall 0-100 ratings to 8,329 FWF grant proposals in 23,414 reviews; the single-rater intraclass correlation of .259 indicates low agreement between two reviews of the same proposal.","vf":"unverified"},{"key":"7EKMP546","au":"Mutz, Rüdiger","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC (single IRR)","estd":"ICC (single/unspec)","v":0.183,"n":"1628","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 to 100 (from poor to excellent)","field":"multi-field","wr":"external reviewers on biosciences grant proposal ratings","conf":"high","self":false,"doi":"10.1371/journal.pone.0048509","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Within the 1,628 bioscience proposals the single-rater ICC was .183, the lowest of the six FWF research areas, indicating especially weak reviewer agreement there.","vf":"unverified"},{"key":"7EKMP546","au":"Mutz, Rüdiger","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC (single IRR)","estd":"ICC (single/unspec)","v":0.229,"n":"1621","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 to 100 (from poor to excellent)","field":"multi-field","wr":"external reviewers on human medicine grant proposal ratings","conf":"high","self":false,"doi":"10.1371/journal.pone.0048509","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Within the 1,621 human medicine proposals the single-rater ICC was .229, somewhat below the overall FWF value.","vf":"unverified"},{"key":"7EKMP546","au":"Mutz, Rüdiger","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC (single IRR)","estd":"ICC (single/unspec)","v":0.319,"n":"1413","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 to 100 (from poor to excellent)","field":"multi-field","wr":"external reviewers on humanities grant proposal ratings","conf":"high","self":false,"doi":"10.1371/journal.pone.0048509","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Within the 1,413 humanities proposals the single-rater ICC was .319, the highest of the six research areas, driven by greater between-proposal quality variance.","vf":"unverified"},{"key":"7EKMP546","au":"Mutz, Rüdiger","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC (single IRR)","estd":"ICC (single/unspec)","v":0.255,"n":"2450","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 to 100 (from poor to excellent)","field":"multi-field","wr":"external reviewers on natural sciences grant proposal ratings","conf":"high","self":false,"doi":"10.1371/journal.pone.0048509","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Within the 2,450 natural science proposals the single-rater ICC was .255, close to the overall FWF value.","vf":"unverified"},{"key":"7EKMP546","au":"Mutz, Rüdiger","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC (single IRR)","estd":"ICC (single/unspec)","v":0.213,"n":"697","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 to 100 (from poor to excellent)","field":"multi-field","wr":"external reviewers on social sciences grant proposal ratings","conf":"high","self":false,"doi":"10.1371/journal.pone.0048509","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Within the 697 social science proposals the single-rater ICC was .213; despite highly heterogeneous ratings the reliability stayed low.","vf":"unverified"},{"key":"7EKMP546","au":"Mutz, Rüdiger","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC (single IRR)","estd":"ICC (single/unspec)","v":0.326,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 to 100 (from poor to excellent)","field":"multi-field","wr":"external reviewers on proposals decided in year 2000","conf":"high","self":false,"doi":"10.1371/journal.pone.0048509","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For proposals with a 2000 final decision the single-rater ICC was .326, significantly higher than in 2004 and 2005; across years the ICCs otherwise varied little.","vf":"unverified"},{"key":"7EKMP546","au":"Mutz, Rüdiger","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC (single IRR)","estd":"ICC (single/unspec)","v":0.215,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 to 100 (from poor to excellent)","field":"multi-field","wr":"external reviewers on proposals decided in year 2004","conf":"high","self":false,"doi":"10.1371/journal.pone.0048509","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For proposals with a 2004 final decision the single-rater ICC was .215, among the lowest across the study years.","vf":"unverified"},{"key":"7EKMP546","au":"Mutz, Rüdiger","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC (single IRR)","estd":"ICC (single/unspec)","v":0.224,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 to 100 (from poor to excellent)","field":"multi-field","wr":"external reviewers on proposals decided in year 2005","conf":"high","self":false,"doi":"10.1371/journal.pone.0048509","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For proposals with a 2005 final decision the single-rater ICC was .224, significantly lower than the year-2000 value.","vf":"unverified"},{"key":"B68EMUVK","au":"Mutz, Rüdiger","y":2012,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraclass correlation coefficient for single ratings (sum of variance components except the residual component divided by the total variance)","estd":"ICC (single/unspec)","v":0.26,"n":"8358","k":"","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 to 100 (from poor to excellent)","field":"multi-field","wr":"external reviewers on grant proposal scores","conf":"med","self":false,"doi":"10.1027/2151-2604/a000103","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"External reviewers scored 8,358 FWF grant proposals on a 1 to 100 scale in 23,977 reviews, about two to three reviews per proposal. The single-rating intraclass correlation of 0.26 means two independent ratings of the same proposal correlate about 0.26, indicating low inter-rater reliability.","vf":"unverified"},{"key":"74M956VR","au":"Mutz, Rüdiger","y":2016,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"interrater reliability of the average of ratings","estd":"ICC (average)","v":0.21,"n":"3011","k":"2.49","samp":"funded-only","blind":"unclear","agg":"average-of-k","scale":"A-D 4-point (2007); A-E 5-point (2008)","field":"chemistry","wr":"averaged referee ratings of journal manuscript importance","conf":"med","self":false,"doi":"10.1002/asi.23701","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"When the referees' importance ratings are averaged per paper, the intraclass correlation rises only from 0.13 to 0.21, still indicating low reliability of the aggregated importance rating.","vf":"unverified"},{"key":"74M956VR","au":"Mutz, Rüdiger","y":2016,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass coefficient for ordinal measures","estd":"ICC (single/unspec)","v":0.13,"n":"3011","k":"2.49","samp":"funded-only","blind":"unclear","agg":"single-rater","scale":"A-D 4-point (2007); A-E 5-point (2008)","field":"chemistry","wr":"referees on journal manuscript importance ratings","conf":"med","self":false,"doi":"10.1002/asi.23701","ciLow":0.1,"ciHigh":0.16,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":true,"he":false,"ms":"Journal referees (on average 2.49 per paper) rated the importance of 3,011 published Angewandte Chemie Communications; the single-rater intraclass correlation of 0.13 (95% CI 0.10 to 0.16) shows very low agreement between referees, computed only on published papers.","vf":"unverified"},{"key":"H9B5WQZN","au":"Mutz, Rüdiger","y":2016,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intra-class correlation for the mean ratings (qM), Spearman-Brown adjusted from single-rater ICC (Equation (8))","estd":"ICC (average)","v":0.495,"n":"8496","k":"2.82","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"continuous scale from 0 to 100, poor to excellent","field":"multi-field","wr":"external referees on grant proposal merit scores","conf":"med","self":false,"doi":"10.1093/reseval/rvw002","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"External referees independently scored 8,496 FWF grant proposals on a 0-100 merit scale, on average 2.82 referees per proposal. The intra-class correlation of 0.495 estimates the reliability of each proposal's mean rating, obtained by Spearman-Brown adjustment of a single-rater ICC of 0.259 the authors reported in an earlier study.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,1) mean of k, Spearman-Brown, continuous, composite of both criteria, k=2.56","estd":"ICC (average)","v":0.31,"n":"875","k":"2.56","samp":"full-pool","blind":"open","agg":"average-of-k","scale":"sum of significance + strength (0-11)","field":"life sciences and medicine","wr":"reviewers on combined significance+strength score","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":0.23,"ciHigh":0.4,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"For a composite indicator summing both review criteria, the reliability of the mean of 2.56 reviewers was an ICC of 0.31, showing no improvement over the individual criteria.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,1) one-way random, single-rater, continuous, composite of both criteria","estd":"ICC (single/unspec)","v":0.15,"n":"875","k":"2.56","samp":"full-pool","blind":"open","agg":"single-rater","scale":"sum of significance + strength (0-11)","field":"life sciences and medicine","wr":"reviewers on combined significance+strength score","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":0.1,"ciHigh":0.2,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"For a composite indicator summing both review criteria, the single-rater intraclass correlation across a mean of 2.56 reviewers was 0.15, showing no improvement over the individual criteria.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"Gwet-AC","form":"Gwet’s chance-corrected AC1","estd":"Gwet AC","v":0.2,"n":"875","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"0=prefer not, useful, valuable, important, fundamental, landmark","field":"life sciences and medicine","wr":"reviewers on manuscript significance ratings","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":0.16,"ciHigh":0.23,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Gwet’s paradox-resistant AC1 for significance of findings was 0.20 between the first two reviewers, a slightly higher but still low chance-corrected agreement.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"Fleiss-kappa","form":"Fleiss kappa (3 raters)","estd":"kappa","v":0.03,"n":"875","k":"3","samp":"full-pool","blind":"open","agg":"single-rater","scale":"0=prefer not, useful, valuable, important, fundamental, landmark","field":"life sciences and medicine","wr":"reviewers on manuscript significance ratings","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":0,"ciHigh":0.06,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Across submissions rated by three reviewers, a Fleiss kappa of 0.03 for significance of findings indicates essentially no chance-corrected agreement among the reviewers.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,1) mean of k raters, Spearman-Brown stepped up to k=2.56, continuous treatment","estd":"ICC (average)","v":0.2,"n":"875","k":"2.56","samp":"full-pool","blind":"open","agg":"average-of-k","scale":"0=prefer not, useful, valuable, important, fundamental, landmark","field":"life sciences and medicine","wr":"reviewers on manuscript significance ratings","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":0.11,"ciHigh":0.3,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Treating significance scores as continuous, the reliability of the mean of 2.56 reviewers per manuscript rose to an ICC of 0.20, still a low reliability for the averaged rating.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC(k) average from multilevel ordinal logistic regression with six covariates, k=2.56","estd":"ICC (average)","v":0.22,"n":"875","k":"2.56","samp":"full-pool","blind":"open","agg":"average-of-k","scale":"0=prefer not, useful, valuable, important, fundamental, landmark","field":"life sciences and medicine","wr":"reviewers on manuscript significance ratings","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":0.12,"ciHigh":0.32,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"In the covariate-adjusted multilevel ordinal model, the reliability of the mean of 2.56 reviewers for significance of findings was an ICC of 0.22, essentially unchanged from the model without covariates.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,1)-type mean of k, Spearman-Brown, from ordinal mixed-effects model, k=2.56","estd":"ICC (average)","v":0.24,"n":"875","k":"2.56","samp":"full-pool","blind":"open","agg":"average-of-k","scale":"0=prefer not, useful, valuable, important, fundamental, landmark","field":"life sciences and medicine","wr":"reviewers on manuscript significance ratings","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":0.14,"ciHigh":0.34,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Treating significance scores as ordinal, the reliability of the mean of 2.56 reviewers per manuscript was an ICC of 0.24, still a low reliability for the averaged rating.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,1) one-way random, single-rater (Shrout & Fleiss), continuous treatment","estd":"ICC (single/unspec)","v":0.09,"n":"875","k":"2.56","samp":"full-pool","blind":"open","agg":"single-rater","scale":"0=prefer not, useful, valuable, important, fundamental, landmark","field":"life sciences and medicine","wr":"reviewers on manuscript significance ratings","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":0.04,"ciHigh":0.14,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Treating significance-of-findings scores as continuous, the single-rater intraclass correlation across a mean of 2.56 reviewers per manuscript was 0.09, a low reliability for one reviewer’s rating.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC(1) single from multilevel ordinal logistic regression with six covariates","estd":"ICC (single/unspec)","v":0.1,"n":"875","k":"2.56","samp":"full-pool","blind":"open","agg":"single-rater","scale":"0=prefer not, useful, valuable, important, fundamental, landmark","field":"life sciences and medicine","wr":"reviewers on manuscript significance ratings","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":0.05,"ciHigh":0.15,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"In the covariate-adjusted multilevel ordinal model, the single-rater intraclass correlation for significance of findings was 0.10, essentially unchanged from the model without covariates.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,1)-type single-rater from mixed-effects multinomial logistic model (Hedeker 2003), ordinal","estd":"ICC (single/unspec)","v":0.11,"n":"875","k":"2.56","samp":"full-pool","blind":"open","agg":"single-rater","scale":"0=prefer not, useful, valuable, important, fundamental, landmark","field":"life sciences and medicine","wr":"reviewers on manuscript significance ratings","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":0.06,"ciHigh":0.16,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Treating significance scores as ordinal via a mixed-effects model, the single-rater intraclass correlation across a mean of 2.56 reviewers was 0.11, a low reliability for one reviewer’s rating.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,1)-type single-rater ordinal, by submission month","estd":"ICC (single/unspec)","v":0.27,"n":"","k":"","samp":"full-pool","blind":"open","agg":"single-rater","scale":"0=prefer not, useful, valuable, important, fundamental, landmark","field":"life sciences and medicine","wr":"reviewers on manuscript significance ratings","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Split by submission month, the single-rater ordinal intraclass correlation for significance of findings reached 0.27 in July to September, above the overall average of 0.11.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"observed percentage agreement (exact agreement, diagonal of first-two-reviewer cross-table)","estd":"percent agreement","v":0.32,"n":"875","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"0=prefer not, useful, valuable, important, fundamental, landmark","field":"life sciences and medicine","wr":"reviewers on manuscript significance ratings","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"First two reviewers of 875 eLife manuscripts rated significance of findings on a 6-category scale; they used the identical category in 32% of manuscripts, an uncorrected agreement.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"single Cohen’s kappa (unweighted), first two reviewers","estd":"kappa","v":0.08,"n":"875","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"0=prefer not, useful, valuable, important, fundamental, landmark","field":"life sciences and medicine","wr":"reviewers on manuscript significance ratings","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":0.04,"ciHigh":0.12,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":false,"ms":"First two reviewers independently rated significance of findings for 875 eLife manuscripts; a chance-corrected Cohen’s kappa of 0.08 indicates a lack of agreement between them.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"weighted Cohen's kappa, Cicchetti-Allison linear weights","estd":"weighted kappa","v":0.09,"n":"875","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"0=prefer not, useful, valuable, important, fundamental, landmark","field":"life sciences and medicine","wr":"reviewers on manuscript significance ratings","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":0.05,"ciHigh":0.14,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"A quadratically weighted Cohen’s kappa of 0.09 for significance of findings, giving partial credit for near disagreements, still shows a lack of chance-corrected agreement between the first two reviewers.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"weighted observed agreement, Cicchetti-Allison linear weights (partial credit for near-diagonal cells)","estd":"percent agreement","v":0.78,"n":"875","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"0=prefer not, useful, valuable, important, fundamental, landmark","field":"life sciences and medicine","wr":"reviewers on manuscript significance ratings","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Giving partial credit to near-diagonal disagreements with Cicchetti-Allison weights, the weighted observed agreement for significance of findings between the first two reviewers was 0.78.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"single Cohen's kappa","estd":"kappa","v":0.1,"n":"875","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"Fundamental versus all other categories","field":"life sciences and medicine","wr":"fundamental significance classification","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Two reviewers classified manuscripts as fundamental or another significance category. The category-specific Cohen's kappa was 0.10.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"single Cohen's kappa","estd":"kappa","v":0.04,"n":"875","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"Important versus all other categories","field":"life sciences and medicine","wr":"important significance classification","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Two reviewers classified manuscripts as important or another significance category. The category-specific Cohen's kappa was 0.04.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"single Cohen's kappa","estd":"kappa","v":-0.01,"n":"875","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"Landmark versus all other categories","field":"life sciences and medicine","wr":"landmark significance classification","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Two reviewers classified manuscripts as landmark or another significance category. The category-specific Cohen's kappa was -0.01.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"single Cohen's kappa","estd":"kappa","v":0.02,"n":"875","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"Prefer not to answer versus all other categories","field":"life sciences and medicine","wr":"significance category classification","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Two reviewers classified manuscripts as prefer not to answer or another significance category. The category-specific Cohen's kappa was 0.02.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"single Cohen's kappa","estd":"kappa","v":0.11,"n":"875","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"Useful versus all other categories","field":"life sciences and medicine","wr":"useful significance classification","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Two reviewers classified manuscripts as useful or another significance category. The category-specific Cohen's kappa was 0.11.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"single Cohen's kappa","estd":"kappa","v":0.1,"n":"875","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"Valuable versus all other categories","field":"life sciences and medicine","wr":"valuable significance classification","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Two reviewers classified manuscripts as valuable or another significance category. The category-specific Cohen's kappa was 0.10.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"Gwet-AC","form":"Gwet’s chance-corrected AC1","estd":"Gwet AC","v":0.17,"n":"875","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"0=prefer not, inadequate, incomplete, solid, convincing, compelling, exceptional","field":"life sciences and medicine","wr":"reviewers on manuscript strength-of-support ratings","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":0.14,"ciHigh":0.21,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Gwet’s paradox-resistant AC1 for strength of support was 0.17 between the first two reviewers, a slightly higher but still low chance-corrected agreement.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"Fleiss-kappa","form":"Fleiss kappa (3 raters)","estd":"kappa","v":0.04,"n":"875","k":"3","samp":"full-pool","blind":"open","agg":"single-rater","scale":"0=prefer not, inadequate, incomplete, solid, convincing, compelling, exceptional","field":"life sciences and medicine","wr":"reviewers on manuscript strength-of-support ratings","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":0.01,"ciHigh":0.07,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Across submissions rated by three reviewers, a Fleiss kappa of 0.04 for strength of support indicates essentially no chance-corrected agreement among the reviewers.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,1) mean of k raters, Spearman-Brown stepped up to k=2.56, continuous treatment","estd":"ICC (average)","v":0.35,"n":"875","k":"2.56","samp":"full-pool","blind":"open","agg":"average-of-k","scale":"0=prefer not, inadequate, incomplete, solid, convincing, compelling, exceptional","field":"life sciences and medicine","wr":"reviewers on manuscript strength-of-support ratings","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":0.27,"ciHigh":0.43,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Treating strength-of-support scores as continuous, the reliability of the mean of 2.56 reviewers per manuscript rose to an ICC of 0.35, still a modest reliability for the averaged rating.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC(k) average from multilevel ordinal logistic regression with six covariates, k=2.56","estd":"ICC (average)","v":0.35,"n":"875","k":"2.56","samp":"full-pool","blind":"open","agg":"average-of-k","scale":"0=prefer not, inadequate, incomplete, solid, convincing, compelling, exceptional","field":"life sciences and medicine","wr":"reviewers on manuscript strength-of-support ratings","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":0.27,"ciHigh":0.43,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"In the covariate-adjusted multilevel ordinal model, the reliability of the mean of 2.56 reviewers for strength of support was an ICC of 0.35, essentially unchanged from the model without covariates.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,1)-type mean of k, Spearman-Brown, from ordinal mixed-effects model, k=2.56","estd":"ICC (average)","v":0.37,"n":"875","k":"2.56","samp":"full-pool","blind":"open","agg":"average-of-k","scale":"0=prefer not, inadequate, incomplete, solid, convincing, compelling, exceptional","field":"life sciences and medicine","wr":"reviewers on manuscript strength-of-support ratings","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":0.29,"ciHigh":0.45,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Treating strength-of-support scores as ordinal, the reliability of the mean of 2.56 reviewers per manuscript was an ICC of 0.37, still a modest reliability for the averaged rating.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,1) one-way random, single-rater (Shrout & Fleiss), continuous treatment","estd":"ICC (single/unspec)","v":0.18,"n":"875","k":"2.56","samp":"full-pool","blind":"open","agg":"single-rater","scale":"0=prefer not, inadequate, incomplete, solid, convincing, compelling, exceptional","field":"life sciences and medicine","wr":"reviewers on manuscript strength-of-support ratings","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":0.13,"ciHigh":0.22,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Treating strength-of-support scores as continuous, the single-rater intraclass correlation across a mean of 2.56 reviewers was 0.18, a low reliability for one reviewer’s rating.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC(1) single from multilevel ordinal logistic regression with six covariates","estd":"ICC (single/unspec)","v":0.17,"n":"875","k":"2.56","samp":"full-pool","blind":"open","agg":"single-rater","scale":"0=prefer not, inadequate, incomplete, solid, convincing, compelling, exceptional","field":"life sciences and medicine","wr":"reviewers on manuscript strength-of-support ratings","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":0.12,"ciHigh":0.23,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"In the covariate-adjusted multilevel ordinal model, the single-rater intraclass correlation for strength of support was 0.17, essentially unchanged from the model without covariates.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,1)-type single-rater from mixed-effects multinomial logistic model (Hedeker 2003), ordinal","estd":"ICC (single/unspec)","v":0.18,"n":"875","k":"2.56","samp":"full-pool","blind":"open","agg":"single-rater","scale":"0=prefer not, inadequate, incomplete, solid, convincing, compelling, exceptional","field":"life sciences and medicine","wr":"reviewers on manuscript strength-of-support ratings","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":0.13,"ciHigh":0.23,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Treating strength-of-support scores as ordinal, the single-rater intraclass correlation across a mean of 2.56 reviewers was 0.18, a low reliability for one reviewer’s rating.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"observed percentage agreement (exact agreement, diagonal of first-two-reviewer cross-table)","estd":"percent agreement","v":0.28,"n":"875","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"0=prefer not, inadequate, incomplete, solid, convincing, compelling, exceptional","field":"life sciences and medicine","wr":"reviewers on manuscript strength-of-support ratings","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"First two reviewers of 875 eLife manuscripts rated strength of support on a 7-category scale; they used the identical category in 28% of manuscripts, an uncorrected agreement.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"single Cohen’s kappa (unweighted), first two reviewers","estd":"kappa","v":0.08,"n":"875","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"0=prefer not, inadequate, incomplete, solid, convincing, compelling, exceptional","field":"life sciences and medicine","wr":"reviewers on manuscript strength-of-support ratings","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":0.04,"ciHigh":0.12,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"First two reviewers independently rated strength of support for 875 eLife manuscripts; a chance-corrected Cohen’s kappa of 0.08 indicates a lack of agreement between them.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"weighted Cohen's kappa, Cicchetti-Allison linear weights","estd":"weighted kappa","v":0.12,"n":"875","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"0=prefer not, inadequate, incomplete, solid, convincing, compelling, exceptional","field":"life sciences and medicine","wr":"reviewers on manuscript strength-of-support ratings","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":0.1,"ciHigh":0.19,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"A quadratically weighted Cohen’s kappa of 0.12 for strength of support, giving partial credit for near disagreements, still shows a lack of chance-corrected agreement between the first two reviewers.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"weighted observed agreement, Cicchetti-Allison linear weights (partial credit for near-diagonal cells)","estd":"percent agreement","v":0.8,"n":"875","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"0=prefer not, inadequate, incomplete, solid, convincing, compelling, exceptional","field":"life sciences and medicine","wr":"reviewers on manuscript strength-of-support ratings","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Giving partial credit to near-diagonal disagreements with Cicchetti-Allison weights, the weighted observed agreement for strength of support between the first two reviewers was 0.80.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"single Cohen's kappa","estd":"kappa","v":0.1,"n":"875","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"Compelling versus all other categories","field":"life sciences and medicine","wr":"compelling strength classification","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Two reviewers classified manuscripts as compelling or another strength category. The category-specific Cohen's kappa was 0.10.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"single Cohen's kappa","estd":"kappa","v":0.04,"n":"875","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"Convincing versus all other categories","field":"life sciences and medicine","wr":"convincing strength classification","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Two reviewers classified manuscripts as convincing or another strength category. The category-specific Cohen's kappa was 0.04.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"single Cohen's kappa","estd":"kappa","v":0.14,"n":"875","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"Exceptional versus all other categories","field":"life sciences and medicine","wr":"exceptional strength classification","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Two reviewers classified manuscripts as exceptional or another strength category. The category-specific Cohen's kappa was 0.14.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"single Cohen's kappa","estd":"kappa","v":0.16,"n":"875","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"Inadequate versus all other categories","field":"life sciences and medicine","wr":"inadequate strength classification","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Two reviewers classified manuscripts as inadequate or another strength category. The category-specific Cohen's kappa was 0.16.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"single Cohen's kappa","estd":"kappa","v":0.12,"n":"875","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"Incomplete versus all other categories","field":"life sciences and medicine","wr":"incomplete strength classification","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Two reviewers classified manuscripts as incomplete or another strength category. The category-specific Cohen's kappa was 0.12.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"single Cohen's kappa","estd":"kappa","v":0.05,"n":"875","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"Prefer not to answer versus all other categories","field":"life sciences and medicine","wr":"strength category classification","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Two reviewers classified manuscripts as prefer not to answer or another strength category. The category-specific Cohen's kappa was 0.05.","vf":"unverified"},{"key":"J2WJBPIT","au":"Mutz, Rüdiger","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"single Cohen's kappa","estd":"kappa","v":0.06,"n":"875","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"Solid versus all other categories","field":"life sciences and medicine","wr":"solid strength classification","conf":"high","self":false,"doi":"10.1007/s11192-025-05422-y","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Two reviewers classified manuscripts as solid or another strength category. The category-specific Cohen's kappa was 0.06.","vf":"unverified"},{"key":"NZSRHRPT","au":"Noble, John H.","y":1974,"cx":"Grant","ob":"other","fam":"other","form":"Kendall's U, an index of interjudge agreement for paired comparisons which varies from zero (complete disagreement) to one (perfect agreement)","estd":"other","v":0.45,"n":"6","k":"15","samp":"funded-only","blind":"unclear","agg":"single-rater","scale":"paired comparisons of six projects' overall methodological adequacy","field":"applied social research (rehabilitation)","wr":"judges on research projects' methodological adequacy","conf":"med","self":false,"doi":"10.1126/science.185.4155.916","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":true,"he":false,"ms":"Fifteen recruited judges each made paired comparisons of six ongoing rehabilitation research projects on overall methodological adequacy, 225 comparisons in total. A Kendall's U of 0.45 indicates as much disagreement as agreement among the judges.","vf":"unverified"},{"key":"8IUUWWR8","au":"Otte, Gunnar","y":2019,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC","estd":"ICC (single/unspec)","v":0.5,"n":"157","k":"5","samp":"full-pool","blind":"double","agg":"unspecified","scale":"1 (Ablehnung) bis 4 (vorbehaltlose Annahme), reduced from finer scheme","field":"sociology","wr":"five editors' votes on manuscripts","conf":"high","self":false,"doi":"10.1515/zfsoz-2019-0001","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"restricted-other","pr":false,"he":false,"ms":"Agreement among the five editors' independent votes on the 157 manuscripts, expressed as an intraclass correlation of 0.50 with the model unspecified.","vf":"unverified"},{"key":"8IUUWWR8","au":"Otte, Gunnar","y":2019,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"mittlere paarweise Korrelation (mean pairwise correlation, r)","estd":"correlation","v":0.52,"n":"157","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 (Ablehnung) bis 4 (vorbehaltlose Annahme), reduced from finer scheme","field":"sociology","wr":"five editors' votes on manuscripts","conf":"high","self":false,"doi":"10.1515/zfsoz-2019-0001","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"restricted-other","pr":true,"he":false,"ms":"The five ZfS editors independently voted on the 157 manuscripts before the editorial meetings; the mean pairwise correlation of their votes was 0.52, indicating moderate agreement between editors.","vf":"unverified"},{"key":"8IUUWWR8","au":"Otte, Gunnar","y":2019,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"proportion of manuscripts where the two reviewers differed by two or more of the four levels (disagreement rate)","estd":"percent agreement","v":0.17,"n":"33","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 (Ablehnung) bis 4 (vorbehaltlose Annahme)","field":"sociology","wr":"two reviewers on qualitative manuscripts","conf":"high","self":false,"doi":"10.1515/zfsoz-2019-0001","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Among the 33 qualitative-methods manuscripts, the two reviewers' ratings differed by two or more of the four levels in 17 per cent of cases, a measure of inter-reviewer disagreement.","vf":"unverified"},{"key":"8IUUWWR8","au":"Otte, Gunnar","y":2019,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"proportion of manuscripts where the two reviewers differed by two or more of the four levels (disagreement rate)","estd":"percent agreement","v":0.25,"n":"69","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 (Ablehnung) bis 4 (vorbehaltlose Annahme)","field":"sociology","wr":"two reviewers on quantitative manuscripts","conf":"high","self":false,"doi":"10.1515/zfsoz-2019-0001","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Among the 69 quantitative-methods manuscripts, the two reviewers' ratings differed by two or more of the four levels in 25 per cent of cases, a measure of inter-reviewer disagreement.","vf":"unverified"},{"key":"8IUUWWR8","au":"Otte, Gunnar","y":2019,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC","estd":"ICC (single/unspec)","v":0.24,"n":"157","k":"2","samp":"full-pool","blind":"double","agg":"unspecified","scale":"1 (Ablehnung) bis 4 (vorbehaltlose Annahme)","field":"sociology","wr":"two reviewers on journal manuscripts","conf":"high","self":false,"doi":"10.1515/zfsoz-2019-0001","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The paper restates the agreement between the two referees' ratings of the 157 manuscripts as an intraclass correlation of 0.24, with the ICC model unspecified.","vf":"unverified"},{"key":"8IUUWWR8","au":"Otte, Gunnar","y":2019,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"Korrelation (r)","estd":"correlation","v":0.24,"n":"157","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1 (Ablehnung) bis 4 (vorbehaltlose Annahme)","field":"sociology","wr":"two reviewers on journal manuscripts","conf":"high","self":false,"doi":"10.1515/zfsoz-2019-0001","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":true,"ms":"For 157 manuscripts first submitted to ZfS in 2015-2016, the two independent referee ratings on the four-level publishability scale correlated at 0.24, indicating low agreement between reviewers.","vf":"unverified"},{"key":"8IUUWWR8","au":"Otte, Gunnar","y":2019,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"correlation between the mean of the two referee ratings and the final editorial decision (Zusammenhang, r)","estd":"correlation","v":0.45,"n":"157","k":"","samp":"full-pool","blind":"double","agg":"unspecified","scale":"mean of 1-4 referee ratings vs 1-4 final decision","field":"sociology","wr":"mean referee rating vs final editorial decision","conf":"high","self":false,"doi":"10.1515/zfsoz-2019-0001","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"restricted-other","pr":false,"he":false,"ms":"The mean of the two referee ratings correlated at 0.45 with the editorial board's final decision on each of the 157 manuscripts, showing that reviews guided but did not determine the decision.","vf":"unverified"},{"key":"ABE5N6J3","au":"Oxman, Andrew D","y":1991,"cx":"General","ob":"journal-manuscript","fam":"ICC","form":"ICC when the between judge variation is not included in the denominator","estd":"ICC (single/unspec)","v":0.51,"n":"36","k":"3","samp":"special","blind":"double","agg":"single-rater","scale":"7-point scale response options, four anchors","field":"biomedical","wr":"three research assistants on item 7, between-judge variation excluded","conf":"med","self":false,"doi":"10.1016/0895-4356(91)90205-n","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Recomputing item 7 for the research assistants with between-judge variation excluded from the denominator raised the ICC to 0.51, still below the other two groups.","vf":"unverified"},{"key":"ABE5N6J3","au":"Oxman, Andrew D","y":1991,"cx":"General","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (ICC), Shrout and Fleiss guidelines, ANOVA model with non-random selection of judges","estd":"ICC (single/unspec)","v":0.31,"n":"36","k":"3","samp":"special","blind":"double","agg":"single-rater","scale":"7-point scale response options, four anchors","field":"biomedical","wr":"three research assistants on item 7 (methods for combining studies)","conf":"med","self":false,"doi":"10.1016/0895-4356(91)90205-n","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Three research assistants rated all 36 review articles on whether the methods used to combine studies were reported; their ICC of 0.31 was the poorest agreement observed.","vf":"unverified"},{"key":"ABE5N6J3","au":"Oxman, Andrew D","y":1991,"cx":"General","ob":"journal-manuscript","fam":"ICC","form":"ICC when the between judge variation is not included in the denominator","estd":"ICC (single/unspec)","v":0.48,"n":"36","k":"3","samp":"special","blind":"double","agg":"single-rater","scale":"7-point scale response options, four anchors","field":"biomedical","wr":"three research assistants on item 8, between-judge variation excluded","conf":"med","self":false,"doi":"10.1016/0895-4356(91)90205-n","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"For item 8 (findings combined appropriately), excluding between-judge variation from the denominator gave the research assistants an ICC of 0.48, below the other two groups.","vf":"unverified"},{"key":"ABE5N6J3","au":"Oxman, Andrew D","y":1991,"cx":"General","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (ICC), Shrout and Fleiss guidelines, ANOVA model with non-random selection of judges","estd":"ICC (single/unspec)","v":0.71,"n":"36","k":"9","samp":"special","blind":"double","agg":"single-rater","scale":"1-7, 7 = minimal flaws (exemplary)","field":"biomedical","wr":"nine judges on review articles' overall scientific quality","conf":"med","self":false,"doi":"10.1016/0895-4356(91)90205-n","ciLow":0.59,"ciHigh":0.81,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":false,"ms":"Nine trained judges each rated the overall scientific quality of 36 anonymised published review articles on a 7-point scale; a single-judge ICC of 0.71 describes agreement across all nine judges.","vf":"unverified"},{"key":"ABE5N6J3","au":"Oxman, Andrew D","y":1991,"cx":"General","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (ICC), Shrout and Fleiss guidelines, ANOVA model with non-random selection of judges","estd":"ICC (single/unspec)","v":0.74,"n":"36","k":"3","samp":"special","blind":"double","agg":"single-rater","scale":"1-7, 7 = minimal flaws (exemplary)","field":"biomedical","wr":"three clinicians with research training on overall quality","conf":"med","self":false,"doi":"10.1016/0895-4356(91)90205-n","ciLow":0.51,"ciHigh":0.79,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Three clinicians with research training each rated the overall scientific quality of 36 review articles; their single-judge ICC was 0.74 (confidence interval OCR-uncertain, transcribed as printed).","vf":"unverified"},{"key":"ABE5N6J3","au":"Oxman, Andrew D","y":1991,"cx":"General","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (ICC), Shrout and Fleiss guidelines, ANOVA model with non-random selection of judges","estd":"ICC (single/unspec)","v":0.77,"n":"36","k":"3","samp":"special","blind":"double","agg":"single-rater","scale":"1-7, 7 = minimal flaws (exemplary)","field":"biomedical","wr":"three methodology experts on articles' overall scientific quality","conf":"med","self":false,"doi":"10.1016/0895-4356(91)90205-n","ciLow":0.69,"ciHigh":0.87,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Three research-methodology experts each rated the overall scientific quality of 36 review articles; their single-judge ICC of 0.77 describes agreement within this judge category.","vf":"unverified"},{"key":"ABE5N6J3","au":"Oxman, Andrew D","y":1991,"cx":"General","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (ICC), Shrout and Fleiss guidelines, ANOVA model with non-random selection of judges","estd":"ICC (single/unspec)","v":0.62,"n":"36","k":"3","samp":"special","blind":"double","agg":"single-rater","scale":"1-7, 7 = minimal flaws (exemplary)","field":"biomedical","wr":"three research assistants on articles' overall scientific quality","conf":"med","self":false,"doi":"10.1016/0895-4356(91)90205-n","ciLow":0.38,"ciHigh":0.78,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Three research assistants each rated the overall scientific quality of 36 review articles; their single-judge ICC of 0.62 was acceptable but lower than the other two groups.","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"kappa","form":"Cohen's kappa","estd":"kappa","v":0.11,"n":"72","k":"6","samp":"special","blind":"single","agg":"average-of-k","scale":"quartiles of simulated panel ranks","field":"astronomy","wr":"simulated panels on central-quartile membership","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For two simulated six-member panels, predicted chance-corrected agreement within central quartiles was Cohen's kappa 0.11.","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement within same quartile","estd":"percent agreement","v":0.33,"n":"72","k":"6","samp":"special","blind":"single","agg":"average-of-k","scale":"quartiles of simulated panel ranks","field":"astronomy","wr":"simulated panels ranking applications","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":0.17,"ciHigh":0.49,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Empirically calibrated simulations compared independent six-member panels ranking 72 applications. Predicted agreement within a central quartile was 33%.","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"panel-panel first-quartile correlation","estd":"correlation","v":0.62,"n":"72","k":"6","samp":"special","blind":"single","agg":"average-of-k","scale":"simulated panel ranks","field":"astronomy","wr":"simulated panels ranking applications","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Empirically calibrated simulations compared rankings from two independent six-member panels. The predicted average panel-panel correlation was 0.62.","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement within first quartile","estd":"percent agreement","v":0.33,"n":"72","k":"1","samp":"special","blind":"single","agg":"single-rater","scale":"first quartile of simulated ranks","field":"astronomy","wr":"simulated reviewers on top-quartile membership","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For simulated one-reviewer panels, predicted agreement on first-quartile membership was 33%. Intermediate panel-size strata are reported in Table 8.","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"kappa","form":"Cohen's k","estd":"kappa","v":0.11,"n":"72","k":"1","samp":"special","blind":"single","agg":"single-rater","scale":"first quartile versus other quartiles","field":"astronomy","wr":"simulated reviewers on top-quartile membership","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For simulated one-reviewer panels, chance-corrected agreement on first-quartile membership was Cohen's kappa 0.11.","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement within first quartile","estd":"percent agreement","v":0.86,"n":"72","k":"100","samp":"special","blind":"single","agg":"average-of-k","scale":"first quartile of simulated ranks","field":"astronomy","wr":"simulated panels on top-quartile membership","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For simulated 100-reviewer panels, predicted agreement on first-quartile membership was 86%. Intermediate panel-size strata are reported in Table 8.","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"kappa","form":"Cohen's k","estd":"kappa","v":0.81,"n":"72","k":"100","samp":"special","blind":"single","agg":"average-of-k","scale":"first quartile versus other quartiles","field":"astronomy","wr":"simulated panels on top-quartile membership","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For simulated 100-reviewer panels, chance-corrected agreement on first-quartile membership was Cohen's kappa 0.81.","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"kappa","form":"Cohen's kappa","estd":"kappa","v":0.4,"n":"72","k":"6","samp":"special","blind":"single","agg":"average-of-k","scale":"quartiles of simulated panel ranks","field":"astronomy","wr":"simulated panels on top-quartile membership","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For two simulated six-member panels, predicted chance-corrected agreement on top-quartile membership was Cohen's kappa 0.40.","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement within same quartile","estd":"percent agreement","v":0.55,"n":"72","k":"6","samp":"special","blind":"single","agg":"average-of-k","scale":"quartiles of simulated panel ranks","field":"astronomy","wr":"simulated panels ranking applications","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":0.383,"ciHigh":0.717,"mt":"inter-rater","tgt":"other","rr":"none","pr":true,"he":false,"ms":"Empirically calibrated simulations compared two independent six-member panels ranking 72 applications. Their predicted top-quartile agreement was 55.0% plus or minus 16.7% at the 95% confidence level.","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement: fraction of runs placed in a central quartile by the panel majority relative to the expected count","estd":"percent agreement","v":0.18,"n":"16091","k":"6","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"quartile rank (1-4) derived from 1-5 grade","field":"astronomy","wr":"panel majority consensus on central-quartile proposals","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Fraction of proposal runs for which a majority of the six-member panel agreed on central-quartile placement, equal to 18%, much lower than for the top quartile.","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement: fraction of runs placed in first quartile by the panel majority relative to the expected count","estd":"percent agreement","v":0.39,"n":"16091","k":"6","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"quartile rank (1-4) derived from 1-5 grade","field":"astronomy","wr":"panel majority consensus on top-quartile proposals","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Fraction of proposal runs for which a majority of the six-member panel agreed on top-quartile placement, relative to the number expected per quartile, equal to 39% (about 15% expected by chance).","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement of a referee with the panel majority (>=4 of 6) on first-quartile membership","estd":"percent agreement","v":0.23,"n":"16091","k":"6","samp":"full-pool","blind":"single","agg":"single-rater","scale":"quartile rank (1-4) derived from 1-5 grade","field":"astronomy","wr":"referee agreeing with panel majority on top-quartile proposals","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Fraction of proposal runs a referee ranked in the top quartile that at least four of the six panel members also ranked in the top quartile, averaging 23%.","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson correlation between a referee's grade and the mean of the other referees' grades","estd":"correlation","v":0.35,"n":"","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1.0-5.0 in 0.5 steps, 1 = best","field":"astronomy","wr":"referee vs mean of rest of panel on proposals","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Pearson correlation between each referee's grade and the average grade given by the other panel members for the same telescope proposal run, equal to 0.35 across runs reviewed by five or six referees.","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement on first-quartile (top 25%) membership between two referees","estd":"percent agreement","v":0.34,"n":"16091","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"quartile rank (1-4) derived from 1-5 grade","field":"astronomy","wr":"referees agreeing on top-quartile proposals","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":0.06,"ciHigh":0.62,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Fraction of proposal runs a referee ranked in the top quartile that another referee of the same panel also ranked in the top quartile, averaging 34% (25% expected by chance); Section 10 restates it as 34% plus or minus 28% at the 95% confidence level.","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement on second-quartile membership between two referees","estd":"percent agreement","v":0.26,"n":"16091","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"quartile rank (1-4) derived from 1-5 grade","field":"astronomy","wr":"referees agreeing on second-quartile proposals","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Diagonal element of the referee-referee quartile agreement matrix: fraction of proposal runs two referees both placed in the second quartile, equal to 0.26, only marginally above the 0.25 chance level.","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement on third-quartile membership between two referees","estd":"percent agreement","v":0.27,"n":"16091","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"quartile rank (1-4) derived from 1-5 grade","field":"astronomy","wr":"referees agreeing on third-quartile proposals","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Diagonal element of the referee-referee quartile agreement matrix: fraction of proposal runs two referees both placed in the third quartile, equal to 0.27, only marginally above the 0.25 chance level.","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement on fourth-quartile (bottom 25%) membership between two referees","estd":"percent agreement","v":0.36,"n":"16091","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"quartile rank (1-4) derived from 1-5 grade","field":"astronomy","wr":"referees agreeing on bottom-quartile proposals","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Diagonal element of the referee-referee quartile agreement matrix: fraction of proposal runs two referees both placed in the bottom quartile, equal to 0.36, higher than in the central quartiles.","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson correlation coefficient between two referees' grade sets","estd":"correlation","v":0.22,"n":"16091","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1.0-5.0 in 0.5 steps, 1 = best","field":"astronomy","wr":"referees on telescope proposal grades","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Pairwise Pearson correlation between the grades of two panel referees who independently scored the same telescope proposal runs; the average across 2445 referee pairs was 0.22, indicating weak inter-referee consistency.","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"kappa","form":"Cohen's kappa coefficient (Cohen 1960), chance-corrected first-quartile agreement","estd":"kappa","v":0.12,"n":"16091","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"quartile rank (1-4) derived from 1-5 grade","field":"astronomy","wr":"referees agreeing on top-quartile proposals (chance-corrected)","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Cohen's kappa expressing the chance-corrected first-quartile agreement between two referees, equal to 0.12, indicating only slight agreement beyond chance.","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement on first-decile (top 10%) membership between two bootstrapped 3-member sub-panels","estd":"percent agreement","v":0.22,"n":"16091","k":"3","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"decile rank (1-10) derived from 1-5 grade","field":"astronomy","wr":"two 3-referee sub-panels agreeing on top-decile proposals","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Bootstrapped agreement fraction between two three-referee sub-panels on top-decile placement of proposal runs, equal to 0.22, above the 0.10 chance level but not exceptional.","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement on last-decile (bottom 10%) membership between two bootstrapped 3-member sub-panels","estd":"percent agreement","v":0.3,"n":"16091","k":"3","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"decile rank (1-10) derived from 1-5 grade","field":"astronomy","wr":"two 3-referee sub-panels agreeing on bottom-decile proposals","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Bootstrapped agreement fraction between two three-referee sub-panels on bottom-decile placement of proposal runs, equal to 0.30, above the 0.10 chance level.","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement on first-quartile membership between two bootstrapped 3-member sub-panels","estd":"percent agreement","v":0.43,"n":"16091","k":"3","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"quartile rank (1-4) derived from 1-5 grade","field":"astronomy","wr":"two 3-referee sub-panels agreeing on top-quartile proposals","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Bootstrapped agreement fraction between two independent three-referee sub-panels on top-quartile placement of proposal runs, equal to 0.43, derived by resampling the real Nr=6 grade data.","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement on second-quartile membership between two bootstrapped 3-member sub-panels","estd":"percent agreement","v":0.3,"n":"16091","k":"3","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"quartile rank (1-4) derived from 1-5 grade","field":"astronomy","wr":"two 3-referee sub-panels agreeing on second-quartile proposals","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Bootstrapped agreement fraction between two three-referee sub-panels on second-quartile placement of proposal runs, equal to 0.30, only marginally above the 0.25 chance level.","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement on third-quartile membership between two bootstrapped 3-member sub-panels","estd":"percent agreement","v":0.29,"n":"16091","k":"3","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"quartile rank (1-4) derived from 1-5 grade","field":"astronomy","wr":"two 3-referee sub-panels agreeing on third-quartile proposals","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Bootstrapped agreement fraction between two three-referee sub-panels on third-quartile placement of proposal runs, equal to 0.29, only marginally above the 0.25 chance level.","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement on fourth-quartile membership between two bootstrapped 3-member sub-panels","estd":"percent agreement","v":0.46,"n":"16091","k":"3","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"quartile rank (1-4) derived from 1-5 grade","field":"astronomy","wr":"two 3-referee sub-panels agreeing on bottom-quartile proposals","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Bootstrapped agreement fraction between two three-referee sub-panels on bottom-quartile placement of proposal runs, equal to 0.46, higher than in the central quartiles.","vf":"unverified"},{"key":"I358ZRGB","au":"Patat, Ferdinando","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"kappa","form":"Cohen's k (chance-corrected central-quartile sub-panel agreement)","estd":"kappa","v":0.07,"n":"16091","k":"3","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"quartile rank (1-4) derived from 1-5 grade","field":"astronomy","wr":"two 3-referee sub-panels agreeing on central-quartile proposals (chance-corrected)","conf":"high","self":false,"doi":"10.1088/1538-3873/aac463","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Cohen's kappa for the chance-corrected central-quartile agreement between two bootstrapped three-referee sub-panels, equal to 0.07, indicating almost no agreement beyond chance.","vf":"unverified"},{"key":"IV59GNII","au":"Pier, Elizabeth L","y":2017,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha across the four panels' consensus scores for the same applications","estd":"Krippendorff","v":-0.052,"n":"12","k":"","samp":"special","blind":"double","agg":"panel-consensus","scale":"panel consensus score, 10-90","field":"biomedical (oncology)","wr":"between-panel agreement on final panel consensus scores","conf":"high","self":false,"doi":"10.1093/reseval/rvw025","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Comparing the four panels' final consensus scores for the same applications, between-panel agreement was slightly negative after discussion (Krippendorff's alpha -0.052).","vf":"unverified"},{"key":"IV59GNII","au":"Pier, Elizabeth L","y":2017,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha across the four panels' scores for the same applications","estd":"Krippendorff","v":0.095,"n":"12","k":"","samp":"special","blind":"double","agg":"average-of-k","scale":"panel mean of 1-9 reverse scores","field":"biomedical (oncology)","wr":"between-panel agreement on final reviewer scores","conf":"high","self":false,"doi":"10.1093/reseval/rvw025","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Comparing the four panels' final reviewer scores for the same applications, between-panel agreement fell after discussion (Krippendorff's alpha 0.095).","vf":"unverified"},{"key":"IV59GNII","au":"Pier, Elizabeth L","y":2017,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha across the four panels' scores for the same applications","estd":"Krippendorff","v":0.231,"n":"25","k":"","samp":"special","blind":"double","agg":"average-of-k","scale":"panel mean of 1-9 reverse scores","field":"biomedical (oncology)","wr":"between-panel agreement on preliminary scores","conf":"high","self":false,"doi":"10.1093/reseval/rvw025","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Comparing the four panels' preliminary scores for the same applications, between-panel agreement was low (Krippendorff's alpha 0.231).","vf":"unverified"},{"key":"IV59GNII","au":"Pier, Elizabeth L","y":2017,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha across panels' preliminary scores, discussed applications only","estd":"Krippendorff","v":-0.012,"n":"","k":"","samp":"special","blind":"double","agg":"average-of-k","scale":"panel mean of 1-9 reverse scores","field":"biomedical (oncology)","wr":"between-panel agreement on preliminary scores, discussed only","conf":"high","self":false,"doi":"10.1093/reseval/rvw025","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Restricting to discussed applications, between-panel agreement on preliminary scores was near zero (Krippendorff's alpha -0.012), even lower than for the full set.","vf":"unverified"},{"key":"IV59GNII","au":"Pier, Elizabeth L","y":2017,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha (accommodates >2 raters, allows missing values)","estd":"Krippendorff","v":0.793,"n":"","k":"10","samp":"special","blind":"double","agg":"single-rater","scale":"1-9 reverse, 1 = best","field":"biomedical (oncology)","wr":"within-CSS1 agreement on all panelists' final scores","conf":"high","self":false,"doi":"10.1093/reseval/rvw025","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Within panel CSS1, agreement among all panelists' final scores was the highest of the four panels (Krippendorff's alpha 0.793).","vf":"unverified"},{"key":"IV59GNII","au":"Pier, Elizabeth L","y":2017,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha (accommodates >2 raters, allows missing values)","estd":"Krippendorff","v":0.477,"n":"","k":"12","samp":"special","blind":"double","agg":"single-rater","scale":"1-9 reverse, 1 = best","field":"biomedical (oncology)","wr":"within-CSS2 agreement on all panelists' final scores","conf":"high","self":false,"doi":"10.1093/reseval/rvw025","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Within panel CSS2, agreement among all panelists' final scores was the lowest of the four panels (Krippendorff's alpha 0.477).","vf":"unverified"},{"key":"IV59GNII","au":"Pier, Elizabeth L","y":2017,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha (accommodates >2 raters, allows missing values)","estd":"Krippendorff","v":0.665,"n":"","k":"","samp":"special","blind":"double","agg":"single-rater","scale":"1-9 reverse, 1 = best","field":"biomedical (oncology)","wr":"within-panel agreement on all panelists' final scores","conf":"high","self":false,"doi":"10.1093/reseval/rvw025","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Averaged across panels, agreement among all panelists' final scores (assigned reviewers plus non-reviewing members) after discussion was a mean Krippendorff's alpha of 0.665, the highest of the three score sets.","vf":"unverified"},{"key":"IV59GNII","au":"Pier, Elizabeth L","y":2017,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha (accommodates >2 raters, allows missing values)","estd":"Krippendorff","v":0.717,"n":"","k":"3","samp":"special","blind":"double","agg":"single-rater","scale":"1-9 reverse, 1 = best","field":"biomedical (oncology)","wr":"within-CSS1 agreement on final reviewer scores","conf":"high","self":false,"doi":"10.1093/reseval/rvw025","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Within panel CSS1, agreement among the three assigned reviewers' final scores after discussion was the highest of the four panels (Krippendorff's alpha 0.717).","vf":"unverified"},{"key":"IV59GNII","au":"Pier, Elizabeth L","y":2017,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha (accommodates >2 raters, allows missing values)","estd":"Krippendorff","v":0.334,"n":"","k":"3","samp":"special","blind":"double","agg":"single-rater","scale":"1-9 reverse, 1 = best","field":"biomedical (oncology)","wr":"within-CSS2 agreement on final reviewer scores","conf":"high","self":false,"doi":"10.1093/reseval/rvw025","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Within panel CSS2, agreement among the three assigned reviewers' final scores after discussion was the lowest of the four panels (Krippendorff's alpha 0.334).","vf":"unverified"},{"key":"IV59GNII","au":"Pier, Elizabeth L","y":2017,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha (accommodates >2 raters, allows missing values)","estd":"Krippendorff","v":0.553,"n":"","k":"3","samp":"special","blind":"double","agg":"single-rater","scale":"1-9 reverse, 1 = best","field":"biomedical (oncology)","wr":"within-panel agreement on final reviewer scores","conf":"high","self":false,"doi":"10.1093/reseval/rvw025","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Averaged across panels, agreement among the three assigned reviewers' final scores after discussion rose to a mean Krippendorff's alpha of 0.553, higher than the preliminary agreement.","vf":"unverified"},{"key":"IV59GNII","au":"Pier, Elizabeth L","y":2017,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha (accommodates >2 raters, allows missing values)","estd":"Krippendorff","v":0.028,"n":"","k":"3","samp":"special","blind":"double","agg":"single-rater","scale":"1-9 reverse, 1 = best","field":"biomedical (oncology)","wr":"within-CSS1 agreement on preliminary scores","conf":"high","self":false,"doi":"10.1093/reseval/rvw025","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Within panel CSS1, agreement among the three assigned reviewers' preliminary scores was the lowest of the four panels (Krippendorff's alpha 0.028).","vf":"unverified"},{"key":"IV59GNII","au":"Pier, Elizabeth L","y":2017,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha (accommodates >2 raters, allows missing values)","estd":"Krippendorff","v":0.135,"n":"","k":"3","samp":"special","blind":"double","agg":"single-rater","scale":"1-9 reverse, 1 = best","field":"biomedical (oncology)","wr":"within-CSS2 agreement on preliminary scores","conf":"high","self":false,"doi":"10.1093/reseval/rvw025","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Within panel CSS2, agreement among the three assigned reviewers' preliminary scores was the highest of the four panels but still very low (Krippendorff's alpha 0.135).","vf":"unverified"},{"key":"IV59GNII","au":"Pier, Elizabeth L","y":2017,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha (accommodates >2 raters, allows missing values)","estd":"Krippendorff","v":-0.304,"n":"","k":"3","samp":"special","blind":"double","agg":"single-rater","scale":"1-9 reverse, 1 = best","field":"biomedical (oncology)","wr":"within-CSS2 agreement on preliminary scores, discussed only","conf":"high","self":false,"doi":"10.1093/reseval/rvw025","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Restricting to discussed applications in panel CSS2, within-panel agreement on preliminary scores was the lowest of the four panels (Krippendorff's alpha -0.304).","vf":"unverified"},{"key":"IV59GNII","au":"Pier, Elizabeth L","y":2017,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha (accommodates >2 raters, allows missing values)","estd":"Krippendorff","v":-0.167,"n":"","k":"3","samp":"special","blind":"double","agg":"single-rater","scale":"1-9 reverse, 1 = best","field":"biomedical (oncology)","wr":"within-CSS4 agreement on preliminary scores, discussed only","conf":"high","self":false,"doi":"10.1093/reseval/rvw025","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Restricting to discussed applications in panel CSS4, within-panel agreement on preliminary scores was the least negative of the four panels (Krippendorff's alpha -0.167).","vf":"unverified"},{"key":"IV59GNII","au":"Pier, Elizabeth L","y":2017,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha (accommodates >2 raters, allows missing values)","estd":"Krippendorff","v":-0.224,"n":"","k":"3","samp":"special","blind":"double","agg":"single-rater","scale":"1-9 reverse, 1 = best","field":"biomedical (oncology)","wr":"within-panel agreement on preliminary scores, discussed only","conf":"high","self":false,"doi":"10.1093/reseval/rvw025","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Restricting to discussed (top 50%) applications, within-panel agreement on preliminary scores averaged Krippendorff's alpha -0.224 across the four panels, even lower than for the full set.","vf":"unverified"},{"key":"IV59GNII","au":"Pier, Elizabeth L","y":2017,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha (accommodates >2 raters, allows missing values)","estd":"Krippendorff","v":0.09,"n":"","k":"3","samp":"special","blind":"double","agg":"single-rater","scale":"1-9 reverse, 1 = best","field":"biomedical (oncology)","wr":"within-panel agreement on preliminary scores","conf":"high","self":false,"doi":"10.1093/reseval/rvw025","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Averaged across the four panels, agreement among the three assigned reviewers' preliminary scores within a panel was very low (mean Krippendorff's alpha 0.090).","vf":"unverified"},{"key":"IV59GNII","au":"Pier, Elizabeth L","y":2017,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha (accommodates >2 raters, allows missing values)","estd":"Krippendorff","v":0.084,"n":"25","k":"","samp":"special","blind":"double","agg":"single-rater","scale":"1-9 reverse, 1 = best","field":"biomedical (oncology)","wr":"reviewers' preliminary scores of grant applications","conf":"high","self":false,"doi":"10.1093/reseval/rvw025","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Assigned reviewers independently gave preliminary scores to 25 NIH grant applications before their panels met; a Krippendorff's alpha of 0.084 indicates extremely poor agreement among independent reviewers.","vf":"unverified"},{"key":"IV59GNII","au":"Pier, Elizabeth L","y":2017,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha (accommodates >2 raters, allows missing values)","estd":"Krippendorff","v":-0.088,"n":"","k":"","samp":"special","blind":"double","agg":"single-rater","scale":"1-9 reverse, 1 = best","field":"biomedical (oncology)","wr":"reviewers' preliminary scores of discussed applications","conf":"high","self":false,"doi":"10.1093/reseval/rvw025","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Restricting to the top-scoring 50% of applications that panels discussed, agreement among independent reviewers' preliminary scores was slightly negative (Krippendorff's alpha -0.088), indicating systematic disagreement.","vf":"unverified"},{"key":"IV59GNII","au":"Pier, Elizabeth L","y":2017,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha on binary discussed/triaged variable","estd":"Krippendorff","v":0.2,"n":"25","k":"","samp":"special","blind":"double","agg":"panel-consensus","scale":"discussed = 1 / triaged = 0","field":"biomedical (oncology)","wr":"panels' decisions on which applications to discuss","conf":"high","self":false,"doi":"10.1093/reseval/rvw025","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":true,"he":false,"ms":"Coding each application as discussed or triaged by each of the four panels, agreement between panels on which applications to discuss was low (Krippendorff's alpha 0.200).","vf":"unverified"},{"key":"85QBBTET","au":"Pier, Elizabeth L","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"random-intercept-for-application model; ICC = variance of random intercept / total variance","estd":"ICC (single/unspec)","v":0,"n":"25","k":"","samp":"funded-only","blind":"double","agg":"single-rater","scale":"reverse 9-point, 1 = exceptional, 9 = poor","field":"biomedical (oncology)","wr":"reviewers' preliminary 9-point ratings of grant applications","conf":"high","self":false,"doi":"10.1073/pnas.1714379115","ciLow":0,"ciHigh":0.14,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":true,"he":true,"ms":"Between two and four primary reviewers independently gave preliminary ratings on the reverse 9-point NIH scale to each of 25 NIH R01 oncology applications (83 ratings from 43 reviewers). The single-rater ICC of 0 means ratings of the same application were no more alike than ratings of different applications.","vf":"unverified"},{"key":"85QBBTET","au":"Pier, Elizabeth L","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"random-intercept-for-application model; ICC = variance of random intercept / total variance","estd":"ICC (single/unspec)","v":0,"n":"25","k":"","samp":"funded-only","blind":"double","agg":"single-rater","scale":"number of strengths mentioned in the critique","field":"biomedical (oncology)","wr":"number of strengths noted per application","conf":"high","self":false,"doi":"10.1073/pnas.1714379115","ciLow":0,"ciHigh":0.15,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"The research team counted the strengths each primary reviewer listed in the written critique of each of the 25 applications. The single-rater ICC of 0 indicates reviewers did not agree on how many strengths a given application had.","vf":"unverified"},{"key":"85QBBTET","au":"Pier, Elizabeth L","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"random-intercept-for-application model; ICC = variance of random intercept / total variance","estd":"ICC (single/unspec)","v":0.017,"n":"25","k":"","samp":"funded-only","blind":"double","agg":"single-rater","scale":"number of weaknesses mentioned in the critique","field":"biomedical (oncology)","wr":"number of weaknesses noted per application","conf":"high","self":false,"doi":"10.1073/pnas.1714379115","ciLow":0,"ciHigh":0.18,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"The research team counted the weaknesses each primary reviewer listed per application. The single-rater ICC of 0.017 indicates essentially no agreement between reviewers on how many weaknesses a given application had.","vf":"unverified"},{"key":"85QBBTET","au":"Pier, Elizabeth L","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha","estd":"Krippendorff","v":0.024,"n":"25","k":"","samp":"funded-only","blind":"double","agg":"single-rater","scale":"reverse 9-point, 1 = exceptional, 9 = poor","field":"biomedical (oncology)","wr":"reviewers' preliminary 9-point ratings of grant applications","conf":"high","self":false,"doi":"10.1073/pnas.1714379115","ciLow":-0.047,"ciHigh":0.093,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Treating the 43 reviewers as raters and the 25 applications as targets, Krippendorff's alpha for the preliminary ratings was 0.024, far below the 0.7 acceptability threshold, indicating no agreement between reviewers. Confidence intervals came from 1,000 bootstrapped samples.","vf":"unverified"},{"key":"85QBBTET","au":"Pier, Elizabeth L","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha","estd":"Krippendorff","v":-0.011,"n":"25","k":"","samp":"funded-only","blind":"double","agg":"single-rater","scale":"number of strengths mentioned in the critique","field":"biomedical (oncology)","wr":"number of strengths noted per application","conf":"high","self":false,"doi":"10.1073/pnas.1714379115","ciLow":-0.094,"ciHigh":0.079,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Krippendorff's alpha for the number of strengths reviewers listed per application was -0.011, indicating no agreement between reviewers on strength counts.","vf":"unverified"},{"key":"85QBBTET","au":"Pier, Elizabeth L","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorff's alpha","estd":"Krippendorff","v":0.004,"n":"25","k":"","samp":"funded-only","blind":"double","agg":"single-rater","scale":"number of weaknesses mentioned in the critique","field":"biomedical (oncology)","wr":"number of weaknesses noted per application","conf":"high","self":false,"doi":"10.1073/pnas.1714379115","ciLow":-0.063,"ciHigh":0.072,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Krippendorff's alpha for the number of weaknesses reviewers listed per application was 0.004, indicating no agreement between reviewers on weakness counts.","vf":"unverified"},{"key":"85QBBTET","au":"Pier, Elizabeth L","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"overall similarity score (mean difference of average absolute rating differences), one-sample t test vs 0","estd":"other","v":0.01,"n":"25","k":"","samp":"funded-only","blind":"double","agg":"single-rater","scale":"reverse 9-point, 1 = exceptional, 9 = poor","field":"biomedical (oncology)","wr":"reviewers' preliminary 9-point ratings of grant applications","conf":"high","self":false,"doi":"10.1073/pnas.1714379115","ciLow":-0.21,"ciHigh":0.22,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"For each application the authors subtracted the average absolute difference among its own ratings from the average absolute difference to other applications' ratings. The mean similarity score of 0.01 was not reliably above zero, so ratings of the same application were no more similar than ratings of different applications.","vf":"unverified"},{"key":"85QBBTET","au":"Pier, Elizabeth L","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"overall similarity score (mean difference of average absolute differences), one-sample t test vs 0","estd":"other","v":0.01,"n":"25","k":"","samp":"funded-only","blind":"double","agg":"single-rater","scale":"number of strengths mentioned in the critique","field":"biomedical (oncology)","wr":"number of strengths noted per application","conf":"high","self":false,"doi":"10.1073/pnas.1714379115","ciLow":-0.23,"ciHigh":0.25,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"The same similarity analysis applied to the strength counts gave a mean of 0.01, not reliably different from zero, showing no greater within-application similarity in strengths.","vf":"unverified"},{"key":"85QBBTET","au":"Pier, Elizabeth L","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"overall similarity score (mean difference of average absolute differences), one-sample t test vs 0","estd":"other","v":0.01,"n":"25","k":"","samp":"funded-only","blind":"double","agg":"single-rater","scale":"number of weaknesses mentioned in the critique","field":"biomedical (oncology)","wr":"number of weaknesses noted per application","conf":"high","self":false,"doi":"10.1073/pnas.1714379115","ciLow":-0.21,"ciHigh":0.22,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"The same similarity analysis applied to the weakness counts gave a mean of 0.01, not reliably different from zero, showing no greater within-application similarity in weaknesses.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"median average deviation (AD) index, points on 0-100 scale","estd":"AD index","v":7.3,"n":"759","k":"3","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-100 final score, weighted composite of criteria","field":"multi-field","wr":"raters on Industry-Academia partnership proposals","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 759 Industry-Academia Partnerships and Pathways proposals, the median average deviation index among the three independent raters was 7.3 points on the 0 to 100 scale, somewhat more disagreement than the overall figure.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"median average deviation (AD) index, points on 0-100 scale","estd":"AD index","v":5.2,"n":"20593","k":"3","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-100 final score, weighted composite of criteria","field":"multi-field","wr":"raters on Intra-European Fellowship proposals","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 20,593 Intra-European Fellowship applications, the median average deviation index among the three independent raters was 5.2 points on the 0 to 100 scale.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"median average deviation (AD) index, points on 0-100 scale","estd":"AD index","v":6.3,"n":"3545","k":"3","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-100 final score, weighted composite of criteria","field":"multi-field","wr":"raters on Initial Training Network proposals","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 3,545 Initial Training Networks proposals, the median average deviation index among the three independent raters was 6.3 points on the 0 to 100 scale.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"median average deviation (AD) index, points on 0-100 scale","estd":"AD index","v":8.8,"n":"6","k":"3","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-100 final score, weighted composite of criteria","field":"multi-field","wr":"raters on IAPP Mathematics panel proposals","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The largest of the 25 stratified median average deviation values in Table 2 was 8.8 points, in the Mathematics panel of the Industry-Academia calls, based on only six proposals.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"median average deviation (AD) index, points on 0-100 scale","estd":"AD index","v":4.7,"n":"2469","k":"3","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-100 final score, weighted composite of criteria","field":"multi-field","wr":"raters on IEF Physics panel proposals","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Table 2 breaks the average deviation index down by action and panel across 25 strata. The smallest median value, 4.7 points, was in the Physics panel of the Intra-European Fellowship calls, indicating the closest agreement of any stratum.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"median average deviation (AD) index, points on 0-100 scale","estd":"AD index","v":5.4,"n":"24897","k":"3","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-100 final score, weighted composite of criteria","field":"multi-field","wr":"raters on grant proposal final scores","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Three external raters independently scored each of 24,897 Marie Curie proposals on a 0 to 100 scale. The median average deviation index of 5.4 points is the paper's main measure of inter-rater agreement and indicates small typical disagreement between raters.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"median average deviation (AD) index after transforming scores to a 5 point scale","estd":"AD index","v":0.27,"n":"24897","k":"3","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"IER scores transformed to a 5 point scale","field":"multi-field","wr":"raters on rescaled grant proposal scores","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"To compare against published significance criteria, the authors rescaled the 0 to 100 individual scores onto a five point scale and recomputed the average deviation index. The median of 0.27 was below all published null distribution cut-offs, indicating statistically significant agreement.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"one-way random intraclass correlation coefficient","estd":"ICC (single/unspec)","v":0.67,"n":"24897","k":"3","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0-100 final score, weighted composite of criteria","field":"multi-field","wr":"raters on grant proposal final scores","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":0.66,"ciHigh":0.68,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"A one-way random intraclass correlation was computed on the independent 0 to 100 scores of the three raters for all 24,897 proposals, because different raters rated different proposals. The value of 0.67 was described as indicating good inter-rater agreement. Panel level intraclass correlations sit in a supplementary table that is not part of this text.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"proportion of proposals where AVIER and CR DIFFERED by 10 or more points on the 0-100 scale","estd":"other","v":0.015,"n":"24897","k":"3","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0-100 final score, weighted composite of criteria","field":"multi-field","wr":"average individual scores versus consensus score","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"other","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"In 368 of the 24,897 proposals, or 1.5 per cent, the consensus score moved 10 or more points away from the average of the three independent scores. Positive and negative moves were about equally common.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"proportion of proposals where AVIER and CR agreed within less than 2 points on the 0-100 scale","estd":"percent agreement","v":0.614,"n":"24897","k":"3","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0-100 final score, weighted composite of criteria","field":"multi-field","wr":"average individual scores versus consensus score","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"other","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"For 61.4 per cent of the 24,897 proposals, the averaged independent score and the later consensus score differed by less than two points on the 0 to 100 scale.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"proportion of proposals agreeing within an AD index below 10 points on the 0-100 scale","estd":"percent agreement","v":0.843,"n":"24897","k":"3","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-100 final score, weighted composite of criteria","field":"multi-field","wr":"raters on grant proposal final scores","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"In 20,988 of the 24,897 proposals, the three raters' independent 0 to 100 scores deviated from their own mean by less than 10 points, an agreement rate of 84.3 per cent.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"proportion of proposals with DISAGREEMENT, all three raters differing 10 or more points from each other","estd":"other","v":0.083,"n":"24897","k":"3","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0-100 final score, weighted composite of criteria","field":"multi-field","wr":"raters on grant proposal final scores","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"In 2,075 of the 24,897 proposals, or 8.3 per cent, every pair of the three raters differed by 10 or more points on the 0 to 100 scale. This is a disagreement rate, so higher values mean worse agreement.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"proportion of proposals with DISAGREEMENT, one rater differing 10 or more points from two raters agreeing within 5 points","estd":"other","v":0.057,"n":"24897","k":"3","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0-100 final score, weighted composite of criteria","field":"multi-field","wr":"raters on grant proposal final scores","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"In 1,424 of the 24,897 proposals, or 5.7 per cent, one rater's independent score sat 10 or more points away from two raters who agreed within 5 points. This is a disagreement rate, so higher values mean worse agreement.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation coefficient","estd":"correlation","v":0.917,"n":"2075","k":"3","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0-100 final score, weighted composite of criteria","field":"multi-field","wr":"average individual scores versus consensus score","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"other","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"In the 2,075 proposals where every pair of raters differed by ten or more points, the averaged independent score still correlated 0.917 with the later consensus score.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation coefficient","estd":"correlation","v":0.97,"n":"759","k":"3","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0-100 final score, weighted composite of criteria","field":"multi-field","wr":"average individual scores versus consensus score","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"other","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"For the 759 Industry-Academia proposals, the average of the three independent scores correlated 0.970 with the consensus score agreed by the same raters.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation coefficient","estd":"correlation","v":0.958,"n":"20593","k":"3","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0-100 final score, weighted composite of criteria","field":"multi-field","wr":"average individual scores versus consensus score","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"other","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"For the 20,593 Intra-European Fellowship applications, the average of the three independent scores correlated 0.958 with the consensus score agreed by the same raters.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation coefficient","estd":"correlation","v":0.946,"n":"3545","k":"3","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0-100 final score, weighted composite of criteria","field":"multi-field","wr":"average individual scores versus consensus score","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"other","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"For the 3,545 Initial Training Networks proposals, the average of the three independent scores correlated 0.946 with the consensus score agreed by the same raters.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation coefficient","estd":"correlation","v":0.994,"n":"6","k":"3","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0-100 final score, weighted composite of criteria","field":"multi-field","wr":"average individual scores versus consensus score","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"other","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"The highest of the 25 stratified correlations between averaged individual scores and consensus scores was 0.994, in the Mathematics panel of the Industry-Academia calls, based on only six proposals.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation coefficient","estd":"correlation","v":0.903,"n":"60","k":"3","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0-100 final score, weighted composite of criteria","field":"multi-field","wr":"average individual scores versus consensus score","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"other","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Table 2 reports this correlation for 25 action and panel strata. The lowest value, 0.903, was in the Mathematics panel of the Initial Training Networks calls, based on 60 proposals.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation coefficient","estd":"correlation","v":0.913,"n":"1424","k":"3","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0-100 final score, weighted composite of criteria","field":"multi-field","wr":"average individual scores versus consensus score","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"other","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"In the 1,424 proposals where one rater differed by ten or more points from two raters who agreed within five points, the averaged independent score still correlated 0.913 with the later consensus score.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation coefficient","estd":"correlation","v":0.957,"n":"24897","k":"3","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0-100 final score, weighted composite of criteria","field":"multi-field","wr":"average individual scores versus consensus score","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"other","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"For all 24,897 proposals the average of the three independent remote scores was correlated with the consensus score the same three raters agreed at the meeting. The Pearson correlation of 0.957 shows the consensus phase reproduced the averaged independent judgement almost exactly.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation of two raters' scores on the same criterion","estd":"correlation","v":0.341,"n":"24897","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 (fail) to 5 (excellent) per criterion","field":"multi-field","wr":"two raters on the impact criterion","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The impact scores of the first and second rater of each proposal correlated 0.341 across 24,897 proposals.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation of two raters' scores on the same criterion","estd":"correlation","v":0.341,"n":"24897","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 (fail) to 5 (excellent) per criterion","field":"multi-field","wr":"two raters on the impact criterion","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The impact scores of the first and third rater of each proposal correlated 0.341 across 24,897 proposals.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation of two raters' scores on the same criterion","estd":"correlation","v":0.342,"n":"24897","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 (fail) to 5 (excellent) per criterion","field":"multi-field","wr":"two raters on the impact criterion","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The impact scores of the second and third rater of each proposal correlated 0.342 across 24,897 proposals.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation of two raters' scores on the same criterion","estd":"correlation","v":0.36,"n":"24897","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 (fail) to 5 (excellent) per criterion","field":"multi-field","wr":"two raters on the implementation criterion","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The implementation scores of the first and second rater of each proposal correlated 0.360 across 24,897 proposals.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation of two raters' scores on the same criterion","estd":"correlation","v":0.367,"n":"24897","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 (fail) to 5 (excellent) per criterion","field":"multi-field","wr":"two raters on the implementation criterion","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The implementation scores of the first and third rater of each proposal correlated 0.367 across 24,897 proposals.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation of two raters' scores on the same criterion","estd":"correlation","v":0.367,"n":"24897","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 (fail) to 5 (excellent) per criterion","field":"multi-field","wr":"two raters on the implementation criterion","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The implementation scores of the second and third rater of each proposal correlated 0.367 across 24,897 proposals.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation of two raters' scores on the same criterion","estd":"correlation","v":0.293,"n":"20593","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 (fail) to 5 (excellent) per criterion","field":"multi-field","wr":"two raters on the researcher criterion","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The researcher criterion was scored only for the 20,593 fellowship applications. The first and second rater of each proposal correlated 0.293 on it.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation of two raters' scores on the same criterion","estd":"correlation","v":0.306,"n":"20593","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 (fail) to 5 (excellent) per criterion","field":"multi-field","wr":"two raters on the researcher criterion","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"On the researcher criterion, scored only for the 20,593 fellowship applications, the first and third rater of each proposal correlated 0.306.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation of two raters' scores on the same criterion","estd":"correlation","v":0.294,"n":"20593","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 (fail) to 5 (excellent) per criterion","field":"multi-field","wr":"two raters on the researcher criterion","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"On the researcher criterion, scored only for the 20,593 fellowship applications, the second and third rater of each proposal correlated 0.294.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation of two raters' scores on the same criterion","estd":"correlation","v":0.291,"n":"24897","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 (fail) to 5 (excellent) per criterion","field":"multi-field","wr":"two raters on the S&T quality criterion","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two of the three raters of the same proposal had their Science and Technology quality scores correlated across 24,897 proposals. The value of 0.291 is the low cross-rater agreement the authors highlight for individual criteria.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation of two raters' scores on the same criterion","estd":"correlation","v":0.296,"n":"24897","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 (fail) to 5 (excellent) per criterion","field":"multi-field","wr":"two raters on the S&T quality criterion","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The Science and Technology quality scores of the first and third rater of each proposal correlated 0.296 across 24,897 proposals.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation of two raters' scores on the same criterion","estd":"correlation","v":0.295,"n":"24897","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 (fail) to 5 (excellent) per criterion","field":"multi-field","wr":"two raters on the S&T quality criterion","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The Science and Technology quality scores of the second and third rater of each proposal correlated 0.295 across 24,897 proposals.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation of two raters' scores on the same criterion","estd":"correlation","v":0.361,"n":"24897","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 (fail) to 5 (excellent) per criterion","field":"multi-field","wr":"two raters on the training criterion","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The training or transfer of knowledge scores of the first and second rater of each proposal correlated 0.361 across 24,897 proposals.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation of two raters' scores on the same criterion","estd":"correlation","v":0.357,"n":"24897","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 (fail) to 5 (excellent) per criterion","field":"multi-field","wr":"two raters on the training criterion","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The training or transfer of knowledge scores of the first and third rater of each proposal correlated 0.357 across 24,897 proposals.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation of two raters' scores on the same criterion","estd":"correlation","v":0.369,"n":"24897","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 (fail) to 5 (excellent) per criterion","field":"multi-field","wr":"two raters on the training criterion","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The training or transfer of knowledge scores of the second and third rater of each proposal correlated 0.369 across 24,897 proposals, the highest cross-rater value for any single criterion.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation of one rater's scores on two different criteria","estd":"correlation","v":0.74,"n":"24897","k":"1","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0 (fail) to 5 (excellent) per criterion","field":"multi-field","wr":"one rater across two evaluation criteria","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"The strongest within-rater correlation between two different criteria was 0.740, reached twice for the first rater. These high within-rater values against low cross-rater values led the authors to conclude that raters scored proposals holistically.","vf":"unverified"},{"key":"YXJVI9WJ","au":"Pina, David G","y":2015,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Pearson's correlation of one rater's scores on two different criteria","estd":"correlation","v":0.564,"n":"20593","k":"1","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0 (fail) to 5 (excellent) per criterion","field":"multi-field","wr":"one rater across two evaluation criteria","conf":"med","self":false,"doi":"10.1371/journal.pone.0130753","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Table 4 also reports how strongly a single rater's scores on different criteria hang together, which the authors read as internal consistency and a holistic scoring style. The weakest of these 30 within-rater correlations was 0.564.","vf":"unverified"},{"key":"2G4INPDL","au":"Pina, David G","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"proportion of proposals with mean AD index 10 points or less on 0-100 scale (agreement within 10 points)","estd":"percent agreement","v":0.787,"n":"75624","k":"3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"0-100 total from 0-5 per criterion, one decimal resolution","field":"multi-field","wr":"reviewers' agreement (AD<=10) on all proposals","conf":"high","self":false,"doi":"10.7554/eLife.59338","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Across all 75,624 proposals, 78.7% had a mean average deviation (AD) index of 10 points or less, which the authors interpret as a high level of agreement between reviewers.","vf":"unverified"},{"key":"2G4INPDL","au":"Pina, David G","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"average deviation (AD) index: sum of absolute pairwise reviewer-score differences divided by number of reviewers (Burke et al. 1999)","estd":"AD index","v":12.86,"n":"3097","k":"3","samp":"special","blind":"single","agg":"single-rater","scale":"0-100 total from 0-5 per criterion, one decimal resolution","field":"multi-field","wr":"reviewers' scores on high-discrepancy proposals","conf":"high","self":false,"doi":"10.7554/eLife.59338","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 3,097 Horizon 2020 proposals where the consensus score differed from the average individual score by more than 10 points, the mean average deviation (AD) index was 12.86, showing markedly greater disagreement among reviewers in this researcher-defined subset.","vf":"unverified"},{"key":"2G4INPDL","au":"Pina, David G","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"proportion of H2020 proposals where all pairwise IER-score differences were 10 points or less (agreement within 10 points on 0-100 scale)","estd":"percent agreement","v":0.253,"n":"50727","k":"3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"0-100 total from 0-5 per criterion, one decimal resolution","field":"multi-field","wr":"reviewers' agreement level on H2020 proposals","conf":"high","self":false,"doi":"10.7554/eLife.59338","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Across the 50,727 Horizon 2020 proposals, 25.3% fell in the 'full agreement' group where every pair of reviewers scored within 10 points of each other; 12.0% were in the 'no agreement' group and 62.7% in between.","vf":"unverified"},{"key":"2G4INPDL","au":"Pina, David G","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"average deviation (AD) index: sum of absolute pairwise reviewer-score differences divided by number of reviewers (Burke et al. 1999)","estd":"AD index","v":7.38,"n":"50727","k":"3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"0-100 total from 0-5 per criterion, one decimal resolution","field":"multi-field","wr":"reviewers' independent scores on H2020 proposals","conf":"high","self":false,"doi":"10.7554/eLife.59338","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 50,727 Horizon 2020 proposals, the mean average deviation (AD) index was 7.38 on a 0 to 100 scale (SD 4.74), very similar to the overall figure and again indicating generally close reviewer scores.","vf":"unverified"},{"key":"2G4INPDL","au":"Pina, David G","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"average deviation (AD) index: sum of absolute pairwise reviewer-score differences divided by number of reviewers (Burke et al. 1999)","estd":"AD index","v":9.4,"n":"283","k":"3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"0-100 total from 0-5 per criterion, one decimal resolution","field":"multi-field","wr":"reviewers' scores on economics/social-sci RISE proposals","conf":"high","self":false,"doi":"10.7554/eLife.59338","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Among the 283 economics and social-sciences RISE proposals from 2014 to 2018, the mean average deviation (AD) index was 9.4 (SD 5.5), the highest value (worst agreement) of the many stratified cells in Table 3.","vf":"unverified"},{"key":"2G4INPDL","au":"Pina, David G","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"average deviation (AD) index: sum of absolute pairwise reviewer-score differences divided by number of reviewers (Burke et al. 1999)","estd":"AD index","v":5.4,"n":"2469","k":"3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"0-100 total from 0-5 per criterion, one decimal resolution","field":"multi-field","wr":"reviewers' scores on physics fellowship proposals","conf":"high","self":false,"doi":"10.7554/eLife.59338","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Among the 2,469 physics Individual Fellowship proposals from 2007 to 2013, the mean average deviation (AD) index was 5.4 (SD 3.5), the lowest value (best agreement) of the many panel-by-action-by-period cells in the heavily stratified Table 3; per the 15-per-family cap only the overall, minimum and maximum cells are coded.","vf":"unverified"},{"key":"2G4INPDL","au":"Pina, David G","y":2021,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"average deviation (AD) index: sum of absolute pairwise reviewer-score differences divided by number of reviewers (Burke et al. 1999)","estd":"AD index","v":7.02,"n":"75624","k":"3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"0-100 total from 0-5 per criterion, one decimal resolution","field":"multi-field","wr":"reviewers' independent scores on grant/fellowship proposals","conf":"high","self":false,"doi":"10.7554/eLife.59338","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Around 75,624 Marie Curie proposals were each scored independently by (typically) three reviewers, and the average deviation (AD) index summarised how far their scores differed. The mean AD index was 7.02 on a 0 to 100 scale (SD 4.56), indicating reviewers' scores were on average close together.","vf":"unverified"},{"key":"IX2FS8HU","au":"Pleskac, Timothy Joseph","y":2025,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"Pearson correlation of standardised ratings (Bayesian 95% credible interval)","estd":"correlation","v":0.45,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-9, 1 = Poor, 9 = Excellent (standardised within reviewer)","field":"judgment and decision making (multi-field)","wr":"reviewers on weaker abstracts (bottom half)","conf":"high","self":false,"doi":"10.31234/osf.io/q2tkw_v2","ciLow":0.35,"ciHigh":0.54,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For the weaker half of submissions (median split on the single-blind rating), the averaged single- and double-blind ratings agreed more strongly, correlating at 0.45.","vf":"unverified"},{"key":"IX2FS8HU","au":"Pleskac, Timothy Joseph","y":2025,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"Pearson correlation of standardised ratings (Bayesian 95% credible interval)","estd":"correlation","v":0.54,"n":"530","k":"","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-9, 1 = Poor, 9 = Excellent (standardised within reviewer)","field":"judgment and decision making (multi-field)","wr":"reviewers on conference abstract scores, both systems","conf":"high","self":false,"doi":"10.31234/osf.io/q2tkw_v2","ciLow":0.48,"ciHigh":0.61,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Across 530 conference abstracts, each reviewed by at least three single-blind and three double-blind reviewers, the averaged single-blind and double-blind ratings correlated at 0.54, showing moderate agreement between the two systems.","vf":"unverified"},{"key":"IX2FS8HU","au":"Pleskac, Timothy Joseph","y":2025,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"Pearson correlation of standardised ratings (Bayesian 95% credible interval)","estd":"correlation","v":0.19,"n":"","k":"","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-9, 1 = Poor, 9 = Excellent (standardised within reviewer)","field":"judgment and decision making (multi-field)","wr":"reviewers on stronger abstracts (top half)","conf":"high","self":false,"doi":"10.31234/osf.io/q2tkw_v2","ciLow":0.06,"ciHigh":0.3,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"For the stronger half of submissions, agreement between the averaged single- and double-blind ratings was low, correlating at only 0.19.","vf":"unverified"},{"key":"IX2FS8HU","au":"Pleskac, Timothy Joseph","y":2025,"cx":"Journal","ob":"conference-abstract","fam":"percent-agreement","form":"percent overlap: submissions ranked in bottom 108 by both single- and double-blind average ratings (exact selection overlap)","estd":"percent agreement","v":0.54,"n":"108","k":"","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"in bottom 108 selection vs not","field":"judgment and decision making (multi-field)","wr":"bottom-108 selection overlap between systems","conf":"high","self":false,"doi":"10.31234/osf.io/q2tkw_v2","ciLow":0.44,"ciHigh":0.63,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Of the bottom 108 submissions, 58 (54%) overlapped between single- and double-blind review, showing more agreement on identifying weaker submissions than stronger ones.","vf":"unverified"},{"key":"IX2FS8HU","au":"Pleskac, Timothy Joseph","y":2025,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"Pearson correlation of two random individual reviews per submission, averaged over 100 iterations (Bayesian 95% credible interval)","estd":"correlation","v":0.28,"n":"530","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1-9, 1 = Poor, 9 = Excellent (standardised within reviewer)","field":"judgment and decision making (multi-field)","wr":"individual double-blind reviewers on abstract scores","conf":"high","self":false,"doi":"10.31234/osf.io/q2tkw_v2","ciLow":0.2,"ciHigh":0.35,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The same procedure for individual double-blind reviews gave an average correlation of 0.28, again indicating low individual-reviewer agreement.","vf":"unverified"},{"key":"IX2FS8HU","au":"Pleskac, Timothy Joseph","y":2025,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"Pearson correlation of two random individual reviews per submission, averaged over 100 iterations (Bayesian 95% credible interval)","estd":"correlation","v":0.34,"n":"530","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1-9, 1 = Poor, 9 = Excellent (standardised within reviewer)","field":"judgment and decision making (multi-field)","wr":"individual single-blind reviewers on abstract scores","conf":"high","self":false,"doi":"10.31234/osf.io/q2tkw_v2","ciLow":0.26,"ciHigh":0.41,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Correlating two randomly chosen individual single-blind reviews per submission across 530 submissions gave an average reliability of 0.34, indicating low individual-reviewer agreement.","vf":"unverified"},{"key":"IX2FS8HU","au":"Pleskac, Timothy Joseph","y":2025,"cx":"Journal","ob":"conference-abstract","fam":"percent-agreement","form":"exact overlap in simulated top-108 selections","estd":"percent agreement","v":0.46,"n":"","k":"","samp":"special","blind":"unclear","agg":"unspecified","scale":"simulated normally distributed values","field":"judgment and decision making (multi-field)","wr":"simulated top-selection overlap, double-blind reliability","conf":"high","self":false,"doi":"10.31234/osf.io/q2tkw_v2","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"The same simulation using the observed double-blind reliability of 0.57 showed that on average 46 per cent of each measure's top 108 overlapped.","vf":"unverified"},{"key":"IX2FS8HU","au":"Pleskac, Timothy Joseph","y":2025,"cx":"Journal","ob":"conference-abstract","fam":"percent-agreement","form":"exact overlap in simulated top-108 selections","estd":"percent agreement","v":0.49,"n":"","k":"","samp":"special","blind":"unclear","agg":"unspecified","scale":"simulated normally distributed values","field":"judgment and decision making (multi-field)","wr":"simulated top-selection overlap, single-blind reliability","conf":"high","self":false,"doi":"10.31234/osf.io/q2tkw_v2","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"The authors simulated 530 pairs of values correlated at the observed single-blind reliability of 0.63. On average, 49 per cent of each measure's top 108 overlapped, illustrating how the observed reliability constrains selection consistency.","vf":"unverified"},{"key":"IX2FS8HU","au":"Pleskac, Timothy Joseph","y":2025,"cx":"Journal","ob":"conference-abstract","fam":"percent-agreement","form":"percent overlap: submissions ranked in top 108 by both single- and double-blind average ratings (exact selection overlap)","estd":"percent agreement","v":0.4,"n":"108","k":"","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"in top 108 selection vs not","field":"judgment and decision making (multi-field)","wr":"top-108 talk selection overlap between systems","conf":"high","self":false,"doi":"10.31234/osf.io/q2tkw_v2","ciLow":0.31,"ciHigh":0.49,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Of the top 108 submissions by average rating, only 43 (40%) would have been selected for talks under both single- and double-blind review, showing weak agreement on the strongest submissions.","vf":"unverified"},{"key":"IX2FS8HU","au":"Pleskac, Timothy Joseph","y":2025,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"Pearson correlation of standardised ratings; observed within-condition agreement","estd":"correlation","v":0.57,"n":"53","k":"6","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-9, 1 = Poor, 9 = Excellent (standardised within reviewer)","field":"judgment and decision making (multi-field)","wr":"double-blind reviewers on abstract scores","conf":"high","self":false,"doi":"10.31234/osf.io/q2tkw_v2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the same 53 submissions given six double-blind reviews, the split-half correlation of the two three-reviewer averages was 0.57, similar to single-blind.","vf":"unverified"},{"key":"IX2FS8HU","au":"Pleskac, Timothy Joseph","y":2025,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"Pearson correlation of standardised ratings; observed within-condition agreement","estd":"correlation","v":0.63,"n":"53","k":"6","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-9, 1 = Poor, 9 = Excellent (standardised within reviewer)","field":"judgment and decision making (multi-field)","wr":"single-blind reviewers on abstract scores","conf":"high","self":false,"doi":"10.31234/osf.io/q2tkw_v2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For 53 submissions each given six single-blind reviews, splitting them into two sets of three and correlating the two three-reviewer averages gave a within-system reliability of 0.63.","vf":"unverified"},{"key":"EK4WFSX9","au":"Pranić, Shelly Melissa","y":2020,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"absolute agreement: exact matching of all reviewers' recommendations","estd":"percent agreement","v":0.59,"n":"395","k":"2","samp":"special","blind":"single","agg":"unspecified","scale":"accepted, revise, rejected","field":"multi-field","wr":"matching reviewer recommendations on manuscripts","conf":"med","self":false,"doi":"10.1002/leap.1344","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Among 395 two-review manuscripts, 59% had both reviewers giving the same recommendation (accepted, revise, rejected), indicating modest exact agreement between reviewers.","vf":"unverified"},{"key":"EK4WFSX9","au":"Pranić, Shelly Melissa","y":2020,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"absolute agreement: exact matching of all reviewers' recommendations","estd":"percent agreement","v":0.08,"n":"53","k":"3","samp":"special","blind":"single","agg":"unspecified","scale":"accepted, revise, rejected","field":"multi-field","wr":"matching reviewer recommendations on manuscripts","conf":"med","self":false,"doi":"10.1002/leap.1344","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Among 53 three-review manuscripts, only 8% had all reviewers giving the same recommendation, indicating very low exact agreement between reviewers.","vf":"unverified"},{"key":"EK4WFSX9","au":"Pranić, Shelly Melissa","y":2020,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"absolute agreement: exact matching of all reviewers' recommendations","estd":"percent agreement","v":0,"n":"9","k":"4","samp":"special","blind":"single","agg":"unspecified","scale":"accepted, revise, rejected","field":"multi-field","wr":"matching reviewer recommendations on manuscripts","conf":"med","self":false,"doi":"10.1002/leap.1344","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Among 9 four-review manuscripts, none had all reviewers giving the same recommendation (0% exact agreement).","vf":"unverified"},{"key":"EK4WFSX9","au":"Pranić, Shelly Melissa","y":2020,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"absolute agreement: exact matching of all reviewers' recommendations","estd":"percent agreement","v":0,"n":"1","k":"6","samp":"special","blind":"single","agg":"unspecified","scale":"accepted, revise, rejected","field":"multi-field","wr":"matching reviewer recommendations on manuscripts","conf":"med","self":false,"doi":"10.1002/leap.1344","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"The single manuscript with six reviews had 0% exact agreement among its reviewers' recommendations; no ICC was computed for this stratum.","vf":"unverified"},{"key":"EK4WFSX9","au":"Pranić, Shelly Melissa","y":2020,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"inter-class correlation coefficient (ICC)","estd":"ICC (single/unspec)","v":0.33,"n":"395","k":"2","samp":"special","blind":"single","agg":"unspecified","scale":"accepted, revise, rejected","field":"multi-field","wr":"reviewers' publication recommendations for manuscripts","conf":"med","self":false,"doi":"10.1002/leap.1344","ciLow":0.184,"ciHigh":0.451,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":false,"ms":"For 395 manuscripts with two reviews each, an ICC of 0.330 indicates low agreement between the two reviewers' recommendations (accepted, revise, rejected).","vf":"unverified"},{"key":"EK4WFSX9","au":"Pranić, Shelly Melissa","y":2020,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"inter-class correlation coefficient (ICC)","estd":"ICC (single/unspec)","v":0.109,"n":"53","k":"3","samp":"special","blind":"single","agg":"unspecified","scale":"accepted, revise, rejected","field":"multi-field","wr":"reviewers' publication recommendations for manuscripts","conf":"med","self":false,"doi":"10.1002/leap.1344","ciLow":-0.403,"ciHigh":0.456,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"For 53 manuscripts with three reviews each, an ICC of 0.109 indicates very low agreement between the reviewers' recommendations.","vf":"unverified"},{"key":"EK4WFSX9","au":"Pranić, Shelly Melissa","y":2020,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"inter-class correlation coefficient (ICC)","estd":"ICC (single/unspec)","v":0.744,"n":"9","k":"4","samp":"special","blind":"single","agg":"unspecified","scale":"accepted, revise, rejected","field":"multi-field","wr":"reviewers' publication recommendations for manuscripts","conf":"med","self":false,"doi":"10.1002/leap.1344","ciLow":0.308,"ciHigh":0.935,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"For 9 manuscripts with four reviews each, an ICC of 0.744 indicates higher agreement between the reviewers' recommendations, though on a very small sample.","vf":"unverified"},{"key":"EK4WFSX9","au":"Pranić, Shelly Melissa","y":2020,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"kappa statistic (weighting unspecified)","estd":"kappa","v":0.65,"n":"100","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1-5 Likert (RQI quality rating)","field":"multi-field","wr":"two assessors rating quality of review reports","conf":"med","self":false,"doi":"10.1002/leap.1344","ciLow":0.5,"ciHigh":0.8,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":true,"he":false,"ms":"Two study assessors independently rated the quality of a random subsample of 100 review reports using the modified RQI; a kappa of 0.65 indicates moderate chance-corrected agreement between them.","vf":"unverified"},{"key":"MHK2CREE","au":"Rakap, Salih","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"percent agreement, exact agreement on accept/revise/reject decision","estd":"percent agreement","v":0.22,"n":"18","k":"3","samp":"special","blind":"double","agg":"single-rater","scale":"Accept / Revise / Reject","field":"special education and psychology","wr":"all reviewers and ChatGPT on publication decision","conf":"high","self":false,"doi":"10.1080/08856257.2025.2533555","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"The two assigned human reviewers and ChatGPT all agreed on the accept/revise/reject decision for only 4 of 18 manuscripts (22%).","vf":"unverified"},{"key":"MHK2CREE","au":"Rakap, Salih","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"percent agreement, exact agreement on accept/revise/reject decision","estd":"percent agreement","v":0.44,"n":"9","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"Accept / Revise / Reject","field":"special education and psychology","wr":"individual reviewer vs ChatGPT on publication decision","conf":"high","self":false,"doi":"10.1080/08856257.2025.2533555","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"Comparing each individual reviewer's publication recommendations with ChatGPT's over the nine manuscripts each reviewed, the highest reviewer-specific exact agreement was 44%.","vf":"unverified"},{"key":"MHK2CREE","au":"Rakap, Salih","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"percent agreement, exact agreement on accept/revise/reject decision","estd":"percent agreement","v":0.22,"n":"9","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"Accept / Revise / Reject","field":"special education and psychology","wr":"individual reviewer vs ChatGPT on publication decision","conf":"high","self":false,"doi":"10.1080/08856257.2025.2533555","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"Comparing each individual reviewer's publication recommendations with ChatGPT's over the nine manuscripts each reviewed, the lowest reviewer-specific exact agreement was 22%.","vf":"unverified"},{"key":"MHK2CREE","au":"Rakap, Salih","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"interobserver agreement percentage, exact agreement on Yes/No checklist items","estd":"percent agreement","v":0.852,"n":"9","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"Yes/No","field":"special education and psychology","wr":"reviewer R1 vs ChatGPT on checklist","conf":"high","self":false,"doi":"10.1080/08856257.2025.2533555","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Reviewer R1's Yes/No checklist ratings agreed with ChatGPT's on a mean 85.2% of items across the nine manuscripts R1 reviewed.","vf":"unverified"},{"key":"MHK2CREE","au":"Rakap, Salih","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"percent agreement, exact agreement on accept/revise/reject decision","estd":"percent agreement","v":0.28,"n":"18","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"Accept / Revise / Reject","field":"special education and psychology","wr":"first reviewer vs ChatGPT on publication decision","conf":"high","self":false,"doi":"10.1080/08856257.2025.2533555","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"ChatGPT and the first assigned reviewer of each manuscript agreed on the accept/revise/reject decision for 5 of 18 manuscripts (28%).","vf":"unverified"},{"key":"MHK2CREE","au":"Rakap, Salih","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"interobserver agreement percentage, exact agreement on Yes/No checklist items","estd":"percent agreement","v":0.868,"n":"9","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"Yes/No","field":"special education and psychology","wr":"reviewer R2 vs ChatGPT on checklist","conf":"high","self":false,"doi":"10.1080/08856257.2025.2533555","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Reviewer R2's Yes/No checklist ratings agreed with ChatGPT's on a mean 86.8% of items across the nine manuscripts R2 reviewed.","vf":"unverified"},{"key":"MHK2CREE","au":"Rakap, Salih","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"percent agreement, exact agreement on accept/revise/reject decision","estd":"percent agreement","v":0.39,"n":"18","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"Accept / Revise / Reject","field":"special education and psychology","wr":"second reviewer vs ChatGPT on publication decision","conf":"high","self":false,"doi":"10.1080/08856257.2025.2533555","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"ChatGPT and the second assigned reviewer of each manuscript agreed on the accept/revise/reject decision for 7 of 18 manuscripts (39%).","vf":"unverified"},{"key":"MHK2CREE","au":"Rakap, Salih","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"interobserver agreement percentage, exact agreement on Yes/No checklist items","estd":"percent agreement","v":0.871,"n":"9","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"Yes/No","field":"special education and psychology","wr":"reviewer R3 vs ChatGPT on checklist","conf":"high","self":false,"doi":"10.1080/08856257.2025.2533555","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Reviewer R3's Yes/No checklist ratings agreed with ChatGPT's on a mean 87.1% of items across the nine manuscripts R3 reviewed.","vf":"unverified"},{"key":"MHK2CREE","au":"Rakap, Salih","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"interobserver agreement percentage, exact agreement on Yes/No checklist items","estd":"percent agreement","v":0.857,"n":"9","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"Yes/No","field":"special education and psychology","wr":"reviewer R4 vs ChatGPT on checklist","conf":"high","self":false,"doi":"10.1080/08856257.2025.2533555","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Reviewer R4's Yes/No checklist ratings agreed with ChatGPT's on a mean 85.7% of items across the nine manuscripts R4 reviewed.","vf":"unverified"},{"key":"MHK2CREE","au":"Rakap, Salih","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"interobserver agreement percentage, exact agreement on Yes/No checklist items","estd":"percent agreement","v":0.929,"n":"1","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"Yes/No","field":"special education and psychology","wr":"one human reviewer vs ChatGPT on checklist","conf":"high","self":false,"doi":"10.1080/08856257.2025.2533555","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The highest pairwise human-ChatGPT checklist agreement was 92.9%, attained on several manuscripts in Table 1 (e.g. manuscript 16); each value is one reviewer and ChatGPT rating one manuscript.","vf":"unverified"},{"key":"MHK2CREE","au":"Rakap, Salih","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"interobserver agreement percentage, exact agreement on Yes/No checklist items","estd":"percent agreement","v":0.786,"n":"1","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"Yes/No","field":"special education and psychology","wr":"one human reviewer vs ChatGPT on checklist","conf":"high","self":false,"doi":"10.1080/08856257.2025.2533555","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The lowest pairwise human-ChatGPT checklist agreement was 78.6%, attained on several manuscripts in Table 1; each value is one reviewer and ChatGPT rating one manuscript.","vf":"unverified"},{"key":"MHK2CREE","au":"Rakap, Salih","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"interobserver agreement percentage, exact agreement on Yes/No checklist items","estd":"percent agreement","v":0.862,"n":"18","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"Yes/No","field":"special education and psychology","wr":"human reviewers vs ChatGPT on checklist","conf":"high","self":false,"doi":"10.1080/08856257.2025.2533555","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Comparing each human reviewer's Yes/No checklist ratings with ChatGPT's across the 18 manuscripts gave a mean interobserver agreement of 86.2%, the study's headline concordance result.","vf":"unverified"},{"key":"MHK2CREE","au":"Rakap, Salih","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"Krippendorff","form":"Krippendorff's Alpha","estd":"Krippendorff","v":0.122,"n":"18","k":"3","samp":"special","blind":"double","agg":"single-rater","scale":"Accept / Revise / Reject","field":"special education and psychology","wr":"reviewers and ChatGPT on publication decisions","conf":"high","self":false,"doi":"10.1080/08856257.2025.2533555","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":true,"he":false,"ms":"Krippendorff's alpha across the two assigned human reviewers and ChatGPT for accept/revise/reject decisions on the 18 manuscripts was 0.122, indicating very low agreement.","vf":"unverified"},{"key":"MHK2CREE","au":"Rakap, Salih","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"interobserver agreement percentage, exact agreement on Yes/No checklist items","estd":"percent agreement","v":1,"n":"2","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"Yes/No","field":"special education and psychology","wr":"human reviewers on manuscript quality checklist","conf":"high","self":false,"doi":"10.1080/08856257.2025.2533555","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two manuscripts (#13 and #18), each rated by a reviewer pair, reached perfect 100% between-reviewer agreement on the Yes/No quality checklist.","vf":"unverified"},{"key":"MHK2CREE","au":"Rakap, Salih","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"interobserver agreement percentage, exact agreement on Yes/No checklist items","estd":"percent agreement","v":0.714,"n":"1","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"Yes/No","field":"special education and psychology","wr":"human reviewers on manuscript quality checklist","conf":"high","self":false,"doi":"10.1080/08856257.2025.2533555","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The lowest between-reviewer checklist agreement across the 18 manuscripts, 71.4%, occurred on manuscript 15, rated by one reviewer pair.","vf":"unverified"},{"key":"MHK2CREE","au":"Rakap, Salih","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"interobserver agreement percentage, exact agreement on Yes/No checklist items","estd":"percent agreement","v":0.857,"n":"18","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"Yes/No","field":"special education and psychology","wr":"human reviewers on manuscript quality checklist","conf":"high","self":false,"doi":"10.1080/08856257.2025.2533555","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Pairs drawn from four expert human reviewers independently scored 18 unpublished single-case design manuscripts on a 42-question Yes/No quality checklist; mean interobserver agreement between the two reviewers was 85.7%.","vf":"unverified"},{"key":"MHK2CREE","au":"Rakap, Salih","y":2025,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"percent agreement, exact agreement on accept/revise/reject decision","estd":"percent agreement","v":0.61,"n":"18","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"Accept / Revise / Reject","field":"special education and psychology","wr":"two human reviewers on publication decision","conf":"high","self":false,"doi":"10.1080/08856257.2025.2533555","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"The two assigned human reviewers agreed on the accept/revise/reject decision for 11 of 18 manuscripts (61%).","vf":"unverified"},{"key":"U4JKK6XI","au":"Rastogi, Charvi","y":2024,"cx":"Journal","ob":"conference-abstract","fam":"other","form":"percent disagreement in pairwise ordering","estd":"other","v":0.34,"n":"","k":"","samp":"special","blind":"double","agg":"unspecified","scale":"strict paper ranking; accept or reject","field":"machine learning","wr":"relative quality of paired conference papers","conf":"med","self":false,"doi":"10.1371/journal.pone.0300710","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"unclear","pr":true,"he":false,"ms":"Authors ranked their own submitted papers by perceived scientific contribution before review outcomes were known. Among 10,171 paired responses, in the subset where the author gave a strict ranking and the two papers received different decisions, the ranking disagreed with the peer-review decision 34% of the time.","vf":"unverified"},{"key":"U4JKK6XI","au":"Rastogi, Charvi","y":2024,"cx":"Journal","ob":"conference-abstract","fam":"other","form":"percent disagreement between strict pairwise rankings","estd":"other","v":0.32,"n":"","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"rank submissions by perceived scientific contribution","field":"machine learning","wr":"relative quality of jointly authored papers","conf":"med","self":false,"doi":"10.1371/journal.pone.0300710","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"unclear","pr":false,"he":false,"ms":"Two co-authors independently ranked a pair of papers they had jointly authored by perceived scientific contribution. Across 1,357 such co-author response pairs, when both gave strict rankings their rankings disagreed 32% of the time.","vf":"unverified"},{"key":"MBCQSMRJ","au":"Research on Research Institute (RoRI)","y":2026,"cx":"Grant","ob":"grant-proposal","fam":"weighted-kappa","form":"weighted kappa, quadratic weights (Cohen, 1968); obtained per proposal, mean of all 41 proposals","estd":"weighted kappa","v":0.38,"n":"","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"5pt scale, 'constructive' to 'non-constructive'","field":"humanities and social sciences","wr":"co-applicants on constructiveness of received reviews","conf":"high","self":false,"doi":"10.6084/m9.figshare.29994841.v4","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":true,"he":false,"ms":"For 41 proposals, the two co-applicants each rated the constructiveness of every review their proposal received on a 5-point scale. A quadratically weighted kappa was computed per proposal and averaged; the mean of 0.38 indicates fair agreement between co-applicants on review constructiveness.","vf":"unverified"},{"key":"MBCQSMRJ","au":"Research on Research Institute (RoRI)","y":2026,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"Intraclass Correlation Coefficient (ICC) calculated for the proposal random effect (mixed-effects model, proposal + reviewer random intercepts)","estd":"ICC (single/unspec)","v":0.0887,"n":"140","k":"9.92","samp":"full-pool","blind":"double","agg":"single-rater","scale":"A+ ('outstanding') to C- ('unsuitable'), 9-point, recoded 1-9","field":"humanities and social sciences","wr":"applicant-reviewers on grant proposal scores (DPR)","conf":"high","self":false,"doi":"10.6084/m9.figshare.29994841.v4","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"In a distributed peer review trial, 323 applicants independently scored 140 grant proposals (1,387 reviews, about ten per proposal) on a 9-point scale. The proposal ICC of 0.0887 means only 8.87% of score variance reflected differences between proposals, indicating low inter-reviewer reliability.","vf":"unverified"},{"key":"MBCQSMRJ","au":"Research on Research Institute (RoRI)","y":2026,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"Intraclass Correlation Coefficient (ICC) calculated for the reviewer random effect (mixed-effects model, proposal + reviewer random intercepts)","estd":"ICC (single/unspec)","v":0.2004,"n":"140","k":"9.92","samp":"full-pool","blind":"double","agg":"unspecified","scale":"A+ ('outstanding') to C- ('unsuitable'), 9-point, recoded 1-9","field":"humanities and social sciences","wr":"share of score variance due to reviewer differences","conf":"high","self":false,"doi":"10.6084/m9.figshare.29994841.v4","ciLow":null,"ciHigh":null,"mt":"other","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"In the same mixed-effects decomposition of 1,387 distributed-review scores of 140 proposals, the reviewer ICC of 0.2004 shows that 20.04% of score variance was due to systematic differences between reviewers, more than twice the share attributable to the proposals themselves. It quantifies reviewer effects rather than agreement.","vf":"unverified"},{"key":"UYD3UVJU","au":"Rohmann, Jessica L.","y":2025,"cx":"Journal","ob":"review-report","fam":"weighted-kappa","form":"weighted Kappa","estd":"weighted kappa","v":0.5,"n":"40","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"1 - “poor” to 5 - “excellent”","field":"biomedical","wr":"editors rating quality of students' peer review reports","conf":"high","self":false,"doi":"10.1101/2025.02.11.25322060","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Two independent BMJ editors scored the overall quality (RQI global item, 1 to 5) of doctoral students' peer review reports; across 40 reports this editor pair's weighted kappa was 0.50, indicating moderate agreement.","vf":"unverified"},{"key":"UYD3UVJU","au":"Rohmann, Jessica L.","y":2025,"cx":"Journal","ob":"review-report","fam":"weighted-kappa","form":"weighted Kappa","estd":"weighted kappa","v":0.49,"n":"42","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"1 - “poor” to 5 - “excellent”","field":"biomedical","wr":"editors rating quality of students' peer review reports","conf":"high","self":false,"doi":"10.1101/2025.02.11.25322060","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Two independent BMJ editors scored the overall quality (RQI global item, 1 to 5) of doctoral students' peer review reports; across 42 reports this editor pair's weighted kappa was 0.49, indicating moderate agreement.","vf":"unverified"},{"key":"UYD3UVJU","au":"Rohmann, Jessica L.","y":2025,"cx":"Journal","ob":"review-report","fam":"weighted-kappa","form":"weighted Kappa","estd":"weighted kappa","v":0.53,"n":"74","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"1 - “poor” to 5 - “excellent”","field":"biomedical","wr":"editors rating quality of students' peer review reports","conf":"high","self":false,"doi":"10.1101/2025.02.11.25322060","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":true,"he":false,"ms":"Two independent BMJ editors scored the overall quality (RQI global item, 1 to 5) of doctoral students' peer review reports; across 74 reports the weighted kappa was 0.53, indicating moderate agreement. This is the largest of three editor-pair kappas (0.53, 0.50, 0.49); no overall value was reported.","vf":"unverified"},{"key":"YBW8PCPY","au":"Rooyen, Susan van","y":1999,"cx":"Journal","ob":"review-report","fam":"correlation","form":"correlation between the mean total scores given by the editors and authors; type unspecified","estd":"correlation","v":0.52,"n":"","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1-5 per item (1=poor, 5=excellent); total = mean of seven items","field":"biomedical","wr":"editors versus corresponding authors rating review reports","conf":"high","self":false,"doi":"10.1136/bmj.318.7175.23","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"The mean of the two editors' total quality scores was compared with the corresponding author's own rating of the same reviews. For reviews by anonymous reviewers the two sets of scores correlated 0.52.","vf":"unverified"},{"key":"YBW8PCPY","au":"Rooyen, Susan van","y":1999,"cx":"Journal","ob":"review-report","fam":"correlation","form":"correlation between the mean total scores given by the editors and authors; type unspecified","estd":"correlation","v":0.35,"n":"","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1-5 per item (1=poor, 5=excellent); total = mean of seven items","field":"biomedical","wr":"editors versus corresponding authors rating review reports","conf":"high","self":false,"doi":"10.1136/bmj.318.7175.23","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"For reviews written by reviewers asked to be identified, the mean of the two editors' total quality scores correlated 0.35 with the corresponding author's rating of the same reviews.","vf":"unverified"},{"key":"YBW8PCPY","au":"Rooyen, Susan van","y":1999,"cx":"Journal","ob":"review-report","fam":"weighted-kappa","form":"weighted kappa, with a maximum difference of 1 in scores between editors representing agreement","estd":"weighted kappa","v":0.67,"n":"226","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"five-point Likert scale (1=poor, 5=excellent)","field":"biomedical","wr":"two editors rating single review quality items","conf":"high","self":false,"doi":"10.1136/bmj.318.7175.23","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"0.67 is the highest of the seven item level weighted kappas between the two editors; the paper reports only the range, so the item is not named.","vf":"unverified"},{"key":"YBW8PCPY","au":"Rooyen, Susan van","y":1999,"cx":"Journal","ob":"review-report","fam":"weighted-kappa","form":"weighted kappa, with a maximum difference of 1 in scores between editors representing agreement","estd":"weighted kappa","v":0.38,"n":"226","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"five-point Likert scale (1=poor, 5=excellent)","field":"biomedical","wr":"two editors rating single review quality items","conf":"high","self":false,"doi":"10.1136/bmj.318.7175.23","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"The two editors' item level agreement was summarised only as a range across the seven instrument items; 0.38 is the lowest weighted kappa reported, item not named.","vf":"unverified"},{"key":"YBW8PCPY","au":"Rooyen, Susan van","y":1999,"cx":"Journal","ob":"review-report","fam":"correlation","form":"correlation between the mean total scores for the two editors; type unspecified","estd":"correlation","v":0.69,"n":"113","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-5 per item (1=poor, 5=excellent); total = mean of seven items","field":"biomedical","wr":"two editors rating anonymous reviewers' review reports","conf":"high","self":false,"doi":"10.1136/bmj.318.7175.23","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"For the 113 reviews written by reviewers randomised to remain anonymous, the two editors' mean total quality scores correlated 0.69. This is the same instrument reliability measured within the control arm.","vf":"unverified"},{"key":"YBW8PCPY","au":"Rooyen, Susan van","y":1999,"cx":"Journal","ob":"review-report","fam":"correlation","form":"correlation between the mean total scores for the two editors; type unspecified","estd":"correlation","v":0.64,"n":"113","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-5 per item (1=poor, 5=excellent); total = mean of seven items","field":"biomedical","wr":"two editors rating identified reviewers' review reports","conf":"high","self":false,"doi":"10.1136/bmj.318.7175.23","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"For the 113 reviews written by reviewers randomised to be asked to be identified to authors, the two editors' mean total quality scores correlated 0.64.","vf":"unverified"},{"key":"YBW8PCPY","au":"Rooyen, Susan van","y":1999,"cx":"Journal","ob":"review-report","fam":"correlation","form":"correlation between the mean total scores for the two editors; type unspecified","estd":"correlation","v":0.66,"n":"226","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1-5 per item (1=poor, 5=excellent); total = mean of seven items","field":"biomedical","wr":"two editors rating quality of review reports","conf":"high","self":false,"doi":"10.1136/bmj.318.7175.23","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":true,"he":false,"ms":"Two BMJ editors independently rated the quality of each of 226 reviews with a seven item review quality instrument. Their mean total scores correlated 0.66, the paper's headline reliability result for the instrument.","vf":"unverified"},{"key":"A7ERVCRX","au":"Rothwell, Peter M.","y":2000,"cx":"Journal","ob":"other","fam":"kappa","form":"kappa statistic (Thompson & Walter 1988)","estd":"kappa","v":0.08,"n":"179","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"accept / accept if revised / reject","field":"clinical neuroscience","wr":"two reviewers on manuscript accept/revise/reject","conf":"high","self":false,"doi":"10.1093/brain/123.9.1964","ciLow":-0.04,"ciHigh":0.2,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"For 179 consecutive manuscripts submitted to Journal A, two independent reviewers each recommended accept, revise or reject; a kappa of 0.08 shows agreement barely above chance.","vf":"unverified"},{"key":"A7ERVCRX","au":"Rothwell, Peter M.","y":2000,"cx":"Journal","ob":"other","fam":"percent-agreement","form":"observed percent agreement, exact category match","estd":"percent agreement","v":0.47,"n":"179","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"accept / accept if revised / reject","field":"clinical neuroscience","wr":"two reviewers on manuscript accept/revise/reject","conf":"high","self":false,"doi":"10.1093/brain/123.9.1964","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers of 179 Journal A manuscripts gave the same accept/revise/reject recommendation for 47% of them (raw observed agreement).","vf":"unverified"},{"key":"A7ERVCRX","au":"Rothwell, Peter M.","y":2000,"cx":"Journal","ob":"other","fam":"kappa","form":"kappa statistic (Thompson & Walter 1988)","estd":"kappa","v":-0.12,"n":"54","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"priority low / medium / high","field":"clinical neuroscience","wr":"two reviewers on publication priority","conf":"high","self":false,"doi":"10.1093/brain/123.9.1964","ciLow":-0.3,"ciHigh":0.11,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Among 54 Journal A manuscripts both reviewers judged publishable, agreement on low/medium/high priority was negative (kappa -0.12), indicating disagreement beyond chance.","vf":"unverified"},{"key":"A7ERVCRX","au":"Rothwell, Peter M.","y":2000,"cx":"Journal","ob":"other","fam":"percent-agreement","form":"observed percent agreement, exact category match","estd":"percent agreement","v":0.35,"n":"54","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"priority low / medium / high","field":"clinical neuroscience","wr":"two reviewers on publication priority","conf":"high","self":false,"doi":"10.1093/brain/123.9.1964","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Among 54 Journal A manuscripts both reviewers judged publishable, the two reviewers gave the same priority rating for 35% of them (raw observed agreement).","vf":"unverified"},{"key":"A7ERVCRX","au":"Rothwell, Peter M.","y":2000,"cx":"Journal","ob":"other","fam":"kappa","form":"kappa statistic (Thompson & Walter 1988)","estd":"kappa","v":0.28,"n":"116","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"accept / accept if revised / reject","field":"clinical neuroscience","wr":"two reviewers on manuscript accept/revise/reject","conf":"high","self":false,"doi":"10.1093/brain/123.9.1964","ciLow":0.12,"ciHigh":0.4,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For 116 manuscripts submitted to Journal B and assessed by two independent reviewers, a kappa of 0.28 indicates poor chance-corrected agreement on the accept/revise/reject recommendation.","vf":"unverified"},{"key":"A7ERVCRX","au":"Rothwell, Peter M.","y":2000,"cx":"Journal","ob":"other","fam":"percent-agreement","form":"observed percent agreement, exact category match","estd":"percent agreement","v":0.61,"n":"116","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"accept / accept if revised / reject","field":"clinical neuroscience","wr":"two reviewers on manuscript accept/revise/reject","conf":"high","self":false,"doi":"10.1093/brain/123.9.1964","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers of 116 Journal B manuscripts gave the same accept/revise/reject recommendation for 61% of them (raw observed agreement).","vf":"unverified"},{"key":"A7ERVCRX","au":"Rothwell, Peter M.","y":2000,"cx":"Journal","ob":"other","fam":"kappa","form":"kappa statistic (Thompson & Walter 1988)","estd":"kappa","v":0.27,"n":"49","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"priority low / medium / high","field":"clinical neuroscience","wr":"two reviewers on publication priority","conf":"high","self":false,"doi":"10.1093/brain/123.9.1964","ciLow":0.01,"ciHigh":0.53,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Among 49 Journal B manuscripts both reviewers judged publishable, agreement on low/medium/high priority was poor (kappa 0.27).","vf":"unverified"},{"key":"A7ERVCRX","au":"Rothwell, Peter M.","y":2000,"cx":"Journal","ob":"other","fam":"percent-agreement","form":"observed percent agreement, exact category match","estd":"percent agreement","v":0.61,"n":"49","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"priority low / medium / high","field":"clinical neuroscience","wr":"two reviewers on publication priority","conf":"high","self":false,"doi":"10.1093/brain/123.9.1964","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Among 49 Journal B manuscripts both reviewers judged publishable, the two reviewers gave the same priority rating for 61% of them (raw observed agreement).","vf":"unverified"},{"key":"A7ERVCRX","au":"Rothwell, Peter M.","y":2000,"cx":"Journal","ob":"other","fam":"other","form":"adjusted r2 from ANOVA, share of total abstract-score variance due to abstract identity","estd":"other","v":0.11,"n":"32","k":"16","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"1 (poor) to 6 (excellent)","field":"clinical neuroscience","wr":"reviewers on conference abstract merit scores","conf":"high","self":false,"doi":"10.1093/brain/123.9.1964","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Sixteen independent reviewers each scored 32 Meeting A abstracts on a 1 to 6 scale. Abstract identity accounted for only 11% of total score variance, an ANOVA analogue of low reliability; 512 ratings derived from stated complete crossing.","vf":"unverified"},{"key":"A7ERVCRX","au":"Rothwell, Peter M.","y":2000,"cx":"Journal","ob":"other","fam":"other","form":"adjusted r2 from ANOVA, share of total abstract-score variance due to reviewer identity","estd":"other","v":0.27,"n":"32","k":"16","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"1 (poor) to 6 (excellent)","field":"clinical neuroscience","wr":"reviewer identity share of abstract score variance","conf":"high","self":false,"doi":"10.1093/brain/123.9.1964","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"At Meeting A, 27% of total abstract-score variance was due to some of the 16 reviewers systematically scoring higher or lower than others, a leniency and severity source of disagreement; 512 ratings derived from stated complete crossing.","vf":"unverified"},{"key":"A7ERVCRX","au":"Rothwell, Peter M.","y":2000,"cx":"Journal","ob":"other","fam":"other","form":"adjusted r2 from ANOVA, share of total abstract-score variance due to abstract identity","estd":"other","v":0.15,"n":"28","k":"14","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"1 (poor) to 6 (excellent)","field":"clinical neuroscience","wr":"reviewers on conference abstract merit scores","conf":"high","self":false,"doi":"10.1093/brain/123.9.1964","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Fourteen independent reviewers each scored 28 Meeting B abstracts on a 1 to 6 scale. Abstract identity accounted for only 15% of total score variance, an ANOVA analogue of low reliability; 392 ratings derived from stated complete crossing.","vf":"unverified"},{"key":"A7ERVCRX","au":"Rothwell, Peter M.","y":2000,"cx":"Journal","ob":"other","fam":"other","form":"adjusted r2 from ANOVA, share of total abstract-score variance due to reviewer identity","estd":"other","v":0.32,"n":"28","k":"14","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"1 (poor) to 6 (excellent)","field":"clinical neuroscience","wr":"reviewer identity share of abstract score variance","conf":"high","self":false,"doi":"10.1093/brain/123.9.1964","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"At Meeting B, 32% of total abstract-score variance was due to some of the 14 reviewers systematically scoring higher or lower than others, a leniency and severity source of disagreement; 392 ratings derived from stated complete crossing.","vf":"unverified"},{"key":"Q9DUXILF","au":"Rubin, Haya R.","y":1993,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"average pairwise kappa (Fleiss 1981) between any two reviewers, on accept/reject classification derived from the rating cutoff","estd":"kappa","v":0.18,"n":"426","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"accept / reject (dichotomised from 1-5 global rating)","field":"general internal medicine","wr":"any two reviewers' accept/reject on abstract global rating","conf":"high","self":false,"doi":"10.1007/bf02600092","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Each reviewer's global rating of an abstract was classified as accept or reject at the selection cutoff; the average kappa between any two reviewers was 0.18, reflecting only slight chance-corrected agreement.","vf":"unverified"},{"key":"Q9DUXILF","au":"Rubin, Haya R.","y":1993,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"average pairwise kappa (Fleiss 1981) between any two reviewers, on accept/reject classification derived from the rating cutoff","estd":"kappa","v":0.11,"n":"426","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"accept / reject (dichotomised from 1-5 interest rating)","field":"general internal medicine","wr":"any two reviewers' accept/reject on abstract interest rating","conf":"high","self":false,"doi":"10.1007/bf02600092","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For reviewers' interest ratings dichotomised to accept or reject, the average kappa between any two reviewers was 0.11, the lowest of the dimensions.","vf":"unverified"},{"key":"Q9DUXILF","au":"Rubin, Haya R.","y":1993,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"average pairwise kappa (Fleiss 1981) between any two reviewers, on accept/reject classification derived from the rating cutoff","estd":"kappa","v":0.16,"n":"426","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"accept / reject (dichotomised from 1-5 methods rating)","field":"general internal medicine","wr":"any two reviewers' accept/reject on abstract methods rating","conf":"high","self":false,"doi":"10.1007/bf02600092","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For reviewers' methods-quality ratings dichotomised to accept or reject, the average kappa between any two reviewers was 0.16.","vf":"unverified"},{"key":"Q9DUXILF","au":"Rubin, Haya R.","y":1993,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"average pairwise kappa (Fleiss 1981) between any two reviewers, on accept/reject classification derived from the rating cutoff","estd":"kappa","v":0.14,"n":"426","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"accept / reject (dichotomised from 1-5 presentation rating)","field":"general internal medicine","wr":"any two reviewers' accept/reject on abstract presentation rating","conf":"high","self":false,"doi":"10.1007/bf02600092","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For reviewers' presentation-quality ratings dichotomised to accept or reject, the average kappa between any two reviewers was 0.14.","vf":"unverified"},{"key":"Q9DUXILF","au":"Rubin, Haya R.","y":1993,"cx":"Journal","ob":"conference-abstract","fam":"percent-agreement","form":"average crude percentage agreement (exact) between any two reviewers on accept/reject classification","estd":"percent agreement","v":0.66,"n":"426","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"accept / reject (dichotomised from 1-5 global rating)","field":"general internal medicine","wr":"any two reviewers' accept/reject on abstract global rating","conf":"high","self":false,"doi":"10.1007/bf02600092","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Any two reviewers gave the same accept-or-reject verdict on an abstract's global rating 66% of the time, though this crude agreement is inflated by chance.","vf":"unverified"},{"key":"Q9DUXILF","au":"Rubin, Haya R.","y":1993,"cx":"Journal","ob":"conference-abstract","fam":"percent-agreement","form":"average crude percentage agreement (exact) between any two reviewers on accept/reject classification","estd":"percent agreement","v":0.56,"n":"426","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"accept / reject (dichotomised from 1-5 interest rating)","field":"general internal medicine","wr":"any two reviewers' accept/reject on abstract interest rating","conf":"high","self":false,"doi":"10.1007/bf02600092","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Any two reviewers agreed on the accept-or-reject verdict for an abstract's interest 56% of the time, the lowest crude agreement among the dimensions.","vf":"unverified"},{"key":"Q9DUXILF","au":"Rubin, Haya R.","y":1993,"cx":"Journal","ob":"conference-abstract","fam":"percent-agreement","form":"average crude percentage agreement (exact) between any two reviewers on accept/reject classification","estd":"percent agreement","v":0.61,"n":"426","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"accept / reject (dichotomised from 1-5 methods rating)","field":"general internal medicine","wr":"any two reviewers' accept/reject on abstract methods rating","conf":"high","self":false,"doi":"10.1007/bf02600092","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Any two reviewers agreed on the accept-or-reject verdict for an abstract's quality of methods 61% of the time.","vf":"unverified"},{"key":"Q9DUXILF","au":"Rubin, Haya R.","y":1993,"cx":"Journal","ob":"conference-abstract","fam":"percent-agreement","form":"average crude percentage agreement (exact) between any two reviewers on accept/reject classification","estd":"percent agreement","v":0.59,"n":"426","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"accept / reject (dichotomised from 1-5 presentation rating)","field":"general internal medicine","wr":"any two reviewers' accept/reject on abstract presentation rating","conf":"high","self":false,"doi":"10.1007/bf02600092","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Any two reviewers agreed on the accept-or-reject verdict for an abstract's quality of presentation 59% of the time.","vf":"unverified"},{"key":"Q9DUXILF","au":"Rubin, Haya R.","y":1993,"cx":"Journal","ob":"conference-abstract","fam":"G-theory","form":"proportion of variance associated with abstract identity ('reliable variance') from analysis of variance / variance-components","estd":"G-theory","v":0.38,"n":"426","k":"","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5 global (1=poor, 5=outstanding)","field":"general internal medicine","wr":"reviewers on abstract global ratings","conf":"high","self":false,"doi":"10.1007/bf02600092","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For reviewers' single-item global ratings of the same abstracts, 38% of the variance reflected the abstract's identity, making global ratings slightly more reproducible than the composite summary scores.","vf":"unverified"},{"key":"Q9DUXILF","au":"Rubin, Haya R.","y":1993,"cx":"Journal","ob":"conference-abstract","fam":"G-theory","form":"proportion of variance associated with abstract identity ('reliable variance') from analysis of variance / variance-components","estd":"G-theory","v":0.27,"n":"426","k":"","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5 interest to SGIM audience (1=poor, 5=outstanding)","field":"general internal medicine","wr":"reviewers on abstract interest ratings","conf":"high","self":false,"doi":"10.1007/bf02600092","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For reviewers' ratings of an abstract's interest to the SGIM audience, only 27% of the variance reflected the abstract's identity, the least reproducible of the rated dimensions.","vf":"unverified"},{"key":"Q9DUXILF","au":"Rubin, Haya R.","y":1993,"cx":"Journal","ob":"conference-abstract","fam":"G-theory","form":"proportion of variance associated with abstract identity ('reliable variance') from analysis of variance / variance-components","estd":"G-theory","v":0.34,"n":"426","k":"","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5 quality of methods (1=poor, 5=outstanding)","field":"general internal medicine","wr":"reviewers on abstract methods-quality ratings","conf":"high","self":false,"doi":"10.1007/bf02600092","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For reviewers' ratings of an abstract's quality of methods, 34% of the variance reflected the abstract's identity.","vf":"unverified"},{"key":"Q9DUXILF","au":"Rubin, Haya R.","y":1993,"cx":"Journal","ob":"conference-abstract","fam":"G-theory","form":"proportion of variance associated with abstract identity ('reliable variance') from analysis of variance / variance-components","estd":"G-theory","v":0.31,"n":"426","k":"","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5 quality of presentation (1=poor, 5=outstanding)","field":"general internal medicine","wr":"reviewers on abstract presentation-quality ratings","conf":"high","self":false,"doi":"10.1007/bf02600092","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For reviewers' ratings of an abstract's quality of presentation, 31% of the variance reflected the abstract's identity.","vf":"unverified"},{"key":"Q9DUXILF","au":"Rubin, Haya R.","y":1993,"cx":"Journal","ob":"conference-abstract","fam":"G-theory","form":"proportion of variance associated with abstract identity ('reliable variance') from analysis of variance / variance-components","estd":"G-theory","v":0.36,"n":"426","k":"","samp":"full-pool","blind":"double","agg":"single-rater","scale":"3-15 summary score (sum of three 1-5 subscales)","field":"general internal medicine","wr":"reviewers on abstract summary scores","conf":"high","self":false,"doi":"10.1007/bf02600092","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Each of 426 abstracts was independently scored by five to seven of 55 reviewers using a 3 to 15 summary score. An analysis of variance found that only 36% of the variance in summary scores reflected the abstract's identity (the reliable variance), indicating substantial reviewer disagreement.","vf":"unverified"},{"key":"G6K4RHJ6","au":"Sattler, David N","y":2015,"cx":"Grant","ob":"other","fam":"ICC","form":"ICC with two-way estimation, for agreement (vs consistency) and for single values; Shrout and Fleiss ICC(2,k)","estd":"ICC (single/unspec)","v":0.61,"n":"4","k":"","samp":"special","blind":"unclear","agg":"single-rater","scale":"1-9, 1 = exceptional to 9 = poor","field":"public health","wr":"NIH scoring-criteria proposal summary descriptions","conf":"high","self":false,"doi":"10.1371/journal.pone.0130450","ciLow":0.32,"ciHigh":0.96,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Public health professors in the no-training control arm independently assigned NIH 9-point scores to the same four proposal summary descriptions. A two-way single-values ICC of 0.61 indicates moderate inter-rater agreement without training.","vf":"unverified"},{"key":"G6K4RHJ6","au":"Sattler, David N","y":2015,"cx":"Grant","ob":"other","fam":"ICC","form":"ICC with two-way estimation, for agreement (vs consistency) and for single values; Shrout and Fleiss ICC(2,k)","estd":"ICC (single/unspec)","v":0.89,"n":"4","k":"","samp":"special","blind":"unclear","agg":"single-rater","scale":"1-9, 1 = exceptional to 9 = poor","field":"public health","wr":"NIH scoring-criteria proposal summary descriptions","conf":"high","self":false,"doi":"10.1371/journal.pone.0130450","ciLow":0.71,"ciHigh":0.99,"mt":"inter-rater","tgt":"other","rr":"none","pr":true,"he":false,"ms":"Public health professors randomised to watch an 11-minute training video independently assigned NIH 9-point scores to four proposal summary descriptions taken from the NIH scoring criteria. A two-way single-values ICC of 0.89 indicates high inter-rater agreement after training.","vf":"unverified"},{"key":"I4ZU3XFU","au":"Scarr, Sandra","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement on collapsed reject vs possibly-accept dichotomy","estd":"percent agreement","v":0.78,"n":"87","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"reject (1-2) vs possibly accept (3-5)","field":"psychology","wr":"two reviewers on collapsed reject/accept recommendation","conf":"high","self":false,"doi":"10.1037/h0078544","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Collapsing the 1-to-5 scale into reject (1-2) versus possibly accept (3-5), the two independent reviewers agreed on 68 of 87 decisions, a raw agreement of about 0.78.","vf":"unverified"},{"key":"I4ZU3XFU","au":"Scarr, Sandra","y":1978,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement on the 5-point acceptability rating","estd":"percent agreement","v":0.66,"n":"87","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = reject to 5 = accept in the present form","field":"psychology","wr":"two reviewers on manuscript acceptability ratings","conf":"high","self":false,"doi":"10.1037/h0078544","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"For 87 editorial decisions on American Psychologist manuscripts (78 manuscripts, 9 rated twice after resubmission), two independent reviewers each rated acceptability on a 1-to-5 scale; they gave exactly the same rating on 57 of 87, a raw agreement of about 0.66.","vf":"unverified"},{"key":"DVDHBSTH","au":"Schmaling, Karen B.","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-measure intraclass correlations (ICC) for absolute agreement with mixed effects (reviewer effects random, component scores fixed)","estd":"ICC (single/unspec)","v":0.614,"n":"1","k":"605","samp":"special","blind":"single","agg":"single-rater","scale":"1-9, 1 = exceptional to 9 = poor","field":"biomedical","wr":"reviewers scoring an outstanding simulated proposal","conf":"med","self":false,"doi":"10.1371/journal.pone.0315567","ciLow":0.542,"ciHigh":0.675,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":true,"ms":"605 experienced reviewers independently scored one outstanding simulated NIH proposal across six criterion scores; a single-measure absolute-agreement ICC of 0.614 indicates good agreement among reviewers.","vf":"unverified"},{"key":"DVDHBSTH","au":"Schmaling, Karen B.","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-measure intraclass correlations (ICC) for absolute agreement with mixed effects (reviewer effects random, component scores fixed)","estd":"ICC (single/unspec)","v":0.276,"n":"1","k":"205","samp":"special","blind":"single","agg":"single-rater","scale":"1-9, 1 = exceptional to 9 = poor","field":"biomedical","wr":"reviewers scoring a proposal with a risky approach","conf":"med","self":false,"doi":"10.1371/journal.pone.0315567","ciLow":0.139,"ciHigh":0.409,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"205 reviewers independently scored one simulated proposal with a risky approach across six criterion scores; the single-measure ICC of 0.276 shows poor agreement among reviewers.","vf":"unverified"},{"key":"DVDHBSTH","au":"Schmaling, Karen B.","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-measure intraclass correlations (ICC) for absolute agreement with mixed effects (reviewer effects random, component scores fixed)","estd":"ICC (single/unspec)","v":0.271,"n":"1","k":"201","samp":"special","blind":"single","agg":"single-rater","scale":"1-9, 1 = exceptional to 9 = poor","field":"biomedical","wr":"reviewers scoring a proposal with both risks","conf":"med","self":false,"doi":"10.1371/journal.pone.0315567","ciLow":0.149,"ciHigh":0.392,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"201 reviewers independently scored one simulated proposal with both a risky approach and a risky investigator across six criterion scores; the single-measure ICC of 0.271 shows poor agreement.","vf":"unverified"},{"key":"DVDHBSTH","au":"Schmaling, Karen B.","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-measure intraclass correlations (ICC) for absolute agreement with mixed effects (reviewer effects random, component scores fixed)","estd":"ICC (single/unspec)","v":0.258,"n":"1","k":"199","samp":"special","blind":"single","agg":"single-rater","scale":"1-9, 1 = exceptional to 9 = poor","field":"biomedical","wr":"reviewers scoring a proposal with a risky investigator","conf":"med","self":false,"doi":"10.1371/journal.pone.0315567","ciLow":0.147,"ciHigh":0.37,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"199 reviewers independently scored one simulated proposal with a risky investigator across six criterion scores; the single-measure ICC of 0.258 shows poor agreement among reviewers.","vf":"unverified"},{"key":"DVDHBSTH","au":"Schmaling, Karen B.","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"percent endorsing the modal score (exact agreement)","estd":"percent agreement","v":0.52,"n":"1","k":"605","samp":"special","blind":"single","agg":"single-rater","scale":"1-9, 1 = exceptional to 9 = poor","field":"biomedical","wr":"reviewers' modal-score agreement on outstanding proposal","conf":"med","self":false,"doi":"10.1371/journal.pone.0315567","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Averaged across the six criteria, 52% of the 605 reviewers gave the outstanding proposal its modal score, the study's exact-agreement measure of interrater agreement.","vf":"unverified"},{"key":"DVDHBSTH","au":"Schmaling, Karen B.","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"percent endorsing the modal score (exact agreement)","estd":"percent agreement","v":0.607,"n":"1","k":"605","samp":"special","blind":"single","agg":"single-rater","scale":"1-9, 1 = exceptional to 9 = poor","field":"biomedical","wr":"reviewers' modal-score agreement on environment, control OIS","conf":"med","self":false,"doi":"10.1371/journal.pone.0315567","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"On the outstanding proposal, the environment criterion showed the highest exact agreement, with 60.7% of the 605 reviewers giving the modal score. This is the maximum percent-agreement stratum.","vf":"unverified"},{"key":"DVDHBSTH","au":"Schmaling, Karen B.","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"percent endorsing the modal score (exact agreement)","estd":"percent agreement","v":0.244,"n":"1","k":"205","samp":"special","blind":"single","agg":"single-rater","scale":"1-9, 1 = exceptional to 9 = poor","field":"biomedical","wr":"reviewers' modal-score agreement on approach, risky-approach OIS","conf":"med","self":false,"doi":"10.1371/journal.pone.0315567","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"On the risky-approach proposal, the approach criterion showed the lowest exact agreement, with 24.4% of the 205 reviewers giving the modal score. This is the minimum percent-agreement stratum.","vf":"unverified"},{"key":"DVDHBSTH","au":"Schmaling, Karen B.","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"percent endorsing the modal score (exact agreement)","estd":"percent agreement","v":0.34,"n":"1","k":"201","samp":"special","blind":"single","agg":"single-rater","scale":"1-9, 1 = exceptional to 9 = poor","field":"biomedical","wr":"reviewers' modal-score agreement on both-risks proposal","conf":"med","self":false,"doi":"10.1371/journal.pone.0315567","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For the proposal with both risks, an average of 34% of the 201 reviewers endorsed the modal criterion score, showing lower exact interrater agreement.","vf":"unverified"},{"key":"DVDHBSTH","au":"Schmaling, Karen B.","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"mean of six single-measure test-retest ICCs between two occasions two weeks apart","estd":"ICC (single/unspec)","v":0.52,"n":"1","k":"83","samp":"special","blind":"single","agg":"single-rater","scale":"1-9 initially; adjectives only at retest","field":"biomedical","wr":"reviewers rescoring the outstanding proposal two weeks later","conf":"med","self":false,"doi":"10.1371/journal.pone.0315567","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"83 reviewers re-scored the same outstanding simulated proposal two weeks later; the mean single-measure test-retest ICC across the six criterion scores was 0.52, indicating fair within-reviewer consistency.","vf":"unverified"},{"key":"DVDHBSTH","au":"Schmaling, Karen B.","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-measure test-retest ICC between two occasions two weeks apart","estd":"ICC (single/unspec)","v":0.71,"n":"1","k":"30","samp":"special","blind":"single","agg":"single-rater","scale":"1-9 initially; adjectives only at retest","field":"biomedical","wr":"reviewers rescoring investigator, both-risks proposal","conf":"med","self":false,"doi":"10.1371/journal.pone.0315567","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"For the proposal with both risks, the investigator criterion showed the highest within-reviewer consistency (test-retest ICC = 0.71; 30 reviewers). This is the maximum test-retest stratum.","vf":"unverified"},{"key":"DVDHBSTH","au":"Schmaling, Karen B.","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-measure test-retest ICC between two occasions two weeks apart","estd":"ICC (single/unspec)","v":0,"n":"1","k":"26","samp":"special","blind":"single","agg":"single-rater","scale":"1-9 initially; adjectives only at retest","field":"biomedical","wr":"reviewers rescoring overall impact, risky-approach proposal","conf":"med","self":false,"doi":"10.1371/journal.pone.0315567","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"For the risky-approach proposal, the overall impact score showed no within-reviewer consistency across the two occasions (test-retest ICC = 0.00; 26 reviewers). This is a minimum test-retest stratum (tied with environment, risky investigator).","vf":"unverified"},{"key":"DVDHBSTH","au":"Schmaling, Karen B.","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"single-measure test-retest ICC between two occasions two weeks apart","estd":"ICC (single/unspec)","v":0,"n":"1","k":"27","samp":"special","blind":"single","agg":"single-rater","scale":"1-9 initially; adjectives only at retest","field":"biomedical","wr":"reviewers rescoring environment, risky-investigator proposal","conf":"med","self":false,"doi":"10.1371/journal.pone.0315567","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"For the risky-investigator proposal, the environment score showed no within-reviewer consistency across the two occasions (test-retest ICC = 0.00; 27 reviewers). This is a minimum test-retest stratum (tied with overall impact, risky approach).","vf":"unverified"},{"key":"DVDHBSTH","au":"Schmaling, Karen B.","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"mean of six single-measure test-retest ICCs between two occasions two weeks apart","estd":"ICC (single/unspec)","v":0.31,"n":"1","k":"26","samp":"special","blind":"single","agg":"single-rater","scale":"1-9 initially; adjectives only at retest","field":"biomedical","wr":"reviewers rescoring the risky-approach proposal","conf":"med","self":false,"doi":"10.1371/journal.pone.0315567","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"26 reviewers re-scored the same risky-approach proposal two weeks later; the mean single-measure test-retest ICC across the six criterion scores was 0.31, indicating poor within-reviewer consistency.","vf":"unverified"},{"key":"DVDHBSTH","au":"Schmaling, Karen B.","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"mean of six single-measure test-retest ICCs between two occasions two weeks apart","estd":"ICC (single/unspec)","v":0.47,"n":"1","k":"30","samp":"special","blind":"single","agg":"single-rater","scale":"1-9 initially; adjectives only at retest","field":"biomedical","wr":"reviewers rescoring the proposal with both risks","conf":"med","self":false,"doi":"10.1371/journal.pone.0315567","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"30 reviewers re-scored the same dual-risk proposal two weeks later; the mean single-measure test-retest ICC across the six criterion scores was 0.47, indicating fair within-reviewer consistency.","vf":"unverified"},{"key":"DVDHBSTH","au":"Schmaling, Karen B.","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"mean of six single-measure test-retest ICCs between two occasions two weeks apart","estd":"ICC (single/unspec)","v":0.33,"n":"1","k":"27","samp":"special","blind":"single","agg":"single-rater","scale":"1-9 initially; adjectives only at retest","field":"biomedical","wr":"reviewers rescoring the risky-investigator proposal","conf":"med","self":false,"doi":"10.1371/journal.pone.0315567","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"27 reviewers re-scored the same risky-investigator proposal two weeks later; the mean single-measure test-retest ICC across the six criterion scores was 0.33, indicating poor within-reviewer consistency.","vf":"unverified"},{"key":"Y26RHK69","au":"Schroter, Sara","y":2004,"cx":"Journal","ob":"review-report","fam":"ICC","form":"intraclass correlation coefficient","estd":"ICC (single/unspec)","v":0.91,"n":"","k":"2","samp":"special","blind":"unclear","agg":"unspecified","scale":"Number of nine major errors identified","field":"biomedical","wr":"researchers counting major errors in review reports","conf":"high","self":false,"doi":"10.1136/bmj.38023.700775.AE","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Two researchers, blind to the reviewer's identity and study group, independently counted how many of the nine deliberate major errors each review reported. An intraclass correlation of 0.91 indicates very close agreement between the two counters.","vf":"unverified"},{"key":"Y26RHK69","au":"Schroter, Sara","y":2004,"cx":"Journal","ob":"review-report","fam":"ICC","form":"intraclass correlation coefficient for total review quality instrument score","estd":"ICC (single/unspec)","v":0.65,"n":"","k":"2","samp":"special","blind":"unclear","agg":"unspecified","scale":"Total scores range from 1 to 5; higher scores reflect higher quality","field":"biomedical","wr":"editors rating quality of reviewers' review reports","conf":"high","self":false,"doi":"10.1136/bmj.38023.700775.AE","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":true,"he":false,"ms":"Two editors independently rated the quality of each review submitted in this trial, using the eight item review quality instrument scored from 1 to 5. An intraclass correlation of 0.65 indicates good agreement between the paired editor raters on total review quality scores.","vf":"unverified"},{"key":"52D3XAIF","au":"Schroter, Sara","y":2006,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"weighted kappa statistic","estd":"weighted kappa","v":0.56,"n":"788","k":"2","samp":"special","blind":"single","agg":"single-rater","scale":"1-5 per RQI item","field":"biomedical","wr":"trained raters on review report quality","conf":"high","self":false,"doi":"10.1001/jama.295.3.314","ciLow":0.49,"ciHigh":0.63,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":true,"he":false,"ms":"Each of 788 peer review reports at 10 biomedical journals was rated independently by two of 16 trained raters using the Review Quality Instrument. The weighted kappa of 0.56 indicates moderate chance-corrected agreement between raters on review quality.","vf":"unverified"},{"key":"UKWAWEDT","au":"Scott, William A.","y":1974,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (Haggard, 1958)","estd":"ICC (single/unspec)","v":0.19,"n":"268","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = sound and thorough; 4 = inadequate","field":"psychology (personality/social)","wr":"two referees on manuscript design ratings","conf":"med","self":false,"doi":"10.1037/h0037631","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Two referees independently rated adequacy of research design and analysis; across 268 manuscripts the intraclass correlation of 0.19 indicates poor, though significant, agreement between referees.","vf":"unverified"},{"key":"UKWAWEDT","au":"Scott, William A.","y":1974,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (Haggard, 1958)","estd":"ICC (single/unspec)","v":0.28,"n":"154","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = very important; 5 = minimal, makes no contribution","field":"psychology (personality/social)","wr":"two referees on manuscript importance ratings","conf":"med","self":false,"doi":"10.1037/h0037631","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Two referees independently rated the importance of each manuscript's contribution; across 154 manuscripts the intraclass correlation of 0.28 indicates weak but statistically significant agreement between referees.","vf":"unverified"},{"key":"UKWAWEDT","au":"Scott, William A.","y":1974,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (Haggard, 1958)","estd":"ICC (single/unspec)","v":0.37,"n":"197","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = thorough and relevant; 4 = ignores relevant studies","field":"psychology (personality/social)","wr":"two referees on manuscript literature-coverage ratings","conf":"med","self":false,"doi":"10.1037/h0037631","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Two referees independently rated attention to relevant literature; across 197 manuscripts the intraclass correlation of 0.37 was the highest of the seven attributes, indicating only modest agreement.","vf":"unverified"},{"key":"UKWAWEDT","au":"Scott, William A.","y":1974,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (Haggard, 1958)","estd":"ICC (single/unspec)","v":0.07,"n":"154","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = major interest to most; 5 = little interest to anyone","field":"psychology (personality/social)","wr":"two referees on manuscript reader-interest ratings","conf":"med","self":false,"doi":"10.1037/h0037631","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Two referees independently rated the probable reader interest of each manuscript; across 154 double-reviewed manuscripts the intraclass correlation was 0.07, showing negligible and non-significant agreement between referees on this attribute.","vf":"unverified"},{"key":"UKWAWEDT","au":"Scott, William A.","y":1974,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (Haggard, 1958)","estd":"ICC (single/unspec)","v":0.26,"n":"286","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = accept; 2 = accept if revised; 3 = reject","field":"psychology (personality/social)","wr":"two referees on manuscript publish/reject recommendation","conf":"med","self":false,"doi":"10.1037/h0037631","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":true,"ms":"Two referees independently recorded a publish recommendation (accept, accept if revised, or reject) for each manuscript; across 286 manuscripts the intraclass correlation of 0.26 indicates weak agreement on the recommendation itself. Marked primary as the study's bottom-line judgment, since no pooled overall ICC across attributes was reported.","vf":"unverified"},{"key":"UKWAWEDT","au":"Scott, William A.","y":1974,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (Haggard, 1958)","estd":"ICC (single/unspec)","v":0.25,"n":"283","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = highly literate, well organised; 4 = not intelligible","field":"psychology (personality/social)","wr":"two referees on manuscript style ratings","conf":"med","self":false,"doi":"10.1037/h0037631","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Two referees independently rated style and organisation; across 283 manuscripts the intraclass correlation of 0.25 indicates weak agreement between referees.","vf":"unverified"},{"key":"UKWAWEDT","au":"Scott, William A.","y":1974,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation coefficient (Haggard, 1958)","estd":"ICC (single/unspec)","v":0.31,"n":"255","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = too condensed; 2 = appropriate; 3 = too leisurely","field":"psychology (personality/social)","wr":"two referees on manuscript succinctness ratings","conf":"med","self":false,"doi":"10.1037/h0037631","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Two referees independently rated succinctness; across 255 manuscripts the intraclass correlation of 0.31 indicates modest agreement between referees.","vf":"unverified"},{"key":"UKWAWEDT","au":"Scott, William A.","y":1974,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"r","estd":"correlation","v":0.58,"n":"287","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"mean of two referees, 1 = accept to 3 = reject","field":"psychology (personality/social)","wr":"pooled referee recommendation vs editorial disposition","conf":"med","self":false,"doi":"10.1037/h0037631","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"restricted-other","pr":false,"he":false,"ms":"The mean of the two referees' recommendations (1 = accept to 3 = reject) correlated 0.58 with the editorial disposition across 287 manuscripts; the disposition was made after reading the reviews, so this reflects decision-following rather than independent inter-rater agreement, hence typed 'other'.","vf":"unverified"},{"key":"JMK4CLUJ","au":"Seeber, Marco","y":2022,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach's alpha of all items","estd":"Cronbach alpha","v":0.92,"n":"1928","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0 to 5 scores for 13 different criteria","field":"multi-field","wr":"internal consistency of the 13-criterion scoring instrument","conf":"high","self":false,"doi":"10.1002/asi.24617","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Cronbach's alpha across the 13 evaluation criteria was 0.92, indicating high internal consistency of the scoring instrument and supporting the use of the summed total score as the dependent variable.","vf":"unverified"},{"key":"JMK4CLUJ","au":"Seeber, Marco","y":2022,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"correlation between mean individual scores and final scores (method unspecified)","estd":"correlation","v":0.91,"n":"1928","k":"3","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"overall evaluation score ranges from 0 to 65","field":"multi-field","wr":"mean individual reviewer scores versus final proposal scores","conf":"high","self":false,"doi":"10.1002/asi.24617","ciLow":null,"ciHigh":null,"mt":"other","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"The mean of the three reviewers' individual scores per proposal correlated at 0.91 with the final proposal scores that determine funding.","vf":"unverified"},{"key":"JMK4CLUJ","au":"Seeber, Marco","y":2022,"cx":"Grant","ob":"grant-proposal","fam":"G-theory","form":"proportion of evaluation-score variance at the proposal level, three-level cross-classified null model","estd":"G-theory","v":0.1439,"n":"1928","k":"3","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-65 total score (sum of 13 criteria)","field":"multi-field","wr":"reviewers on grant proposal scores","conf":"high","self":false,"doi":"10.1002/asi.24617","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"In a three-level cross-classified model of 5,330 individual evaluations of 1,928 COST proposals by 3,050 reviewers, 14.4% of the total score variance lay at the proposal level, so a single reviewer's score carries limited reliability for distinguishing proposals.","vf":"unverified"},{"key":"JMK4CLUJ","au":"Seeber, Marco","y":2022,"cx":"Grant","ob":"grant-proposal","fam":"G-theory","form":"proportion of evaluation-score variance at the evaluation (residual) level, three-level cross-classified null model","estd":"G-theory","v":0.687,"n":"1928","k":"3","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-65 total score (sum of 13 criteria)","field":"multi-field","wr":"residual disagreement between reviewers on the same proposal","conf":"high","self":false,"doi":"10.1002/asi.24617","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"In the same decomposition, 68.7% of the total variance sat at the evaluation level, the residual disagreement between reviewers scoring the same proposal that is not explained by proposal or reviewer.","vf":"unverified"},{"key":"JMK4CLUJ","au":"Seeber, Marco","y":2022,"cx":"Grant","ob":"grant-proposal","fam":"G-theory","form":"proportion of evaluation-score variance at the reviewer level, three-level cross-classified null model","estd":"G-theory","v":0.1692,"n":"1928","k":"3","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0-65 total score (sum of 13 criteria)","field":"multi-field","wr":"systematic between-reviewer differences in proposal scores","conf":"high","self":false,"doi":"10.1002/asi.24617","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"In the same variance decomposition, 16.9% of the total evaluation-score variance was attributable to systematic differences between reviewers, reflecting reviewer leniency or severity effects.","vf":"unverified"},{"key":"QR68P67D","au":"Seeber, Marco","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"Average ICC (ICC (1, k))","estd":"ICC (average)","v":0.22,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1 (poor) to 7 (excellent)","field":"health research (multi-field)","wr":"reviewers scoring long grant proposals","conf":"high","self":true,"doi":"10.1093/reseval/rvae020","ciLow":0.08,"ciHigh":0.24,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Average-measures ICC for the mean of the two reviewers rating long grant proposals in the one-stage procedure; 0.22 reflects the reliability of the two reviewers' averaged score.","vf":"unverified"},{"key":"QR68P67D","au":"Seeber, Marco","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"Single Intra-Class Correlation coefficient (ICC (1,1))","estd":"ICC (single/unspec)","v":0.12,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1 (poor) to 7 (excellent)","field":"health research (multi-field)","wr":"reviewers scoring long grant proposals","conf":"high","self":true,"doi":"10.1093/reseval/rvae020","ciLow":0.04,"ciHigh":0.2,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Pairs of reviewers independently scored long grant proposals on a 1-7 scale in Foundation Dam's one-stage procedure; a single-rater ICC of 0.12 indicates poor reliability of one reviewer's judgment.","vf":"unverified"},{"key":"QR68P67D","au":"Seeber, Marco","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"absolute value of the score difference between all pairs of reviewers rating the same proposal","estd":"other","v":1.304,"n":"542","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1 (poor) to 7 (excellent)","field":"health research (multi-field)","wr":"reviewer pairs on long grant proposals","conf":"high","self":true,"doi":"10.1093/reseval/rvae020","ciLow":1.21,"ciHigh":1.399,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Mean absolute difference between the scores of the two reviewers rating the same long proposal in the one-stage procedure; 1.304 points on a 1-7 scale indicates substantial disagreement.","vf":"unverified"},{"key":"QR68P67D","au":"Seeber, Marco","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"Standard Deviation (SD) in the reviewers' scores","estd":"other","v":1.115,"n":"542","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1 (poor) to 7 (excellent)","field":"health research (multi-field)","wr":"reviewers scoring long grant proposals","conf":"high","self":true,"doi":"10.1093/reseval/rvae020","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Standard deviation of the reviewers' scores for the same one-stage long proposals, used by the authors as an agreement proxy; the larger SD of 1.115 indicates more spread between reviewers.","vf":"unverified"},{"key":"QR68P67D","au":"Seeber, Marco","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"mean absolute change in score from the first to the second stage","estd":"other","v":0.39,"n":"184","k":"5","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1 (poor) to 7 (excellent)","field":"health research (multi-field)","wr":"proposal average scores, short vs long stage","conf":"high","self":true,"doi":"10.1093/reseval/rvae020","ciLow":null,"ciHigh":null,"mt":"other","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"For the 184 proposals evaluated at both stages, the mean absolute change between the stage-one short-proposal average score and the stage-two long-proposal average score was 0.39 points on a 1-7 scale.","vf":"unverified"},{"key":"QR68P67D","au":"Seeber, Marco","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Spearman rank correlation","estd":"correlation","v":-0.05,"n":"90","k":"5","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1 (poor) to 7 (excellent)","field":"health research (multi-field)","wr":"proposal rankings, short (stage one) vs long (stage two)","conf":"high","self":true,"doi":"10.1093/reseval/rvae020","ciLow":null,"ciHigh":null,"mt":"other","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Spearman correlation between a proposal's average-score ranking as a short proposal at stage one and as a long proposal at stage two in 2020; -0.05 shows the rankings changed substantially between formats.","vf":"unverified"},{"key":"QR68P67D","au":"Seeber, Marco","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Spearman rank correlation","estd":"correlation","v":0.16,"n":"94","k":"5","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1 (poor) to 7 (excellent)","field":"health research (multi-field)","wr":"proposal rankings, short (stage one) vs long (stage two)","conf":"high","self":true,"doi":"10.1093/reseval/rvae020","ciLow":null,"ciHigh":null,"mt":"other","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Spearman correlation between a proposal's stage-one short-proposal ranking and stage-two long-proposal ranking in 2021; 0.16 shows only weak stability of the ranking across formats.","vf":"unverified"},{"key":"QR68P67D","au":"Seeber, Marco","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"Average ICC (ICC (1, k))","estd":"ICC (average)","v":0.37,"n":"","k":"5","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1 (poor) to 7 (excellent)","field":"health research (multi-field)","wr":"reviewers scoring long grant proposals (stage two)","conf":"high","self":true,"doi":"10.1093/reseval/rvae020","ciLow":0.22,"ciHigh":0.51,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Average-measures ICC for the mean of five reviewers scoring long proposals at stage two, on a restricted top subsample; 0.37 reflects the averaged reliability under a restricted quality range.","vf":"unverified"},{"key":"QR68P67D","au":"Seeber, Marco","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"Single Intra-Class Correlation coefficient (ICC (1,1))","estd":"ICC (single/unspec)","v":0.11,"n":"","k":"5","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1 (poor) to 7 (excellent)","field":"health research (multi-field)","wr":"reviewers scoring long grant proposals (stage two)","conf":"high","self":true,"doi":"10.1093/reseval/rvae020","ciLow":0.05,"ciHigh":0.17,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Single-rater ICC for five reviewers scoring long proposals at stage two, on a restricted top subsample of high-quality proposals; 0.11 is attenuated by the restricted quality range.","vf":"unverified"},{"key":"QR68P67D","au":"Seeber, Marco","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"absolute value of the score difference between all pairs of reviewers rating the same proposal","estd":"other","v":0.907,"n":"184","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1 (poor) to 7 (excellent)","field":"health research (multi-field)","wr":"reviewer pairs on long grant proposals (stage two)","conf":"high","self":true,"doi":"10.1093/reseval/rvae020","ciLow":0.872,"ciHigh":0.942,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Mean absolute difference between reviewer pairs' scores of the same long proposal at stage two, on a restricted top subsample; 0.907 is the smallest average distance.","vf":"unverified"},{"key":"QR68P67D","au":"Seeber, Marco","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"Standard Deviation (SD) in the reviewers' scores","estd":"other","v":0.69,"n":"184","k":"5","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1 (poor) to 7 (excellent)","field":"health research (multi-field)","wr":"reviewers scoring long grant proposals (stage two)","conf":"high","self":true,"doi":"10.1093/reseval/rvae020","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Standard deviation of the reviewers' scores for the same two-stage long proposals, on a restricted top subsample; 0.690 is the smallest SD, indicating the least spread between reviewers.","vf":"unverified"},{"key":"QR68P67D","au":"Seeber, Marco","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"Average ICC (ICC (1, k))","estd":"ICC (average)","v":0.55,"n":"","k":"5","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"1 (poor) to 7 (excellent)","field":"health research (multi-field)","wr":"reviewers scoring short grant proposals","conf":"high","self":true,"doi":"10.1093/reseval/rvae020","ciLow":0.49,"ciHigh":0.6,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Average-measures ICC for the mean of five reviewers rating short proposals in the two-stage procedure; 0.55 partly reflects the greater number of reviewers averaged.","vf":"unverified"},{"key":"QR68P67D","au":"Seeber, Marco","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"Single Intra-Class Correlation coefficient (ICC (1,1))","estd":"ICC (single/unspec)","v":0.2,"n":"","k":"5","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1 (poor) to 7 (excellent)","field":"health research (multi-field)","wr":"reviewers scoring short grant proposals","conf":"high","self":true,"doi":"10.1093/reseval/rvae020","ciLow":0.16,"ciHigh":0.23,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Groups of five reviewers independently scored short grant proposals in the two-stage procedure's first stage; the single-rater ICC of 0.20 is the paper's headline result, higher than for one-stage long proposals.","vf":"unverified"},{"key":"QR68P67D","au":"Seeber, Marco","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"absolute value of the score difference between all pairs of reviewers rating the same proposal","estd":"other","v":0.944,"n":"662","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1 (poor) to 7 (excellent)","field":"health research (multi-field)","wr":"reviewer pairs on short grant proposals","conf":"high","self":true,"doi":"10.1093/reseval/rvae020","ciLow":0.925,"ciHigh":0.964,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Mean absolute difference between reviewer pairs' scores of the same short proposal in the two-stage procedure; 0.944 is smaller than the one-stage value, showing greater agreement.","vf":"unverified"},{"key":"QR68P67D","au":"Seeber, Marco","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"Standard Deviation (SD) in the reviewers' scores","estd":"other","v":0.725,"n":"662","k":"5","samp":"full-pool","blind":"single","agg":"single-rater","scale":"1 (poor) to 7 (excellent)","field":"health research (multi-field)","wr":"reviewers scoring short grant proposals","conf":"high","self":true,"doi":"10.1093/reseval/rvae020","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Standard deviation of the reviewers' scores for the same two-stage short proposals, used as an agreement proxy; 0.725 is smaller than the one-stage SD, indicating less spread. Three to five reviewers per proposal after conflict exemptions.","vf":"unverified"},{"key":"HPGFXDLZ","au":"Severin, Anna","y":2023,"cx":"Journal","ob":"review-report","fam":"Krippendorff","form":"Krippendorff's alpha, average across the 8 content categories","estd":"Krippendorff","v":0.7,"n":"","k":"2","samp":"special","blind":"unclear","agg":"unspecified","scale":"1 for yes, 0 for no (per category)","field":"biomedical","wr":"two coders labelling review-report sentences by content category","conf":"med","self":false,"doi":"10.1371/journal.pbio.3002238","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":true,"he":false,"ms":"Two trained coders independently classified sentences from biomedical peer review reports into eight thoroughness and helpfulness categories derived from a review-quality instrument; the average Krippendorff's alpha across the eight categories was 0.70, indicating acceptable to good agreement between the coders.","vf":"unverified"},{"key":"NB4IK4KH","au":"Shah, Nihar B.","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"other","form":"agreement-to-disagreement ratio","estd":"other","v":5,"n":"2425","k":"","samp":"re-reviewed-subset","blind":"double","agg":"unspecified","scale":"total ordering compared with accept or reject","field":"machine learning","wr":"reviewer rankings versus final acceptance decisions","conf":"med","self":false,"doi":"10.48550/arxiv.1708.09794","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Reviewer rankings of accepted-rejected paper pairs were compared with the final conference decisions. The paper reports roughly five concordant orderings for every discordant ordering.","vf":"unverified"},{"key":"NB4IK4KH","au":"Shah, Nihar B.","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"other","form":"relative-order disagreement with final decisions","estd":"other","v":0.17,"n":"2425","k":"","samp":"re-reviewed-subset","blind":"double","agg":"unspecified","scale":"total ordering compared with accept or reject","field":"machine learning","wr":"reviewer rankings versus final acceptance decisions","conf":"med","self":false,"doi":"10.48550/arxiv.1708.09794","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Across all papers, ordinal rankings disagreed with final decisions in 16 to 17 per cent of comparisons across the overall and reviewer-pool categories. This row records the maximum.","vf":"unverified"},{"key":"NB4IK4KH","au":"Shah, Nihar B.","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"other","form":"relative-order disagreement with final decisions","estd":"other","v":0.16,"n":"2425","k":"","samp":"re-reviewed-subset","blind":"double","agg":"unspecified","scale":"total ordering compared with accept or reject","field":"machine learning","wr":"reviewer rankings versus final acceptance decisions","conf":"med","self":false,"doi":"10.48550/arxiv.1708.09794","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Across all papers, ordinal rankings disagreed with final decisions in 16 to 17 per cent of comparisons across the overall and reviewer-pool categories. This row records the minimum.","vf":"unverified"},{"key":"NB4IK4KH","au":"Shah, Nihar B.","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"other","form":"relative-order disagreement with final decisions","estd":"other","v":0.28,"n":"","k":"","samp":"re-reviewed-subset","blind":"double","agg":"unspecified","scale":"total ordering compared with accept or reject","field":"machine learning","wr":"reviewer rankings versus final acceptance decisions","conf":"med","self":false,"doi":"10.48550/arxiv.1708.09794","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"restricted-other","pr":false,"he":false,"ms":"Within the restricted top 2k set, ordinal rankings disagreed with final decisions in 27 to 28 per cent of comparisons. This row records the maximum.","vf":"unverified"},{"key":"NB4IK4KH","au":"Shah, Nihar B.","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"other","form":"relative-order disagreement with final decisions","estd":"other","v":0.27,"n":"","k":"","samp":"re-reviewed-subset","blind":"double","agg":"unspecified","scale":"total ordering compared with accept or reject","field":"machine learning","wr":"reviewer rankings versus final acceptance decisions","conf":"med","self":false,"doi":"10.48550/arxiv.1708.09794","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"restricted-other","pr":false,"he":false,"ms":"Within the restricted top 2k set, ordinal rankings disagreed with final decisions in 27 to 28 per cent of comparisons. This row records the minimum.","vf":"unverified"},{"key":"NB4IK4KH","au":"Shah, Nihar B.","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"percent-agreement","form":"exact agreement between median cardinal score and ordinal ranking","estd":"percent agreement","v":0.9,"n":"","k":"1","samp":"re-reviewed-subset","blind":"double","agg":"unspecified","scale":"total ordering and median of four 1-5 scores","field":"machine learning","wr":"reviewers' ordinal and median cardinal paper judgments","conf":"med","self":false,"doi":"10.48550/arxiv.1708.09794","ciLow":null,"ciHigh":null,"mt":"other","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Reviewers later ranked papers they had previously scored on four features. Within the same reviewer, the ordering implied by the median cardinal score agreed with the ordinal ranking in about 90 per cent of paper pairs.","vf":"unverified"},{"key":"NB4IK4KH","au":"Shah, Nihar B.","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"percent-agreement","form":"fraction of reviewer-pair agreements on relative ordering of paper-pairs by mean score; exact directional agreement, ties discarded","estd":"percent agreement","v":0.68,"n":"2425","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5, 5=award level; agreement on mean across 4 features","field":"machine learning","wr":"reviewers on pairwise ranking of paper mean scores","conf":"med","self":false,"doi":"10.48550/arxiv.1708.09794","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":false,"ms":"Across all NIPS 2016 submissions, pairs of reviewers who reviewed the same two papers agreed on which paper was better, by mean cardinal score, in 68 per cent of the 1087 comparison pairs; agreement fell towards chance for middle-ranked papers.","vf":"unverified"},{"key":"NB4IK4KH","au":"Shah, Nihar B.","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"percent-agreement","form":"fraction of reviewer-pair agreements on relative ordering of paper-pairs by mean score; exact directional agreement, ties discarded","estd":"percent agreement","v":0.8,"n":"","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5, 5=award level; agreement on mean across 4 features","field":"machine learning","wr":"reviewers on relative ordering of middle papers","conf":"med","self":false,"doi":"10.48550/arxiv.1708.09794","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Pairs of reviewers compared the papers remaining after removing the top 85 per cent and bottom 5 per cent by mean score. Their relative ordering agreed in 80 per cent of ten comparisons, the maximum cell of Figure 16.","vf":"unverified"},{"key":"NB4IK4KH","au":"Shah, Nihar B.","y":2017,"cx":"Journal","ob":"conference-abstract","fam":"percent-agreement","form":"fraction of reviewer-pair agreements on relative ordering of paper-pairs by mean score; exact directional agreement, ties discarded","estd":"percent agreement","v":0.2,"n":"","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"1-5, 5=award level; agreement on mean across 4 features","field":"machine learning","wr":"reviewers on relative ordering of middle papers","conf":"med","self":false,"doi":"10.48550/arxiv.1708.09794","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Pairs of reviewers compared the papers remaining after removing the top 50 per cent and bottom 40 per cent by mean score. Their relative ordering agreed in 20 per cent of five comparisons, the minimum cell of Figure 16.","vf":"unverified"},{"key":"M72NE68B","au":"Shcherbiak, Anna","y":2024,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"r (type unspecified; t-test reported)","estd":"correlation","v":0.08,"n":"245","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"0-10, 'very unlikely' to 'very likely'","field":"judgment and decision-making","wr":"AI and human reviewer on likelihood abstract is AI-generated","conf":"med","self":false,"doi":"10.1017/jdm.2024.24","ciLow":-0.04,"ciHigh":0.21,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"The AI reviewer and one randomly selected human reviewer's AI-likelihood ratings correlated at only r = 0.08 across 245 abstracts, a very small and non-significant association.","vf":"unverified"},{"key":"M72NE68B","au":"Shcherbiak, Anna","y":2024,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"Cohen's Kappa (reported as percent agreement; labelling ambiguous)","estd":"kappa","v":0.176,"n":"245","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"0-10, 'very unlikely' to 'very likely'","field":"judgment and decision-making","wr":"AI and human reviewer on likelihood abstract is AI-generated","conf":"med","self":false,"doi":"10.1017/jdm.2024.24","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"A Cohen's kappa between the AI reviewer and one randomly selected human reviewer's AI-likelihood assessments of 245 abstracts was reported as 17.6% agreement, similar to the human-human figure.","vf":"unverified"},{"key":"M72NE68B","au":"Shcherbiak, Anna","y":2024,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"r (type unspecified; t-test reported)","estd":"correlation","v":0.23,"n":"245","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"0-10, 'very poor quality' to 'very high quality'","field":"judgment and decision-making","wr":"AI and human reviewer on abstract quality","conf":"med","self":false,"doi":"10.1017/jdm.2024.24","ciLow":0.11,"ciHigh":0.35,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The AI reviewer (GPT-4) and a human reviewer independently rated abstract quality; across 245 abstracts their scores correlated at r = 0.23, a small-to-medium agreement comparable to human-human pairs.","vf":"unverified"},{"key":"M72NE68B","au":"Shcherbiak, Anna","y":2024,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"Cohen's Kappa (reported as percent agreement; labelling ambiguous)","estd":"kappa","v":0.22,"n":"245","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"0-10, 'very poor quality' to 'very high quality'","field":"judgment and decision-making","wr":"AI and human reviewer on abstract quality","conf":"med","self":false,"doi":"10.1017/jdm.2024.24","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"A Cohen's kappa between the AI reviewer and one randomly selected human reviewer's quality assessments of 245 abstracts was reported as 22% agreement, very similar to the human-human figure.","vf":"unverified"},{"key":"M72NE68B","au":"Shcherbiak, Anna","y":2024,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"r (type unspecified; t-test reported)","estd":"correlation","v":0.33,"n":"231","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"0-10, 'very unlikely' to 'very likely'","field":"judgment and decision-making","wr":"human reviewers on likelihood abstract is AI-generated","conf":"med","self":false,"doi":"10.1017/jdm.2024.24","ciLow":0.21,"ciHigh":0.44,"mt":"inter-rater","tgt":"other","rr":"none","pr":true,"he":false,"ms":"Two human reviewers independently rated how likely each abstract was AI-generated on a 0-10 scale; across 231 abstracts their ratings correlated at r = 0.33, a medium level of agreement on this authorship judgment.","vf":"unverified"},{"key":"M72NE68B","au":"Shcherbiak, Anna","y":2024,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"Cohen's Kappa (reported as percent agreement; labelling ambiguous)","estd":"kappa","v":0.16,"n":"231","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"0-10, 'very unlikely' to 'very likely'","field":"judgment and decision-making","wr":"human reviewers on likelihood abstract is AI-generated","conf":"med","self":false,"doi":"10.1017/jdm.2024.24","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"A Cohen's kappa between paired human reviewers' AI-likelihood assessments of 231 abstracts was reported as 16% agreement, indicating low chance-corrected agreement on this authorship judgment.","vf":"unverified"},{"key":"M72NE68B","au":"Shcherbiak, Anna","y":2024,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"r (type unspecified; t-test reported)","estd":"correlation","v":0.43,"n":"","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"0-10, 'very poor quality' to 'very high quality'","field":"judgment and decision-making","wr":"junior human reviewers on abstract quality","conf":"med","self":false,"doi":"10.1017/jdm.2024.24","ciLow":0.23,"ciHigh":0.59,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Among pairs of junior reviewers, independent quality ratings of abstracts correlated at r = 0.43, similar to the overall human-human agreement.","vf":"unverified"},{"key":"M72NE68B","au":"Shcherbiak, Anna","y":2024,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"r (type unspecified; t-test reported)","estd":"correlation","v":0.38,"n":"231","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"0-10, 'very poor quality' to 'very high quality'","field":"judgment and decision-making","wr":"human reviewers on conference abstract quality","conf":"med","self":false,"doi":"10.1017/jdm.2024.24","ciLow":0.26,"ciHigh":0.48,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Pairs of human reviewers independently rated the research quality of 231 conference abstracts on a 0-10 scale; their scores correlated at r = 0.38, indicating limited agreement between reviewers.","vf":"unverified"},{"key":"M72NE68B","au":"Shcherbiak, Anna","y":2024,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"r (type unspecified; t-test reported)","estd":"correlation","v":0.46,"n":"","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"0-10, 'very poor quality' to 'very high quality'","field":"judgment and decision-making","wr":"senior human reviewers on abstract quality","conf":"med","self":false,"doi":"10.1017/jdm.2024.24","ciLow":0.31,"ciHigh":0.59,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Among pairs of senior reviewers, independent quality ratings of abstracts correlated at r = 0.46, similar to the overall human-human agreement.","vf":"unverified"},{"key":"M72NE68B","au":"Shcherbiak, Anna","y":2024,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"Cohen's Kappa (reported as percent agreement; labelling ambiguous)","estd":"kappa","v":0.208,"n":"231","k":"2","samp":"special","blind":"double","agg":"single-rater","scale":"0-10, 'very poor quality' to 'very high quality'","field":"judgment and decision-making","wr":"human reviewers on abstract quality","conf":"med","self":false,"doi":"10.1017/jdm.2024.24","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"A Cohen's kappa between paired human reviewers' quality assessments of 231 abstracts was reported as 20.8% agreement, indicating low chance-corrected agreement; the paper's wording mixes kappa and percentage labels.","vf":"unverified"},{"key":"3KXX595H","au":"Snell, Richard R","y":2015,"cx":"Grant","ob":"fellowship","fam":"percent-agreement","form":"exact agreement (final score unchanged from initial pre-score)","estd":"percent agreement","v":0.7,"n":"","k":"","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"0 (least competitive) to 4.9 (most competitive)","field":"biomedical","wr":"reviewers' pre vs final scores on fellowship applications","conf":"med","self":false,"doi":"10.1371/journal.pone.0120838","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"Committee reviewers in a CIHR post-doctoral fellowship competition submitted independent pre-scores and could revise them after optional electronic discussion; about 70 per cent of final scores were unchanged from the initial pre-scores, an approximate figure indicating high within-reviewer score stability.","vf":"unverified"},{"key":"QPIKNW58","au":"Solans‐Domènech, Maite","y":2017,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"exact agreement (unchanged category between first and second assessment)","estd":"percent agreement","v":0.815,"n":"2256","k":"1","samp":"full-pool","blind":"double","agg":"single-rater","scale":"recommended / recommended with reservations / questionable / not recommended","field":"biomedical","wr":"reviewers on grant proposal recommendation categories","conf":"high","self":false,"doi":"10.1093/reseval/rvx021","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Across 5,002 evaluations of 2,256 proposals, reviewers gave the same categorical recommendation on the blinded and the later unblinded assessment in 81.5% of cases, an exact agreement rate between a reviewer's two assessments of the same proposal.","vf":"unverified"},{"key":"QPIKNW58","au":"Solans‐Domènech, Maite","y":2017,"cx":"Grant","ob":"grant-proposal","fam":"weighted-kappa","form":"weighted Kappa statistic","estd":"weighted kappa","v":0.75,"n":"2256","k":"1","samp":"full-pool","blind":"double","agg":"single-rater","scale":"recommended / recommended with reservations / questionable / not recommended","field":"biomedical","wr":"reviewers on grant proposal recommendation categories","conf":"high","self":false,"doi":"10.1093/reseval/rvx021","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"none","pr":true,"he":false,"ms":"The same international reviewers assessed 2,256 biomedical grant proposals twice, first with the applicant's identity concealed and then revealed about one to two weeks later, across 5,002 evaluations. A weighted kappa of 0.75 shows very good agreement between a reviewer's blinded and unblinded categorical recommendations.","vf":"unverified"},{"key":"WNJ55J6C","au":"Sorrell, Lexy","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraclass correlation (ICC); Table 2 column headed 'Average ICC'","estd":"ICC (single/unspec)","v":0.35,"n":"40","k":"4","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"1-6, 1 = extremely poor and unsupportable, 6 = excellent","field":"applied health research","wr":"external reviewers scoring NIHR funding applications","conf":"high","self":false,"doi":"10.1136/bmjopen-2018-022547","ciLow":-0.05,"ciHigh":0.63,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Forty NIHR funding applications were each scored by four external reviewers on a one to six scale. The intraclass correlation of 0.35 indicates low agreement between the reviewers of an application.","vf":"unverified"},{"key":"WNJ55J6C","au":"Sorrell, Lexy","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraclass correlation (ICC); Table 2 column headed 'Average ICC'","estd":"ICC (single/unspec)","v":0.35,"n":"90","k":"5","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"1-6, 1 = extremely poor and unsupportable, 6 = excellent","field":"applied health research","wr":"external reviewers scoring NIHR funding applications","conf":"high","self":false,"doi":"10.1136/bmjopen-2018-022547","ciLow":0.11,"ciHigh":0.54,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":true,"ms":"Ninety NIHR funding applications, the largest group in the study, were each scored by five external reviewers on a one to six scale. The intraclass correlation of 0.35 indicates low agreement between the reviewers of an application.","vf":"unverified"},{"key":"WNJ55J6C","au":"Sorrell, Lexy","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraclass correlation (ICC); Table 2 column headed 'Average ICC'","estd":"ICC (single/unspec)","v":0.18,"n":"82","k":"6","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"1-6, 1 = extremely poor and unsupportable, 6 = excellent","field":"applied health research","wr":"external reviewers scoring NIHR funding applications","conf":"high","self":false,"doi":"10.1136/bmjopen-2018-022547","ciLow":-0.13,"ciHigh":0.43,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Eighty-two NIHR funding applications were each scored by six external reviewers on a one to six scale. The intraclass correlation of 0.18 is the lowest reported and indicates very little agreement between the reviewers of an application.","vf":"unverified"},{"key":"WNJ55J6C","au":"Sorrell, Lexy","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraclass correlation (ICC); Table 2 column headed 'Average ICC'","estd":"ICC (single/unspec)","v":0.41,"n":"51","k":"7","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"1-6, 1 = extremely poor and unsupportable, 6 = excellent","field":"applied health research","wr":"external reviewers scoring NIHR funding applications","conf":"high","self":false,"doi":"10.1136/bmjopen-2018-022547","ciLow":0.12,"ciHigh":0.63,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Fifty-one NIHR funding applications were each scored by seven external reviewers on a one to six scale. The intraclass correlation of 0.41 is the highest reported in the study but still only fair agreement between reviewers.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"Cronbach-alpha","form":"Cronbach alpha coefficient for the average ratings of the 10 reviewers","estd":"Cronbach alpha","v":0.69,"n":"35","k":"10","samp":"full-pool","blind":"double","agg":"average-of-k","scale":"-6 (unacceptable) to +6 (excellent), sum of six criteria","field":"plastic surgery","wr":"ten reviewers on aesthetic surgery abstract scores","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Cronbach alpha across the 10 reviewers' scores for the 35 aesthetic surgery abstracts was 0.69, the highest of the three categories, describing the reliability of the averaged rating.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"Cronbach-alpha","form":"Cronbach alpha coefficient for the average ratings of the 10 reviewers","estd":"Cronbach alpha","v":0.6,"n":"220","k":"10","samp":"full-pool","blind":"double","agg":"average-of-k","scale":"-6 (unacceptable) to +6 (excellent), sum of six criteria","field":"plastic surgery","wr":"ten reviewers on all 220 abstract scores","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Cronbach alpha was computed across the 10 reviewers' summed scores for all 220 abstracts, treating each reviewer as an item. The value of 0.60 describes the reliability of the average of the 10 reviewers' ratings.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"Cronbach-alpha","form":"Cronbach alpha coefficient for the average ratings of the 10 reviewers","estd":"Cronbach alpha","v":0.53,"n":"65","k":"10","samp":"full-pool","blind":"double","agg":"average-of-k","scale":"-6 (unacceptable) to +6 (excellent), sum of six criteria","field":"plastic surgery","wr":"ten reviewers on basic research abstract scores","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Cronbach alpha across the 10 reviewers' scores for the 65 basic research abstracts was 0.53, the lowest of the three categories.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"Cronbach-alpha","form":"Cronbach alpha coefficient for the average ratings of the 10 reviewers","estd":"Cronbach alpha","v":0.59,"n":"120","k":"10","samp":"full-pool","blind":"double","agg":"average-of-k","scale":"-6 (unacceptable) to +6 (excellent), sum of six criteria","field":"plastic surgery","wr":"ten reviewers on clinical study abstract scores","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Cronbach alpha across the 10 reviewers' scores for the 120 clinical study abstracts was 0.59, describing the reliability of the averaged rating.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"Pearson's coefficient of correlation between one reviewer's scores and the total scores by the other nine reviewers","estd":"correlation","v":0.42,"n":"220","k":"10","samp":"full-pool","blind":"double","agg":"unspecified","scale":"-6 (unacceptable) to +6 (excellent), sum of six criteria","field":"plastic surgery","wr":"one reviewer versus nine co-reviewers on abstract scores","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Reviewer A's scores for all 220 abstracts correlated 0.42 with the summed scores of the other nine reviewers, one of the two highest such correlations in the panel.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"Pearson's coefficient of correlation between one reviewer's scores and the total scores by the other nine reviewers","estd":"correlation","v":0.34,"n":"220","k":"10","samp":"full-pool","blind":"double","agg":"unspecified","scale":"-6 (unacceptable) to +6 (excellent), sum of six criteria","field":"plastic surgery","wr":"one reviewer versus nine co-reviewers on abstract scores","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Reviewer B's scores for all 220 abstracts correlated 0.34 with the summed scores of the other nine reviewers.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"Pearson's coefficient of correlation between one reviewer's scores and the total scores by the other nine reviewers","estd":"correlation","v":0.38,"n":"220","k":"10","samp":"full-pool","blind":"double","agg":"unspecified","scale":"-6 (unacceptable) to +6 (excellent), sum of six criteria","field":"plastic surgery","wr":"one reviewer versus nine co-reviewers on abstract scores","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Reviewer C's scores for all 220 abstracts correlated 0.38 with the summed scores of the other nine reviewers.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"Pearson's coefficient of correlation between one reviewer's scores and the total scores by the other nine reviewers","estd":"correlation","v":0.21,"n":"220","k":"10","samp":"full-pool","blind":"double","agg":"unspecified","scale":"-6 (unacceptable) to +6 (excellent), sum of six criteria","field":"plastic surgery","wr":"one reviewer versus nine co-reviewers on abstract scores","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Reviewer D's scores for all 220 abstracts correlated 0.21 with the summed scores of the other nine reviewers, the weakest agreement in the panel.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"Pearson's coefficient of correlation between one reviewer's scores and the total scores by the other nine reviewers","estd":"correlation","v":0.29,"n":"220","k":"10","samp":"full-pool","blind":"double","agg":"unspecified","scale":"-6 (unacceptable) to +6 (excellent), sum of six criteria","field":"plastic surgery","wr":"one reviewer versus nine co-reviewers on abstract scores","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Reviewer E's scores for all 220 abstracts correlated 0.29 with the summed scores of the other nine reviewers.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"Pearson's coefficient of correlation between one reviewer's scores and the total scores by the other nine reviewers","estd":"correlation","v":0.4,"n":"220","k":"10","samp":"full-pool","blind":"double","agg":"unspecified","scale":"-6 (unacceptable) to +6 (excellent), sum of six criteria","field":"plastic surgery","wr":"one reviewer versus nine co-reviewers on abstract scores","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Reviewer F's scores for all 220 abstracts correlated 0.40 with the summed scores of the other nine reviewers. Reviewer F was the only one whose score distribution was uniform.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"Pearson's coefficient of correlation between one reviewer's scores and the total scores by the other nine reviewers","estd":"correlation","v":0.26,"n":"220","k":"10","samp":"full-pool","blind":"double","agg":"unspecified","scale":"-6 (unacceptable) to +6 (excellent), sum of six criteria","field":"plastic surgery","wr":"one reviewer versus nine co-reviewers on abstract scores","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Reviewer G's scores for all 220 abstracts correlated 0.26 with the summed scores of the other nine reviewers.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"Pearson's coefficient of correlation between one reviewer's scores and the total scores by the other nine reviewers","estd":"correlation","v":0.38,"n":"220","k":"10","samp":"full-pool","blind":"double","agg":"unspecified","scale":"-6 (unacceptable) to +6 (excellent), sum of six criteria","field":"plastic surgery","wr":"one reviewer versus nine co-reviewers on abstract scores","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Reviewer H's scores for all 220 abstracts correlated 0.38 with the summed scores of the other nine reviewers.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"Pearson's coefficient of correlation between one reviewer's scores and the total scores by the other nine reviewers","estd":"correlation","v":0.32,"n":"220","k":"10","samp":"full-pool","blind":"double","agg":"unspecified","scale":"-6 (unacceptable) to +6 (excellent), sum of six criteria","field":"plastic surgery","wr":"one reviewer versus nine co-reviewers on abstract scores","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Reviewer I's scores for all 220 abstracts correlated 0.32 with the summed scores of the other nine reviewers.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"Pearson's coefficient of correlation between one reviewer's scores and the total scores by the other nine reviewers","estd":"correlation","v":0.43,"n":"220","k":"10","samp":"full-pool","blind":"double","agg":"unspecified","scale":"-6 (unacceptable) to +6 (excellent), sum of six criteria","field":"plastic surgery","wr":"one reviewer versus nine co-reviewers on abstract scores","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Reviewer J's scores for all 220 abstracts correlated 0.43 with the summed scores of the other nine reviewers, the highest such correlation in the panel.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"Pearson's coefficient of correlation between one reviewer's scores and the total scores by the other nine reviewers","estd":"correlation","v":0.62,"n":"35","k":"10","samp":"full-pool","blind":"double","agg":"unspecified","scale":"-6 (unacceptable) to +6 (excellent), sum of six criteria","field":"plastic surgery","wr":"one reviewer versus nine co-reviewers on abstract scores","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Table III holds 40 reviewer by category correlations, so only the extremes are recorded here. The highest, 0.62, is reviewer A on the 35 aesthetic surgery abstracts, one of two exceeding 0.60.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"correlation","form":"Pearson's coefficient of correlation between one reviewer's scores and the total scores by the other nine reviewers","estd":"correlation","v":0.05,"n":"65","k":"10","samp":"full-pool","blind":"double","agg":"unspecified","scale":"-6 (unacceptable) to +6 (excellent), sum of six criteria","field":"plastic surgery","wr":"one reviewer versus nine co-reviewers on abstract scores","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The lowest of the 40 correlations in Table III, 0.05, is reviewer C on the 65 basic research abstracts, showing almost no relation to the other nine reviewers' summed scores.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"kappa statistic between two individual reviewers","estd":"kappa","v":0.6,"n":"35","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"'accept' or 'reject', dichotomised from the -6 to +6 score","field":"plastic surgery","wr":"one reviewer pair on aesthetic surgery abstracts","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The best agreeing of the 45 reviewer pairs on the 35 aesthetic surgery abstracts reached a kappa of 0.60, which the authors classify as moderate agreement.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"kappa statistic between two individual reviewers","estd":"kappa","v":-0.25,"n":"35","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"'accept' or 'reject', dichotomised from the -6 to +6 score","field":"plastic surgery","wr":"one reviewer pair on aesthetic surgery abstracts","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Table II tabulates all 45 pairwise kappas for the 35 aesthetic surgery abstracts; only the extremes are recorded here. The worst agreeing pair reached -0.25, that is worse than chance agreement.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"average kappa statistic over all 45 combinations of two of 10 reviewers","estd":"kappa","v":0.14,"n":"35","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"'accept' or 'reject', dichotomised from the -6 to +6 score","field":"plastic surgery","wr":"reviewers on aesthetic surgery abstract accept/reject","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 35 aesthetic surgery abstracts, the accept or reject judgments of all 45 reviewer pairs gave a mean kappa of 0.14 (SD 0.18), which the authors read as poor agreement.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"kappa statistic between two individual reviewers","estd":"kappa","v":0.26,"n":"220","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"'accept' or 'reject', dichotomised from the -6 to +6 score","field":"plastic surgery","wr":"one reviewer pair on all 220 abstracts","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The best agreeing of the 45 reviewer pairs on all 220 abstracts reached a kappa of 0.26, still only fair agreement.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"kappa statistic between two individual reviewers","estd":"kappa","v":-0.01,"n":"220","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"'accept' or 'reject', dichotomised from the -6 to +6 score","field":"plastic surgery","wr":"one reviewer pair on all 220 abstracts","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Across the 45 reviewer pairs judging all 220 abstracts, the lowest pairwise kappa was -0.01, that is no better than chance. Only the extremes of the distribution are recorded here.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"average kappa statistic over all 45 combinations of two of 10 reviewers","estd":"kappa","v":0.12,"n":"220","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"'accept' or 'reject', dichotomised from the -6 to +6 score","field":"plastic surgery","wr":"reviewers on conference abstract accept/reject decisions","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Ten scientific committee members each rated all 220 submitted abstracts, and each reviewer's total score was dichotomised into accept or reject. Averaged over all 45 reviewer pairs, kappa was 0.12 (SD across pairs 0.07), showing poor chance corrected agreement.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"kappa statistic between two individual reviewers","estd":"kappa","v":0.34,"n":"65","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"'accept' or 'reject', dichotomised from the -6 to +6 score","field":"plastic surgery","wr":"one reviewer pair on basic research abstracts","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The best agreeing of the 45 reviewer pairs on the 65 basic research abstracts reached a kappa of 0.34, still only fair agreement.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"kappa statistic between two individual reviewers","estd":"kappa","v":-0.22,"n":"65","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"'accept' or 'reject', dichotomised from the -6 to +6 score","field":"plastic surgery","wr":"one reviewer pair on basic research abstracts","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Across the 45 reviewer pairs judging the 65 basic research abstracts, the lowest pairwise kappa was -0.22, that is worse than chance agreement. Only the extremes of the distribution are recorded here.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"average kappa statistic over all 45 combinations of two of 10 reviewers","estd":"kappa","v":0.09,"n":"65","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"'accept' or 'reject', dichotomised from the -6 to +6 score","field":"plastic surgery","wr":"reviewers on basic research abstract accept/reject","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 65 basic research abstracts, the mean kappa across all 45 reviewer pairs on the dichotomised accept or reject judgment was 0.09 (SD 0.14), the lowest of the three categories.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"kappa statistic between two individual reviewers","estd":"kappa","v":0.28,"n":"120","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"'accept' or 'reject', dichotomised from the -6 to +6 score","field":"plastic surgery","wr":"one reviewer pair on clinical study abstracts","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The best agreeing of the 45 reviewer pairs on the 120 clinical study abstracts reached a kappa of 0.28, still only fair agreement.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"kappa statistic between two individual reviewers","estd":"kappa","v":-0.05,"n":"120","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"'accept' or 'reject', dichotomised from the -6 to +6 score","field":"plastic surgery","wr":"one reviewer pair on clinical study abstracts","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Across the 45 reviewer pairs judging the 120 clinical study abstracts, the lowest pairwise kappa was -0.05, marginally worse than chance. Only the extremes of the distribution are recorded here.","vf":"unverified"},{"key":"ZGLDA6ZW","au":"Steen, Lydia P. E. van der","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"average kappa statistic over all 45 combinations of two of 10 reviewers","estd":"kappa","v":0.12,"n":"120","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"'accept' or 'reject', dichotomised from the -6 to +6 score","field":"plastic surgery","wr":"reviewers on clinical study abstract accept/reject","conf":"high","self":false,"doi":"10.1097/01.prs.0000061092.88629.82","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 120 clinical study abstracts, the mean kappa across all 45 reviewer pairs on the dichotomised accept or reject judgment was 0.12 (SD 0.09).","vf":"unverified"},{"key":"PEMI66JZ","au":"Strayhorn, J","y":1993,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"weighted kappa, quadratic weights (Fleiss equations)","estd":"weighted kappa","v":0.17,"n":"268","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 rejection, 1 major rewrite, 2 minor revisions, 3 acceptance","field":"child and adolescent psychiatry","wr":"reviewers on 4-step reject-to-accept rating","conf":"high","self":false,"doi":"10.1176/ajp.150.6.947","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers placed each of 268 manuscripts on a four-step reject-to-accept scale before the intervention; a quadratic-weighted kappa of 0.17 indicates poor agreement.","vf":"unverified"},{"key":"PEMI66JZ","au":"Strayhorn, J","y":1993,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"Cohen's kappa (unweighted)","estd":"kappa","v":0.12,"n":"268","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"rejection coded 0, any other recommendation coded 1","field":"child and adolescent psychiatry","wr":"reviewers on manuscript accept/reject decision","conf":"high","self":false,"doi":"10.1176/ajp.150.6.947","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers classified each of 268 manuscripts as reject versus any other recommendation in the pre-intervention year; a Cohen kappa of 0.12 indicates poor chance-corrected agreement.","vf":"unverified"},{"key":"PEMI66JZ","au":"Strayhorn, J","y":1993,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation","estd":"ICC (single/unspec)","v":0.27,"n":"289","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1-10 quality, mark on a line, nearest tenth","field":"child and adolescent psychiatry","wr":"reviewers on 1-10 manuscript quality rating","conf":"high","self":false,"doi":"10.1176/ajp.150.6.947","ciLow":0.16,"ciHigh":0.38,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated the overall quality of each of 289 manuscripts on a 1-10 line scale; the single-measures intraclass correlation of 0.27 was the most reliable pre-intervention metric.","vf":"unverified"},{"key":"PEMI66JZ","au":"Strayhorn, J","y":1993,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"stepped up by the Spearman-Brown formula","estd":"ICC (average)","v":0.43,"n":"289","k":"2","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-10 quality, mark on a line, nearest tenth","field":"child and adolescent psychiatry","wr":"averaged two reviewers on 1-10 quality rating","conf":"high","self":false,"doi":"10.1176/ajp.150.6.947","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Applying the Spearman-Brown formula to the pre-intervention single-measures quality reliability of 0.27, the predicted reliability of the average of two reviewers was 0.43.","vf":"unverified"},{"key":"PEMI66JZ","au":"Strayhorn, J","y":1993,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"stepped up by the Spearman-Brown formula","estd":"ICC (average)","v":0.53,"n":"289","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"1-10 quality, mark on a line, nearest tenth","field":"child and adolescent psychiatry","wr":"averaged three reviewers on 1-10 quality rating","conf":"high","self":false,"doi":"10.1176/ajp.150.6.947","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Applying the Spearman-Brown formula to the pre-intervention single-measures quality reliability of 0.27, the predicted reliability of the average of three reviewers was 0.53.","vf":"unverified"},{"key":"PEMI66JZ","au":"Strayhorn, J","y":1993,"cx":"Journal","ob":"journal-manuscript","fam":"Cronbach-alpha","form":"coefficient alpha (first rater, listwise deletion, SPSS)","estd":"Cronbach alpha","v":0.88,"n":"36","k":"1","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0-5 per item, fails greatly to succeeds greatly","field":"child and adolescent psychiatry","wr":"internal consistency of case-study rating scale items","conf":"high","self":false,"doi":"10.1176/ajp.150.6.947","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Cronbach's alpha across the items of the case-study version of the rating scale, computed on the first rater's ratings of 36 manuscripts, was 0.88, indicating high internal consistency.","vf":"unverified"},{"key":"PEMI66JZ","au":"Strayhorn, J","y":1993,"cx":"Journal","ob":"journal-manuscript","fam":"Cronbach-alpha","form":"coefficient alpha (first rater, listwise deletion, SPSS)","estd":"Cronbach alpha","v":0.92,"n":"143","k":"1","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0-5 per item, fails greatly to succeeds greatly","field":"child and adolescent psychiatry","wr":"internal consistency of research-article rating scale items","conf":"high","self":false,"doi":"10.1176/ajp.150.6.947","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Cronbach's alpha across the items of the research-article rating scale, computed on the first rater's ratings of 143 manuscripts, was 0.92, indicating high internal consistency of the new instrument.","vf":"unverified"},{"key":"PEMI66JZ","au":"Strayhorn, J","y":1993,"cx":"Journal","ob":"journal-manuscript","fam":"Cronbach-alpha","form":"coefficient alpha (first rater, listwise deletion, SPSS)","estd":"Cronbach alpha","v":0.85,"n":"10","k":"1","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0-5 per item, fails greatly to succeeds greatly","field":"child and adolescent psychiatry","wr":"internal consistency of review-article rating scale items","conf":"high","self":false,"doi":"10.1176/ajp.150.6.947","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Cronbach's alpha across the items of the review-article version of the rating scale, computed on the first rater's ratings of 10 manuscripts, was 0.85, indicating high internal consistency.","vf":"unverified"},{"key":"PEMI66JZ","au":"Strayhorn, J","y":1993,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"weighted kappa, quadratic weights (Fleiss equations)","estd":"weighted kappa","v":0.26,"n":"253","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"0 rejection, 1 major rewrite, 2 minor revisions, 3 acceptance","field":"child and adolescent psychiatry","wr":"reviewers on 4-step reject-to-accept rating","conf":"high","self":false,"doi":"10.1176/ajp.150.6.947","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"In the post-intervention year, two independent reviewers rated 253 manuscripts on the four-step reject-to-accept scale; the quadratic-weighted kappa was 0.26.","vf":"unverified"},{"key":"PEMI66JZ","au":"Strayhorn, J","y":1993,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"Cohen's kappa (unweighted)","estd":"kappa","v":0.27,"n":"253","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"rejection coded 0, any other recommendation coded 1","field":"child and adolescent psychiatry","wr":"reviewers on manuscript accept/reject decision","conf":"high","self":false,"doi":"10.1176/ajp.150.6.947","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"In the post-intervention year, two independent reviewers classified 253 manuscripts as reject versus any other recommendation; the Cohen kappa rose to 0.27, still poor agreement.","vf":"unverified"},{"key":"PEMI66JZ","au":"Strayhorn, J","y":1993,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation","estd":"ICC (single/unspec)","v":0.43,"n":"265","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"average item rating, actual range 0 to 5","field":"child and adolescent psychiatry","wr":"reviewers on averaged new multi-item quality scale","conf":"high","self":false,"doi":"10.1176/ajp.150.6.947","ciLow":0.32,"ciHigh":0.52,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":false,"ms":"Two independent reviewers scored 265 manuscripts on the new multi-item rating scale, averaged across items; the single-measures intraclass correlation of 0.43 was the study's headline post-intervention reliability.","vf":"unverified"},{"key":"PEMI66JZ","au":"Strayhorn, J","y":1993,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"stepped up by the Spearman-Brown formula","estd":"ICC (average)","v":0.6,"n":"265","k":"2","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"average item rating, actual range 0 to 5","field":"child and adolescent psychiatry","wr":"averaged two reviewers on new multi-item scale","conf":"high","self":false,"doi":"10.1176/ajp.150.6.947","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Applying the Spearman-Brown formula to the new scale's single-measures reliability of 0.43, the predicted reliability of the average of two reviewers was 0.60.","vf":"unverified"},{"key":"PEMI66JZ","au":"Strayhorn, J","y":1993,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"stepped up by the Spearman-Brown formula","estd":"ICC (average)","v":0.69,"n":"265","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"average item rating, actual range 0 to 5","field":"child and adolescent psychiatry","wr":"averaged three reviewers on new multi-item scale","conf":"high","self":false,"doi":"10.1176/ajp.150.6.947","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Applying the Spearman-Brown formula to the new scale's single-measures reliability of 0.43, the predicted reliability of the average of three reviewers was 0.69.","vf":"unverified"},{"key":"PEMI66JZ","au":"Strayhorn, J","y":1993,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation","estd":"ICC (single/unspec)","v":0.36,"n":"265","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1-10 quality, mark on a line, nearest tenth","field":"child and adolescent psychiatry","wr":"reviewers on 1-10 manuscript quality rating","conf":"high","self":false,"doi":"10.1176/ajp.150.6.947","ciLow":0.25,"ciHigh":0.46,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"In the post-intervention year, two independent reviewers rated the 1-10 overall quality of 265 manuscripts; the single-measures intraclass correlation was 0.36.","vf":"unverified"},{"key":"QA4UL4N5","au":"Sueur, Helen Le","y":2020,"cx":"Journal","ob":"review-report","fam":"Fleiss-kappa","form":"Fleiss' extended Kappa statistic","estd":"kappa","v":0.336,"n":"","k":"4","samp":"special","blind":"open","agg":"single-rater","scale":"request for additional analysis, yes/no","field":"biomedical","wr":"four evaluators judging whether a report requested additional analysis","conf":"med","self":false,"doi":"10.1080/0142159x.2020.1774527","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Four evaluators independently judged whether each review report requested additional analysis (yes/no); a Fleiss extended kappa of 0.336 indicates fair chance-corrected agreement between them.","vf":"unverified"},{"key":"QA4UL4N5","au":"Sueur, Helen Le","y":2020,"cx":"Journal","ob":"review-report","fam":"ICC","form":"absolute-agreement, two-way random-effects model","estd":"ICC (single/unspec)","v":0.279,"n":"","k":"4","samp":"special","blind":"open","agg":"unspecified","scale":"0 to 5, constructiveness","field":"biomedical","wr":"four evaluators scoring review reports' constructiveness","conf":"med","self":false,"doi":"10.1080/0142159x.2020.1774527","ciLow":0.065,"ciHigh":0.527,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Four evaluators independently scored the constructiveness of review reports drawn from 10 papers; an ICC of 0.279 under an absolute-agreement two-way random-effects model indicates poor agreement between them.","vf":"unverified"},{"key":"QA4UL4N5","au":"Sueur, Helen Le","y":2020,"cx":"Journal","ob":"review-report","fam":"ICC","form":"absolute-agreement, two-way random-effects model","estd":"ICC (single/unspec)","v":0.439,"n":"","k":"4","samp":"special","blind":"open","agg":"unspecified","scale":"0 to 5, harshness","field":"biomedical","wr":"four evaluators scoring review reports' harshness","conf":"med","self":false,"doi":"10.1080/0142159x.2020.1774527","ciLow":0.24,"ciHigh":0.65,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":true,"he":false,"ms":"Four evaluators independently scored the harshness of review reports drawn from 10 papers; an ICC of 0.439 under an absolute-agreement two-way random-effects model indicates fair agreement on the study's focal measure.","vf":"unverified"},{"key":"QA4UL4N5","au":"Sueur, Helen Le","y":2020,"cx":"Journal","ob":"review-report","fam":"ICC","form":"absolute-agreement, two-way random-effects model","estd":"ICC (single/unspec)","v":0.585,"n":"","k":"4","samp":"special","blind":"open","agg":"unspecified","scale":"0 to 5, level of detail","field":"biomedical","wr":"four evaluators scoring review reports' level of detail","conf":"med","self":false,"doi":"10.1080/0142159x.2020.1774527","ciLow":0.37,"ciHigh":0.766,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Four evaluators independently scored the level of detail of review reports drawn from 10 papers; an ICC of 0.585 under an absolute-agreement two-way random-effects model indicates moderate agreement between them.","vf":"unverified"},{"key":"QA4UL4N5","au":"Sueur, Helen Le","y":2020,"cx":"Journal","ob":"review-report","fam":"ICC","form":"absolute-agreement, two-way random-effects model","estd":"ICC (single/unspec)","v":0.479,"n":"","k":"4","samp":"special","blind":"open","agg":"unspecified","scale":"0 to 5, positiveness","field":"biomedical","wr":"four evaluators scoring review reports' positiveness","conf":"med","self":false,"doi":"10.1080/0142159x.2020.1774527","ciLow":0.146,"ciHigh":0.73,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Four evaluators independently scored the positiveness of review reports drawn from 10 papers; an ICC of 0.479 under an absolute-agreement two-way random-effects model indicates fair agreement between them.","vf":"unverified"},{"key":"5PN6GE4C","au":"Sun, Zhuanlan","y":2026,"cx":"Journal","ob":"review-report","fam":"percent-agreement","form":"observed agreement, difference of <=1 point counted as agreement","estd":"percent agreement","v":0.98,"n":"","k":"1","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"mean of seven 1-5 RQI items (overall score)","field":"biomedical","wr":"Gemini re-rating overall RQI score on review reports","conf":"high","self":false,"doi":"10.1016/j.joi.2026.101801","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Gemini 2.0 Flash scored the review reports twice, and the two repetitions agreed at 0.98 under a one-point tolerance, indicating high internal consistency of the model's overall RQI scores. The sample size for this figure is not explicitly stated.","vf":"unverified"},{"key":"5PN6GE4C","au":"Sun, Zhuanlan","y":2026,"cx":"Journal","ob":"review-report","fam":"percent-agreement","form":"observed agreement, difference of <=1 point counted as agreement","estd":"percent agreement","v":0.99,"n":"300","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1-5 Likert, 1 = poor to 5 = excellent","field":"biomedical","wr":"two humans rating RQI constructive comments item on review reports","conf":"high","self":false,"doi":"10.1016/j.joi.2026.101801","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Two trained evaluators independently applied the constructive comments item of the RQI to 300 eLife peer review reports; observed agreement was 0.99 under a one-point tolerance. n_ratings_total derived from stated complete crossing.","vf":"unverified"},{"key":"5PN6GE4C","au":"Sun, Zhuanlan","y":2026,"cx":"Journal","ob":"review-report","fam":"percent-agreement","form":"observed agreement, difference of <=1 point counted as agreement","estd":"percent agreement","v":0.99,"n":"300","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1-5 Likert, 1 = poor to 5 = excellent","field":"biomedical","wr":"two humans rating RQI importance item on review reports","conf":"high","self":false,"doi":"10.1016/j.joi.2026.101801","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Two trained evaluators independently applied the importance item of the RQI to 300 eLife peer review reports; observed agreement was 0.99, treating score differences of one point or less as agreement. n_ratings_total derived from stated complete crossing (2 evaluators by 300 reports).","vf":"unverified"},{"key":"5PN6GE4C","au":"Sun, Zhuanlan","y":2026,"cx":"Journal","ob":"review-report","fam":"percent-agreement","form":"observed agreement, difference of <=1 point counted as agreement","estd":"percent agreement","v":0.99,"n":"300","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1-5 Likert, 1 = poor to 5 = excellent","field":"biomedical","wr":"two humans rating RQI interpretation item on review reports","conf":"high","self":false,"doi":"10.1016/j.joi.2026.101801","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Two trained evaluators independently applied the interpretation item of the RQI to 300 eLife peer review reports; observed agreement was 0.99 under a one-point tolerance. n_ratings_total derived from stated complete crossing.","vf":"unverified"},{"key":"5PN6GE4C","au":"Sun, Zhuanlan","y":2026,"cx":"Journal","ob":"review-report","fam":"percent-agreement","form":"observed agreement, difference of <=1 point counted as agreement","estd":"percent agreement","v":0.99,"n":"300","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1-5 Likert, 1 = poor to 5 = excellent","field":"biomedical","wr":"two humans rating RQI methodology item on review reports","conf":"high","self":false,"doi":"10.1016/j.joi.2026.101801","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Two trained evaluators independently applied the methodology item of the RQI to 300 eLife peer review reports; observed agreement was 0.99 under a one-point tolerance. n_ratings_total derived from stated complete crossing.","vf":"unverified"},{"key":"5PN6GE4C","au":"Sun, Zhuanlan","y":2026,"cx":"Journal","ob":"review-report","fam":"percent-agreement","form":"observed agreement, difference of <=1 point counted as agreement","estd":"percent agreement","v":0.99,"n":"300","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1-5 Likert, 1 = poor to 5 = excellent","field":"biomedical","wr":"two humans rating RQI originality item on review reports","conf":"high","self":false,"doi":"10.1016/j.joi.2026.101801","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Two trained evaluators independently applied the originality item of the RQI to 300 eLife peer review reports; observed agreement was 0.99 under a one-point tolerance. n_ratings_total derived from stated complete crossing.","vf":"unverified"},{"key":"5PN6GE4C","au":"Sun, Zhuanlan","y":2026,"cx":"Journal","ob":"review-report","fam":"percent-agreement","form":"observed agreement, difference of <=1 point counted as agreement","estd":"percent agreement","v":0.99,"n":"300","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1-5 Likert, 1 = poor to 5 = excellent","field":"biomedical","wr":"two humans rating RQI presentation item on review reports","conf":"high","self":false,"doi":"10.1016/j.joi.2026.101801","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Two trained evaluators independently applied the presentation item of the RQI to 300 eLife peer review reports; observed agreement was 0.99 under a one-point tolerance. n_ratings_total derived from stated complete crossing.","vf":"unverified"},{"key":"5PN6GE4C","au":"Sun, Zhuanlan","y":2026,"cx":"Journal","ob":"review-report","fam":"percent-agreement","form":"observed agreement, difference of <=1 point counted as agreement","estd":"percent agreement","v":0.99,"n":"300","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1-5 Likert, 1 = poor to 5 = excellent","field":"biomedical","wr":"two humans rating RQI substantiated comments item on review reports","conf":"high","self":false,"doi":"10.1016/j.joi.2026.101801","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Two trained evaluators independently applied the substantiated comments item of the RQI to 300 eLife peer review reports; observed agreement was 0.99 under a one-point tolerance. n_ratings_total derived from stated complete crossing.","vf":"unverified"},{"key":"5PN6GE4C","au":"Sun, Zhuanlan","y":2026,"cx":"Journal","ob":"review-report","fam":"percent-agreement","form":"observed agreement, difference of <=1 point counted as agreement","estd":"percent agreement","v":0.86,"n":"300","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"mean of seven 1-5 RQI items (overall score)","field":"biomedical","wr":"human vs Gemini overall RQI score on review reports","conf":"high","self":false,"doi":"10.1016/j.joi.2026.101801","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":true,"he":false,"ms":"Agreement between the overall (mean of seven RQI items) review-quality scores from human Evaluator 1 and Gemini 2.0 Flash across 300 eLife peer review reports was 0.86 under a one-point tolerance, the study's headline evidence of human-LLM comparability.","vf":"unverified"},{"key":"5PN6GE4C","au":"Sun, Zhuanlan","y":2026,"cx":"Journal","ob":"review-report","fam":"percent-agreement","form":"observed agreement, difference of <=1 point counted as agreement","estd":"percent agreement","v":0.79,"n":"300","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"mean of seven 1-5 RQI items (overall score)","field":"biomedical","wr":"human vs Gemini overall RQI score on review reports","conf":"high","self":false,"doi":"10.1016/j.joi.2026.101801","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Agreement between the overall review-quality scores from human Evaluator 2 and Gemini 2.0 Flash across 300 eLife peer review reports was 0.79 under a one-point tolerance.","vf":"unverified"},{"key":"KARAZVW9","au":"Tamblyn, Robyn","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraclass correlation coefficient (ICC)","estd":"ICC (single/unspec)","v":0.33,"n":"4174","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (poor) to 4.9 (excellent)","field":"health research","wr":"reviewers on applied-science grant application scores","conf":"high","self":false,"doi":"10.1503/cmaj.170901","ciLow":0.3,"ciHigh":0.36,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For 4174 applied-science applications, the two reviewers' independent scores yielded an ICC of 0.33, indicating poor inter-rater reliability.","vf":"unverified"},{"key":"KARAZVW9","au":"Tamblyn, Robyn","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraclass correlation coefficient (ICC)","estd":"ICC (single/unspec)","v":0.41,"n":"7450","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (poor) to 4.9 (excellent)","field":"health research","wr":"reviewers on basic-science grant application scores","conf":"high","self":false,"doi":"10.1503/cmaj.170901","ciLow":0.39,"ciHigh":0.44,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For 7450 basic-science applications, the two reviewers' independent scores yielded an ICC of 0.41, indicating fair inter-rater reliability.","vf":"unverified"},{"key":"KARAZVW9","au":"Tamblyn, Robyn","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraclass correlation coefficient (ICC)","estd":"ICC (single/unspec)","v":0.33,"n":"1924","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (poor) to 4.9 (excellent)","field":"health research","wr":"reviewers on clinical grant application scores","conf":"high","self":false,"doi":"10.1503/cmaj.170901","ciLow":0.29,"ciHigh":0.38,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For 1924 clinical applications, the two reviewers' independent scores yielded an ICC of 0.33, indicating poor inter-rater reliability.","vf":"unverified"},{"key":"KARAZVW9","au":"Tamblyn, Robyn","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraclass correlation coefficient (ICC)","estd":"ICC (single/unspec)","v":0.32,"n":"941","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (poor) to 4.9 (excellent)","field":"health research","wr":"reviewers on health-services-and-policy grant application scores","conf":"high","self":false,"doi":"10.1503/cmaj.170901","ciLow":0.26,"ciHigh":0.38,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For 941 health-services-and-policy applications, the two reviewers' independent scores yielded an ICC of 0.32, indicating poor inter-rater reliability.","vf":"unverified"},{"key":"KARAZVW9","au":"Tamblyn, Robyn","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraclass correlation coefficient (ICC)","estd":"ICC (single/unspec)","v":0.32,"n":"1309","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (poor) to 4.9 (excellent)","field":"health research","wr":"reviewers on population-health grant application scores","conf":"high","self":false,"doi":"10.1503/cmaj.170901","ciLow":0.26,"ciHigh":0.37,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For 1309 population-health applications, the two reviewers' independent scores yielded an ICC of 0.32, indicating poor inter-rater reliability.","vf":"unverified"},{"key":"KARAZVW9","au":"Tamblyn, Robyn","y":2018,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraclass correlation coefficient (ICC)","estd":"ICC (single/unspec)","v":0.41,"n":"11624","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"1 (poor) to 4.9 (excellent)","field":"health research","wr":"first and second reviewers on grant application scores","conf":"high","self":false,"doi":"10.1503/cmaj.170901","ciLow":0.39,"ciHigh":0.43,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"For 11 624 CIHR operating-grant applications, the first and second reviewers independently scored each application; an ICC of 0.41 indicates fair inter-rater reliability of their ratings.","vf":"unverified"},{"key":"V7CCKM45","au":"Tamblyn, Robyn","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach alpha (reported as 'item to total correlations for criterion rating')","estd":"Cronbach alpha","v":0.91,"n":"3156","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"letter-grade criterion scores (0-28 each)","field":"health research","wr":"internal consistency of phase-1 rating criteria","conf":"high","self":false,"doi":"10.1371/journal.pone.0292306","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Cronbach's alpha across the four phase 1 applicant criteria (vision, impact, productivity, leadership) over 3,156 applications was 0.91, indicating high internal consistency of the rating instrument. The alpha's items are criteria, not raters.","vf":"unverified"},{"key":"V7CCKM45","au":"Tamblyn, Robyn","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intra-class correlation coefficient (ICC)","estd":"ICC (single/unspec)","v":0.59,"n":"3156","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"reviewer-adjusted ranks converted to percentiles","field":"health research","wr":"reviewers ranking grant applications on applicant calibre","conf":"high","self":false,"doi":"10.1371/journal.pone.0292306","ciLow":0.58,"ciHigh":0.61,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The same reviewers ranked the applications they had rated, and ranks were converted to percentiles. Across 3,156 applications the ICC of 0.59 shows ranking was slightly more reliable than rating.","vf":"unverified"},{"key":"V7CCKM45","au":"Tamblyn, Robyn","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intra-class correlation coefficient (ICC)","estd":"ICC (single/unspec)","v":0.46,"n":"3156","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"letter grades poor to outstanding++, scored 0-28","field":"health research","wr":"reviewers rating grant applications on impact","conf":"high","self":false,"doi":"10.1371/journal.pone.0292306","ciLow":0.44,"ciHigh":0.48,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Reviewers rated the impact criterion for each of 3,156 phase 1 applications. The ICC of 0.46 measures agreement between reviewers on this criterion.","vf":"unverified"},{"key":"V7CCKM45","au":"Tamblyn, Robyn","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intra-class correlation coefficient (ICC)","estd":"ICC (single/unspec)","v":0.48,"n":"3156","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"letter grades poor to outstanding++, scored 0-28","field":"health research","wr":"reviewers rating grant applications on leadership","conf":"high","self":false,"doi":"10.1371/journal.pone.0292306","ciLow":0.47,"ciHigh":0.5,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Reviewers rated the leadership criterion for each of 3,156 phase 1 applications. The ICC of 0.48 measures agreement between reviewers on this criterion.","vf":"unverified"},{"key":"V7CCKM45","au":"Tamblyn, Robyn","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intra-class correlation coefficient (ICC)","estd":"ICC (single/unspec)","v":0.54,"n":"3156","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"letter grades converted to scores, 0-28 per criterion; maximum 112","field":"health research","wr":"reviewers rating grant applications on applicant calibre","conf":"high","self":false,"doi":"10.1371/journal.pone.0292306","ciLow":0.52,"ciHigh":0.55,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"In phase 1, four to five expert reviewers, most often five, independently rated each of 3,156 CIHR foundation applications on applicant calibre. The ICC of 0.54 indicates moderate inter-reviewer reliability of the summed rating.","vf":"unverified"},{"key":"V7CCKM45","au":"Tamblyn, Robyn","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intra-class correlation coefficient (ICC)","estd":"ICC (single/unspec)","v":0.5,"n":"3156","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"letter grades poor to outstanding++, scored 0-28","field":"health research","wr":"reviewers rating grant applications on productivity","conf":"high","self":false,"doi":"10.1371/journal.pone.0292306","ciLow":0.48,"ciHigh":0.51,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Reviewers rated the productivity criterion for each of 3,156 phase 1 applications. The ICC of 0.50 was the highest of the phase 1 criteria.","vf":"unverified"},{"key":"V7CCKM45","au":"Tamblyn, Robyn","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intra-class correlation coefficient (ICC)","estd":"ICC (single/unspec)","v":0.35,"n":"3156","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"letter grades poor to outstanding++, scored 0-28","field":"health research","wr":"reviewers rating grant applications on vision","conf":"high","self":false,"doi":"10.1371/journal.pone.0292306","ciLow":0.33,"ciHigh":0.37,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Reviewers rated the vision criterion for each of 3,156 phase 1 applications. The ICC of 0.35 was the lowest of the phase 1 criteria.","vf":"unverified"},{"key":"V7CCKM45","au":"Tamblyn, Robyn","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intra-class correlation coefficient (ICC)","estd":"ICC (single/unspec)","v":0.6,"n":"3156","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"reviewer-adjusted rank from best to worst","field":"health research","wr":"reviewers' raw ranks of phase 1 applications","conf":"high","self":false,"doi":"10.1371/journal.pone.0292306","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Before ranks were standardised to percentiles, the ICC of reviewers' adjusted raw ranks of the 3,156 phase 1 applications was 0.60. The percentile conversion added noise and lowered it to 0.59.","vf":"unverified"},{"key":"V7CCKM45","au":"Tamblyn, Robyn","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach alpha (reported as 'item to total correlations for criterion rating')","estd":"Cronbach alpha","v":0.86,"n":"1096","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"letter-grade criterion scores (0-28 each)","field":"health research","wr":"internal consistency of phase-2 rating criteria","conf":"high","self":false,"doi":"10.1371/journal.pone.0292306","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Cronbach's alpha across the five phase 2 programme criteria (research concept, research approach, expertise, mentorship, environmental support) over 1,096 applications was 0.86. The alpha's items are criteria, not raters.","vf":"unverified"},{"key":"V7CCKM45","au":"Tamblyn, Robyn","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intra-class correlation coefficient (ICC)","estd":"ICC (single/unspec)","v":0.38,"n":"1096","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"reviewer-adjusted ranks converted to percentiles","field":"health research","wr":"reviewers ranking grant applications on research programme","conf":"high","self":false,"doi":"10.1371/journal.pone.0292306","ciLow":0.35,"ciHigh":0.4,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Phase 2 reviewers also ranked their assigned applications, with ranks converted to percentiles. The ICC of 0.38 across 1,096 applications again exceeded the rating ICC.","vf":"unverified"},{"key":"V7CCKM45","au":"Tamblyn, Robyn","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intra-class correlation coefficient (ICC)","estd":"ICC (single/unspec)","v":0.19,"n":"1096","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"letter grades poor to outstanding++, scored 0-28","field":"health research","wr":"reviewers rating grant applications on expertise","conf":"high","self":false,"doi":"10.1371/journal.pone.0292306","ciLow":0.16,"ciHigh":0.22,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Reviewers rated the expertise criterion for each of 1,096 phase 2 applications. The ICC of 0.19 measures agreement between reviewers on this criterion.","vf":"unverified"},{"key":"V7CCKM45","au":"Tamblyn, Robyn","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intra-class correlation coefficient (ICC)","estd":"ICC (single/unspec)","v":0.17,"n":"1096","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"letter grades poor to outstanding++, scored 0-28","field":"health research","wr":"reviewers rating grant applications on mentorship","conf":"high","self":false,"doi":"10.1371/journal.pone.0292306","ciLow":0.14,"ciHigh":0.21,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Reviewers rated the mentorship criterion for each of 1,096 phase 2 applications. The ICC of 0.17 measures agreement between reviewers on this criterion.","vf":"unverified"},{"key":"V7CCKM45","au":"Tamblyn, Robyn","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intra-class correlation coefficient (ICC)","estd":"ICC (single/unspec)","v":0.25,"n":"1096","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"same letter grades as phase 1, weighted; maximum score 140","field":"health research","wr":"reviewers rating grant applications on research programme","conf":"high","self":false,"doi":"10.1371/journal.pone.0292306","ciLow":0.22,"ciHigh":0.28,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"In phase 2, a different set of four to five reviewers rated the research programmes of the 1,096 top-ranked applications. The rating ICC fell to 0.25, which the authors partly attribute to the more homogeneous shortlisted sample.","vf":"unverified"},{"key":"V7CCKM45","au":"Tamblyn, Robyn","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intra-class correlation coefficient (ICC)","estd":"ICC (single/unspec)","v":0.23,"n":"1096","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"letter grades poor to outstanding++, scored 0-28","field":"health research","wr":"reviewers rating grant applications on research approach","conf":"high","self":false,"doi":"10.1371/journal.pone.0292306","ciLow":0.2,"ciHigh":0.26,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Reviewers rated the research approach criterion for each of 1,096 phase 2 applications. The ICC of 0.23 was the highest of the phase 2 criteria.","vf":"unverified"},{"key":"V7CCKM45","au":"Tamblyn, Robyn","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intra-class correlation coefficient (ICC)","estd":"ICC (single/unspec)","v":0.21,"n":"1096","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"letter grades poor to outstanding++, scored 0-28","field":"health research","wr":"reviewers rating grant applications on research concept","conf":"high","self":false,"doi":"10.1371/journal.pone.0292306","ciLow":0.18,"ciHigh":0.24,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Reviewers rated the research concept criterion for each of 1,096 phase 2 applications. The ICC of 0.21 measures agreement between reviewers on this criterion.","vf":"unverified"},{"key":"V7CCKM45","au":"Tamblyn, Robyn","y":2023,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intra-class correlation coefficient (ICC)","estd":"ICC (single/unspec)","v":0.12,"n":"1096","k":"","samp":"full-pool","blind":"single","agg":"unspecified","scale":"letter grades poor to outstanding++, scored 0-28","field":"health research","wr":"reviewers rating grant applications on environmental support","conf":"high","self":false,"doi":"10.1371/journal.pone.0292306","ciLow":0.09,"ciHigh":0.15,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Reviewers rated the environmental support criterion, labelled Support in Table 3 and Environment in Table 2, for each of 1,096 phase 2 applications. The ICC of 0.12 was the lowest criterion reliability.","vf":"unverified"},{"key":"7S4GVWVN","au":"Tan, Erwin","y":2016,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach's a (consistency)","estd":"Cronbach alpha","v":0.88,"n":"241","k":"3","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"0-3: does not meet / fair / good / excellent","field":"community service grants","wr":"reviewers on individual grant criteria scores","conf":"high","self":false,"doi":"10.1007/s11266-015-9602-2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Averaged across the 22 individual selection criteria, reviewer consistency on grant applications, measured by Cronbach's alpha, was 0.88.","vf":"unverified"},{"key":"7S4GVWVN","au":"Tan, Erwin","y":2016,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach's a (consistency)","estd":"Cronbach alpha","v":0.97,"n":"241","k":"3","samp":"full-pool","blind":"single","agg":"average-of-k","scale":"0-88 weighted total score","field":"community service grants","wr":"reviewers on grant application overall scores","conf":"high","self":false,"doi":"10.1007/s11266-015-9602-2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"A Cronbach's alpha of 0.97 indicates very high consistency among the three reviewers' overall scores across the 241 grant applications.","vf":"unverified"},{"key":"7S4GVWVN","au":"Tan, Erwin","y":2016,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorf's a (consensus)","estd":"Krippendorff","v":0.71,"n":"241","k":"3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"0-3: does not meet / fair / good / excellent","field":"community service grants","wr":"reviewers on individual grant criteria scores","conf":"high","self":false,"doi":"10.1007/s11266-015-9602-2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Averaged across the 22 individual selection criteria, reviewer consensus on grant applications, measured by Krippendorff's alpha, was 0.71.","vf":"unverified"},{"key":"7S4GVWVN","au":"Tan, Erwin","y":2016,"cx":"Grant","ob":"grant-proposal","fam":"Krippendorff","form":"Krippendorf's a (consensus)","estd":"Krippendorff","v":0.91,"n":"241","k":"3","samp":"full-pool","blind":"single","agg":"single-rater","scale":"0-88 weighted total score","field":"community service grants","wr":"reviewers on grant application overall scores","conf":"high","self":false,"doi":"10.1007/s11266-015-9602-2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":false,"ms":"Three reviewers (two internal staff and one external peer) rated each of 241 grant applications; a Krippendorff's alpha of 0.91 for the overall 0-88 score indicates high consensus.","vf":"unverified"},{"key":"7S4GVWVN","au":"Tan, Erwin","y":2016,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"x (omega; glyph renders as 'x' in extraction)","estd":"other","v":0.78,"n":"241","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"0 = does not meet to 3 = excellent","field":"community service grants","wr":"cost effectiveness and budget adequacy criteria","conf":"high","self":false,"doi":"10.1007/s11266-015-9602-2","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Omega measurement reliability of the Cost Effectiveness and Budget Adequacy domain of the selection instrument was 0.78, the only domain below the 0.8 level.","vf":"unverified"},{"key":"7S4GVWVN","au":"Tan, Erwin","y":2016,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"x (omega; glyph renders as 'x' in extraction)","estd":"other","v":0.91,"n":"241","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"0 = does not meet to 3 = excellent","field":"community service grants","wr":"Organizational Capacity selection criteria","conf":"high","self":false,"doi":"10.1007/s11266-015-9602-2","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Omega measurement reliability of the Organizational Capacity domain of the selection instrument, computed from three reviewers' ratings of 241 applications, was 0.91.","vf":"unverified"},{"key":"7S4GVWVN","au":"Tan, Erwin","y":2016,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"x (omega; glyph renders as 'x' in extraction)","estd":"other","v":0.94,"n":"241","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"0 = does not meet to 3 = excellent","field":"community service grants","wr":"Strengthening Communities selection criteria","conf":"high","self":false,"doi":"10.1007/s11266-015-9602-2","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Omega measurement reliability of the Strengthening Communities domain of the selection instrument, computed from three reviewers' ratings of 241 applications, was 0.94.","vf":"unverified"},{"key":"7S4GVWVN","au":"Tan, Erwin","y":2016,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"multidimensional x (omega; glyph renders as 'x' in extraction)","estd":"other","v":0.97,"n":"241","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"0 = does not meet to 3 = excellent","field":"community service grants","wr":"complete four-domain grant selection instrument","conf":"high","self":false,"doi":"10.1007/s11266-015-9602-2","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Multidimensional omega for the complete four-domain selection instrument was 0.97, indicating high measurement reliability of the final grant scores.","vf":"unverified"},{"key":"7S4GVWVN","au":"Tan, Erwin","y":2016,"cx":"Grant","ob":"grant-proposal","fam":"other","form":"x (omega; glyph renders as 'x' in extraction)","estd":"other","v":0.93,"n":"241","k":"3","samp":"full-pool","blind":"single","agg":"unspecified","scale":"0 = does not meet to 3 = excellent","field":"community service grants","wr":"volunteer recruitment and programme management criteria","conf":"high","self":false,"doi":"10.1007/s11266-015-9602-2","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Omega measurement reliability of the merged Recruitment and Development of Volunteers/Program Management domain of the selection instrument was 0.93.","vf":"unverified"},{"key":"SCBRUUIQ","au":"Teplitskiy, Misha","y":2018,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intra-class correlation coefficient (ICC), decisions converted to numeric 1.0-4.0, first two reviewers","estd":"ICC (single/unspec)","v":0.192,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"Reject / Major revision / Minor revision / Accept, coded 1-4","field":"neuroscience","wr":"reviewers on neuroscience manuscript recommendations","conf":"high","self":false,"doi":"10.1016/j.respol.2018.06.014","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For comparability with earlier studies, the authors also computed an ICC on the first two reviewers' first-round recommendations (recoded 1 to 4); it was 0.192, matching the near-zero alpha.","vf":"unverified"},{"key":"SCBRUUIQ","au":"Teplitskiy, Misha","y":2018,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intra-class correlation coefficient (ICC), decisions converted to numeric 1.0-4.0, first two reviewers","estd":"ICC (single/unspec)","v":0.194,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"Reject / Major revision / Minor revision / Accept, coded 1-4","field":"neuroscience","wr":"reviewers on neuroscience manuscript recommendations","conf":"high","self":false,"doi":"10.1016/j.respol.2018.06.014","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"Restricted to the first two reviewers, the ICC for second-round recommendations was 0.194; from the second round onward reviewers could see each other's decisions, so independence no longer holds.","vf":"unverified"},{"key":"SCBRUUIQ","au":"Teplitskiy, Misha","y":2018,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intra-class correlation coefficient (ICC), decisions converted to numeric 1.0-4.0, first two reviewers","estd":"ICC (single/unspec)","v":0.204,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"Reject / Major revision / Minor revision / Accept, coded 1-4","field":"neuroscience","wr":"reviewers on neuroscience manuscript recommendations","conf":"high","self":false,"doi":"10.1016/j.respol.2018.06.014","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"The third-round ICC on the first two reviewers' recommendations was 0.204, still very low, with reviewers no longer judging independently.","vf":"unverified"},{"key":"SCBRUUIQ","au":"Teplitskiy, Misha","y":2018,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intra-class correlation coefficient (ICC), decisions converted to numeric 1.0-4.0, first two reviewers","estd":"ICC (single/unspec)","v":0.555,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"unspecified","scale":"Reject / Major revision / Minor revision / Accept, coded 1-4","field":"neuroscience","wr":"reviewers on neuroscience manuscript recommendations","conf":"high","self":false,"doi":"10.1016/j.respol.2018.06.014","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"By the fourth round the ICC on the first two reviewers rose to 0.555, but this reflects surviving manuscripts and non-independent reviewers rather than genuinely higher agreement.","vf":"unverified"},{"key":"SCBRUUIQ","au":"Teplitskiy, Misha","y":2018,"cx":"Journal","ob":"journal-manuscript","fam":"Krippendorff","form":"Krippendorff's alpha, ordinal level of measurement","estd":"Krippendorff","v":0.193,"n":"","k":"","samp":"full-pool","blind":"single","agg":"single-rater","scale":"Reject / Major revision / Minor revision / Accept, coded 1-4","field":"neuroscience","wr":"reviewers on neuroscience manuscript recommendations","conf":"high","self":false,"doi":"10.1016/j.respol.2018.06.014","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"PLOS ONE reviewers independently recommended accept, revise, or reject on 2011-2012 neuroscience manuscripts. The first-round Krippendorff's alpha of 0.193, treating the scale as ordinal, shows agreement barely above chance and is the paper's headline reliability figure.","vf":"unverified"},{"key":"SCBRUUIQ","au":"Teplitskiy, Misha","y":2018,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"Cohen's weighted Kappa, squared (quadratic) weighting, first two reviewers","estd":"weighted kappa","v":0.192,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"Reject / Major revision / Minor revision / Accept, coded 1-4","field":"neuroscience","wr":"reviewers on neuroscience manuscript recommendations","conf":"high","self":false,"doi":"10.1016/j.respol.2018.06.014","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Cohen's quadratically weighted kappa on the first two reviewers' first-round recommendations was 0.192, again indicating agreement close to chance.","vf":"unverified"},{"key":"SCBRUUIQ","au":"Teplitskiy, Misha","y":2018,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"Cohen's weighted Kappa, squared (quadratic) weighting, first two reviewers","estd":"weighted kappa","v":0.194,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"Reject / Major revision / Minor revision / Accept, coded 1-4","field":"neuroscience","wr":"reviewers on neuroscience manuscript recommendations","conf":"high","self":false,"doi":"10.1016/j.respol.2018.06.014","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"The quadratically weighted kappa for the first two second-round reviewers was 0.194, computed where reviewers could already see each other's decisions.","vf":"unverified"},{"key":"SCBRUUIQ","au":"Teplitskiy, Misha","y":2018,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"Cohen's weighted Kappa, squared (quadratic) weighting, first two reviewers","estd":"weighted kappa","v":0.206,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"Reject / Major revision / Minor revision / Accept, coded 1-4","field":"neuroscience","wr":"reviewers on neuroscience manuscript recommendations","conf":"high","self":false,"doi":"10.1016/j.respol.2018.06.014","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"The third-round quadratically weighted kappa on the first two reviewers was 0.206, still very low.","vf":"unverified"},{"key":"SCBRUUIQ","au":"Teplitskiy, Misha","y":2018,"cx":"Journal","ob":"journal-manuscript","fam":"weighted-kappa","form":"Cohen's weighted Kappa, squared (quadratic) weighting, first two reviewers","estd":"weighted kappa","v":0.545,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"Reject / Major revision / Minor revision / Accept, coded 1-4","field":"neuroscience","wr":"reviewers on neuroscience manuscript recommendations","conf":"high","self":false,"doi":"10.1016/j.respol.2018.06.014","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":false,"ms":"The fourth-round quadratically weighted kappa on the first two reviewers rose to 0.545, among surviving manuscripts and non-independent reviewers.","vf":"unverified"},{"key":"4RYFCKV2","au":"Thelwall, Mike","y":2023,"cx":"General","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement on four-point 1*-4* scale","estd":"percent agreement","v":0.589,"n":"8015","k":"","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"1* recognised nationally to 4* world-leading","field":"multi-field","wr":"UoA panels scoring same journal article's quality","conf":"high","self":false,"doi":"10.1108/jd-01-2023-0012","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the 8,015 journal articles scored by two or more different subject panels, the raw exact agreement rate between panel scores for the same article on the four-point scale was 58.9%; the number of copies per article varied.","vf":"unverified"},{"key":"4RYFCKV2","au":"Thelwall, Mike","y":2023,"cx":"General","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement on four-point 1*-4* scale, projected to single submission","estd":"percent agreement","v":0.53,"n":"8015","k":"2","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"1* recognised nationally to 4* world-leading","field":"multi-field","wr":"UoA panels scoring same journal article's quality","conf":"high","self":false,"doi":"10.1108/jd-01-2023-0012","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"The paper's headline estimate: for a journal article scored by two different subject panels, the projected chance that the two panels give the same grade on the four-point REF scale is about 53%, extrapolated to singly submitted outputs from 8,015 multiply submitted articles.","vf":"unverified"},{"key":"4RYFCKV2","au":"Thelwall, Mike","y":2023,"cx":"General","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement between two randomly selected scores, four-point scale","estd":"percent agreement","v":0.798,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"1* recognised nationally to 4* world-leading","field":"multi-field","wr":"panels scoring same multiply-submitted journal article","conf":"high","self":false,"doi":"10.1108/jd-01-2023-0012","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Across all articles submitted more than once (about a quarter of the sample), two randomly chosen scores for the same article matched on the four-point scale 79.8% of the time, pooling within-panel and between-panel copies.","vf":"unverified"},{"key":"4RYFCKV2","au":"Thelwall, Mike","y":2023,"cx":"General","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement on four-point 1*-4* scale","estd":"percent agreement","v":0.85,"n":"40","k":"2","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"1* recognised nationally to 4* world-leading","field":"multi-field","wr":"Archaeology panel scoring duplicate journal articles","conf":"high","self":false,"doi":"10.1108/jd-01-2023-0012","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two REF2021 Archaeology scores were available for each of 40 duplicate articles; the scores agreed exactly for 85% of the articles.","vf":"unverified"},{"key":"4RYFCKV2","au":"Thelwall, Mike","y":2023,"cx":"General","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement on four-point 1*-4* scale","estd":"percent agreement","v":1,"n":"4","k":"3","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"1* recognised nationally to 4* world-leading","field":"multi-field","wr":"Archaeology panel scoring triplicate journal articles","conf":"high","self":false,"doi":"10.1108/jd-01-2023-0012","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Three REF2021 Archaeology scores were available for each of four triplicate articles; all scores agreed exactly for every article.","vf":"unverified"},{"key":"4RYFCKV2","au":"Thelwall, Mike","y":2023,"cx":"General","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement on four-point 1*-4* scale, per article single-counted","estd":"percent agreement","v":0.864,"n":"44","k":"","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"1* recognised nationally to 4* world-leading","field":"multi-field","wr":"Archaeology panel scoring same journal article twice","conf":"high","self":false,"doi":"10.1108/jd-01-2023-0012","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"In Archaeology (UoA 15), which had 44 duplicate or triplicate articles and no apparent systematic cross-check procedure, scores for the same article matched 86.4% of the time; the paper treats this as its best estimate of genuine within-field agreement.","vf":"unverified"},{"key":"4RYFCKV2","au":"Thelwall, Mike","y":2023,"cx":"General","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement on four-point 1*-4* scale, per article single-counted","estd":"percent agreement","v":1,"n":"2112","k":"","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"1* recognised nationally to 4* world-leading","field":"multi-field","wr":"Business and Management panel scoring same journal article twice","conf":"high","self":false,"doi":"10.1108/jd-01-2023-0012","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Maximum within-UoA agreement rate in Table 1: several panels including Business and Management (UoA 17, 2112 duplicate articles) reached 100% duplicate-score agreement, attributed to systematic cross-checking or single assessment of duplicates.","vf":"unverified"},{"key":"4RYFCKV2","au":"Thelwall, Mike","y":2023,"cx":"General","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement on four-point 1*-4* scale, per article single-counted","estd":"percent agreement","v":0.667,"n":"24","k":"","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"1* recognised nationally to 4* world-leading","field":"multi-field","wr":"Art and Design panel scoring same journal article twice","conf":"high","self":false,"doi":"10.1108/jd-01-2023-0012","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Minimum within-UoA agreement rate in Table 1: in Art and Design (UoA 32, 24 duplicate articles), scores for the same article matched 66.7% of the time.","vf":"unverified"},{"key":"4RYFCKV2","au":"Thelwall, Mike","y":2023,"cx":"General","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement between two randomly selected scores, four-point scale","estd":"percent agreement","v":0.989,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"1* recognised nationally to 4* world-leading","field":"multi-field","wr":"same UoA panel scoring same journal article twice","conf":"high","self":false,"doi":"10.1108/jd-01-2023-0012","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"When the same journal article was submitted more than once within one subject panel, two randomly chosen scores matched 98.9% of the time, a rate inflated because many panels cross-checked duplicate scores for discrepancies.","vf":"unverified"},{"key":"4RYFCKV2","au":"Thelwall, Mike","y":2023,"cx":"General","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement on four-point 1*-4* scale, projected to single submission","estd":"percent agreement","v":0.7,"n":"44","k":"2","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"1* recognised nationally to 4* world-leading","field":"multi-field","wr":"single UoA panel scoring same journal article twice","conf":"high","self":false,"doi":"10.1108/jd-01-2023-0012","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Extrapolating Archaeology's 44 duplicate and triplicate articles to singly submitted outputs gives a projected within-field agreement rate of 70%, the paper's within-UoA counterpart to its 53% between-UoA figure.","vf":"unverified"},{"key":"8LVJQ749","au":"Thelwall, Mike","y":2023,"cx":"General","ob":"journal-manuscript","fam":"percent-agreement","form":"raw percent agreement, exact ('reviewing team scores agreed 85% of the time')","estd":"percent agreement","v":0.85,"n":"","k":"","samp":"full-pool","blind":"single","agg":"panel-consensus","scale":"0, 1*, 2*, 3*, or 4*","field":"multi-field","wr":"REF teams scoring duplicate journal articles","conf":"low","self":false,"doi":"10.1162/qss_a_00258","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":false,"ms":"In one REF2021 Unit of Assessment that accidentally reviewed duplicate copies of the same journal articles, the reviewing teams' provisional quality scores agreed 85 per cent of the time. Each score is a two-reviewer consensus reached after discussion and wider group norm referencing, and the authors present the figure as a very crude natural-experiment estimate of the consistency of the overall REF assessment process.","vf":"unverified"},{"key":"C9INWUTB","au":"Thelwall, Mike","y":2023,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,1)","estd":"ICC (single/unspec)","v":0.36,"n":"500","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"poor / low / ok / good / high / top","field":"physics","wr":"first two reviewers on originality of physics articles","conf":"high","self":false,"doi":"10.1093/reseval/rvad018","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"The first two reviewers of 500 SciPost Physics theoretical articles scored each article's originality on a six-point scale; an ICC(1,1) of 0.36 indicates moderate single-rater agreement.","vf":"unverified"},{"key":"C9INWUTB","au":"Thelwall, Mike","y":2023,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,1)","estd":"ICC (single/unspec)","v":0.39,"n":"502","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"poor / low / ok / good / high / top","field":"physics","wr":"first two reviewers on significance of physics articles","conf":"high","self":false,"doi":"10.1093/reseval/rvad018","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"The first two reviewers of 502 SciPost Physics theoretical articles scored each article's significance on a six-point scale; an ICC(1,1) of 0.39 indicates moderate single-rater agreement.","vf":"unverified"},{"key":"C9INWUTB","au":"Thelwall, Mike","y":2023,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,1)","estd":"ICC (single/unspec)","v":0.45,"n":"505","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"sum of three six-point facets (validity+significance+originality)","field":"physics","wr":"first two reviewers on summed quality of physics articles","conf":"high","self":false,"doi":"10.1093/reseval/rvad018","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":false,"ms":"The first two reviewers of the 505 SciPost Physics theoretical articles scored validity, significance and originality; the ICC(1,1) of the summed score, 0.45, is the study's overall reviewer-consistency figure and indicates moderate single-rater agreement.","vf":"unverified"},{"key":"C9INWUTB","au":"Thelwall, Mike","y":2023,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,1)","estd":"ICC (single/unspec)","v":0.4,"n":"497","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"poor / low / ok / good / high / top","field":"physics","wr":"first two reviewers on validity of physics articles","conf":"high","self":false,"doi":"10.1093/reseval/rvad018","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"The first two reviewers of 497 SciPost Physics theoretical articles scored each article's validity (rigour) on a six-point scale; an ICC(1,1) of 0.40 indicates moderate single-rater agreement.","vf":"unverified"},{"key":"C9INWUTB","au":"Thelwall, Mike","y":2023,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"agreeing or differing by at most 1 point (within-1 tolerance)","estd":"percent agreement","v":0.86,"n":"505","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"poor / low / ok / good / high / top","field":"physics","wr":"first two reviewers agreeing within one point on facet scores","conf":"high","self":false,"doi":"10.1093/reseval/rvad018","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Across the first two reviewers of the 505 SciPost Physics theoretical articles, 86% of paired facet scores were identical or differed by only one point on the six-point scale, indicating that large disagreements were rare.","vf":"unverified"},{"key":"C9INWUTB","au":"Thelwall, Mike","y":2023,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC(1,1)","estd":"ICC (single/unspec)","v":0.4,"n":"1008","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"1-10 each facet, summed to 3-30","field":"physics","wr":"two VQR reviewers on summed quality of physics articles","conf":"high","self":false,"doi":"10.1093/reseval/rvad018","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":true,"ms":"Two independent VQR field-specialist reviewers scored 1008 Italian physics articles on significance, originality and rigour (1-10 each); the ICC(1,1) of the summed score is 0.40, a reanalysis used to contextualise the SciPost result.","vf":"unverified"},{"key":"C9INWUTB","au":"Thelwall, Mike","y":2023,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"exact agreement (identical total scores)","estd":"percent agreement","v":0.09,"n":"1008","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"sum of three 1-10 facets, 3-30","field":"physics","wr":"two VQR reviewers giving identical total scores","conf":"high","self":false,"doi":"10.1093/reseval/rvad018","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":true,"ms":"Two VQR reviewers gave exactly the same total (significance+originality+rigour) score for only 9% of 1008 Italian physics articles.","vf":"unverified"},{"key":"C9INWUTB","au":"Thelwall, Mike","y":2023,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"within 3 points of each other (within-3 tolerance)","estd":"percent agreement","v":0.51,"n":"1008","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"sum of three 1-10 facets, 3-30","field":"physics","wr":"two VQR reviewers agreeing within three total-score points","conf":"high","self":false,"doi":"10.1093/reseval/rvad018","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":true,"ms":"Two VQR reviewers gave total scores within three points of each other for 51% of 1008 Italian physics articles on the 3-30 summed scale.","vf":"unverified"},{"key":"C9INWUTB","au":"Thelwall, Mike","y":2023,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"within 9 points of each other (within-9 tolerance)","estd":"percent agreement","v":0.89,"n":"1008","k":"2","samp":"full-pool","blind":"open","agg":"single-rater","scale":"sum of three 1-10 facets, 3-30","field":"physics","wr":"two VQR reviewers agreeing within nine total-score points","conf":"high","self":false,"doi":"10.1093/reseval/rvad018","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":true,"ms":"Two VQR reviewers gave total scores within nine points of each other for 89% of 1008 Italian physics articles on the 3-30 summed scale.","vf":"unverified"},{"key":"5ZF6XD7M","au":"Thelwall, Mike","y":2024,"cx":"General","ob":"journal-manuscript","fam":"correlation","form":"Pearson correlation, author self-scores vs mean of 15 ChatGPT-4 rounds","estd":"correlation","v":0.2,"n":"34","k":"2","samp":"special","blind":"open","agg":"average-of-k","scale":"1* to 4* (4* = world-leading)","field":"information science","wr":"author vs ChatGPT on higher-quality article scores","conf":"med","self":false,"doi":"10.2478/jdis-2024-0013","ciLow":-0.148,"ciHigh":0.504,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Restricting to the 34 articles the author rated at least 2.5 stars, the correlation between his scores and the averaged ChatGPT scores fell to 0.20 and was not statistically significant, suggesting weaker discrimination among higher-quality articles.","vf":"unverified"},{"key":"5ZF6XD7M","au":"Thelwall, Mike","y":2024,"cx":"General","ob":"journal-manuscript","fam":"correlation","form":"Pearson correlation, author self-scores vs mean of 15 ChatGPT-4 rounds","estd":"correlation","v":0.246,"n":"24","k":"2","samp":"special","blind":"open","agg":"average-of-k","scale":"1* to 4* (4* = world-leading)","field":"information science","wr":"author vs ChatGPT on top-quality article scores","conf":"med","self":false,"doi":"10.2478/jdis-2024-0013","ciLow":-0.175,"ciHigh":0.59,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"For the 24 articles the author rated at least 3 stars, the correlation between his scores and averaged ChatGPT scores was 0.25 and not statistically significant.","vf":"unverified"},{"key":"5ZF6XD7M","au":"Thelwall, Mike","y":2024,"cx":"General","ob":"journal-manuscript","fam":"correlation","form":"Pearson correlation, author self-scores vs mean of 15 ChatGPT-4 rounds","estd":"correlation","v":0.509,"n":"51","k":"2","samp":"special","blind":"open","agg":"average-of-k","scale":"1* to 4* (4* = world-leading)","field":"information science","wr":"author vs ChatGPT on article REF quality scores","conf":"med","self":false,"doi":"10.2478/jdis-2024-0013","ciLow":0.271,"ciHigh":0.688,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"The author scored 51 of his own articles for quality using REF 2021 criteria, and each article was also scored by ChatGPT-4 across fifteen rounds. A Pearson correlation of 0.51 between the author's scores and the mean ChatGPT score indicates a moderate association between the human and averaged machine evaluations.","vf":"unverified"},{"key":"5ZF6XD7M","au":"Thelwall, Mike","y":2024,"cx":"General","ob":"journal-manuscript","fam":"correlation","form":"Pearson correlation, mean over 15 single-round ChatGPT-vs-author pairs","estd":"correlation","v":0.102,"n":"34","k":"2","samp":"special","blind":"open","agg":"single-rater","scale":"1* to 4* (4* = world-leading)","field":"information science","wr":"single-round ChatGPT vs author, higher-quality subset","conf":"med","self":false,"doi":"10.2478/jdis-2024-0013","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Across the 34 higher-rated articles, the mean single-round ChatGPT-versus-author correlation was 0.10, with only one of fifteen rounds statistically significant.","vf":"unverified"},{"key":"5ZF6XD7M","au":"Thelwall, Mike","y":2024,"cx":"General","ob":"journal-manuscript","fam":"correlation","form":"Pearson correlation, mean over 15 single-round ChatGPT-vs-author pairs","estd":"correlation","v":0.128,"n":"24","k":"2","samp":"special","blind":"open","agg":"single-rater","scale":"1* to 4* (4* = world-leading)","field":"information science","wr":"single-round ChatGPT vs author, top-quality subset","conf":"med","self":false,"doi":"10.2478/jdis-2024-0013","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Across the 24 articles rated at least 3 stars, the mean single-round ChatGPT-versus-author correlation was 0.13.","vf":"unverified"},{"key":"5ZF6XD7M","au":"Thelwall, Mike","y":2024,"cx":"General","ob":"journal-manuscript","fam":"correlation","form":"Pearson correlation, mean over 15 single-round ChatGPT-vs-author pairs","estd":"correlation","v":0.281,"n":"51","k":"2","samp":"special","blind":"open","agg":"single-rater","scale":"1* to 4* (4* = world-leading)","field":"information science","wr":"single-round ChatGPT vs author on article scores","conf":"med","self":false,"doi":"10.2478/jdis-2024-0013","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Averaging the Pearson correlations from the fifteen individual ChatGPT rounds against the author's scores gave a mean correlation of 0.28 across the 51 articles, with eight of the fifteen rounds significantly different from zero.","vf":"unverified"},{"key":"5ZF6XD7M","au":"Thelwall, Mike","y":2024,"cx":"General","ob":"journal-manuscript","fam":"correlation","form":"Pearson correlation, mean over 105 pairwise ChatGPT-round comparisons","estd":"correlation","v":0.194,"n":"34","k":"2","samp":"special","blind":"open","agg":"single-rater","scale":"1* to 4* (4* = world-leading)","field":"information science","wr":"ChatGPT round-to-round consistency, higher-quality subset","conf":"med","self":false,"doi":"10.2478/jdis-2024-0013","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"For the 34 higher-rated articles the mean ChatGPT round-to-round correlation was 0.19.","vf":"unverified"},{"key":"5ZF6XD7M","au":"Thelwall, Mike","y":2024,"cx":"General","ob":"journal-manuscript","fam":"correlation","form":"Pearson correlation, mean over 105 pairwise ChatGPT-round comparisons","estd":"correlation","v":0.215,"n":"24","k":"2","samp":"special","blind":"open","agg":"single-rater","scale":"1* to 4* (4* = world-leading)","field":"information science","wr":"ChatGPT round-to-round consistency, top-quality subset","conf":"med","self":false,"doi":"10.2478/jdis-2024-0013","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"For the 24 articles rated at least 3 stars the mean ChatGPT round-to-round correlation was 0.22.","vf":"unverified"},{"key":"5ZF6XD7M","au":"Thelwall, Mike","y":2024,"cx":"General","ob":"journal-manuscript","fam":"correlation","form":"Pearson correlation, mean over 105 pairwise ChatGPT-round comparisons","estd":"correlation","v":0.245,"n":"51","k":"2","samp":"special","blind":"open","agg":"single-rater","scale":"1* to 4* (4* = world-leading)","field":"information science","wr":"ChatGPT round-to-round consistency on article scores","conf":"med","self":false,"doi":"10.2478/jdis-2024-0013","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Consistency of ChatGPT with itself: the mean of 105 pairwise Pearson correlations between the fifteen scoring rounds of the same 51 articles was 0.25, indicating low agreement of the model across repeated evaluations.","vf":"unverified"},{"key":"5ZF6XD7M","au":"Thelwall, Mike","y":2024,"cx":"General","ob":"journal-manuscript","fam":"other","form":"Mean Absolute Difference (MAD), author vs mean of 15 ChatGPT rounds; 0 = full agreement","estd":"other","v":0.802,"n":"51","k":"2","samp":"special","blind":"open","agg":"average-of-k","scale":"1* to 4* (4* = world-leading)","field":"information science","wr":"absolute gap between author and averaged ChatGPT scores","conf":"med","self":false,"doi":"10.2478/jdis-2024-0013","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Averaging ChatGPT scores over fifteen rounds did not reduce the mean absolute deviation from the author's scores, which remained 0.80 stars.","vf":"unverified"},{"key":"5ZF6XD7M","au":"Thelwall, Mike","y":2024,"cx":"General","ob":"journal-manuscript","fam":"other","form":"Mean Absolute Difference (MAD), author vs single-round ChatGPT scores; 0 = full agreement","estd":"other","v":0.802,"n":"51","k":"2","samp":"special","blind":"open","agg":"single-rater","scale":"1* to 4* (4* = world-leading)","field":"information science","wr":"absolute gap between author and ChatGPT article scores","conf":"med","self":false,"doi":"10.2478/jdis-2024-0013","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"The mean absolute deviation between the author's scores and single-round ChatGPT scores was 0.80 stars, showing the two disagreed by nearly one quality level on average.","vf":"unverified"},{"key":"QCWPTK8N","au":"Thelwall, Mike","y":2026,"cx":"General","ob":"journal-manuscript","fam":"correlation","form":"Spearman rank correlation","estd":"correlation","v":0.339,"n":"2780","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"LLM 1*-4* stars; departmental mean REF proxy (continuous)","field":"health and life sciences (6 REF2021 UoAs)","wr":"best single LLM vs departmental-mean REF score proxy","conf":"med","self":false,"doi":"10.1007/s11192-026-05585-2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"The single best-performing LLM configuration's scores correlated with the departmental-average REF score proxy at a Spearman rho of 0.34 across all six fields.","vf":"unverified"},{"key":"QCWPTK8N","au":"Thelwall, Mike","y":2026,"cx":"General","ob":"journal-manuscript","fam":"correlation","form":"Spearman rank correlation","estd":"correlation","v":0.369,"n":"2780","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"average-of-k","scale":"LLM 1*-4* stars; departmental mean REF proxy (continuous)","field":"health and life sciences (6 REF2021 UoAs)","wr":"LLM-ensemble vs departmental-mean REF score proxy for journal articles","conf":"med","self":false,"doi":"10.1007/s11192-026-05585-2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"The mean of 23 LLM configurations' scores for the articles correlated with the departmental-average REF score proxy at a Spearman rho of 0.37 across all six fields; the proxy dampens correlations relative to individual scores.","vf":"unverified"},{"key":"QCWPTK8N","au":"Thelwall, Mike","y":2026,"cx":"General","ob":"journal-manuscript","fam":"correlation","form":"Spearman rank correlation","estd":"correlation","v":0.347,"n":"2780","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"average-of-k","scale":"LLM 1*-4* stars; departmental mean REF proxy (continuous)","field":"health and life sciences (6 REF2021 UoAs)","wr":"LLM-ensemble (median) vs departmental-mean REF score proxy","conf":"med","self":false,"doi":"10.1007/s11192-026-05585-2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"The median of 23 LLM configurations' scores for the articles correlated with the departmental-average REF score proxy at a Spearman rho of 0.35 across all six fields.","vf":"unverified"},{"key":"QCWPTK8N","au":"Thelwall, Mike","y":2026,"cx":"General","ob":"journal-manuscript","fam":"correlation","form":"Spearman rank correlation","estd":"correlation","v":0.374,"n":"2780","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"average-of-k","scale":"LLM 1*-4* stars; departmental mean REF proxy (continuous)","field":"health and life sciences (6 REF2021 UoAs)","wr":"LLM-ensemble (rank average) vs departmental-mean REF score proxy","conf":"med","self":false,"doi":"10.1007/s11192-026-05585-2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"The mean rank of the articles across the 23 LLM configurations correlated with the departmental-average REF score proxy at a Spearman rho of 0.37 across all six fields.","vf":"unverified"},{"key":"QCWPTK8N","au":"Thelwall, Mike","y":2026,"cx":"General","ob":"journal-manuscript","fam":"correlation","form":"Spearman rank correlation (average of 10 folds)","estd":"correlation","v":0.552,"n":"500","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"average-of-k","scale":"LLM 1*-4* stars (weighted sum); departmental mean REF proxy (continuous)","field":"health and life sciences (6 REF2021 UoAs)","wr":"weighted-sum LLM ensemble vs departmental-mean REF proxy, Psychology/Neuroscience","conf":"med","self":false,"doi":"10.1007/s11192-026-05585-2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For 500 Psychology, Psychiatry and Neuroscience articles, the tenfold-cross-validated weighted LLM ensemble correlated with the departmental-average REF proxy at a Spearman rho of 0.55, the highest stratum value reported in the paper.","vf":"unverified"},{"key":"QCWPTK8N","au":"Thelwall, Mike","y":2026,"cx":"General","ob":"journal-manuscript","fam":"correlation","form":"Spearman rank correlation (average of 10 folds)","estd":"correlation","v":0.118,"n":"500","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"average-of-k","scale":"LLM 1*-4* stars (weighted sum); departmental mean REF proxy (continuous)","field":"health and life sciences (6 REF2021 UoAs)","wr":"weighted-sum LLM ensemble vs departmental-mean REF proxy, Clinical Medicine","conf":"med","self":false,"doi":"10.1007/s11192-026-05585-2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"For 500 Clinical Medicine articles, the tenfold-cross-validated weighted LLM ensemble correlated with the departmental-average REF proxy at a Spearman rho of 0.12, the lowest stratum value reported in the paper.","vf":"unverified"},{"key":"QCWPTK8N","au":"Thelwall, Mike","y":2026,"cx":"General","ob":"journal-manuscript","fam":"correlation","form":"Spearman rank correlation (average of 10 folds)","estd":"correlation","v":0.4,"n":"2780","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"average-of-k","scale":"LLM 1*-4* stars (weighted sum); departmental mean REF proxy (continuous)","field":"health and life sciences (6 REF2021 UoAs)","wr":"optimised weighted-sum LLM ensemble vs departmental-mean REF score proxy","conf":"med","self":false,"doi":"10.1007/s11192-026-05585-2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"A tenfold-cross-validated weighted sum of the 23 LLM configurations' scores correlated with the departmental-average REF score proxy at a Spearman rho of 0.40 across all six fields.","vf":"unverified"},{"key":"QCWPTK8N","au":"Thelwall, Mike","y":2026,"cx":"General","ob":"journal-manuscript","fam":"correlation","form":"Spearman rank correlation","estd":"correlation","v":0.427,"n":"2780","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"LLM 1*-4* stars; author nine-point (averaged, fractional)","field":"health and life sciences (6 REF2021 UoAs)","wr":"best single LLM vs expert research-quality scores of journal articles","conf":"med","self":false,"doi":"10.1007/s11192-026-05585-2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"The single best-performing LLM configuration's scores for the articles correlated with one expert's individual quality scores at a Spearman rho of 0.43 across all six fields.","vf":"unverified"},{"key":"QCWPTK8N","au":"Thelwall, Mike","y":2026,"cx":"General","ob":"journal-manuscript","fam":"correlation","form":"Spearman rank correlation","estd":"correlation","v":0.463,"n":"2780","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"average-of-k","scale":"LLM 1*-4* stars; author nine-point (averaged, fractional)","field":"health and life sciences (6 REF2021 UoAs)","wr":"LLM-ensemble vs expert research-quality scores of journal articles","conf":"med","self":false,"doi":"10.1007/s11192-026-05585-2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":true,"he":false,"ms":"The mean of 23 LLM configurations' quality scores for 2780 journal articles correlated with one expert's individual quality scores at a Spearman rho of 0.46 across all six health and life science fields.","vf":"unverified"},{"key":"QCWPTK8N","au":"Thelwall, Mike","y":2026,"cx":"General","ob":"journal-manuscript","fam":"correlation","form":"Spearman rank correlation","estd":"correlation","v":0.442,"n":"2780","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"average-of-k","scale":"LLM 1*-4* stars; author nine-point (averaged, fractional)","field":"health and life sciences (6 REF2021 UoAs)","wr":"LLM-ensemble (median) vs expert research-quality scores of journal articles","conf":"med","self":false,"doi":"10.1007/s11192-026-05585-2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"The median of 23 LLM configurations' scores for 2780 journal articles correlated with one expert's individual quality scores at a Spearman rho of 0.44 across all six fields.","vf":"unverified"},{"key":"QCWPTK8N","au":"Thelwall, Mike","y":2026,"cx":"General","ob":"journal-manuscript","fam":"correlation","form":"Spearman rank correlation","estd":"correlation","v":0.465,"n":"2780","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"average-of-k","scale":"LLM 1*-4* stars; author nine-point (averaged, fractional)","field":"health and life sciences (6 REF2021 UoAs)","wr":"LLM-ensemble (rank average) vs expert research-quality scores of journal articles","conf":"med","self":false,"doi":"10.1007/s11192-026-05585-2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"The mean rank of the articles across the 23 LLM configurations correlated with one expert's individual quality scores at a Spearman rho of 0.47 across all six fields.","vf":"unverified"},{"key":"QCWPTK8N","au":"Thelwall, Mike","y":2026,"cx":"General","ob":"journal-manuscript","fam":"correlation","form":"Spearman rank correlation (average of 10 folds)","estd":"correlation","v":0.497,"n":"2780","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"average-of-k","scale":"LLM 1*-4* stars (weighted sum); author nine-point (averaged, fractional)","field":"health and life sciences (6 REF2021 UoAs)","wr":"optimised weighted-sum LLM ensemble vs expert research-quality scores","conf":"med","self":false,"doi":"10.1007/s11192-026-05585-2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"A differential-evolution weighted sum of the 23 LLM configurations' scores, fitted with tenfold cross-validation, correlated with one expert's individual quality scores at a Spearman rho of 0.50 across all six fields.","vf":"unverified"},{"key":"RRPR8N5D","au":"Timmer, Antje","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"intra class correlation coefficient (Ri), ratio of between-abstract variance to overall variance","estd":"ICC (single/unspec)","v":0.6,"n":"19","k":"2","samp":"re-reviewed-subset","blind":"double","agg":"single-rater","scale":"0 to 1 summary quality score","field":"biomedical (gastroenterology)","wr":"two raters on abstract quality summary scores","conf":"high","self":false,"doi":"10.1186/1471-2288-3-2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two trained raters independently scored the formal quality of 19 basic-science abstracts submitted to the AGA meeting; an intraclass correlation of 0.60 indicates moderate agreement between them. n_ratings_total derived from stated complete crossing.","vf":"unverified"},{"key":"RRPR8N5D","au":"Timmer, Antje","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"intra class correlation coefficient (Ri), ratio of between-abstract variance to overall variance","estd":"ICC (single/unspec)","v":0.81,"n":"42","k":"2","samp":"re-reviewed-subset","blind":"double","agg":"single-rater","scale":"0 to 1 summary quality score","field":"biomedical (gastroenterology)","wr":"two raters on abstract quality summary scores","conf":"high","self":false,"doi":"10.1186/1471-2288-3-2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Two trained raters independently scored the formal quality of 42 controlled-trial abstracts submitted to the AGA meeting; an intraclass correlation of 0.81 indicates good agreement between them. n_ratings_total derived from stated complete crossing (both raters scored all abstracts).","vf":"unverified"},{"key":"RRPR8N5D","au":"Timmer, Antje","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"intra class correlation coefficient (Ri), ratio of between-abstract variance to overall variance","estd":"ICC (single/unspec)","v":0.67,"n":"39","k":"2","samp":"re-reviewed-subset","blind":"double","agg":"single-rater","scale":"0 to 1 summary quality score","field":"biomedical (gastroenterology)","wr":"two raters on abstract quality summary scores","conf":"high","self":false,"doi":"10.1186/1471-2288-3-2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two trained raters independently scored the formal quality of 39 other-clinical-research abstracts submitted to the AGA meeting; an intraclass correlation of 0.67 indicates moderate agreement between them. n_ratings_total derived from stated complete crossing.","vf":"unverified"},{"key":"RRPR8N5D","au":"Timmer, Antje","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"ICC","form":"intra class correlation coefficient","estd":"ICC (single/unspec)","v":0.85,"n":"85","k":"1","samp":"re-reviewed-subset","blind":"double","agg":"single-rater","scale":"0 to 1 summary quality score","field":"biomedical (gastroenterology)","wr":"single rater re-scoring abstract quality summary scores","conf":"high","self":false,"doi":"10.1186/1471-2288-3-2","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"One rater re-scored the quality of 85 abstracts after four to six weeks; a test-retest intraclass correlation of 0.85 indicates good stability of the summary scores over time. n_ratings_total derived from the stated two-occasion design.","vf":"unverified"},{"key":"RRPR8N5D","au":"Timmer, Antje","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"kappa coefficient per item (test-retest); agreement definition not stated","estd":"kappa","v":0.85,"n":"85","k":"1","samp":"re-reviewed-subset","blind":"double","agg":"single-rater","scale":"per-item categorical rating","field":"biomedical (gastroenterology)","wr":"single rater re-scoring individual instrument items","conf":"high","self":false,"doi":"10.1186/1471-2288-3-2","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"For the individual instrument items, one rater re-scored 85 abstracts after four to six weeks; the highest per-item test-retest kappa reported was 0.85 (only the range endpoints are given, not per-item values).","vf":"unverified"},{"key":"RRPR8N5D","au":"Timmer, Antje","y":2003,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"kappa coefficient per item (test-retest); agreement definition not stated","estd":"kappa","v":0.54,"n":"85","k":"1","samp":"re-reviewed-subset","blind":"double","agg":"single-rater","scale":"per-item categorical rating","field":"biomedical (gastroenterology)","wr":"single rater re-scoring individual instrument items","conf":"high","self":false,"doi":"10.1186/1471-2288-3-2","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"For the individual instrument items, one rater re-scored 85 abstracts after four to six weeks; the lowest per-item test-retest kappa reported was 0.54 (only the range endpoints are given, not per-item values).","vf":"unverified"},{"key":"A42AQLJA","au":"Tran, The Phi","y":2025,"cx":"General","ob":"other","fam":"weighted-kappa","form":"Cohen quadratic weighted kappa, averaged over three design criteria","estd":"weighted kappa","v":0.507,"n":"40","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, very unsatisfactory to very satisfactory","field":"English language education","wr":"raters scoring research proposals' design criteria","conf":"med","self":false,"doi":"10.1007/978-3-032-01348-4_3","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"ChatGPT-4o scored 40 proposals twice on the three research-design criteria; an average quadratic weighted kappa of 0.507 indicates moderate consistency between its two rounds.","vf":"unverified"},{"key":"A42AQLJA","au":"Tran, The Phi","y":2025,"cx":"General","ob":"other","fam":"weighted-kappa","form":"Cohen quadratic weighted kappa, averaged over three design criteria","estd":"weighted kappa","v":0.35,"n":"40","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, very unsatisfactory to very satisfactory","field":"English language education","wr":"human and AI raters on proposal design criteria","conf":"med","self":false,"doi":"10.1007/978-3-032-01348-4_3","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Human consensus scores were compared with ChatGPT's first run across 40 proposals on the research-design criteria; an average weighted kappa of 0.350 indicates fair agreement.","vf":"unverified"},{"key":"A42AQLJA","au":"Tran, The Phi","y":2025,"cx":"General","ob":"other","fam":"weighted-kappa","form":"Cohen quadratic weighted kappa, averaged over three design criteria","estd":"weighted kappa","v":0.389,"n":"40","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, very unsatisfactory to very satisfactory","field":"English language education","wr":"human and AI raters on proposal design criteria","conf":"med","self":false,"doi":"10.1007/978-3-032-01348-4_3","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Human consensus scores were compared with ChatGPT's second run across 40 proposals on the research-design criteria; an average weighted kappa of 0.389 indicates fair agreement.","vf":"unverified"},{"key":"A42AQLJA","au":"Tran, The Phi","y":2025,"cx":"General","ob":"other","fam":"weighted-kappa","form":"Cohen quadratic weighted kappa, averaged over two paradigm criteria","estd":"weighted kappa","v":0.509,"n":"40","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, very unsatisfactory to very satisfactory","field":"English language education","wr":"raters scoring research proposals' paradigm criteria","conf":"med","self":false,"doi":"10.1007/978-3-032-01348-4_3","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"ChatGPT-4o scored 40 proposals twice on the two research-paradigm criteria; an average quadratic weighted kappa of 0.509 indicates moderate consistency between its two rounds.","vf":"unverified"},{"key":"A42AQLJA","au":"Tran, The Phi","y":2025,"cx":"General","ob":"other","fam":"weighted-kappa","form":"Cohen quadratic weighted kappa, averaged over two paradigm criteria","estd":"weighted kappa","v":0.461,"n":"40","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, very unsatisfactory to very satisfactory","field":"English language education","wr":"human and AI raters on proposal paradigm criteria","conf":"med","self":false,"doi":"10.1007/978-3-032-01348-4_3","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Human consensus scores were compared with ChatGPT's first run across 40 proposals on the research-paradigm criteria; an average weighted kappa of 0.461 indicates moderate agreement.","vf":"unverified"},{"key":"A42AQLJA","au":"Tran, The Phi","y":2025,"cx":"General","ob":"other","fam":"weighted-kappa","form":"Cohen quadratic weighted kappa, averaged over two paradigm criteria","estd":"weighted kappa","v":0.585,"n":"40","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, very unsatisfactory to very satisfactory","field":"English language education","wr":"human and AI raters on proposal paradigm criteria","conf":"med","self":false,"doi":"10.1007/978-3-032-01348-4_3","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":false,"ms":"Human consensus scores were compared with ChatGPT's second run across 40 proposals on the research-paradigm criteria; an average weighted kappa of 0.585 indicates moderate agreement, the study's most highlighted inter-rater result (no single grand overall kappa is reported).","vf":"unverified"},{"key":"A42AQLJA","au":"Tran, The Phi","y":2025,"cx":"General","ob":"other","fam":"weighted-kappa","form":"Cohen quadratic weighted kappa, averaged over six research-question criteria","estd":"weighted kappa","v":0.356,"n":"40","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, very unsatisfactory to very satisfactory","field":"English language education","wr":"raters scoring research proposals' question criteria","conf":"med","self":false,"doi":"10.1007/978-3-032-01348-4_3","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"ChatGPT-4o scored 40 proposals twice on the six research-question criteria; an average quadratic weighted kappa of 0.356 indicates fair consistency between its two rounds.","vf":"unverified"},{"key":"A42AQLJA","au":"Tran, The Phi","y":2025,"cx":"General","ob":"other","fam":"weighted-kappa","form":"Cohen quadratic weighted kappa, averaged over six research-question criteria","estd":"weighted kappa","v":0.248,"n":"40","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, very unsatisfactory to very satisfactory","field":"English language education","wr":"human and AI raters on proposal question criteria","conf":"med","self":false,"doi":"10.1007/978-3-032-01348-4_3","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Human consensus scores were compared with ChatGPT's first run across 40 proposals on the research-question criteria; an average weighted kappa of 0.248 indicates fair agreement.","vf":"unverified"},{"key":"A42AQLJA","au":"Tran, The Phi","y":2025,"cx":"General","ob":"other","fam":"weighted-kappa","form":"Cohen quadratic weighted kappa, averaged over six research-question criteria","estd":"weighted kappa","v":0.135,"n":"40","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, very unsatisfactory to very satisfactory","field":"English language education","wr":"human and AI raters on proposal question criteria","conf":"med","self":false,"doi":"10.1007/978-3-032-01348-4_3","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Human consensus scores were compared with ChatGPT's second run across 40 proposals on the research-question criteria; an average weighted kappa of 0.135 indicates only slight agreement.","vf":"unverified"},{"key":"A42AQLJA","au":"Tran, The Phi","y":2025,"cx":"General","ob":"other","fam":"weighted-kappa","form":"Cohen quadratic weighted kappa, averaged over five techniques criteria","estd":"weighted kappa","v":0.526,"n":"40","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, very unsatisfactory to very satisfactory","field":"English language education","wr":"raters scoring research proposals' techniques criteria","conf":"med","self":false,"doi":"10.1007/978-3-032-01348-4_3","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"ChatGPT-4o scored 40 proposals twice on the five research-techniques criteria; an average quadratic weighted kappa of 0.526 indicates moderate consistency between its two rounds.","vf":"unverified"},{"key":"A42AQLJA","au":"Tran, The Phi","y":2025,"cx":"General","ob":"other","fam":"weighted-kappa","form":"Cohen quadratic weighted kappa, averaged over five techniques criteria","estd":"weighted kappa","v":0.318,"n":"40","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, very unsatisfactory to very satisfactory","field":"English language education","wr":"human and AI raters on proposal techniques criteria","conf":"med","self":false,"doi":"10.1007/978-3-032-01348-4_3","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Human consensus scores were compared with ChatGPT's first run across 40 proposals on the research-techniques criteria; an average weighted kappa of 0.318 indicates fair agreement.","vf":"unverified"},{"key":"A42AQLJA","au":"Tran, The Phi","y":2025,"cx":"General","ob":"other","fam":"weighted-kappa","form":"Cohen quadratic weighted kappa, averaged over five techniques criteria","estd":"weighted kappa","v":0.419,"n":"40","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, very unsatisfactory to very satisfactory","field":"English language education","wr":"human and AI raters on proposal techniques criteria","conf":"med","self":false,"doi":"10.1007/978-3-032-01348-4_3","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Human consensus scores were compared with ChatGPT's second run across 40 proposals on the research-techniques criteria; an average weighted kappa of 0.419 indicates fair-to-moderate agreement.","vf":"unverified"},{"key":"A42AQLJA","au":"Tran, The Phi","y":2025,"cx":"General","ob":"other","fam":"weighted-kappa","form":"Cohen quadratic weighted kappa, averaged over five title criteria","estd":"weighted kappa","v":0.497,"n":"40","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, very unsatisfactory to very satisfactory","field":"English language education","wr":"raters scoring research proposals' title criteria","conf":"med","self":false,"doi":"10.1007/978-3-032-01348-4_3","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"ChatGPT-4o scored 40 student research proposals twice on the five research-title criteria; an average quadratic weighted kappa of 0.497 indicates moderate consistency between its two rounds.","vf":"unverified"},{"key":"A42AQLJA","au":"Tran, The Phi","y":2025,"cx":"General","ob":"other","fam":"weighted-kappa","form":"Cohen quadratic weighted kappa, averaged over five title criteria","estd":"weighted kappa","v":0.159,"n":"40","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, very unsatisfactory to very satisfactory","field":"English language education","wr":"human and AI raters on proposal title criteria","conf":"med","self":false,"doi":"10.1007/978-3-032-01348-4_3","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Human lecturers' consensus scores were compared with ChatGPT's first run across 40 proposals on the research-title criteria; an average weighted kappa of 0.159 indicates only slight agreement.","vf":"unverified"},{"key":"A42AQLJA","au":"Tran, The Phi","y":2025,"cx":"General","ob":"other","fam":"weighted-kappa","form":"Cohen quadratic weighted kappa, averaged over five title criteria","estd":"weighted kappa","v":0.238,"n":"40","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, very unsatisfactory to very satisfactory","field":"English language education","wr":"human and AI raters on proposal title criteria","conf":"med","self":false,"doi":"10.1007/978-3-032-01348-4_3","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Human lecturers' consensus scores were compared with ChatGPT's second run across 40 proposals on the research-title criteria; an average weighted kappa of 0.238 indicates slight-to-fair agreement.","vf":"unverified"},{"key":"E7836ZR3","au":"Tran, The Phi","y":2025,"cx":"Analogue","ob":"other","fam":"weighted-kappa","form":"Quadratic Cohen's weighted Kappa","estd":"weighted kappa","v":0.269,"n":"37","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, 1=Very Unsatisfactory, 5=Very Satisfactory","field":"English-language education / applied linguistics","wr":"lecturers vs ChatGPT on proposal design scores","conf":"med","self":false,"doi":"10.54855/callej.252637","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Two lecturers' consensus scores and ChatGPT-4o's first-round scores on 37 proposals' design criteria agreed at a mean quadratic weighted kappa of 0.27, a fair level of inter-rater agreement.","vf":"unverified"},{"key":"E7836ZR3","au":"Tran, The Phi","y":2025,"cx":"Analogue","ob":"other","fam":"weighted-kappa","form":"Quadratic Cohen's weighted Kappa","estd":"weighted kappa","v":0.223,"n":"37","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, 1=Very Unsatisfactory, 5=Very Satisfactory","field":"English-language education / applied linguistics","wr":"lecturers vs ChatGPT on proposal design scores","conf":"med","self":false,"doi":"10.54855/callej.252637","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Two lecturers' consensus scores and ChatGPT-4o's second-round scores on 37 proposals' design criteria agreed at a mean quadratic weighted kappa of 0.22, a fair level of inter-rater agreement.","vf":"unverified"},{"key":"E7836ZR3","au":"Tran, The Phi","y":2025,"cx":"Analogue","ob":"other","fam":"weighted-kappa","form":"Quadratic Cohen's weighted Kappa","estd":"weighted kappa","v":0.42,"n":"37","k":"1","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, 1=Very Unsatisfactory, 5=Very Satisfactory","field":"English-language education / applied linguistics","wr":"ChatGPT self-consistency on proposal design scores","conf":"med","self":false,"doi":"10.54855/callej.252637","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"ChatGPT-4o scored 37 proposals on research-design criteria twice; a mean quadratic weighted kappa of 0.42 indicates moderate self-consistency across its two rounds.","vf":"unverified"},{"key":"E7836ZR3","au":"Tran, The Phi","y":2025,"cx":"Analogue","ob":"other","fam":"weighted-kappa","form":"Quadratic Cohen's weighted Kappa","estd":"weighted kappa","v":0.292,"n":"","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, 1=Very Unsatisfactory, 5=Very Satisfactory","field":"English-language education / applied linguistics","wr":"lecturers vs ChatGPT on proposal hypothesis scores","conf":"med","self":false,"doi":"10.54855/callej.252637","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Two lecturers' consensus scores and ChatGPT-4o's first-round scores on hypothesis criteria, for proposals with explicit hypotheses, agreed at a mean quadratic weighted kappa of 0.29, a fair level of inter-rater agreement.","vf":"unverified"},{"key":"E7836ZR3","au":"Tran, The Phi","y":2025,"cx":"Analogue","ob":"other","fam":"weighted-kappa","form":"Quadratic Cohen's weighted Kappa","estd":"weighted kappa","v":0.431,"n":"","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, 1=Very Unsatisfactory, 5=Very Satisfactory","field":"English-language education / applied linguistics","wr":"lecturers vs ChatGPT on proposal hypothesis scores","conf":"med","self":false,"doi":"10.54855/callej.252637","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Two lecturers' consensus scores and ChatGPT-4o's second-round scores on hypothesis criteria, for proposals with explicit hypotheses, agreed at a mean quadratic weighted kappa of 0.43, a moderate level of inter-rater agreement.","vf":"unverified"},{"key":"E7836ZR3","au":"Tran, The Phi","y":2025,"cx":"Analogue","ob":"other","fam":"weighted-kappa","form":"Quadratic Cohen's weighted Kappa","estd":"weighted kappa","v":0.758,"n":"","k":"1","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, 1=Very Unsatisfactory, 5=Very Satisfactory","field":"English-language education / applied linguistics","wr":"ChatGPT self-consistency on proposal hypothesis scores","conf":"med","self":false,"doi":"10.54855/callej.252637","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"ChatGPT-4o twice scored the hypothesis criteria of proposals with explicitly stated hypotheses; a mean quadratic weighted kappa of 0.76 indicates high self-consistency across its two rounds.","vf":"unverified"},{"key":"E7836ZR3","au":"Tran, The Phi","y":2025,"cx":"Analogue","ob":"other","fam":"weighted-kappa","form":"Quadratic Cohen's weighted Kappa","estd":"weighted kappa","v":1,"n":"","k":"1","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, 1=Very Unsatisfactory, 5=Very Satisfactory","field":"English-language education / applied linguistics","wr":"ChatGPT self-consistency on hypothesis testability scores","conf":"med","self":false,"doi":"10.54855/callej.252637","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"ChatGPT-4o twice judged whether explicitly stated hypotheses were empirically testable; the quadratic weighted kappa of 1.000 is the highest item-level value in the paper, indicating perfect repeatability. Included as the maximum item-level stratum.","vf":"unverified"},{"key":"E7836ZR3","au":"Tran, The Phi","y":2025,"cx":"Analogue","ob":"other","fam":"weighted-kappa","form":"Quadratic Cohen's weighted Kappa","estd":"weighted kappa","v":-0.5,"n":"","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, 1=Very Unsatisfactory, 5=Very Satisfactory","field":"English-language education / applied linguistics","wr":"lecturers vs ChatGPT on hypothesis accessibility scores","conf":"med","self":false,"doi":"10.54855/callej.252637","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Lecturers' consensus scores for hypothesis accessibility were compared with ChatGPT-4o's first-round scores; the quadratic weighted kappa of -0.500 is the lowest item-level value in the paper. Included as the minimum item-level stratum.","vf":"unverified"},{"key":"E7836ZR3","au":"Tran, The Phi","y":2025,"cx":"Analogue","ob":"other","fam":"weighted-kappa","form":"Quadratic Cohen's weighted Kappa","estd":"weighted kappa","v":0.361,"n":"37","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, 1=Very Unsatisfactory, 5=Very Satisfactory","field":"English-language education / applied linguistics","wr":"lecturers vs ChatGPT on proposal paradigm scores","conf":"med","self":false,"doi":"10.54855/callej.252637","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Two lecturers' consensus scores and ChatGPT-4o's first-round scores on 37 proposals' paradigm criteria agreed at a mean quadratic weighted kappa of 0.36, a fair level of inter-rater agreement.","vf":"unverified"},{"key":"E7836ZR3","au":"Tran, The Phi","y":2025,"cx":"Analogue","ob":"other","fam":"weighted-kappa","form":"Quadratic Cohen's weighted Kappa","estd":"weighted kappa","v":0.326,"n":"37","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, 1=Very Unsatisfactory, 5=Very Satisfactory","field":"English-language education / applied linguistics","wr":"lecturers vs ChatGPT on proposal paradigm scores","conf":"med","self":false,"doi":"10.54855/callej.252637","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Two lecturers' consensus scores and ChatGPT-4o's second-round scores on 37 proposals' paradigm criteria agreed at a mean quadratic weighted kappa of 0.33, a fair level of inter-rater agreement.","vf":"unverified"},{"key":"E7836ZR3","au":"Tran, The Phi","y":2025,"cx":"Analogue","ob":"other","fam":"weighted-kappa","form":"Quadratic Cohen's weighted Kappa","estd":"weighted kappa","v":0.703,"n":"37","k":"1","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, 1=Very Unsatisfactory, 5=Very Satisfactory","field":"English-language education / applied linguistics","wr":"ChatGPT self-consistency on proposal paradigm scores","conf":"med","self":false,"doi":"10.54855/callej.252637","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"ChatGPT-4o scored 37 proposals on research-paradigm criteria twice; a mean quadratic weighted kappa of 0.70 indicates substantial self-consistency across its two rounds.","vf":"unverified"},{"key":"E7836ZR3","au":"Tran, The Phi","y":2025,"cx":"Analogue","ob":"other","fam":"weighted-kappa","form":"Quadratic Cohen's weighted Kappa","estd":"weighted kappa","v":0.126,"n":"37","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, 1=Very Unsatisfactory, 5=Very Satisfactory","field":"English-language education / applied linguistics","wr":"lecturers vs ChatGPT on proposal question scores","conf":"med","self":false,"doi":"10.54855/callej.252637","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Two lecturers' consensus scores and ChatGPT-4o's first-round scores on 37 proposals' question criteria agreed at a mean quadratic weighted kappa of 0.13, a slight level of inter-rater agreement.","vf":"unverified"},{"key":"E7836ZR3","au":"Tran, The Phi","y":2025,"cx":"Analogue","ob":"other","fam":"weighted-kappa","form":"Quadratic Cohen's weighted Kappa","estd":"weighted kappa","v":0.092,"n":"37","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, 1=Very Unsatisfactory, 5=Very Satisfactory","field":"English-language education / applied linguistics","wr":"lecturers vs ChatGPT on proposal question scores","conf":"med","self":false,"doi":"10.54855/callej.252637","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Two lecturers' consensus scores and ChatGPT-4o's second-round scores on 37 proposals' question criteria agreed at a mean quadratic weighted kappa of 0.09, a slight level of inter-rater agreement.","vf":"unverified"},{"key":"E7836ZR3","au":"Tran, The Phi","y":2025,"cx":"Analogue","ob":"other","fam":"weighted-kappa","form":"Quadratic Cohen's weighted Kappa","estd":"weighted kappa","v":0.341,"n":"37","k":"1","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, 1=Very Unsatisfactory, 5=Very Satisfactory","field":"English-language education / applied linguistics","wr":"ChatGPT self-consistency on proposal question scores","conf":"med","self":false,"doi":"10.54855/callej.252637","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"ChatGPT-4o scored 37 proposals on research-question criteria twice; a mean quadratic weighted kappa of 0.34 indicates only fair self-consistency across its two rounds.","vf":"unverified"},{"key":"E7836ZR3","au":"Tran, The Phi","y":2025,"cx":"Analogue","ob":"other","fam":"weighted-kappa","form":"Quadratic Cohen's weighted Kappa","estd":"weighted kappa","v":0.186,"n":"37","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, 1=Very Unsatisfactory, 5=Very Satisfactory","field":"English-language education / applied linguistics","wr":"lecturers vs ChatGPT on proposal technique scores","conf":"med","self":false,"doi":"10.54855/callej.252637","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Two lecturers' consensus scores and ChatGPT-4o's first-round scores on 37 proposals' technique criteria agreed at a mean quadratic weighted kappa of 0.19, a slight level of inter-rater agreement.","vf":"unverified"},{"key":"E7836ZR3","au":"Tran, The Phi","y":2025,"cx":"Analogue","ob":"other","fam":"weighted-kappa","form":"Quadratic Cohen's weighted Kappa","estd":"weighted kappa","v":0.117,"n":"37","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, 1=Very Unsatisfactory, 5=Very Satisfactory","field":"English-language education / applied linguistics","wr":"lecturers vs ChatGPT on proposal technique scores","conf":"med","self":false,"doi":"10.54855/callej.252637","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Two lecturers' consensus scores and ChatGPT-4o's second-round scores on 37 proposals' technique criteria agreed at a mean quadratic weighted kappa of 0.12, a slight level of inter-rater agreement.","vf":"unverified"},{"key":"E7836ZR3","au":"Tran, The Phi","y":2025,"cx":"Analogue","ob":"other","fam":"weighted-kappa","form":"Quadratic Cohen's weighted Kappa","estd":"weighted kappa","v":0.446,"n":"37","k":"1","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, 1=Very Unsatisfactory, 5=Very Satisfactory","field":"English-language education / applied linguistics","wr":"ChatGPT self-consistency on proposal technique scores","conf":"med","self":false,"doi":"10.54855/callej.252637","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"ChatGPT-4o scored 37 proposals on research-technique criteria twice; a mean quadratic weighted kappa of 0.45 indicates moderate self-consistency across its two rounds.","vf":"unverified"},{"key":"E7836ZR3","au":"Tran, The Phi","y":2025,"cx":"Analogue","ob":"other","fam":"weighted-kappa","form":"Quadratic Cohen's weighted Kappa","estd":"weighted kappa","v":0.248,"n":"37","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, 1=Very Unsatisfactory, 5=Very Satisfactory","field":"English-language education / applied linguistics","wr":"lecturers vs ChatGPT on proposal title scores","conf":"med","self":false,"doi":"10.54855/callej.252637","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":false,"ms":"Two lecturers' consensus scores and ChatGPT-4o's first-round scores on 37 proposals' title criteria agreed at a mean quadratic weighted kappa of 0.25, a low level of inter-rater agreement.","vf":"unverified"},{"key":"E7836ZR3","au":"Tran, The Phi","y":2025,"cx":"Analogue","ob":"other","fam":"weighted-kappa","form":"Quadratic Cohen's weighted Kappa","estd":"weighted kappa","v":0.099,"n":"37","k":"2","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, 1=Very Unsatisfactory, 5=Very Satisfactory","field":"English-language education / applied linguistics","wr":"lecturers vs ChatGPT on proposal title scores","conf":"med","self":false,"doi":"10.54855/callej.252637","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Two lecturers' consensus scores and ChatGPT-4o's second-round scores on 37 proposals' title criteria agreed at a mean quadratic weighted kappa of 0.10, a slight level of inter-rater agreement.","vf":"unverified"},{"key":"E7836ZR3","au":"Tran, The Phi","y":2025,"cx":"Analogue","ob":"other","fam":"weighted-kappa","form":"Quadratic Cohen's weighted Kappa","estd":"weighted kappa","v":0.354,"n":"37","k":"1","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1-5, 1=Very Unsatisfactory, 5=Very Satisfactory","field":"English-language education / applied linguistics","wr":"ChatGPT self-consistency on proposal title scores","conf":"med","self":false,"doi":"10.54855/callej.252637","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"ChatGPT-4o scored 37 student research proposals on title criteria twice; a mean quadratic weighted kappa of 0.35 shows only fair self-consistency between its two rounds.","vf":"unverified"},{"key":"GPQAWJJ9","au":"Turner, D. P.","y":2017,"cx":"Journal","ob":"journal-manuscript","fam":"G-theory","form":"variance component as % of total variance; item facet; two-facet partially nested generalizability study","estd":"G-theory","v":0.1522,"n":"","k":"","samp":"unclear","blind":"unclear","agg":"unspecified","scale":"five-item rating scale","field":"unclear","wr":"rating items in the five-item publishability instrument","conf":"low","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"In the same generalisability study, the rating items accounted for 15.22% of the total variance in publishability scores.","vf":"unverified"},{"key":"GPQAWJJ9","au":"Turner, D. P.","y":2017,"cx":"Journal","ob":"journal-manuscript","fam":"G-theory","form":"variance component as % of total variance; manuscript (object of measurement) facet; two-facet partially nested generalizability study","estd":"G-theory","v":0.1221,"n":"","k":"","samp":"unclear","blind":"unclear","agg":"unspecified","scale":"five-item rating scale","field":"unclear","wr":"true between-manuscript variance in publishability scores","conf":"low","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"In the same generalisability study, true differences between manuscripts (the object variance) accounted for only 12.21% of the total variance in publishability scores.","vf":"unverified"},{"key":"GPQAWJJ9","au":"Turner, D. P.","y":2017,"cx":"Journal","ob":"journal-manuscript","fam":"G-theory","form":"variance component as % of total variance; reviewers nested within manuscripts; two-facet partially nested generalizability study","estd":"G-theory","v":0.3548,"n":"","k":"","samp":"unclear","blind":"unclear","agg":"unspecified","scale":"five-item rating scale","field":"unclear","wr":"reviewers on manuscript publishability scores","conf":"low","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":true,"he":false,"ms":"In a two-facet partially nested generalisability study of 635 journal peer reviews scored on a five-item scale, the reviewer facet (reviewers nested within manuscripts) accounted for 35.48% of the variance in publishability scores, indicating substantial rater-related error variance.","vf":"unverified"},{"key":"85KL54GC","au":"Tymms, Peter","y":2017,"cx":"General","ob":"other","fam":"other","form":"proportion of paper-measure variance located at the academic level (multilevel model)","estd":"other","v":0.668,"n":"42","k":"","samp":"special","blind":"unclear","agg":"unspecified","scale":"0 (unclassified) to 4* (5-point REF scale)","field":"education","wr":"share of paper-quality variance between academics","conf":"med","self":false,"doi":"10.1080/03075079.2016.1266609","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"In a multilevel model with papers nested within academics, 66.8% of the variance in paper measures lay between academics rather than within them.","vf":"unverified"},{"key":"85KL54GC","au":"Tymms, Peter","y":2017,"cx":"General","ob":"other","fam":"other","form":"reliability of academic-level measure via Goldstein shrinkage formula (three papers)","estd":"other","v":0.86,"n":"42","k":"","samp":"special","blind":"unclear","agg":"average-of-k","scale":"0 (unclassified) to 4* (5-point REF scale)","field":"education","wr":"reliability of academics' aggregate quality measure","conf":"med","self":false,"doi":"10.1080/03075079.2016.1266609","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"The reliability of an academic's aggregate quality measure dropped to 0.86 when based on three papers rather than four.","vf":"unverified"},{"key":"85KL54GC","au":"Tymms, Peter","y":2017,"cx":"General","ob":"other","fam":"other","form":"reliability of academic-level measure via Goldstein shrinkage formula (four papers)","estd":"other","v":0.89,"n":"42","k":"","samp":"special","blind":"unclear","agg":"average-of-k","scale":"0 (unclassified) to 4* (5-point REF scale)","field":"education","wr":"reliability of academics' aggregate quality measure","conf":"med","self":false,"doi":"10.1080/03075079.2016.1266609","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"Using a multilevel shrinkage formula, the reliability of an academic's aggregate quality measure across 42 academics was 0.89 when based on four papers.","vf":"unverified"},{"key":"85KL54GC","au":"Tymms, Peter","y":2017,"cx":"General","ob":"other","fam":"correlation","form":"point-biserial correlation of a rater's ratings with the overall Rasch measure","estd":"correlation","v":0.82,"n":"","k":"","samp":"special","blind":"unclear","agg":"single-rater","scale":"0 (unclassified) to 4* (5-point REF scale)","field":"education","wr":"typical raters' agreement with consensus measure","conf":"med","self":false,"doi":"10.1080/03075079.2016.1266609","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Excluding two misfitting raters, the remaining 21 raters' ratings correlated on average 0.82 (SD 0.10) with the overall Rasch paper measure, indicating high agreement with the consensus.","vf":"unverified"},{"key":"85KL54GC","au":"Tymms, Peter","y":2017,"cx":"General","ob":"other","fam":"other","form":"Rasch (Winsteps) paper/person separation reliability","estd":"other","v":0.76,"n":"224","k":"","samp":"special","blind":"unclear","agg":"unspecified","scale":"0 (unclassified) to 4* (5-point REF scale)","field":"education","wr":"senior staff rating research papers","conf":"med","self":false,"doi":"10.1080/03075079.2016.1266609","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":true,"ms":"Twenty-three raters scored 224 research papers (710 ratings) on the 0-4* REF scale; a Rasch paper separation reliability of 0.76 indicates the paper measures were moderately reliably separated by the raters.","vf":"unverified"},{"key":"85KL54GC","au":"Tymms, Peter","y":2017,"cx":"General","ob":"other","fam":"percent-agreement","form":"exact agreement: percentage of papers given identical ratings by all their raters","estd":"percent agreement","v":0.25,"n":"","k":"","samp":"special","blind":"unclear","agg":"unspecified","scale":"0 (unclassified) to 4* (5-point REF scale)","field":"education","wr":"raters agreeing exactly on paper ratings","conf":"med","self":false,"doi":"10.1080/03075079.2016.1266609","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Among papers rated by two or more raters, 25% received identical ratings from all their raters (exact agreement).","vf":"unverified"},{"key":"85KL54GC","au":"Tymms, Peter","y":2017,"cx":"General","ob":"other","fam":"percent-agreement","form":"agreement within one point on the 5-point REF scale","estd":"percent agreement","v":0.72,"n":"","k":"","samp":"special","blind":"unclear","agg":"unspecified","scale":"0 (unclassified) to 4* (5-point REF scale)","field":"education","wr":"raters agreeing within one point on paper ratings","conf":"med","self":false,"doi":"10.1080/03075079.2016.1266609","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"For 72% of papers rated by two or more raters, the ratings differed by no more than one point on the 5-point REF scale.","vf":"unverified"},{"key":"85KL54GC","au":"Tymms, Peter","y":2017,"cx":"General","ob":"other","fam":"correlation","form":"point-biserial correlation of rater J's ratings with the overall Rasch measure","estd":"correlation","v":0.44,"n":"","k":"","samp":"special","blind":"unclear","agg":"single-rater","scale":"0 (unclassified) to 4* (5-point REF scale)","field":"education","wr":"misfitting rater J agreement with consensus","conf":"med","self":false,"doi":"10.1080/03075079.2016.1266609","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Misfitting rater J's ratings correlated 0.44 with the overall Rasch measure, a low value indicating poor agreement with the other raters.","vf":"unverified"},{"key":"85KL54GC","au":"Tymms, Peter","y":2017,"cx":"General","ob":"other","fam":"correlation","form":"point-biserial correlation of rater L's ratings with the overall Rasch measure","estd":"correlation","v":0.36,"n":"","k":"","samp":"special","blind":"unclear","agg":"single-rater","scale":"0 (unclassified) to 4* (5-point REF scale)","field":"education","wr":"misfitting rater L agreement with consensus","conf":"med","self":false,"doi":"10.1080/03075079.2016.1266609","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Misfitting rater L's ratings correlated only 0.36 with the overall Rasch measure, a low value indicating poor agreement with the other raters.","vf":"unverified"},{"key":"85KL54GC","au":"Tymms, Peter","y":2017,"cx":"General","ob":"other","fam":"other","form":"Rasch (Winsteps) rater/item separation reliability","estd":"other","v":0.87,"n":"224","k":"","samp":"special","blind":"unclear","agg":"unspecified","scale":"0 (unclassified) to 4* (5-point REF scale)","field":"education","wr":"separation of the 23 raters' severities","conf":"med","self":false,"doi":"10.1080/03075079.2016.1266609","ciLow":null,"ciHigh":null,"mt":"other","tgt":"other","rr":"restricted-other","pr":false,"he":false,"ms":"A Rasch rater separation reliability of 0.87 indicates the 23 raters were reliably distinguished in their severity/leniency, meaning they differed systematically in how harshly they rated papers.","vf":"unverified"},{"key":"E7IAN24K","au":"UK Metascience Unit","y":2025,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"ICC One-Way Random (Absolute Agreement, Average Measures)","estd":"ICC (average)","v":0.4,"n":"80","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"scoring scale of 1-6","field":"metascience","wr":"resampled 3-reviewer scores of fellowship proposals","conf":"high","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":true,"he":false,"ms":"By randomly resampling three of the nine real reviews per application over 10,000 iterations, the mean ICC (same one-way random, absolute agreement, average measures form) fell to 0.4, showing that a typical three-reviewer design captures much less of the between-application quality signal. The value is simulated from the paper's own empirical review scores.","vf":"unverified"},{"key":"E7IAN24K","au":"UK Metascience Unit","y":2025,"cx":"Grant","ob":"fellowship","fam":"ICC","form":"ICC One-Way Random (Absolute Agreement, Average Measures)","estd":"ICC (average)","v":0.7,"n":"80","k":"9","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"scoring scale of 1-6","field":"metascience","wr":"applicants scoring each other's fellowship proposals","conf":"high","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"In a distributed peer review pilot, applicants independently scored each other's fellowship proposals on a 1-6 scale, with nine reviews per application across 80 proposals. An ICC (one-way random, absolute agreement, average measures) of 0.7 shows the averaged nine reviewers capture a substantial share of the between-application quality signal. Confidence intervals were shown graphically but not stated numerically.","vf":"unverified"},{"key":"E7IAN24K","au":"UK Metascience Unit","y":2025,"cx":"Grant","ob":"fellowship","fam":"other","form":"Smallest Detectable Difference (SDD)","estd":"other","v":1.8,"n":"80","k":"3","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"scoring scale of 1-6","field":"metascience","wr":"resampled 3-reviewer scores of fellowship proposals","conf":"high","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"When three of the nine real reviews per application were resampled over 10,000 iterations, the smallest detectable difference rose to 1.8 points on the 1-6 scale, meaning nearly twice the score change is needed to detect a real difference with three reviewers. The value is simulated from the paper's own empirical review scores.","vf":"unverified"},{"key":"E7IAN24K","au":"UK Metascience Unit","y":2025,"cx":"Grant","ob":"fellowship","fam":"other","form":"Smallest Detectable Difference (SDD)","estd":"other","v":1.1,"n":"80","k":"8","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"scoring scale of 1-6","field":"metascience","wr":"resampled 8-reviewer scores of fellowship proposals","conf":"high","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"When eight of the nine real reviews per application were resampled over 10,000 iterations, the smallest detectable difference was 1.1 points on the 1-6 scale, approaching the 1.0-point precision of the full nine-reviewer dataset. The value is simulated from the paper's own empirical review scores.","vf":"unverified"},{"key":"E7IAN24K","au":"UK Metascience Unit","y":2025,"cx":"Grant","ob":"fellowship","fam":"other","form":"Smallest Detectable Difference (SDD)","estd":"other","v":1,"n":"80","k":"9","samp":"full-pool","blind":"unclear","agg":"average-of-k","scale":"scoring scale of 1-6","field":"metascience","wr":"applicants scoring each other's fellowship proposals","conf":"high","self":false,"doi":"","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"With nine reviews per application across 80 proposals, the smallest detectable difference was 1.0 point on the 1-6 scale. Two applications whose mean scores differ by less than one point cannot be reliably distinguished, an agreement-based precision limit of the averaged nine reviewers. Confidence intervals were shown graphically but not stated numerically.","vf":"unverified"},{"key":"K5N2TIMX","au":"Verharen, Jeroen P. H.","y":2023,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"R2","estd":"correlation","v":0.859,"n":"","k":"1","samp":"funded-only","blind":"single","agg":"single-rater","scale":"-100 (rude) to 0 (neutral) to +100 (polite)","field":"neuroscience","wr":"politeness of peer-review reports","conf":"med","self":false,"doi":"10.7554/elife.90230.2","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"ChatGPT (a later model version) rated the politeness of the same review reports in two subsequent iterations. The two iterations shared 85.9% of their variance.","vf":"unverified"},{"key":"K5N2TIMX","au":"Verharen, Jeroen P. H.","y":2023,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"R2","estd":"correlation","v":0.992,"n":"","k":"1","samp":"funded-only","blind":"single","agg":"single-rater","scale":"-100 (negative) to 0 (neutral) to +100 (positive) sentiment","field":"neuroscience","wr":"sentiment of peer-review reports","conf":"med","self":false,"doi":"10.7554/elife.90230.2","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"ChatGPT (a later model version) rated the sentiment of the same review reports in two subsequent iterations. The two iterations shared 99.2% of their variance, indicating highly reproducible algorithmic scoring.","vf":"unverified"},{"key":"K5N2TIMX","au":"Verharen, Jeroen P. H.","y":2023,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"linear regression R2","estd":"correlation","v":0.7,"n":"7","k":"","samp":"funded-only","blind":"single","agg":"unspecified","scale":"human 1-5 (converted); ChatGPT -100 to +100","field":"neuroscience","wr":"politeness of peer-review reports","conf":"med","self":false,"doi":"10.7554/elife.90230.2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"Seven blinded human scorers and ChatGPT rated the politeness of seven review reports. The average human score shared 70% of its variance with ChatGPT's score.","vf":"unverified"},{"key":"K5N2TIMX","au":"Verharen, Jeroen P. H.","y":2023,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"linear regression R2","estd":"correlation","v":0.91,"n":"7","k":"","samp":"funded-only","blind":"single","agg":"unspecified","scale":"human 1-5 (converted); ChatGPT -100 to +100","field":"neuroscience","wr":"sentiment of peer-review reports","conf":"med","self":false,"doi":"10.7554/elife.90230.2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"restricted-top","pr":true,"he":false,"ms":"Seven blinded human scorers and ChatGPT rated the sentiment of seven review reports. The average human score shared 91% of its variance with ChatGPT's score, validating the algorithmic sentiment measure.","vf":"unverified"},{"key":"K5N2TIMX","au":"Verharen, Jeroen P. H.","y":2023,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"ICC type ICC1,1 (log-transformed sentiment scores)","estd":"ICC (single/unspec)","v":0.055,"n":"200","k":"","samp":"funded-only","blind":"single","agg":"single-rater","scale":"-100 (negative) to 0 (neutral) to +100 (positive) sentiment","field":"neuroscience","wr":"reviewers' favourability (sentiment) toward manuscripts","conf":"med","self":false,"doi":"10.7554/elife.90230.2","ciLow":-0.025,"ciHigh":0.144,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":true,"he":true,"ms":"Each of 200 accepted Nature Communications neuroscience manuscripts received two or more first-round reviews (572 in total), and ChatGPT-derived sentiment scores were compared across the reviewers of each paper. An ICC of 0.055 indicates poor agreement between reviewers on how favourable the same paper is.","vf":"unverified"},{"key":"K5N2TIMX","au":"Verharen, Jeroen P. H.","y":2023,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"linear regression variance explained (R2); maximum across pairwise reviewer comparisons","estd":"correlation","v":0.055,"n":"","k":"2","samp":"funded-only","blind":"single","agg":"single-rater","scale":"-100 (negative) to 0 (neutral) to +100 (positive) sentiment","field":"neuroscience","wr":"reviewers' favourability (sentiment) toward manuscripts","conf":"med","self":false,"doi":"10.7554/elife.90230.2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"A linear regression of ChatGPT-derived sentiment scores between reviewer 1 and reviewer 3 of the same manuscripts explained only 5.5% of the variance, the largest and only significant of the pairwise reviewer comparisons, again indicating very low agreement between reviewers.","vf":"unverified"},{"key":"K5N2TIMX","au":"Verharen, Jeroen P. H.","y":2023,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"linear regression R2","estd":"correlation","v":0.13,"n":"7","k":"","samp":"funded-only","blind":"single","agg":"unspecified","scale":"unspecified","field":"neuroscience","wr":"sentiment of peer-review reports","conf":"med","self":false,"doi":"10.7554/elife.90230.2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"TextBlob sentiment scores for seven review reports were compared with the average ratings of seven blinded human scorers. The scores shared 13% of their variance, a non-significant relation.","vf":"unverified"},{"key":"K5N2TIMX","au":"Verharen, Jeroen P. H.","y":2023,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"linear regression R2","estd":"correlation","v":0.07,"n":"7","k":"","samp":"funded-only","blind":"single","agg":"unspecified","scale":"unspecified","field":"neuroscience","wr":"sentiment of peer-review reports","conf":"med","self":false,"doi":"10.7554/elife.90230.2","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"VADER sentiment scores for seven review reports were compared with the average ratings of seven blinded human scorers. The scores shared 7% of their variance, a non-significant relation.","vf":"unverified"},{"key":"JX8TSFKY","au":"Watkins, Marley W.","y":1979,"cx":"Journal","ob":"journal-manuscript","fam":"Fleiss-kappa","form":"kappa (Cohen 1960, extended by Light 1971 and Fleiss 1971), computed with Fleiss's (1971) formulas","estd":"kappa","v":0.49,"n":"","k":"","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"Category 1 \"reject\" to Category 5 \"accept in present form\"","field":"psychology","wr":"reviewers on journal manuscript acceptability","conf":"high","self":false,"doi":"10.1037/0003-066x.34.9.796","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"Applying Fleiss's kappa to Scarr and Weber's (1978) American Psychologist reviewer ratings on a five-category scale, a kappa of 0.49 indicates moderate chance-corrected agreement between reviewers.","vf":"unverified"},{"key":"JX8TSFKY","au":"Watkins, Marley W.","y":1979,"cx":"Journal","ob":"journal-manuscript","fam":"Fleiss-kappa","form":"kappa (Cohen 1960, extended by Light 1971 and Fleiss 1971), computed with Fleiss's (1971) formulas","estd":"kappa","v":0.53,"n":"","k":"","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"dichotomy: ratings 1-2 \"reject\", 3-5 \"accept\"","field":"psychology","wr":"reviewers on journal manuscript accept/reject decision","conf":"high","self":false,"doi":"10.1037/0003-066x.34.9.796","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"Collapsing the American Psychologist 5-point ratings to an accept/reject dichotomy, a Fleiss kappa of 0.53 indicates reviewer agreement substantially beyond chance on the broad publishability decision.","vf":"unverified"},{"key":"JX8TSFKY","au":"Watkins, Marley W.","y":1979,"cx":"Journal","ob":"journal-manuscript","fam":"Fleiss-kappa","form":"kappa (Cohen 1960, extended by Light 1971 and Fleiss 1971), computed with Fleiss's (1971) formulas","estd":"kappa","v":0.15,"n":"","k":"","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"Category 1 \"definitely accept\" to Category 5 \"definitely reject\"","field":"psychology","wr":"reviewers on journal manuscript acceptability","conf":"high","self":false,"doi":"10.1037/0003-066x.34.9.796","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":true,"he":false,"ms":"Applying Fleiss's kappa to Hendrick's (1977) PSPB reviewer ratings on a five-category accept-reject scale, a kappa of 0.15 indicates only marginal chance-corrected agreement between manuscript reviewers.","vf":"unverified"},{"key":"JX8TSFKY","au":"Watkins, Marley W.","y":1979,"cx":"Journal","ob":"journal-manuscript","fam":"Fleiss-kappa","form":"kappa (Cohen 1960, extended by Light 1971 and Fleiss 1971), computed with Fleiss's (1971) formulas","estd":"kappa","v":0.091,"n":"","k":"","samp":"unclear","blind":"unclear","agg":"single-rater","scale":"dichotomy: ratings 1-3 \"possibly accept\", 4-5 \"reject\"","field":"psychology","wr":"reviewers on journal manuscript accept/reject decision","conf":"high","self":false,"doi":"10.1037/0003-066x.34.9.796","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"unclear","pr":false,"he":false,"ms":"Collapsing the PSPB 5-point ratings to an accept/reject dichotomy, a Fleiss kappa of 0.091 shows reviewer agreement no greater than chance on the broad publishability decision.","vf":"unverified"},{"key":"FAWCL6KR","au":"Whaley, Arthur L","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach's alpha, unstandardized, all raters","estd":"Cronbach alpha","v":0.75,"n":"20","k":"5","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"not at all (0), somewhat (1), or definitely (2)","field":"mental health","wr":"internal consistency of the Agency subscale items","conf":"high","self":false,"doi":"10.1177/0193841X05275586","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Unstandardized Cronbach's alpha of the Agency subscale, pooled across the four raters, was 0.75, indicating acceptable internal consistency. k is the number of items forming the subscale (5).","vf":"unverified"},{"key":"FAWCL6KR","au":"Whaley, Arthur L","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach's alpha, unstandardized, all raters","estd":"Cronbach alpha","v":0.71,"n":"20","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"not at all (0), somewhat (1), or definitely (2)","field":"mental health","wr":"internal consistency of the Funding subscale items","conf":"high","self":false,"doi":"10.1177/0193841X05275586","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Unstandardized Cronbach's alpha of the Funding subscale, pooled across the four raters, was 0.71, indicating acceptable internal consistency. k is the number of items forming the subscale (2).","vf":"unverified"},{"key":"FAWCL6KR","au":"Whaley, Arthur L","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach's alpha, standardized, single rater","estd":"Cronbach alpha","v":0.93,"n":"20","k":"6","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"not at all (0), somewhat (1), or definitely (2)","field":"mental health","wr":"internal consistency of Project Evaluation items, Rater 1","conf":"high","self":false,"doi":"10.1177/0193841X05275586","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"The maximum per-rater standardized alpha in Table 3: Rater 1's standardized Cronbach's alpha for the six-item Project Evaluation subscale across the 20 proposals was 0.93 (Rater 4 tied). Included as the maximum stratum under the per-family cap.","vf":"unverified"},{"key":"FAWCL6KR","au":"Whaley, Arthur L","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach's alpha, unstandardized, single rater","estd":"Cronbach alpha","v":0.93,"n":"20","k":"6","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"not at all (0), somewhat (1), or definitely (2)","field":"mental health","wr":"internal consistency of Project Evaluation items, Rater 1","conf":"high","self":false,"doi":"10.1177/0193841X05275586","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"The maximum per-rater alpha in Table 3: Rater 1's unstandardized Cronbach's alpha for the six-item Project Evaluation subscale across the 20 proposals was 0.93 (Rater 4 tied). Included as the maximum stratum under the per-family cap.","vf":"unverified"},{"key":"FAWCL6KR","au":"Whaley, Arthur L","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach's alpha, unstandardized, all raters","estd":"Cronbach alpha","v":0.8,"n":"20","k":"3","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"not at all (0), somewhat (1), or definitely (2)","field":"mental health","wr":"internal consistency of the Mental Health subscale items","conf":"high","self":false,"doi":"10.1177/0193841X05275586","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Unstandardized Cronbach's alpha of the Mental Health subscale, pooled across the four raters, was 0.80, indicating good internal consistency. k is the number of items forming the subscale (3).","vf":"unverified"},{"key":"FAWCL6KR","au":"Whaley, Arthur L","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach's alpha, standardized, single rater","estd":"Cronbach alpha","v":0.53,"n":"20","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"not at all (0), somewhat (1), or definitely (2)","field":"mental health","wr":"internal consistency of Funding subscale items, Rater 1","conf":"high","self":false,"doi":"10.1177/0193841X05275586","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"The minimum per-rater standardized alpha in Table 3: Rater 1's standardized Cronbach's alpha for the two-item Funding subscale across the 20 proposals was 0.53. Included as the minimum stratum under the per-family cap.","vf":"unverified"},{"key":"FAWCL6KR","au":"Whaley, Arthur L","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach's alpha, unstandardized, single rater","estd":"Cronbach alpha","v":0.53,"n":"20","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"not at all (0), somewhat (1), or definitely (2)","field":"mental health","wr":"internal consistency of Funding subscale items, Rater 1","conf":"high","self":false,"doi":"10.1177/0193841X05275586","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"The minimum per-rater alpha in Table 3: Rater 1's unstandardized Cronbach's alpha for the two-item Funding subscale across the 20 proposals was 0.53. Included as the minimum stratum under the per-family cap.","vf":"unverified"},{"key":"FAWCL6KR","au":"Whaley, Arthur L","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach's alpha, unstandardized, all raters","estd":"Cronbach alpha","v":0.89,"n":"20","k":"6","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"not at all (0), somewhat (1), or definitely (2)","field":"mental health","wr":"internal consistency of the Project Evaluation subscale items","conf":"high","self":false,"doi":"10.1177/0193841X05275586","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Unstandardized Cronbach's alpha of the Project Evaluation subscale, pooled across the four raters, was 0.89, indicating good internal consistency. k is the number of items forming the subscale (6).","vf":"unverified"},{"key":"FAWCL6KR","au":"Whaley, Arthur L","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach's alpha, unstandardized, all raters","estd":"Cronbach alpha","v":0.79,"n":"20","k":"5","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"not at all (0), somewhat (1), or definitely (2)","field":"mental health","wr":"internal consistency of the Target Population subscale items","conf":"high","self":false,"doi":"10.1177/0193841X05275586","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Unstandardized Cronbach's alpha of the Target Population subscale, pooled across the four raters, was 0.79, indicating acceptable internal consistency. k is the number of items forming the subscale (5).","vf":"unverified"},{"key":"FAWCL6KR","au":"Whaley, Arthur L","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach's alpha, standardized, all raters","estd":"Cronbach alpha","v":0.87,"n":"20","k":"22","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"not at all (0), somewhat (1), or definitely (2)","field":"mental health","wr":"internal consistency of the 22-item GPRF total scale","conf":"high","self":false,"doi":"10.1177/0193841X05275586","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Standardized Cronbach's alpha of the 22-item GPRF total scale, pooled across the four raters' scoring of 20 proposals (80 cases), was 0.87, indicating good internal consistency. k is the number of items forming the scale (22).","vf":"unverified"},{"key":"FAWCL6KR","au":"Whaley, Arthur L","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach's alpha, unstandardized, all raters","estd":"Cronbach alpha","v":0.87,"n":"20","k":"22","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"not at all (0), somewhat (1), or definitely (2)","field":"mental health","wr":"internal consistency of the 22-item GPRF total scale","conf":"high","self":false,"doi":"10.1177/0193841X05275586","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Unstandardized Cronbach's alpha of the 22-item GPRF total scale, pooled across the four raters' scoring of 20 proposals (80 cases), was 0.87, indicating good internal consistency. k is the number of items forming the scale (22).","vf":"unverified"},{"key":"FAWCL6KR","au":"Whaley, Arthur L","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC (3, k), model for test raters being the only raters of interest, rating of the average of multiple measurements","estd":"ICC (average)","v":0.97,"n":"20","k":"4","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"not at all (0), somewhat (1), or definitely (2)","field":"mental health","wr":"four raters scoring 20 grant proposals","conf":"high","self":false,"doi":"10.1177/0193841X05275586","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Inter-rater agreement of the four raters on the Agency subscale scores was excellent, ICC(3,k) = 0.97 across the 20 grant proposals.","vf":"unverified"},{"key":"FAWCL6KR","au":"Whaley, Arthur L","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC (3, k), model for test raters being the only raters of interest, rating of the average of multiple measurements","estd":"ICC (average)","v":0.85,"n":"20","k":"4","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"not at all (0), somewhat (1), or definitely (2)","field":"mental health","wr":"four raters scoring 20 grant proposals","conf":"high","self":false,"doi":"10.1177/0193841X05275586","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Inter-rater agreement of the four raters on the Funding subscale scores was good, ICC(3,k) = 0.85 across the 20 grant proposals.","vf":"unverified"},{"key":"FAWCL6KR","au":"Whaley, Arthur L","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC (3, k), model for test raters being the only raters of interest, rating of the average of multiple measurements","estd":"ICC (average)","v":0.67,"n":"20","k":"4","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"not at all (0), somewhat (1), or definitely (2)","field":"mental health","wr":"four raters scoring 20 grant proposals","conf":"high","self":false,"doi":"10.1177/0193841X05275586","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Inter-rater agreement of the four raters on the Mental Health subscale scores was moderate, ICC(3,k) = 0.67 across the 20 grant proposals.","vf":"unverified"},{"key":"FAWCL6KR","au":"Whaley, Arthur L","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC (3, k), model for test raters being the only raters of interest, rating of the average of multiple measurements","estd":"ICC (average)","v":0.33,"n":"20","k":"4","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"not at all (0), somewhat (1), or definitely (2)","field":"mental health","wr":"four raters scoring 20 grant proposals","conf":"high","self":false,"doi":"10.1177/0193841X05275586","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Inter-rater agreement of the four raters on the Project Evaluation subscale scores was poor, ICC(3,k) = 0.33 across the 20 grant proposals.","vf":"unverified"},{"key":"FAWCL6KR","au":"Whaley, Arthur L","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC (3, k), model for test raters being the only raters of interest, rating of the average of multiple measurements","estd":"ICC (average)","v":0.83,"n":"20","k":"4","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"not at all (0), somewhat (1), or definitely (2)","field":"mental health","wr":"four raters scoring 20 grant proposals","conf":"high","self":false,"doi":"10.1177/0193841X05275586","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Inter-rater agreement of the four raters on the Target Population subscale scores was good, ICC(3,k) = 0.83 across the 20 grant proposals.","vf":"unverified"},{"key":"FAWCL6KR","au":"Whaley, Arthur L","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"ICC (3, k), model for test raters being the only raters of interest, rating of the average of multiple measurements","estd":"ICC (average)","v":0.81,"n":"20","k":"4","samp":"re-reviewed-subset","blind":"single","agg":"average-of-k","scale":"not at all (0), somewhat (1), or definitely (2)","field":"mental health","wr":"four raters scoring 20 grant proposals","conf":"high","self":false,"doi":"10.1177/0193841X05275586","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Four master's-level raters independently scored 20 grant proposals on the 22-item GPRF in a mock review; an ICC(3,k) of 0.81 for the total scale indicates good agreement on the averaged rating. The 80 total ratings are stated in the text (the 20 proposals rated by four raters yielded 80 cases).","vf":"unverified"},{"key":"IXVXHAXS","au":"Whaley, Arthur L.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Spearman’s rank order correlation (rho)","estd":"correlation","v":0.77,"n":"27","k":"","samp":"re-reviewed-subset","blind":"unclear","agg":"unspecified","scale":"percentage of NA responses per item","field":"mental health","wr":"EC vs fellows NA response patterns across brief-form items","conf":"high","self":false,"doi":"10.1177/0193841x06288737","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"The proportion of Not Applicable responses each rater group gave to each brief-form item was compared between the Executive Committee and the fellows; the Spearman correlation between the two groups' item-level response patterns was 0.77.","vf":"unverified"},{"key":"IXVXHAXS","au":"Whaley, Arthur L.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Spearman’s rank order correlation (rho)","estd":"correlation","v":0.92,"n":"27","k":"","samp":"re-reviewed-subset","blind":"unclear","agg":"unspecified","scale":"percentage of NMI responses per item","field":"mental health","wr":"EC vs fellows NMI response patterns across brief-form items","conf":"high","self":false,"doi":"10.1177/0193841x06288737","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"The proportion of Need More Information responses each rater group gave to each brief-form item was compared between the Executive Committee and the fellows; the Spearman correlation between the two groups' item-level response patterns was 0.92.","vf":"unverified"},{"key":"IXVXHAXS","au":"Whaley, Arthur L.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach’s alpha (Cronbach 1951)","estd":"Cronbach alpha","v":0.59,"n":"8","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0 Not at all, 1 Somewhat, 2 Definitely; NA/NMI scored zero","field":"mental health","wr":"internal consistency of 5-item agency subscale, EC","conf":"high","self":false,"doi":"10.1177/0193841x06288737","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Within the Executive Committee's 14 long-form ratings of eight full proposals, Cronbach's alpha across the 5 agency subscale items was 0.59.","vf":"unverified"},{"key":"IXVXHAXS","au":"Whaley, Arthur L.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach’s alpha (Cronbach 1951)","estd":"Cronbach alpha","v":0.78,"n":"27","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0 No, 1 Yes; NA/NMI scored zero","field":"mental health","wr":"internal consistency of 12-item brief GPRF, EC ratings","conf":"high","self":false,"doi":"10.1177/0193841x06288737","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Across the 203 brief-form ratings the Executive Committee produced for 27 applications, Cronbach's alpha over the 12 items was 0.78, indicating good internal consistency of the total score.","vf":"unverified"},{"key":"IXVXHAXS","au":"Whaley, Arthur L.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraclass correlation coefficient (Shrout and Fleiss 1979)","estd":"ICC (single/unspec)","v":0.84,"n":"27","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"total percentage score, 0-1","field":"mental health","wr":"EC members' total scores on grant applications","conf":"high","self":false,"doi":"10.1177/0193841x06288737","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"All eight Executive Committee members were asked to score 27 grant applications on the brief form; 203 of 216 ratings were completed, so raters per application varied. The intraclass correlation for total percentage scores was 0.84.","vf":"unverified"},{"key":"IXVXHAXS","au":"Whaley, Arthur L.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach’s alpha (Cronbach 1951)","estd":"Cronbach alpha","v":0.87,"n":"8","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0 Not at all, 1 Somewhat, 2 Definitely; NA/NMI scored zero","field":"mental health","wr":"internal consistency of 20-item long GPRF, EC ratings","conf":"high","self":false,"doi":"10.1177/0193841x06288737","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Across the 14 long-form ratings the Executive Committee produced for eight full proposals, Cronbach's alpha over the 20 items was 0.87, indicating good internal consistency of the total score.","vf":"unverified"},{"key":"IXVXHAXS","au":"Whaley, Arthur L.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraclass correlation coefficient (Shrout and Fleiss 1979)","estd":"ICC (single/unspec)","v":0.79,"n":"8","k":"2","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"total percentage score, 0-1","field":"mental health","wr":"EC members' total scores on full grant proposals","conf":"high","self":false,"doi":"10.1177/0193841x06288737","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":true,"he":true,"ms":"Executive Committee members independently scored eight full grant proposals on the 20-item long form before the review meeting, producing 14 ratings; the intraclass correlation for total percentage scores was 0.79. The number of raters per proposal varied and is not clearly stated.","vf":"unverified"},{"key":"IXVXHAXS","au":"Whaley, Arthur L.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach’s alpha (Cronbach 1951)","estd":"Cronbach alpha","v":0.56,"n":"8","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0 Not at all, 1 Somewhat, 2 Definitely; NA/NMI scored zero","field":"mental health","wr":"internal consistency of 3-item funding subscale, EC","conf":"high","self":false,"doi":"10.1177/0193841x06288737","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Within the Executive Committee's 14 long-form ratings of eight full proposals, Cronbach's alpha across the 3 funding subscale items was 0.56.","vf":"unverified"},{"key":"IXVXHAXS","au":"Whaley, Arthur L.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach’s alpha (Cronbach 1951)","estd":"Cronbach alpha","v":0.56,"n":"8","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0 Not at all, 1 Somewhat, 2 Definitely; NA/NMI scored zero","field":"mental health","wr":"internal consistency of 3-item mental health subscale, EC","conf":"high","self":false,"doi":"10.1177/0193841x06288737","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Within the Executive Committee's 14 long-form ratings of eight full proposals, Cronbach's alpha across the 3 mental health subscale items was 0.56.","vf":"unverified"},{"key":"IXVXHAXS","au":"Whaley, Arthur L.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach’s alpha (Cronbach 1951)","estd":"Cronbach alpha","v":0.79,"n":"8","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0 Not at all, 1 Somewhat, 2 Definitely; NA/NMI scored zero","field":"mental health","wr":"internal consistency of 5-item project evaluation subscale, EC","conf":"high","self":false,"doi":"10.1177/0193841x06288737","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Within the Executive Committee's 14 long-form ratings of eight full proposals, Cronbach's alpha across the 5 project evaluation subscale items was 0.79.","vf":"unverified"},{"key":"IXVXHAXS","au":"Whaley, Arthur L.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach’s alpha (Cronbach 1951)","estd":"Cronbach alpha","v":0.59,"n":"8","k":"","samp":"full-pool","blind":"unclear","agg":"unspecified","scale":"0 Not at all, 1 Somewhat, 2 Definitely; NA/NMI scored zero","field":"mental health","wr":"internal consistency of 4-item target population subscale, EC","conf":"high","self":false,"doi":"10.1177/0193841x06288737","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Within the Executive Committee's 14 long-form ratings of eight full proposals, Cronbach's alpha across the 4 target population subscale items was 0.59.","vf":"unverified"},{"key":"IXVXHAXS","au":"Whaley, Arthur L.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach’s alpha (Cronbach 1951)","estd":"Cronbach alpha","v":0.72,"n":"7","k":"4","samp":"re-reviewed-subset","blind":"unclear","agg":"unspecified","scale":"0 Not at all, 1 Somewhat, 2 Definitely; NA/NMI scored zero","field":"mental health","wr":"internal consistency of 5-item agency subscale, fellows","conf":"high","self":false,"doi":"10.1177/0193841x06288737","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Within the four fellows' 28 long-form ratings of seven full proposals, Cronbach's alpha across the 5 agency subscale items was 0.72.","vf":"unverified"},{"key":"IXVXHAXS","au":"Whaley, Arthur L.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach’s alpha (Cronbach 1951)","estd":"Cronbach alpha","v":0.75,"n":"27","k":"4","samp":"re-reviewed-subset","blind":"unclear","agg":"unspecified","scale":"0 No, 1 Yes; NA/NMI scored zero","field":"mental health","wr":"internal consistency of 12-item brief GPRF, fellows' ratings","conf":"high","self":false,"doi":"10.1177/0193841x06288737","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Across the 108 brief-form ratings the four fellows produced for 27 applications, Cronbach's alpha over the 12 items was 0.75, indicating moderate to good internal consistency of the total score.","vf":"unverified"},{"key":"IXVXHAXS","au":"Whaley, Arthur L.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraclass correlation coefficient (Shrout and Fleiss 1979)","estd":"ICC (single/unspec)","v":0.81,"n":"27","k":"4","samp":"re-reviewed-subset","blind":"unclear","agg":"unspecified","scale":"total percentage score, 0-1","field":"mental health","wr":"student fellows' total scores on grant applications","conf":"high","self":false,"doi":"10.1177/0193841x06288737","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Four doctoral-student fellows independently scored all 27 grant applications on the brief form (108 ratings); the intraclass correlation for total percentage scores was 0.81.","vf":"unverified"},{"key":"IXVXHAXS","au":"Whaley, Arthur L.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach’s alpha (Cronbach 1951)","estd":"Cronbach alpha","v":0.88,"n":"7","k":"4","samp":"re-reviewed-subset","blind":"unclear","agg":"unspecified","scale":"0 Not at all, 1 Somewhat, 2 Definitely; NA/NMI scored zero","field":"mental health","wr":"internal consistency of 20-item long GPRF, fellows' ratings","conf":"high","self":false,"doi":"10.1177/0193841x06288737","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Across the 28 long-form ratings the four fellows produced for seven full proposals, Cronbach's alpha over the 20 items was 0.88, indicating good internal consistency of the total score.","vf":"unverified"},{"key":"IXVXHAXS","au":"Whaley, Arthur L.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"ICC","form":"intraclass correlation coefficient (Shrout and Fleiss 1979)","estd":"ICC (single/unspec)","v":0.82,"n":"7","k":"4","samp":"re-reviewed-subset","blind":"unclear","agg":"unspecified","scale":"total percentage score, 0-1","field":"mental health","wr":"student fellows' total scores on full grant proposals","conf":"high","self":false,"doi":"10.1177/0193841x06288737","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":true,"ms":"Four doctoral-student fellows, naive raters working in parallel to the committee, independently scored seven of the eight full proposals (28 ratings); the intraclass correlation for total percentage scores was 0.82.","vf":"unverified"},{"key":"IXVXHAXS","au":"Whaley, Arthur L.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach’s alpha (Cronbach 1951)","estd":"Cronbach alpha","v":0.6,"n":"7","k":"4","samp":"re-reviewed-subset","blind":"unclear","agg":"unspecified","scale":"0 Not at all, 1 Somewhat, 2 Definitely; NA/NMI scored zero","field":"mental health","wr":"internal consistency of 3-item funding subscale, fellows","conf":"high","self":false,"doi":"10.1177/0193841x06288737","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Within the four fellows' 28 long-form ratings of seven full proposals, Cronbach's alpha across the 3 funding subscale items was 0.60.","vf":"unverified"},{"key":"IXVXHAXS","au":"Whaley, Arthur L.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach’s alpha (Cronbach 1951)","estd":"Cronbach alpha","v":0.44,"n":"7","k":"4","samp":"re-reviewed-subset","blind":"unclear","agg":"unspecified","scale":"0 Not at all, 1 Somewhat, 2 Definitely; NA/NMI scored zero","field":"mental health","wr":"internal consistency of 3-item mental health subscale, fellows","conf":"high","self":false,"doi":"10.1177/0193841x06288737","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Within the four fellows' 28 long-form ratings of seven full proposals, Cronbach's alpha across the 3 mental health subscale items was 0.44.","vf":"unverified"},{"key":"IXVXHAXS","au":"Whaley, Arthur L.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach’s alpha (Cronbach 1951)","estd":"Cronbach alpha","v":0.83,"n":"7","k":"4","samp":"re-reviewed-subset","blind":"unclear","agg":"unspecified","scale":"0 Not at all, 1 Somewhat, 2 Definitely; NA/NMI scored zero","field":"mental health","wr":"internal consistency of 5-item project evaluation subscale, fellows","conf":"high","self":false,"doi":"10.1177/0193841x06288737","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Within the four fellows' 28 long-form ratings of seven full proposals, Cronbach's alpha across the 5 project evaluation subscale items was 0.83.","vf":"unverified"},{"key":"IXVXHAXS","au":"Whaley, Arthur L.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"Cronbach-alpha","form":"Cronbach’s alpha (Cronbach 1951)","estd":"Cronbach alpha","v":0.71,"n":"7","k":"4","samp":"re-reviewed-subset","blind":"unclear","agg":"unspecified","scale":"0 Not at all, 1 Somewhat, 2 Definitely; NA/NMI scored zero","field":"mental health","wr":"internal consistency of 4-item target population subscale, fellows","conf":"high","self":false,"doi":"10.1177/0193841x06288737","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"submission-scores","rr":"restricted-top","pr":false,"he":false,"ms":"Within the four fellows' 28 long-form ratings of seven full proposals, Cronbach's alpha across the 4 target population subscale items was 0.71.","vf":"unverified"},{"key":"IXVXHAXS","au":"Whaley, Arthur L.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Spearman’s rank order correlation (rho)","estd":"correlation","v":0.52,"n":"8","k":"","samp":"re-reviewed-subset","blind":"unclear","agg":"unspecified","scale":"percentage of NA responses per item","field":"mental health","wr":"EC vs fellows NA response patterns across long-form items","conf":"high","self":false,"doi":"10.1177/0193841x06288737","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"The proportion of Not Applicable responses each rater group gave to each long-form item was compared between the Executive Committee and the fellows; the Spearman correlation between the two groups' item-level response patterns was 0.52.","vf":"unverified"},{"key":"IXVXHAXS","au":"Whaley, Arthur L.","y":2006,"cx":"Grant","ob":"grant-proposal","fam":"correlation","form":"Spearman’s rank order correlation (rho)","estd":"correlation","v":0.56,"n":"8","k":"","samp":"re-reviewed-subset","blind":"unclear","agg":"unspecified","scale":"percentage of NMI responses per item","field":"mental health","wr":"EC vs fellows NMI response patterns across long-form items","conf":"high","self":false,"doi":"10.1177/0193841x06288737","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"The proportion of Need More Information responses each rater group gave to each long-form item was compared between the Executive Committee and the fellows; the Spearman correlation between the two groups' item-level response patterns was 0.56.","vf":"unverified"},{"key":"836GCTFJ","au":"Whitehurst, Grover J.","y":1983,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Finn's r (Finn, 1970)","estd":"other","v":0.4,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"seven-point rating scale, higher numbers more favorable","field":"developmental psychology","wr":"nine manuscript evaluation criteria","conf":"med","self":false,"doi":"10.1016/0273-2297(83)90009-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Across the nine analysed seven-point evaluation scales, the mean Finn's r was 0.40, summarising the per-scale coefficients rather than a new pooled computation.","vf":"unverified"},{"key":"836GCTFJ","au":"Whitehurst, Grover J.","y":1983,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation using the one-way ANOVA model (Bartko, 1976)","estd":"ICC (single/unspec)","v":0.19,"n":"","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"seven-point rating scale, higher numbers more favorable","field":"developmental psychology","wr":"nine manuscript evaluation criteria","conf":"med","self":false,"doi":"10.1016/0273-2297(83)90009-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Across the nine analysed seven-point evaluation scales, the mean one-way intraclass correlation was 0.19, summarising the per-scale coefficients rather than a new pooled computation.","vf":"unverified"},{"key":"836GCTFJ","au":"Whitehurst, Grover J.","y":1983,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Finn's r (Finn, 1970)","estd":"other","v":0.46,"n":"63","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"seven-point rating scale, higher numbers more favorable","field":"developmental psychology","wr":"reviewers on overall manuscript quality (7-point)","conf":"med","self":false,"doi":"10.1016/0273-2297(83)90009-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Two independent external reviewers rated overall manuscript quality on a seven-point scale for 63 manuscripts; Finn's r was 0.46.","vf":"unverified"},{"key":"836GCTFJ","au":"Whitehurst, Grover J.","y":1983,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation using the one-way ANOVA model (Bartko, 1976)","estd":"ICC (single/unspec)","v":0.23,"n":"63","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"seven-point rating scale, higher numbers more favorable","field":"developmental psychology","wr":"reviewers on overall manuscript quality (7-point)","conf":"med","self":false,"doi":"10.1016/0273-2297(83)90009-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The same seven-point overall-quality ratings by two reviewers over 63 manuscripts gave a one-way intraclass correlation of 0.23.","vf":"unverified"},{"key":"836GCTFJ","au":"Whitehurst, Grover J.","y":1983,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Finn's r (Finn, 1970)","estd":"other","v":0.34,"n":"62","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"seven-point rating scale, higher numbers more favorable","field":"developmental psychology","wr":"reviewers on wide interest (7-point)","conf":"med","self":false,"doi":"10.1016/0273-2297(83)90009-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Two independent external reviewers rated breadth of interest on a seven-point scale for 62 manuscripts; Finn's r was 0.34.","vf":"unverified"},{"key":"836GCTFJ","au":"Whitehurst, Grover J.","y":1983,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation using the one-way ANOVA model (Bartko, 1976)","estd":"ICC (single/unspec)","v":0.17,"n":"62","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"seven-point rating scale, higher numbers more favorable","field":"developmental psychology","wr":"reviewers on wide interest (7-point)","conf":"med","self":false,"doi":"10.1016/0273-2297(83)90009-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The same seven-point breadth-of-interest ratings by two reviewers over 62 manuscripts gave a one-way intraclass correlation of 0.17.","vf":"unverified"},{"key":"836GCTFJ","au":"Whitehurst, Grover J.","y":1983,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Finn's r (Finn, 1970)","estd":"other","v":0.36,"n":"64","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"seven-point rating scale, higher numbers more favorable","field":"developmental psychology","wr":"reviewers on likely impact (7-point)","conf":"med","self":false,"doi":"10.1016/0273-2297(83)90009-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Two independent external reviewers rated likely impact on a seven-point scale for 64 manuscripts; Finn's r was 0.36.","vf":"unverified"},{"key":"836GCTFJ","au":"Whitehurst, Grover J.","y":1983,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation using the one-way ANOVA model (Bartko, 1976)","estd":"ICC (single/unspec)","v":0.17,"n":"64","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"seven-point rating scale, higher numbers more favorable","field":"developmental psychology","wr":"reviewers on likely impact (7-point)","conf":"med","self":false,"doi":"10.1016/0273-2297(83)90009-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The same seven-point likely-impact ratings by two reviewers over 64 manuscripts gave a one-way intraclass correlation of 0.17 (table prints '.I7', an OCR rendering of .17).","vf":"unverified"},{"key":"836GCTFJ","au":"Whitehurst, Grover J.","y":1983,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Finn's r (Finn, 1970)","estd":"other","v":0.39,"n":"62","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"seven-point rating scale, higher numbers more favorable","field":"developmental psychology","wr":"reviewers on problem conceptualisation (7-point)","conf":"med","self":false,"doi":"10.1016/0273-2297(83)90009-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Two independent external reviewers rated how well the problem was conceptualised on a seven-point scale for 62 manuscripts; Finn's r was 0.39.","vf":"unverified"},{"key":"836GCTFJ","au":"Whitehurst, Grover J.","y":1983,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation using the one-way ANOVA model (Bartko, 1976)","estd":"ICC (single/unspec)","v":0.24,"n":"62","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"seven-point rating scale, higher numbers more favorable","field":"developmental psychology","wr":"reviewers on problem conceptualisation (7-point)","conf":"med","self":false,"doi":"10.1016/0273-2297(83)90009-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The same seven-point conceptualisation ratings by two reviewers over 62 manuscripts gave a one-way intraclass correlation of 0.24.","vf":"unverified"},{"key":"836GCTFJ","au":"Whitehurst, Grover J.","y":1983,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Finn's r (Finn, 1970)","estd":"other","v":0.41,"n":"63","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"seven-point rating scale, higher numbers more favorable","field":"developmental psychology","wr":"reviewers on topic importance (7-point)","conf":"med","self":false,"doi":"10.1016/0273-2297(83)90009-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Two independent external reviewers rated topic importance on a seven-point scale for 63 manuscripts; Finn's r was 0.41 despite a near-zero intraclass correlation.","vf":"unverified"},{"key":"836GCTFJ","au":"Whitehurst, Grover J.","y":1983,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation using the one-way ANOVA model (Bartko, 1976)","estd":"ICC (single/unspec)","v":-0.1,"n":"63","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"seven-point rating scale, higher numbers more favorable","field":"developmental psychology","wr":"reviewers on topic importance (7-point)","conf":"med","self":false,"doi":"10.1016/0273-2297(83)90009-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The same seven-point topic-importance ratings gave a one-way intraclass correlation of -0.10; the paper attributes this to very low between-manuscript variance in judged importance.","vf":"unverified"},{"key":"836GCTFJ","au":"Whitehurst, Grover J.","y":1983,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Finn's r (Finn, 1970)","estd":"other","v":0.26,"n":"60","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"seven-point rating scale, higher numbers more favorable","field":"developmental psychology","wr":"reviewers on originality (7-point)","conf":"med","self":false,"doi":"10.1016/0273-2297(83)90009-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Two independent external reviewers rated originality on a seven-point scale for 60 manuscripts; Finn's r was 0.26.","vf":"unverified"},{"key":"836GCTFJ","au":"Whitehurst, Grover J.","y":1983,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation using the one-way ANOVA model (Bartko, 1976)","estd":"ICC (single/unspec)","v":0.11,"n":"60","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"seven-point rating scale, higher numbers more favorable","field":"developmental psychology","wr":"reviewers on originality (7-point)","conf":"med","self":false,"doi":"10.1016/0273-2297(83)90009-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The same seven-point originality ratings by two reviewers over 60 manuscripts gave a one-way intraclass correlation of 0.11 (table prints '.I1', an OCR rendering of .11).","vf":"unverified"},{"key":"836GCTFJ","au":"Whitehurst, Grover J.","y":1983,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Finn's r (Finn, 1970)","estd":"other","v":0.47,"n":"63","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"seven-point rating scale, higher numbers more favorable","field":"developmental psychology","wr":"reviewers on writing quality (7-point)","conf":"med","self":false,"doi":"10.1016/0273-2297(83)90009-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Two independent external reviewers rated how well the manuscript was written on a seven-point scale for 63 manuscripts; Finn's r was 0.47.","vf":"unverified"},{"key":"836GCTFJ","au":"Whitehurst, Grover J.","y":1983,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation using the one-way ANOVA model (Bartko, 1976)","estd":"ICC (single/unspec)","v":0.28,"n":"63","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"seven-point rating scale, higher numbers more favorable","field":"developmental psychology","wr":"reviewers on writing quality (7-point)","conf":"med","self":false,"doi":"10.1016/0273-2297(83)90009-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The same seven-point writing-quality ratings by two reviewers over 63 manuscripts gave a one-way intraclass correlation of 0.28.","vf":"unverified"},{"key":"836GCTFJ","au":"Whitehurst, Grover J.","y":1983,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Finn's r (Finn, 1970)","estd":"other","v":0.49,"n":"60","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"seven-point rating scale, higher numbers more favorable","field":"developmental psychology","wr":"reviewers on theoretical character (7-point)","conf":"med","self":false,"doi":"10.1016/0273-2297(83)90009-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Two independent external reviewers rated how theoretical the presentation was on a seven-point scale for 60 manuscripts; Finn's r was 0.49, the highest of the individual scales.","vf":"unverified"},{"key":"836GCTFJ","au":"Whitehurst, Grover J.","y":1983,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation using the one-way ANOVA model (Bartko, 1976)","estd":"ICC (single/unspec)","v":0.43,"n":"60","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"seven-point rating scale, higher numbers more favorable","field":"developmental psychology","wr":"reviewers on theoretical character (7-point)","conf":"med","self":false,"doi":"10.1016/0273-2297(83)90009-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The same seven-point theoretical-character ratings by two reviewers over 60 manuscripts gave a one-way intraclass correlation of 0.43, the highest scale ICC.","vf":"unverified"},{"key":"836GCTFJ","au":"Whitehurst, Grover J.","y":1983,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Finn's r (Finn, 1970)","estd":"other","v":0.38,"n":"41","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"seven-point rating scale, higher numbers more favorable","field":"developmental psychology","wr":"reviewers on validity of conclusions (7-point)","conf":"med","self":false,"doi":"10.1016/0273-2297(83)90009-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"Two independent external reviewers rated validity of conclusions on a seven-point scale for 41 manuscripts; Finn's r was 0.38.","vf":"unverified"},{"key":"836GCTFJ","au":"Whitehurst, Grover J.","y":1983,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation using the one-way ANOVA model (Bartko, 1976)","estd":"ICC (single/unspec)","v":0.21,"n":"41","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"seven-point rating scale, higher numbers more favorable","field":"developmental psychology","wr":"reviewers on validity of conclusions (7-point)","conf":"med","self":false,"doi":"10.1016/0273-2297(83)90009-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The same seven-point validity-of-conclusions ratings by two reviewers over 41 manuscripts gave a one-way intraclass correlation of 0.21.","vf":"unverified"},{"key":"836GCTFJ","au":"Whitehurst, Grover J.","y":1983,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Finn's r (Finn, 1970)","estd":"other","v":0.62,"n":"72","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept as is / accept with revisions / reject but encourage resubmission / reject","field":"developmental psychology","wr":"reviewers on four-category publishability recommendation","conf":"med","self":false,"doi":"10.1016/0273-2297(83)90009-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":true,"he":true,"ms":"Two independent external reviewers each rated the publishability of 72 manuscripts on a four-category recommendation; a Finn's r of 0.62 indicates agreement above chance, the paper's headline result.","vf":"unverified"},{"key":"836GCTFJ","au":"Whitehurst, Grover J.","y":1983,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"intraclass correlation using the one-way ANOVA model (Bartko, 1976)","estd":"ICC (single/unspec)","v":0.44,"n":"72","k":"2","samp":"full-pool","blind":"single","agg":"single-rater","scale":"accept as is / accept with revisions / reject but encourage resubmission / reject","field":"developmental psychology","wr":"reviewers on four-category publishability recommendation","conf":"med","self":false,"doi":"10.1016/0273-2297(83)90009-6","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"restricted-other","pr":false,"he":true,"ms":"The same four-category publishability recommendations by two reviewers over 72 manuscripts gave a one-way intraclass correlation of 0.44, an alternative index of the same agreement.","vf":"unverified"},{"key":"T3A36PFH","au":"Whitehurst, Grover J.","y":1984,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Finn's r (Finn, 1970)","estd":"other","v":0.56,"n":"87","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"5-point scale","field":"psychology","wr":"reviewers rating journal manuscript publishability","conf":"med","self":false,"doi":"10.1037/0003-066x.39.1.22","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"For the same 87 American Psychologist manuscripts, Finn's r was 0.56, the proportion of interrater correspondence not due to chance. This is the paper's headline result, arguing that agreement is better than intraclass correlations suggest because Finn's r is unaffected by low manuscript variance.","vf":"unverified"},{"key":"T3A36PFH","au":"Whitehurst, Grover J.","y":1984,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"one-way random-effects analysis of variance (ANOVA) intraclass correlation (R_I)","estd":"ICC (single/unspec)","v":0.54,"n":"87","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"5-point scale","field":"psychology","wr":"reviewers rating journal manuscript publishability","conf":"med","self":false,"doi":"10.1037/0003-066x.39.1.22","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated each of 87 manuscripts submitted to the American Psychologist on a five-point scale. The one-way random-effects intraclass correlation for a single rating, recomputed from Scarr and Weber's published agreement matrix, was 0.54.","vf":"unverified"},{"key":"T3A36PFH","au":"Whitehurst, Grover J.","y":1984,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Finn's r (Finn, 1970)","estd":"other","v":0.49,"n":"73","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"4-category rating scale (from Table 2)","field":"psychology","wr":"reviewers rating journal manuscript publishability","conf":"med","self":false,"doi":"10.1037/0003-066x.39.1.22","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the same 73 Developmental Review and Merrill-Palmer Quarterly manuscripts, Finn's r was 0.49, indicating about half of the interrater correspondence was due to nonchance factors, substantially above the intraclass correlation of 0.27.","vf":"unverified"},{"key":"T3A36PFH","au":"Whitehurst, Grover J.","y":1984,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"one-way random-effects analysis of variance (ANOVA) intraclass correlation (R_I)","estd":"ICC (single/unspec)","v":0.27,"n":"73","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"4-category rating scale (from Table 2)","field":"psychology","wr":"reviewers rating journal manuscript publishability","conf":"med","self":false,"doi":"10.1037/0003-066x.39.1.22","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated each of 73 manuscripts from Developmental Review and the Merrill-Palmer Quarterly on a four-category scale. The one-way random-effects intraclass correlation for a single rating, computed from previously unreported raw data, was 0.27.","vf":"unverified"},{"key":"T3A36PFH","au":"Whitehurst, Grover J.","y":1984,"cx":"Journal","ob":"journal-manuscript","fam":"other","form":"Finn's r (Finn, 1970)","estd":"other","v":0.3,"n":"177","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"5-point scale","field":"psychology","wr":"reviewers rating journal manuscript publishability","conf":"med","self":false,"doi":"10.1037/0003-066x.39.1.22","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"For the same 177 Personality and Social Psychology Bulletin manuscripts, Finn's r was 0.30, higher than the intraclass correlation of 0.22 but lower than for the other journals. The printed value is split by table garbling but is confirmed exactly by the stated within-manuscript error variance of 1.40.","vf":"unverified"},{"key":"T3A36PFH","au":"Whitehurst, Grover J.","y":1984,"cx":"Journal","ob":"journal-manuscript","fam":"ICC","form":"one-way random-effects analysis of variance (ANOVA) intraclass correlation (R_I)","estd":"ICC (single/unspec)","v":0.22,"n":"177","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"5-point scale","field":"psychology","wr":"reviewers rating journal manuscript publishability","conf":"med","self":false,"doi":"10.1037/0003-066x.39.1.22","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two independent reviewers rated each of 177 manuscripts handled by the Personality and Social Psychology Bulletin on a five-point scale. The one-way random-effects intraclass correlation for a single rating, recomputed from Crandall's published matrix, was 0.22; the table is garbled in extraction but the value is confirmed by the stated MSB and MSW.","vf":"unverified"},{"key":"HUD5RNL8","au":"Wicherts, Jelte M","y":2016,"cx":"Journal","ob":"other","fam":"Cronbach-alpha","form":"Cronbach's alpha across the 15 items","estd":"Cronbach alpha","v":0.9,"n":"","k":"","samp":"special","blind":"unclear","agg":"unspecified","scale":"1-5 Likert (1=completely fails to apply, 5=applies very well)","field":"multi-field","wr":"internal consistency of tool (Chemistry journals)","conf":"med","self":false,"doi":"10.1371/journal.pone.0147913","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Within Analytical Chemistry journals, the 15-item transparency tool showed Cronbach's alpha of 0.90.","vf":"unverified"},{"key":"HUD5RNL8","au":"Wicherts, Jelte M","y":2016,"cx":"Journal","ob":"other","fam":"Cronbach-alpha","form":"Cronbach's alpha across the 15 items","estd":"Cronbach alpha","v":0.62,"n":"","k":"","samp":"special","blind":"unclear","agg":"unspecified","scale":"1-5 Likert (1=completely fails to apply, 5=applies very well)","field":"multi-field","wr":"internal consistency of tool (Economics journals)","conf":"med","self":false,"doi":"10.1371/journal.pone.0147913","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Within Economics journals, the 15-item transparency tool showed Cronbach's alpha of 0.62.","vf":"unverified"},{"key":"HUD5RNL8","au":"Wicherts, Jelte M","y":2016,"cx":"Journal","ob":"other","fam":"Cronbach-alpha","form":"Cronbach's alpha across the 15 administered items","estd":"Cronbach alpha","v":0.81,"n":"92","k":"2.49","samp":"special","blind":"unclear","agg":"unspecified","scale":"1-5 Likert (1=completely fails to apply, 5=applies very well)","field":"multi-field","wr":"internal consistency of 15-item transparency tool","conf":"med","self":false,"doi":"10.1371/journal.pone.0147913","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"other","rr":"none","pr":false,"he":false,"ms":"The 15-item transparency tool (scored 1-5 per item) showed internal consistency of Cronbach's alpha 0.81 computed across the 92 journals' mean scores (mean 2.49 authors per journal).","vf":"unverified"},{"key":"HUD5RNL8","au":"Wicherts, Jelte M","y":2016,"cx":"Journal","ob":"other","fam":"Cronbach-alpha","form":"Cronbach's alpha across the 15 items","estd":"Cronbach alpha","v":0.73,"n":"","k":"","samp":"special","blind":"unclear","agg":"unspecified","scale":"1-5 Likert (1=completely fails to apply, 5=applies very well)","field":"multi-field","wr":"internal consistency of tool (Physics journals)","conf":"med","self":false,"doi":"10.1371/journal.pone.0147913","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Within Condensed Matter Physics journals, the 15-item transparency tool showed Cronbach's alpha of 0.73.","vf":"unverified"},{"key":"HUD5RNL8","au":"Wicherts, Jelte M","y":2016,"cx":"Journal","ob":"other","fam":"Cronbach-alpha","form":"Cronbach's alpha across the 15 items","estd":"Cronbach alpha","v":0.81,"n":"","k":"","samp":"special","blind":"unclear","agg":"unspecified","scale":"1-5 Likert (1=completely fails to apply, 5=applies very well)","field":"multi-field","wr":"internal consistency of tool (Psychiatry journals)","conf":"med","self":false,"doi":"10.1371/journal.pone.0147913","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Within Psychiatry journals, the 15-item transparency tool showed Cronbach's alpha of 0.81.","vf":"unverified"},{"key":"HUD5RNL8","au":"Wicherts, Jelte M","y":2016,"cx":"Journal","ob":"other","fam":"Cronbach-alpha","form":"Cronbach's alpha across the 15 items","estd":"Cronbach alpha","v":0.82,"n":"","k":"","samp":"special","blind":"unclear","agg":"unspecified","scale":"1-5 Likert (1=completely fails to apply, 5=applies very well)","field":"multi-field","wr":"internal consistency of tool (Radiology journals)","conf":"med","self":false,"doi":"10.1371/journal.pone.0147913","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Within Radiology & Nuclear Medicine journals, the 15-item transparency tool showed Cronbach's alpha of 0.82.","vf":"unverified"},{"key":"HUD5RNL8","au":"Wicherts, Jelte M","y":2016,"cx":"Journal","ob":"other","fam":"Cronbach-alpha","form":"Cronbach's alpha of the 3-item criterion (fairness, rigour, recommendation)","estd":"Cronbach alpha","v":0.86,"n":"92","k":"2.49","samp":"special","blind":"unclear","agg":"unspecified","scale":"1-5 Likert (1=completely fails to apply, 5=applies very well)","field":"multi-field","wr":"authors rating quality of peer review of their own paper","conf":"med","self":false,"doi":"10.1371/journal.pone.0147913","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"other","rr":"restricted-top","pr":false,"he":false,"ms":"The three-item criterion measuring authors' rated quality of the peer review of their own paper (fairness, rigour, recommendation) showed Cronbach's alpha of 0.86 across the 92 journals.","vf":"unverified"},{"key":"HUD5RNL8","au":"Wicherts, Jelte M","y":2016,"cx":"Journal","ob":"other","fam":"other","form":"reliability of the mean score per journal from one-way ANOVA eta-squared","estd":"other","v":0.53,"n":"","k":"","samp":"special","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1=completely fails to apply, 5=applies very well)","field":"multi-field","wr":"authors rating journals' peer-review transparency (Chemistry)","conf":"med","self":false,"doi":"10.1371/journal.pone.0147913","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Within Analytical Chemistry journals, the reliability of the mean author-rated transparency score per journal was 0.53.","vf":"unverified"},{"key":"HUD5RNL8","au":"Wicherts, Jelte M","y":2016,"cx":"Journal","ob":"other","fam":"other","form":"reliability of the mean score per journal from one-way ANOVA eta-squared","estd":"other","v":0.23,"n":"","k":"","samp":"special","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1=completely fails to apply, 5=applies very well)","field":"multi-field","wr":"authors rating journals' peer-review transparency (Economics)","conf":"med","self":false,"doi":"10.1371/journal.pone.0147913","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Within Economics journals, the reliability of the mean author-rated transparency score per journal was 0.23, the lowest of the five fields.","vf":"unverified"},{"key":"HUD5RNL8","au":"Wicherts, Jelte M","y":2016,"cx":"Journal","ob":"other","fam":"other","form":"reliability of the mean score per journal estimated from eta-squared from a standard ANOVA with journal as sole factor","estd":"other","v":0.48,"n":"92","k":"2.49","samp":"special","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1=completely fails to apply, 5=applies very well)","field":"multi-field","wr":"authors rating journals' peer-review transparency","conf":"med","self":false,"doi":"10.1371/journal.pone.0147913","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Authors of published papers rated the peer-review transparency of the 92 journals in which they published (mean 2.49 authors per journal, range 1 to 11); the reliability of the mean transparency score per journal across authors was 0.48 overall.","vf":"unverified"},{"key":"HUD5RNL8","au":"Wicherts, Jelte M","y":2016,"cx":"Journal","ob":"other","fam":"other","form":"reliability of the mean score per journal from one-way ANOVA eta-squared","estd":"other","v":0.46,"n":"","k":"","samp":"special","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1=completely fails to apply, 5=applies very well)","field":"multi-field","wr":"authors rating journals' peer-review transparency (Physics)","conf":"med","self":false,"doi":"10.1371/journal.pone.0147913","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Within Condensed Matter Physics journals, the reliability of the mean author-rated transparency score per journal was 0.46.","vf":"unverified"},{"key":"HUD5RNL8","au":"Wicherts, Jelte M","y":2016,"cx":"Journal","ob":"other","fam":"other","form":"reliability of the mean score per journal from one-way ANOVA eta-squared","estd":"other","v":0.37,"n":"","k":"","samp":"special","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1=completely fails to apply, 5=applies very well)","field":"multi-field","wr":"authors rating journals' peer-review transparency (Psychiatry)","conf":"med","self":false,"doi":"10.1371/journal.pone.0147913","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Within Psychiatry journals, the reliability of the mean author-rated transparency score per journal was 0.37.","vf":"unverified"},{"key":"HUD5RNL8","au":"Wicherts, Jelte M","y":2016,"cx":"Journal","ob":"other","fam":"other","form":"reliability of the mean score per journal from one-way ANOVA eta-squared","estd":"other","v":0.61,"n":"","k":"","samp":"special","blind":"unclear","agg":"average-of-k","scale":"1-5 Likert (1=completely fails to apply, 5=applies very well)","field":"multi-field","wr":"authors rating journals' peer-review transparency (Radiology)","conf":"med","self":false,"doi":"10.1371/journal.pone.0147913","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Within Radiology & Nuclear Medicine journals, the reliability of the mean author-rated transparency score per journal was 0.61.","vf":"unverified"},{"key":"HUD5RNL8","au":"Wicherts, Jelte M","y":2016,"cx":"Journal","ob":"other","fam":"Cronbach-alpha","form":"Cronbach's alpha (ratings averaged across raters per journal)","estd":"Cronbach alpha","v":0.89,"n":"31","k":"","samp":"special","blind":"unclear","agg":"unspecified","scale":"1-5 Likert (1=completely fails to apply, 5=applies very well)","field":"multi-field","wr":"internal consistency of transparency tool (Study 2)","conf":"med","self":false,"doi":"10.1371/journal.pone.0147913","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"other","rr":"none","pr":false,"he":false,"ms":"Across the 31 journals rated in Study 2, the transparency tool showed internal consistency of Cronbach's alpha 0.89 (the abstract reports .91 for this study).","vf":"unverified"},{"key":"HUD5RNL8","au":"Wicherts, Jelte M","y":2016,"cx":"Journal","ob":"other","fam":"ICC","form":"intra-class correlation (percentage of variance in ratings due to journals) from multilevel model with rater and journal random","estd":"ICC (single/unspec)","v":0.52,"n":"31","k":"","samp":"special","blind":"unclear","agg":"single-rater","scale":"1-5 Likert (1=completely fails to apply, 5=applies very well)","field":"multi-field","wr":"stakeholders rating journals' peer-review transparency","conf":"med","self":false,"doi":"10.1371/journal.pone.0147913","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":true,"he":false,"ms":"Conference attendees (editors, librarians, funders, publishers) independently rated the transparency of 31 heterogeneous journals; 52% of the variance was between journals, giving an inter-rater intra-class correlation of 0.52.","vf":"unverified"},{"key":"HUD5RNL8","au":"Wicherts, Jelte M","y":2016,"cx":"Journal","ob":"other","fam":"Cronbach-alpha","form":"Cronbach's alpha across the 14 items (revised tool)","estd":"Cronbach alpha","v":0.8,"n":"140","k":"1","samp":"special","blind":"unclear","agg":"unspecified","scale":"1-5 Likert (1=completely fails to apply, 5=applies very well)","field":"multi-field","wr":"internal consistency of revised tool (DOAJ sample)","conf":"med","self":false,"doi":"10.1371/journal.pone.0147913","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"other","rr":"none","pr":false,"he":false,"ms":"In a random sample of 140 DOAJ journals rated by academic librarians and experts (one rater per journal), the revised 14-item transparency tool showed internal consistency of Cronbach's alpha 0.80.","vf":"unverified"},{"key":"HUD5RNL8","au":"Wicherts, Jelte M","y":2016,"cx":"Journal","ob":"other","fam":"Cronbach-alpha","form":"Cronbach's alpha across the 14 items (revised tool)","estd":"Cronbach alpha","v":0.9,"n":"54","k":"","samp":"special","blind":"unclear","agg":"unspecified","scale":"1-5 Likert (1=completely fails to apply, 5=applies very well)","field":"multi-field","wr":"internal consistency of revised tool (Bohannon sample)","conf":"med","self":false,"doi":"10.1371/journal.pone.0147913","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"other","rr":"none","pr":false,"he":false,"ms":"In the 54 journals from Bohannon's hoax-paper study rated by eight librarians and two experts, the revised 14-item transparency tool showed internal consistency of Cronbach's alpha 0.90.","vf":"unverified"},{"key":"RP9WIA7I","au":"Wing, Deborah A.","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"kappa statistic","estd":"kappa","v":0.4,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"acceptance or rejection (four-tier recommendations collapsed)","field":"biomedical (obstetrics and gynecology)","wr":"female board-member recommendation vs editor's final decision","conf":"high","self":false,"doi":"10.1089/jwh.2009.1904","ciLow":0.36,"ciHigh":0.44,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"Restricting to the 1990 reviews by the 13 female editorial board members, agreement between their collapsed accept or reject recommendations and the editors' final decisions was a kappa of 0.40, identical to the overall value.","vf":"unverified"},{"key":"RP9WIA7I","au":"Wing, Deborah A.","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"kappa statistic","estd":"kappa","v":0.4,"n":"","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"acceptance or rejection (four-tier recommendations collapsed)","field":"biomedical (obstetrics and gynecology)","wr":"male board-member recommendation vs editor's final decision","conf":"high","self":false,"doi":"10.1089/jwh.2009.1904","ciLow":0.37,"ciHigh":0.43,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"Restricting to the 4072 reviews by the 25 male editorial board members, agreement between their collapsed accept or reject recommendations and the editors' final decisions was a kappa of 0.40, identical to the overall value.","vf":"unverified"},{"key":"RP9WIA7I","au":"Wing, Deborah A.","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"kappa","form":"kappa statistic","estd":"kappa","v":0.4,"n":"5958","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"acceptance or rejection (four-tier recommendations collapsed)","field":"biomedical (obstetrics and gynecology)","wr":"board-member recommendation vs editor's final decision","conf":"high","self":false,"doi":"10.1089/jwh.2009.1904","ciLow":0.37,"ciHigh":0.42,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":true,"he":false,"ms":"Editorial board members' accept or reject recommendations on manuscripts were compared with the editors' final accept or reject decisions across 6062 reviews of 5958 manuscripts; a kappa of 0.40 indicates moderate chance-corrected agreement. The rater pool of 38 counts board members only; 3 editors made the final decisions.","vf":"unverified"},{"key":"RP9WIA7I","au":"Wing, Deborah A.","y":2010,"cx":"Journal","ob":"journal-manuscript","fam":"percent-agreement","form":"decision congruence, exact agreement on collapsed accept/reject categories","estd":"percent agreement","v":0.72,"n":"5958","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"acceptance or rejection (four-tier recommendations collapsed)","field":"biomedical (obstetrics and gynecology)","wr":"board-member recommendation vs editor's final decision","conf":"high","self":false,"doi":"10.1089/jwh.2009.1904","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"editor-decisions","rr":"none","pr":false,"he":false,"ms":"Board members' collapsed accept or reject recommendations matched the editors' final decisions on 72% of the 6062 reviews (the abstract rounds this to approximately 73%), an uncorrected agreement rate paralleling the kappa of 0.40.","vf":"unverified"},{"key":"E4ZVSVLA","au":"Wolff, Wirt M.","y":1970,"cx":"Journal","ob":"journal-manuscript","fam":"Kendall-W","form":"coefficient of concordance","estd":"Kendall W","v":0.59,"n":"15","k":"66","samp":"special","blind":"unclear","agg":"unspecified","scale":"rank 1-15, 1 = most important","field":"clinical and personality psychology","wr":"editors ranking 15 manuscript criteria by importance","conf":"med","self":false,"doi":"10.1037/h0029770","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":true,"he":false,"ms":"Sixty-six clinical and personality journal editors each independently ranked 15 manuscript-evaluation criteria by importance; a Kendall coefficient of concordance of 0.59 indicates a significant but moderate level of agreement about the criterion hierarchy. n_ratings_total 990 derived from stated complete crossing, each of the 66 complete questionnaires ranking all 15 items.","vf":"unverified"},{"key":"E4ZVSVLA","au":"Wolff, Wirt M.","y":1970,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"rho","estd":"correlation","v":0.91,"n":"15","k":"","samp":"special","blind":"unclear","agg":"average-of-k","scale":"rank criteria in order of importance","field":"clinical and personality psychology","wr":"two editor cohorts' mean importance rankings of criteria","conf":"med","self":false,"doi":"10.1037/h0029770","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"The 66 clinical and personality journal editors' mean importance rankings of the criteria were correlated with the rankings from Frantz's 1968 survey of counselling and guidance journal editors; the cross-study rho was 0.91. The comparison cohort's data come from the published Frantz study; the paper states 121 editorial-board members across the two studies combined.","vf":"unverified"},{"key":"E4ZVSVLA","au":"Wolff, Wirt M.","y":1970,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"corrected (Spearman-Brown) split-half reliabilities for item intercorrelations, odd-even dichotomy","estd":"correlation","v":0.49,"n":"15","k":"66","samp":"special","blind":"unclear","agg":"average-of-k","scale":"rank 1-15, 1 = most important","field":"clinical and personality psychology","wr":"item intercorrelations across odd-even editor halves","conf":"med","self":false,"doi":"10.1037/h0029770","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"The intercorrelations among the 15 criterion items were computed in the odd-numbered and even-numbered editor halves and compared; the Spearman-Brown corrected reliability was 0.49, again used to argue against reducing the criteria to a simpler factor structure. Objects are the 15 criterion items; interpretation is somewhat ambiguous in the source text.","vf":"unverified"},{"key":"E4ZVSVLA","au":"Wolff, Wirt M.","y":1970,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"odd-even dichotomy of editors' item rankings, intercorrelated (uncorrected)","estd":"correlation","v":0.97,"n":"15","k":"33","samp":"special","blind":"unclear","agg":"average-of-k","scale":"rank 1-15, 1 = most important","field":"clinical and personality psychology","wr":"two odd-even 33-editor halves' mean rankings of 15 criteria","conf":"med","self":false,"doi":"10.1037/h0029770","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"An odd-even split of the 66 editors produced two 33-editor mean rankings of the 15 criteria that correlated 0.97, again indicating highly reliable group-mean rankings.","vf":"unverified"},{"key":"E4ZVSVLA","au":"Wolff, Wirt M.","y":1970,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"corrected (Spearman-Brown) split-half reliabilities for item intercorrelations","estd":"correlation","v":0.26,"n":"15","k":"66","samp":"special","blind":"unclear","agg":"average-of-k","scale":"rank 1-15, 1 = most important","field":"clinical and personality psychology","wr":"item intercorrelations across random editor split-halves","conf":"med","self":false,"doi":"10.1037/h0029770","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"The intercorrelations among the 15 criterion items were computed in each random 33-editor half and compared; the Spearman-Brown corrected split-half reliability was 0.26, a low value the author used to argue against reducing the criteria to a simpler factor structure. Objects are the 15 criterion items; interpretation is somewhat ambiguous in the source text.","vf":"unverified"},{"key":"E4ZVSVLA","au":"Wolff, Wirt M.","y":1970,"cx":"Journal","ob":"journal-manuscript","fam":"correlation","form":"split-half pairing of editors' item rankings, intercorrelated (uncorrected)","estd":"correlation","v":0.94,"n":"15","k":"33","samp":"special","blind":"unclear","agg":"average-of-k","scale":"rank 1-15, 1 = most important","field":"clinical and personality psychology","wr":"two random 33-editor halves' mean rankings of 15 criteria","conf":"med","self":false,"doi":"10.1037/h0029770","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"other","rr":"none","pr":false,"he":false,"ms":"The 66 editors were split into two random halves of 33; the two halves' mean rankings of the 15 criteria correlated 0.94, indicating highly reliable group-mean rankings.","vf":"unverified"},{"key":"GRDE4DDS","au":"Wong, Victoria S S","y":2017,"cx":"Journal","ob":"review-report","fam":"ICC","form":"intra-rater correlation coefficient (ICC)","estd":"ICC (single/unspec)","v":0.7,"n":"","k":"2","samp":"special","blind":"unclear","agg":"unspecified","scale":"5-point Likert scale","field":"biomedical","wr":"two study authors rated quality of residents' manuscript reviews","conf":"high","self":false,"doi":"10.1186/s41073-017-0032-0","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":true,"he":false,"ms":"Two study authors, blinded to group and identity, independently rated the quality of neurology residents' reviews of standardised manuscripts using the Review Quality Instrument. An ICC of 0.7 indicates good inter-rater reliability between the two raters.","vf":"unverified"},{"key":"2F6KAXNL","au":"Wood, Michael","y":2004,"cx":"Journal","ob":"conference-abstract","fam":"percent-agreement","form":"exact agreement on collapsed 2-point (good/bad) scale","estd":"percent agreement","v":0.5,"n":"58","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"good / bad (four-point collapsed to two)","field":"information systems","wr":"reviewers' good/bad verdicts on conference papers","conf":"high","self":false,"doi":"10.1177/0165551504041673","ciLow":0.37,"ciHigh":0.63,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"After collapsing the four-point scale to good/bad, two reviewers agreed on 50% of 58 papers. The paper reports a 95% CI of 37% to 63% for the complementary 50% disagreement rate, which is numerically identical for the 50% agreement rate.","vf":"unverified"},{"key":"2F6KAXNL","au":"Wood, Michael","y":2004,"cx":"Journal","ob":"conference-abstract","fam":"percent-agreement","form":"exact agreement on collapsed 2-point (good/bad) scale","estd":"percent agreement","v":0.68,"n":"68","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"good / bad (four-point collapsed to two)","field":"information systems","wr":"reviewers' good/bad verdicts on conference papers","conf":"high","self":false,"doi":"10.1177/0165551504041673","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"After collapsing to good/bad, two reviewers agreed on 68% of 68 papers at UKAIS 2002; the corresponding disagreement rate was 32%.","vf":"unverified"},{"key":"2F6KAXNL","au":"Wood, Michael","y":2004,"cx":"Journal","ob":"conference-abstract","fam":"percent-agreement","form":"exact agreement (identical grade) on 4-point scale","estd":"percent agreement","v":0.28,"n":"58","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"accept / accept minor rev / accept major rev / reject","field":"information systems","wr":"reviewers' 4-point grades on conference papers","conf":"high","self":false,"doi":"10.1177/0165551504041673","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Two reviewers independently graded each of 58 conference papers on a four-point scale; they gave the identical grade for 28% of papers (exact agreement).","vf":"unverified"},{"key":"2F6KAXNL","au":"Wood, Michael","y":2004,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"Cohen's kappa, unweighted, on collapsed 2-point scale","estd":"kappa","v":-0.04,"n":"58","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"good / bad (four-point collapsed to two)","field":"information systems","wr":"reviewers' good/bad verdicts on conference papers","conf":"high","self":false,"doi":"10.1177/0165551504041673","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":true,"ms":"Cohen's kappa for two reviewers' collapsed good/bad verdicts on 58 papers was -0.04, indicating agreement marginally below chance level.","vf":"unverified"},{"key":"2F6KAXNL","au":"Wood, Michael","y":2004,"cx":"Journal","ob":"conference-abstract","fam":"kappa","form":"Cohen's kappa, unweighted, on collapsed 2-point scale","estd":"kappa","v":0.3,"n":"68","k":"2","samp":"full-pool","blind":"double","agg":"single-rater","scale":"good / bad (four-point collapsed to two)","field":"information systems","wr":"reviewers' good/bad verdicts on conference papers","conf":"high","self":false,"doi":"10.1177/0165551504041673","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":true,"ms":"Cohen's kappa for two reviewers' collapsed good/bad verdicts on 68 papers was 0.30, described by the paper as a 'poor' level of agreement.","vf":"unverified"},{"key":"R4QME8PM","au":"Younan J.","y":2019,"cx":"Analogue","ob":"other","fam":"ICC","form":"intraclass correlation (ICC) in a 2-way random-effects model","estd":"ICC (single/unspec)","v":0.55,"n":"7","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"numeric file score, scale not stated","field":"surgical education (residency selection)","wr":"committee members scoring residency applicant files","conf":"med","self":false,"doi":"10.1503/cjs.011719","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Ten file review committee members were split between two scoring tools and evaluated the same seven residency applicant files. The group using the element-based tool reached an intraclass correlation of 0.55, indicating moderate inter-rater reliability.","vf":"unverified"},{"key":"R4QME8PM","au":"Younan J.","y":2019,"cx":"Analogue","ob":"other","fam":"ICC","form":"intraclass correlation (ICC) in a 2-way random-effects model","estd":"ICC (single/unspec)","v":0.82,"n":"7","k":"","samp":"re-reviewed-subset","blind":"single","agg":"unspecified","scale":"numeric file score, scale not stated","field":"surgical education (residency selection)","wr":"committee members scoring residency applicant files","conf":"med","self":false,"doi":"10.1503/cjs.011719","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":false,"ms":"Ten file review committee members were split between two scoring tools and evaluated the same seven residency applicant files. The group using the trait-based tool reached an intraclass correlation of 0.82, indicating good inter-rater reliability.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"Cronbach-alpha","form":"Cronbach's alpha if item deleted","estd":"Cronbach alpha","v":0.78,"n":"11","k":"11","samp":"special","blind":"unclear","agg":"unspecified","scale":"seven items, 1-5 Likert","field":"biomedical","wr":"internal consistency of RQI with constructiveness item removed","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"review-quality","rr":"unclear","pr":false,"he":false,"ms":"For Version 3.1 (11 reviews, 11 raters), Cronbach's alpha would be 0.78 if the constructiveness item were deleted.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"Cronbach-alpha","form":"Cronbach's alpha if item deleted","estd":"Cronbach alpha","v":0.87,"n":"11","k":"11","samp":"special","blind":"unclear","agg":"unspecified","scale":"seven items, 1-5 Likert","field":"biomedical","wr":"internal consistency of RQI with importance item removed","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"review-quality","rr":"unclear","pr":false,"he":false,"ms":"For Version 3.1 (11 reviews, 11 raters), Cronbach's alpha would be 0.87 if the importance item were deleted.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"Cronbach-alpha","form":"Cronbach's alpha if item deleted","estd":"Cronbach alpha","v":0.8,"n":"11","k":"11","samp":"special","blind":"unclear","agg":"unspecified","scale":"seven items, 1-5 Likert","field":"biomedical","wr":"internal consistency of RQI with interpretation item removed","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"review-quality","rr":"unclear","pr":false,"he":false,"ms":"For Version 3.1 (11 reviews, 11 raters), Cronbach's alpha would be 0.80 if the interpretation item were deleted.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"Cronbach-alpha","form":"Cronbach's alpha if item deleted","estd":"Cronbach alpha","v":0.78,"n":"11","k":"11","samp":"special","blind":"unclear","agg":"unspecified","scale":"seven items, 1-5 Likert","field":"biomedical","wr":"internal consistency of RQI with method item removed","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"review-quality","rr":"unclear","pr":false,"he":false,"ms":"For Version 3.1 (11 reviews, 11 raters), Cronbach's alpha would be 0.78 if the method item were deleted.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"Cronbach-alpha","form":"Cronbach's alpha if item deleted","estd":"Cronbach alpha","v":0.87,"n":"11","k":"11","samp":"special","blind":"unclear","agg":"unspecified","scale":"seven items, 1-5 Likert","field":"biomedical","wr":"internal consistency of RQI with originality item removed","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"review-quality","rr":"unclear","pr":false,"he":false,"ms":"For Version 3.1 (11 reviews, 11 raters), Cronbach's alpha would be 0.87 if the originality item were deleted.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"Cronbach-alpha","form":"Cronbach's alpha if item deleted","estd":"Cronbach alpha","v":0.8,"n":"11","k":"11","samp":"special","blind":"unclear","agg":"unspecified","scale":"seven items, 1-5 Likert","field":"biomedical","wr":"internal consistency of RQI with presentation item removed","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"review-quality","rr":"unclear","pr":false,"he":false,"ms":"For Version 3.1 (11 reviews, 11 raters), Cronbach's alpha would be 0.80 if the presentation item were deleted.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"Cronbach-alpha","form":"Cronbach's alpha if item deleted","estd":"Cronbach alpha","v":0.78,"n":"11","k":"11","samp":"special","blind":"unclear","agg":"unspecified","scale":"seven items, 1-5 Likert","field":"biomedical","wr":"internal consistency of RQI with substantiation item removed","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"review-quality","rr":"unclear","pr":false,"he":false,"ms":"For Version 3.1 (11 reviews, 11 raters), Cronbach's alpha would be 0.78 if the substantiation item were deleted.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"Cronbach-alpha","form":"Cronbach's alpha","estd":"Cronbach alpha","v":0.84,"n":"11","k":"11","samp":"special","blind":"unclear","agg":"unspecified","scale":"seven items, 1-5 Likert","field":"biomedical","wr":"internal consistency of RQI Version 3.2 across items","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"review-quality","rr":"unclear","pr":false,"he":false,"ms":"Eleven BMJ editors rated 11 reviews with the final Version 3.2; the overall Cronbach's alpha was 0.84, varying from 0.65 to 0.91 across raters.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"Cronbach-alpha","form":"Cronbach's alpha","estd":"Cronbach alpha","v":0.81,"n":"15","k":"2","samp":"special","blind":"unclear","agg":"unspecified","scale":"pre-existing review-quality instrument","field":"biomedical","wr":"internal consistency of the pre-existing instrument on own data","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"review-quality","rr":"unclear","pr":false,"he":false,"ms":"The pre-existing (McNutt) review-quality instrument, administered by the authors to two editors on 15 reviews, had a Cronbach's alpha of 0.81 (own result, not a cited value).","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"Cronbach-alpha","form":"Cronbach's alpha","estd":"Cronbach alpha","v":0.83,"n":"20","k":"7","samp":"special","blind":"unclear","agg":"unspecified","scale":"eight items, 1-5 Likert (1 = poor, 5 = excellent)","field":"biomedical","wr":"internal consistency of RQI Version 1 across items","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"review-quality","rr":"unclear","pr":false,"he":false,"ms":"Version 1, tested by seven researchers and editors rating 20 neurology-journal reviews, had a Cronbach's alpha of 0.83, though inter-rater reliability of some items was poor.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"Cronbach-alpha","form":"Cronbach's alpha","estd":"Cronbach alpha","v":0.78,"n":"15","k":"2","samp":"special","blind":"unclear","agg":"unspecified","scale":"seven items, 1-5 Likert","field":"biomedical","wr":"internal consistency of RQI Version 2 across items","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"review-quality","rr":"unclear","pr":false,"he":false,"ms":"Version 2, tested by two BMJ editors on 15 reviews, had a Cronbach's alpha of 0.78.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"Cronbach-alpha","form":"Cronbach's alpha","estd":"Cronbach alpha","v":0.91,"n":"11","k":"11","samp":"special","blind":"unclear","agg":"unspecified","scale":"seven items, 1-5 Likert","field":"biomedical","wr":"internal consistency of RQI Version 3.1 across items","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"internal-consistency","tgt":"review-quality","rr":"unclear","pr":false,"he":false,"ms":"Version 3.1, tested on 11 reviews by 11 BMJ editors, had a reported Cronbach's alpha of 0.91 across its items.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"kappa","form":"Kappa statistic (K)","estd":"kappa","v":0.17,"n":"934","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = not at all, 5 = very constructive","field":"biomedical","wr":"BMJ editors rating the constructiveness item of reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"For the constructiveness item, the unweighted kappa between two independent editors over 934 reviews was 0.17.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"kappa","form":"Kappa statistic (K)","estd":"kappa","v":0.25,"n":"934","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = poor, 5 = excellent","field":"biomedical","wr":"BMJ editors rating the global overall-quality item of reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"For the global overall-quality item, the unweighted kappa between two independent editors over 934 reviews was 0.25.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"kappa","form":"Kappa statistic (K)","estd":"kappa","v":0.17,"n":"934","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = not at all, 5 = extensively","field":"biomedical","wr":"BMJ editors rating the importance item of reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"For the importance item, the unweighted kappa between two independent editors over 934 reviews was 0.17.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"kappa","form":"Kappa statistic (K)","estd":"kappa","v":0.19,"n":"934","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = not at all, 5 = extensively","field":"biomedical","wr":"BMJ editors rating the interpretation item of reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"For the interpretation item, the unweighted kappa between two independent editors over 934 reviews was 0.19.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"kappa","form":"Kappa statistic (K)","estd":"kappa","v":0.31,"n":"934","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"mean of seven items scored 1 to 5","field":"biomedical","wr":"BMJ editors rating quality of manuscript peer reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"Across the same 934 reviews rated by two independent editors, the unweighted kappa for the mean total RQI score was 0.31.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"kappa","form":"Kappa statistic (K)","estd":"kappa","v":0.25,"n":"934","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = not at all, 5 = comprehensive","field":"biomedical","wr":"BMJ editors rating the method item of reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"For the method item, the unweighted kappa between two independent editors over 934 reviews was 0.25.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"kappa","form":"Kappa statistic (K)","estd":"kappa","v":0.33,"n":"934","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = not at all, 5 = extensively with references","field":"biomedical","wr":"BMJ editors rating the originality item of reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"For the originality item, the unweighted kappa between two independent editors over 934 reviews was 0.33.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"kappa","form":"Kappa statistic (K)","estd":"kappa","v":0.22,"n":"934","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = not at all, 5 = extensive","field":"biomedical","wr":"BMJ editors rating the presentation item of reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"For the presentation item, the unweighted kappa between two independent editors over 934 reviews was 0.22.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"kappa","form":"Kappa statistic (K)","estd":"kappa","v":0.24,"n":"934","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = none, 5 = all comments substantiated","field":"biomedical","wr":"BMJ editors rating the substantiation item of reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"For the substantiation item, the unweighted kappa between two independent editors over 934 reviews was 0.24.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"weighted-kappa","form":"weighted kappa statistic (Kw)","estd":"weighted kappa","v":0.53,"n":"934","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = not at all, 5 = very constructive","field":"biomedical","wr":"BMJ editors rating the constructiveness item of reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"For the constructiveness item, two editors' independent ratings of 934 reviews gave a weighted kappa of 0.53 (unweighted kappa 0.17).","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"weighted-kappa","form":"weighted kappa statistic (Kw)","estd":"weighted kappa","v":0.73,"n":"934","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = poor, 5 = excellent","field":"biomedical","wr":"BMJ editors rating the global overall-quality item of reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"For the global overall-quality item, two editors' independent ratings of 934 reviews gave a weighted kappa of 0.73 (unweighted kappa 0.25).","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"weighted-kappa","form":"weighted kappa statistic (Kw)","estd":"weighted kappa","v":0.49,"n":"934","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = not at all, 5 = extensively","field":"biomedical","wr":"BMJ editors rating the importance item of reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"For the importance item, two editors' independent ratings of 934 reviews gave a weighted kappa of 0.49 (unweighted kappa 0.17).","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"weighted-kappa","form":"weighted kappa statistic (Kw)","estd":"weighted kappa","v":0.52,"n":"934","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = not at all, 5 = extensively","field":"biomedical","wr":"BMJ editors rating the interpretation item of reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"For the interpretation item, two editors' independent ratings of 934 reviews gave a weighted kappa of 0.52 (unweighted kappa 0.19).","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"weighted-kappa","form":"weighted kappa statistic (Kw)","estd":"weighted kappa","v":0.83,"n":"934","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"mean of seven items scored 1 to 5","field":"biomedical","wr":"BMJ editors rating quality of manuscript peer reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":true,"he":false,"ms":"Two BMJ editors independently rated the quality of 934 manuscript reviews from a randomised trial; the mean total RQI score had a weighted kappa of 0.83, the study's headline inter-rater reliability.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"weighted-kappa","form":"weighted kappa statistic (Kw)","estd":"weighted kappa","v":0.66,"n":"934","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = not at all, 5 = comprehensive","field":"biomedical","wr":"BMJ editors rating the method item of reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"For the method item, two editors' independent ratings of 934 reviews gave a weighted kappa of 0.66 (unweighted kappa 0.25).","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"weighted-kappa","form":"weighted kappa statistic (Kw)","estd":"weighted kappa","v":0.7,"n":"934","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = not at all, 5 = extensively with references","field":"biomedical","wr":"BMJ editors rating the originality item of reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"For the originality item, two editors' independent ratings of 934 reviews gave a weighted kappa of 0.70 (unweighted kappa 0.33).","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"weighted-kappa","form":"weighted kappa statistic (Kw)","estd":"weighted kappa","v":0.58,"n":"934","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = not at all, 5 = extensive","field":"biomedical","wr":"BMJ editors rating the presentation item of reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"For the presentation item, two editors' independent ratings of 934 reviews gave a weighted kappa of 0.58 (unweighted kappa 0.22).","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"weighted-kappa","form":"weighted kappa statistic (Kw)","estd":"weighted kappa","v":0.54,"n":"934","k":"2","samp":"full-pool","blind":"unclear","agg":"single-rater","scale":"1 = none, 5 = all comments substantiated","field":"biomedical","wr":"BMJ editors rating the substantiation item of reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"none","pr":false,"he":false,"ms":"For the substantiation item, two editors' independent ratings of 934 reviews gave a weighted kappa of 0.54 (unweighted kappa 0.24).","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"weighted-kappa","form":"weighted kappa statistic (Kw)","estd":"weighted kappa","v":0.95,"n":"11","k":"11","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1 = not at all, 5 = very constructive","field":"biomedical","wr":"BMJ editors re-rating the constructiveness item of reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"review-quality","rr":"unclear","pr":false,"he":false,"ms":"For the constructiveness item, 11 editors re-rating 11 reviews after two months gave a test-retest weighted kappa of 0.95.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"weighted-kappa","form":"weighted kappa statistic (Kw)","estd":"weighted kappa","v":0.96,"n":"11","k":"11","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1 = poor, 5 = excellent","field":"biomedical","wr":"BMJ editors re-rating the global overall-quality item of reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"review-quality","rr":"unclear","pr":false,"he":false,"ms":"For the global overall-quality item, 11 editors re-rating 11 reviews after two months gave a test-retest weighted kappa of 0.96.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"weighted-kappa","form":"weighted kappa statistic (Kw)","estd":"weighted kappa","v":0.7,"n":"11","k":"11","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1 = not at all, 5 = extensively","field":"biomedical","wr":"BMJ editors re-rating the importance item of reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"review-quality","rr":"unclear","pr":false,"he":false,"ms":"For the importance item, 11 editors re-rating 11 reviews after two months gave a test-retest weighted kappa of 0.70 (last column of Table 2).","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"weighted-kappa","form":"weighted kappa statistic (Kw)","estd":"weighted kappa","v":0.71,"n":"11","k":"11","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1 = not at all, 5 = extensively","field":"biomedical","wr":"BMJ editors re-rating the interpretation item of reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"review-quality","rr":"unclear","pr":false,"he":false,"ms":"For the interpretation item, 11 editors re-rating 11 reviews after two months gave a test-retest weighted kappa of 0.71.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"weighted-kappa","form":"weighted kappa statistic (Kw)","estd":"weighted kappa","v":1,"n":"11","k":"11","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"mean of seven items scored 1 to 5","field":"biomedical","wr":"BMJ editors re-rating quality of the same reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"review-quality","rr":"unclear","pr":false,"he":false,"ms":"Eleven BMJ editors re-rated the same 11 reviews after two months with Version 3.2; the mean total score had a test-retest weighted kappa of 1.00.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"weighted-kappa","form":"weighted kappa statistic (Kw)","estd":"weighted kappa","v":0.89,"n":"11","k":"11","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1 = not at all, 5 = comprehensive","field":"biomedical","wr":"BMJ editors re-rating the method item of reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"review-quality","rr":"unclear","pr":false,"he":false,"ms":"For the method item, 11 editors re-rating 11 reviews after two months gave a test-retest weighted kappa of 0.89.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"weighted-kappa","form":"weighted kappa statistic (Kw)","estd":"weighted kappa","v":0.88,"n":"11","k":"11","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1 = not at all, 5 = extensively with references","field":"biomedical","wr":"BMJ editors re-rating the originality item of reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"review-quality","rr":"unclear","pr":false,"he":false,"ms":"For the originality item, 11 editors re-rating 11 reviews after two months gave a test-retest weighted kappa of 0.88.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"weighted-kappa","form":"weighted kappa statistic (Kw)","estd":"weighted kappa","v":0.85,"n":"11","k":"11","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1 = not at all, 5 = extensive","field":"biomedical","wr":"BMJ editors re-rating the presentation item of reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"review-quality","rr":"unclear","pr":false,"he":false,"ms":"For the presentation item, 11 editors re-rating 11 reviews after two months gave a test-retest weighted kappa of 0.85.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"weighted-kappa","form":"weighted kappa statistic (Kw)","estd":"weighted kappa","v":0.94,"n":"11","k":"11","samp":"re-reviewed-subset","blind":"unclear","agg":"single-rater","scale":"1 = none, 5 = all comments substantiated","field":"biomedical","wr":"BMJ editors re-rating the substantiation item of reviews","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"test-retest","tgt":"review-quality","rr":"unclear","pr":false,"he":false,"ms":"For the substantiation item, 11 editors re-rating 11 reviews after two months gave a test-retest weighted kappa of 0.94.","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"Kendall-W","form":"Kendall coefficient of concordance","estd":"Kendall W","v":0.83,"n":"11","k":"11","samp":"special","blind":"unclear","agg":"unspecified","scale":"1-5 Likert, 1 = poor, 5 = excellent","field":"biomedical","wr":"eleven editors rating reviews (Version 3.1 items)","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"unclear","pr":false,"he":false,"ms":"For Version 3.1, the reported Kendall concordance across six items and the global item ranged up to 0.83, the upper bound (individual item values not given).","vf":"unverified"},{"key":"E5YQG6G8","au":"van Rooyen, S","y":1999,"cx":"Journal","ob":"review-report","fam":"Kendall-W","form":"Kendall coefficient of concordance","estd":"Kendall W","v":0.65,"n":"11","k":"11","samp":"special","blind":"unclear","agg":"unspecified","scale":"1-5 Likert, 1 = poor, 5 = excellent","field":"biomedical","wr":"eleven editors rating reviews (Version 3.1 items)","conf":"high","self":false,"doi":"10.1016/s0895-4356(99)00047-5","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"review-quality","rr":"unclear","pr":false,"he":false,"ms":"For Version 3.1, 11 editors rating 11 reviews reached Kendall concordance of 0.65 to 0.83 across six items and the global item; 0.65 is the reported lower bound (individual item values not given).","vf":"unverified"},{"key":"WZVF5TQJ","au":"Şimşek, Müge","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"percentage of concordant pairs: exact agreement on which of two applications is better; ties broken randomly","estd":"percent agreement","v":0.62,"n":"133","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1 (excellent) to 9 (bad), CV score","field":"social sciences and humanities","wr":"panelists ranking applications on CV quality","conf":"high","self":false,"doi":"10.1007/s11192-024-04968-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Using only the score for the quality of the researcher, panelists in the P–N group ranked a pair of applications in the same order in 62.0% of 355 paired comparisons.","vf":"unverified"},{"key":"WZVF5TQJ","au":"Şimşek, Müge","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"percentage of concordant pairs: exact agreement on which of two applications is better; ties broken randomly","estd":"percent agreement","v":0.53,"n":"133","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1 (excellent) to 9 (bad), knowledge utilization score","field":"social sciences and humanities","wr":"panelists ranking applications on knowledge utilization","conf":"high","self":false,"doi":"10.1007/s11192-024-04968-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"On the score for the potential societal and economic use of the knowledge, panelists in the P–N group ranked a pair of applications in the same order in 53.0% of 355 paired comparisons.","vf":"unverified"},{"key":"WZVF5TQJ","au":"Şimşek, Müge","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"percentage of concordant pairs: exact agreement on which of two applications is better; ties broken randomly","estd":"percent agreement","v":0.552,"n":"133","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1 (excellent) to 9 (bad); overall = weighted sum of three","field":"social sciences and humanities","wr":"panelists ranking pairs of grant applications","conf":"high","self":false,"doi":"10.1007/s11192-024-04968-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"For applications where only one of the two panelists could read the proposal text, a regular panelist ranked a pair of applications the same way as the shadow panelist in 55.2% of 355 paired comparisons, against 50% expected by chance. The comparisons covered 133 applications in the first round of NWO's 2018 Vidi competition.","vf":"unverified"},{"key":"WZVF5TQJ","au":"Şimşek, Müge","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"percentage of concordant pairs: exact agreement on which of two applications is better; ties broken randomly","estd":"percent agreement","v":0.501,"n":"133","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1 (excellent) to 9 (bad), proposal score","field":"social sciences and humanities","wr":"panelists ranking applications on proposal quality","conf":"high","self":false,"doi":"10.1007/s11192-024-04968-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"On the score for the quality of the proposed research, panelists in the P–N group ranked a pair of applications in the same order in 50.1% of 355 paired comparisons, which is what random scoring would produce.","vf":"unverified"},{"key":"WZVF5TQJ","au":"Şimşek, Müge","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"percentage of concordant pairs: exact agreement on which of two applications is better; ties broken randomly","estd":"percent agreement","v":0.597,"n":"130","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1 (excellent) to 9 (bad), CV score","field":"social sciences and humanities","wr":"panelists ranking applications on CV quality","conf":"high","self":false,"doi":"10.1007/s11192-024-04968-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Using only the score for the quality of the researcher, panelists in the P–P group, where both had read the proposals, ranked a pair of applications in the same order in 59.7% of 367 paired comparisons.","vf":"unverified"},{"key":"WZVF5TQJ","au":"Şimşek, Müge","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"percentage of concordant pairs: exact agreement on which of two applications is better; ties broken randomly","estd":"percent agreement","v":0.526,"n":"130","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1 (excellent) to 9 (bad), knowledge utilization score","field":"social sciences and humanities","wr":"panelists ranking applications on knowledge utilization","conf":"high","self":false,"doi":"10.1007/s11192-024-04968-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"On the score for the potential societal and economic use of the knowledge, panelists in the P–P group ranked a pair of applications in the same order in 52.6% of 367 paired comparisons.","vf":"unverified"},{"key":"WZVF5TQJ","au":"Şimşek, Müge","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"percentage of concordant pairs: exact agreement on which of two applications is better; ties broken randomly","estd":"percent agreement","v":0.589,"n":"130","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1 (excellent) to 9 (bad); overall = weighted sum of three","field":"social sciences and humanities","wr":"panelists ranking pairs of grant applications","conf":"high","self":false,"doi":"10.1007/s11192-024-04968-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":true,"he":false,"ms":"Under ordinary conditions, where both panelists read the full proposals, a regular panelist and a shadow panelist agreed on which of two applications was better in 58.9% of 367 paired comparisons, only marginally above the 50% expected from random scoring. This is the study's headline agreement figure for standard practice.","vf":"unverified"},{"key":"WZVF5TQJ","au":"Şimşek, Müge","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"percent-agreement","form":"percentage of concordant pairs: exact agreement on which of two applications is better; ties broken randomly","estd":"percent agreement","v":0.534,"n":"130","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1 (excellent) to 9 (bad), proposal score","field":"social sciences and humanities","wr":"panelists ranking applications on proposal quality","conf":"high","self":false,"doi":"10.1007/s11192-024-04968-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"On the score for the quality of the proposed research, two panelists who had both read the proposals agreed on which of two applications was better in only 53.4% of 367 paired comparisons, barely above chance. The authors read this as evidence that proposal evaluation is not informative.","vf":"unverified"},{"key":"WZVF5TQJ","au":"Şimşek, Müge","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"mean absolute difference in standardized CV scores between two panelists reviewing the same application","estd":"AD index","v":0.79,"n":"133","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1 (excellent) to 9 (bad), standardized within panelists","field":"social sciences and humanities","wr":"panelist pairs' gap on CV scores","conf":"high","self":false,"doi":"10.1007/s11192-024-04968-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"On the CV score, the two panelists rating the same application in the P–N group differed by 0.79 standard deviations on average, with a standard deviation of 0.61.","vf":"unverified"},{"key":"WZVF5TQJ","au":"Şimşek, Müge","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"mean absolute difference in standardized knowledge utilization scores between two panelists reviewing the same application","estd":"AD index","v":0.99,"n":"133","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1 (excellent) to 9 (bad), standardized within panelists","field":"social sciences and humanities","wr":"panelist pairs' gap on knowledge utilization scores","conf":"high","self":false,"doi":"10.1007/s11192-024-04968-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"On the knowledge utilization score, the two panelists rating the same application in the P–N group differed by 0.99 standard deviations on average, with a standard deviation of 0.70.","vf":"unverified"},{"key":"WZVF5TQJ","au":"Şimşek, Müge","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"mean absolute difference in standardized overall scores between two panelists reviewing the same application","estd":"AD index","v":0.89,"n":"133","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1 (excellent) to 9 (bad), standardized within panelists","field":"social sciences and humanities","wr":"panelist pairs' score gap on applications","conf":"high","self":false,"doi":"10.1007/s11192-024-04968-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Where only one of the two panelists had read the proposal, the two panelists' standardised overall scores for the same application differed by 0.89 standard deviations on average across 342 pairs, with a standard deviation of 0.67. This is a disagreement measure, so higher values mean less agreement.","vf":"unverified"},{"key":"WZVF5TQJ","au":"Şimşek, Müge","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"mean absolute difference in standardized proposal scores between two panelists reviewing the same application","estd":"AD index","v":0.98,"n":"133","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1 (excellent) to 9 (bad), standardized within panelists","field":"social sciences and humanities","wr":"panelist pairs' gap on proposal scores","conf":"high","self":false,"doi":"10.1007/s11192-024-04968-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"On the proposal score, the two panelists rating the same application in the P–N group differed by 0.98 standard deviations on average, with a standard deviation of 0.75.","vf":"unverified"},{"key":"WZVF5TQJ","au":"Şimşek, Müge","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"mean absolute difference in standardized CV scores between two panelists reviewing the same application","estd":"AD index","v":0.81,"n":"130","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1 (excellent) to 9 (bad), standardized within panelists","field":"social sciences and humanities","wr":"panelist pairs' gap on CV scores","conf":"high","self":false,"doi":"10.1007/s11192-024-04968-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"On the CV score, the two panelists rating the same application in the P–P group differed by 0.81 standard deviations on average, with a standard deviation of 0.58.","vf":"unverified"},{"key":"WZVF5TQJ","au":"Şimşek, Müge","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"mean absolute difference in standardized knowledge utilization scores between two panelists reviewing the same application","estd":"AD index","v":0.99,"n":"130","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1 (excellent) to 9 (bad), standardized within panelists","field":"social sciences and humanities","wr":"panelist pairs' gap on knowledge utilization scores","conf":"high","self":false,"doi":"10.1007/s11192-024-04968-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"On the knowledge utilization score, the two panelists rating the same application in the P–P group differed by 0.99 standard deviations on average, with a standard deviation of 0.68, identical to the no-proposal group.","vf":"unverified"},{"key":"WZVF5TQJ","au":"Şimşek, Müge","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"mean absolute difference in standardized overall scores between two panelists reviewing the same application","estd":"AD index","v":0.93,"n":"130","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1 (excellent) to 9 (bad), standardized within panelists","field":"social sciences and humanities","wr":"panelist pairs' score gap on applications","conf":"high","self":false,"doi":"10.1007/s11192-024-04968-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"Where both panelists had read the proposal, their standardised overall scores for the same application differed by 0.93 standard deviations on average across 314 pairs, with a standard deviation of 0.63. Disagreement was thus no smaller than in the no-proposal group.","vf":"unverified"},{"key":"WZVF5TQJ","au":"Şimşek, Müge","y":2024,"cx":"Grant","ob":"grant-proposal","fam":"AD-index","form":"mean absolute difference in standardized proposal scores between two panelists reviewing the same application","estd":"AD index","v":1.01,"n":"130","k":"2","samp":"re-reviewed-subset","blind":"single","agg":"single-rater","scale":"1 (excellent) to 9 (bad), standardized within panelists","field":"social sciences and humanities","wr":"panelist pairs' gap on proposal scores","conf":"high","self":false,"doi":"10.1007/s11192-024-04968-7","ciLow":null,"ciHigh":null,"mt":"inter-rater","tgt":"submission-scores","rr":"none","pr":false,"he":false,"ms":"On the proposal score, the two panelists rating the same application in the P–P group differed by 1.01 standard deviations on average, with a standard deviation of 0.70, the largest disagreement of any scoring dimension.","vf":"unverified"}],"summaries":{"6754BVQM":"A study of a US Department of Defense breast cancer research programme finding that including patient-advocate consumers as voting panel members changed relatively few proposal scores and was viewed positively by both consumers and scientists.","RVJKZWBR":"Analysis of nearly 900 grant applications for career-development awards found that interviews and committee discussion, more than the initial written peer reviews, shaped decisions about which applicants were judged to show research talent.","W5SHTMNM":"An empirical study of 897 career grant applications examining how review panels select talent, finding no clear boundaries of excellence, a dominant role for the interview phase, little influence of external peer reviews, and no gender bias in decisions.","7JG4HLD9":"A study of reviewer recommendations at a German medical journal reporting that standard kappa statistics showed poor agreement, but overall raw agreement was high and an alternative statistic (Gwet's AC1) indicated substantial agreement, with no link found between reviewer recommendations and later citations.","ITEBWPJP":"Analyzes a two-phase European Commission funding scheme in which remote reviewers first assess highly interdisciplinary proposals and a separate on-site panel later reinforces or revises scores, examining how much the panel stage adds to reviewer agreement and proposal selection.","RJEGBDFX":"Models the grant peer-review process used by Turkish regional development agencies as a Bayesian game among three reviewers balancing false-acceptance and false-rejection risks, comparing its outcomes to two simpler review procedures.","BSI9PZ5P":"Uses large language models to automatically classify peer reviews as thorough or superficial, testing whether AI-based sorting of review quality aligns with human judgments.","VPW3P5QI":"Simulation study at the Swiss National Science Foundation comparing remote score-based evaluation of postdoctoral fellowship applications with face-to-face panel review, finding the simplified procedure agreed with official funding decisions in over 80% of cases while halving costs.","HPC6SYSG":"Analyzing nearly 2,000 conference poster reviews, this study found reviewers who were also authors rated posters more critically, while posters with reviewer-authors received more favorable ratings, consistent with self-serving bias effects.","PSBEMHHH":"Proposes a science-funding model where each researcher receives equal base funding but must redistribute part of it to peers, aiming to reach a crowd-informed funding distribution without formal proposals or review.","FDNSXEWP":"Large-scale study of a biomedical fellowship funder's selection procedure assessing reviewer agreement, fairness across applicant characteristics, and whether review scores predicted later research success.","5CTCSBI3":"Study of a European fellowship program's peer review finding that reviewers' own experience, country, and gender had no significant effect on their ratings, which instead tracked applicants' scientific achievement.","49ZXXL9X":"Latent Markov model analysis of a three-stage fellowship review process finding that a positive first-stage external review was essentially a precondition for eventual funding, supporting the value of early-stage screening.","MAPYT9ZR":"Examining referee agreement and the link between publication decisions and later citation rates at a chemistry journal, the study concludes its peer-review process maintained consistent quality despite a sharp rise in submissions and falling acceptance rates.","SHHLHEPD":"A study of one open-access journal's public peer review process found reviewer agreement was still low to only moderate, similar to findings previously reported for traditional closed peer review.","8AAUI9F9":"Analyzing about 100,000 papers, this study found low interrater agreement among F1000Prime post-publication reviewers, but showed that papers rated more highly were more often highly cited, supporting the system's validity.","TE7UT3JG":"Randomized field experiment assigning over 2,000 evaluator-proposal pairs at a research university finds evaluators give lower scores to proposals closer to their own expertise and to highly novel proposals, consistent with bounded rationality rather than random noise.","PB3TYFEA":"An older commentary discusses longstanding dissatisfaction among journal editors and authors with manuscript evaluation procedures, describing gatekeeping challenges and reviewer disagreement from an editor's perspective.","73TNGMX8":"Analyzing eight years of editorial decisions across four computer science journals, the authors found that authors with higher network-based reputation were less likely to be rejected even after receiving negative reviews.","BGW9Q5PA":"A mixed-methods evaluation of a foundation's trial of \"distributed peer review,\" where grant applicants reviewed each other's proposals, finds moderate overlap with a parallel panel's funding choices and substantial variability between individual reviewers.","SA7AMBC5":"A 3.5-year observational study of a peer-reviewed journal finds editors' subjective quality ratings of individual reviewers are moderately reliable and correlate with reviewers' ability to catch errors in a test manuscript, more so than with their acceptance recommendation rates.","W2VF7UBP":"An analysis of journal peer review at Biological Conservation examining agreement between reviewer recommendations and editors’ decisions, and whether manuscripts by Chinese authors receive fair treatment.","KNJE6VS9":"This paper develops a Bayesian hierarchical model for ranking items rated by multiple raters that separately estimates each rater's bias, discrimination, and error, applying the approach to grant review data.","TZXYKKCC":"A retrospective analysis of NIH grant panels found that discussion produced only small average score changes but shifted more than one in ten applications across the funding threshold, regardless of whether the panel met by teleconference or in person.","Q3AG2M2K":"Experimental study with 24 human-computer interaction reviewers testing ChatGPT-assisted peer review, finding it reduced perceived workload but did not meaningfully speed up reviews or improve their quality.","UYGWPDNF":"Analyzing manuscripts submitted to a major anthropology journal, this study finds reviewer agreement is only modestly better than chance despite an initial impression of good agreement, alongside a fairly strong link between reviewer recommendations and editors' final decisions.","RRJ5XCFM":"A statistical assessment of reviewer evaluations submitted to a major psychology journal, situated within a broader literature critical of reviewer competence, consistency, and fairness in the review process.","QQRFQCNU":"Commentary disputes an earlier author's critique of a study on inter-rater agreement in journal manuscript review, addressing disagreement over the framing of politics and science in diagnostic classification debates.","SBVJDN7K":"An experiment duplicating applications across two independent Australian early-career fellowship review panels finds relatively high agreement on funding outcomes (about 83 percent, with substantial chance-adjusted agreement), higher than reported for some other grant schemes.","WU4XSEWV":"Masked randomised trial at a","N8M2AUCU":"Assessing interrater reliability among 11 reviewers rating abstracts for an anesthesiology subspecialty meeting, the study finds only fair agreement (kappa around 0.2 to 0.4), consistent with similarly low reliability reported for other medical society abstract reviews.","JN8MZ48Y":"A study of abstract selection at several anesthesiology subspecialty meetings finds reviewer agreement ranging from poor to moderate, only slightly better than chance, and suggests clearer criteria and interval scoring scales could improve consistency.","II5SPEMP":"A classic experiment reviewing 150 US National Science Foundation proposals with a fresh set of reviewers finds substantial disagreement among eligible reviewers, indicating funding outcomes depend considerably on which reviewers happen to be assigned, without evidence of systematic selection bias.","G6BW4RFI":"A study asked early-career humanities researchers to rate the societal relevance of hundreds of published articles, finding reviewer identity mattered more than article content in shaping ratings, and suggesting much published humanities work may not register as societally relevant.","RJLFSWWX":"Describes development of a scoring tool tailored to implementation-science grant proposals, applied to 30 pilot proposals, finding most scored poorly on criteria like theoretical grounding and readiness for adoption.","QWNRA5FG":"A small randomized trial comparing a structured critical appraisal tool to informal appraisal of health research papers found the structured tool produced higher interrater agreement and reduced variance attributable to individual raters.","YWD3GJ8R":"An empirical evaluation of the journal peer-review system at Angewandte Chemie, reporting selected findings from Daniel's anonymised study of referees' judgements on submitted manuscripts.","FSW6MJXT":"Comparing review outcomes for large multi-investigator grants versus individual research grants in immunology, this study found high consistency in approval rates, scores, and reviewer consensus between the two review formats.","4BPNLHNQ":"Validation study of a structured review questionnaire (CoRE) for a medical journal, finding its concrete, expertise-sensitive items produced strong agreement between reviewer recommendations and their scored answers.","HEYTAC8Z":"Comparing large language model evaluations of 110 health-economic studies against two human reviewers using a standard reporting checklist, this study found substantial human-human disagreement and moderate-to-high agreement between the AI and human consensus scores.","2NTZHLW6":"An early study of a university ethics review committee found reviewers agreed on about two-thirds of decisions, but this was barely above what chance alone would produce.","V5LUUSHS":"A methodological paper uses NIH and biological-sciences grant review data to show that inter-rater reliability estimates can misleadingly equal zero when calculated from a narrow range of top-quality proposals, even though full-range agreement is good.","SA84R68P":"Two studies asking groups of raters to assess methodological rigor of psychology papers using structured criteria found good to excellent inter-rater reliability, even among minimally trained raters, though a recent paper sample showed generally low methodological rigor.","3MVDPR4T":"Analysing reviewer forms and a survey at an Irish funding agency, the study finds reviewers vary in which proposal topics they emphasize and how they map them to scoring criteria, but simulations suggest this variation barely affects overall reviewer disagreement.","GNW32DYJ":"A content analysis of hundreds of reviews submitted to psychology journal editors finding that different reviewers of the same paper rarely raised the same criticisms, instead each highlighting distinct, non-overlapping concerns.","Y5JCHZCD":"An empirical analysis of PCORI's first grant merit review round, finding that scientist, patient and stakeholder reviewers initially scored proposals quite differently but converged substantially after in-person panel discussion, which markedly changed the final funding rankings.","JCNY992P":"Compares scores from two independent expert panels reviewing the same medical research grant proposals, finding panel discussion did not improve agreement compared with simply averaging individual reviewers' scores.","A42HFS8G":"Analysing an NIH experiment with hundreds of reviewers scoring dozens of grant proposals, the authors find low agreement on scientific merit and estimate that roughly a dozen reviewers would be needed to achieve moderately reliable overall-impact ratings.","NNIEFSDW":"Analysis of over 1,300 applications to a patient-centered research funder found that scientists' technical-merit ratings were the strongest predictor of final funding scores, and that scoring agreement between scientist, patient, and stakeholder reviewers improved after panel discussion.","WWP3Q9HN":"A Bayesian reanalysis of the 2014 NIPS conference experiment, in which duplicated reviews showed 60% disagreement on accepted papers, estimating that around 56% of submissions met basic quality criteria and suggesting a higher acceptance rate would reduce arbitrariness.","T67A4S8E":"Studying grant review panels evaluating early-stage technologies, the authors find panellists tend to converge toward consensus after discussion, with technical experts most resistant to changing their views, though shared information did not clearly improve predictions of commercial success.","8JXS4WJA":"A retrospective analysis of roughly 1,600 grant applications found that switching from face-to-face to teleconference panel review had little effect on scores, score distributions, or interrater reliability, apart from discussion time.","9EBX8PWV":"A large observational study of biomedical grant reviews finds that reviewers with greater self-assessed subject expertise tend to score applications more harshly, replicating earlier experimental findings and pointing to a possible role for social networks in scoring.","N99QJWN2":"Experiment testing how manipulated descriptions of risk affected reviewers' scoring of mock grant reviews, finding perceived proposal risk influenced scores more than proposal strengths did.","WJIKJ2M5":"A journal peer-review study re-reviewing accepted manuscripts with new referees, finding frequent but often dissimilar criticisms, while 80% were still recommended for publication and editors rarely disagreed on acceptance decisions.","HH53MPTF":"Survey of psychology journal editors identifying shared dimensions of manuscript quality and demonstrating that structuring reviews around these dimensions can raise reviewer reliability, though quality ratings correlated only weakly with later citations.","AKUVGJ36":"Describes and evaluates a fast, applicant-involved \"community review\" system used to distribute small emergency grants across 147 pandemic-related projects, finding it quick, scalable, and resistant to bias from reviewer removal.","LJCKBTRE":"A retrospective analysis of manuscripts submitted to an Indian pediatrics journal in 2002 examines acceptance patterns and reports low inter-reviewer agreement (kappa around 0.2-0.35) among assigned reviewer pairs.","BCKFAVIX":"Analyzes long-term stability in journal rejection rates across disciplines, concluding that variation is better explained by differing levels of scholarly consensus than by space constraints on publishing.","LJ3DDLY7":"Statistical analysis of referee recommendations across five journals testing whether a single underlying \"publishability\" dimension explains reviewer judgments, finding rejection recommendations more reliable than more favorable ones.","BHZQYW9A":"Compares three scoring systems for evaluating conference abstracts, finding a proposed national evaluation system produced consistent inter-reviewer agreement comparable to established international tools.","SEPHFSGV":"A prospective study compares two simplified grant review processes, shorter panels and paired independent review, against Australia's official longer process, finding agreement just below an acceptable threshold while offering substantial potential cost savings.","ZPZW8PR3":"An analysis of conference peer review in anthrozoology finding poor to fair agreement among reviewers, but substantially improved reliability in acceptance decisions when ratings from three reviewers were averaged.","LVKA3ZWD":"A randomized controlled trial with 42 reviewers at a Norwegian funder compares individual versus general feedback reports on subsequent reviewer agreement, finding the general-feedback group showed higher eligibility agreement, while overall score agreement stayed low in both groups.","A58BN9MR":"Describes a Bayesian ranking method the Swiss National Science Foundation uses to identify proposals near the funding threshold for entry into a lottery, aiming to account for both estimated quality and uncertainty.","YXXSYUJR":"A case study testing whether a Bayesian ranking of individual reviewer scores could replace consensus meetings in Marie Skłodowska-Curie grant evaluation, finding large discrepancies with panel outcomes and recommending pre-scores be used only to triage weak proposals.","S5ZNPKAX":"Comparing scores given to the same grant proposals when submitted simultaneously to two similar review agencies, this study finds a moderate positive correlation between the two systems' scores and only moderate agreement on which proposals were fundable.","XTSCEH5T":"An empirical study of 100 reviews from the 19th Mental Measurements Yearbook, finding that quality judgments ranged uniformly from very good to very bad and that agreement between independent reviewers of the same test was positive but weak.","FPDF4D3R":"An analysis of manuscripts submitted to a psychiatry journal finding poor interrater reliability among assessors' rankings, though assessors and editors agreed well on which papers to clearly reject.","QUNM7567":"Discusses persistent problems in biomedical journal peer review, including low reviewer agreement approaching chance levels in one large journal, and proposes several possible reforms to reviewer training and incentives.","LEGPWSWB":"A year-long study at a general medicine journal found reviewer quality ratings predicted later citation impact only weakly and inconsistently, though editorial accept/reject decisions distinguished high- and low-impact articles fairly well, and agreement between individual reviewers was low.","XZ8J64TH":"Empirical evaluation of grant peer review at the Australian Research Council (2,989 proposals, 6,233 reviewers), finding low inter-rater reliability (.53), higher and less reliable ratings from researcher-nominated reviewers, and recommending fewer, better-selected reviewers and more reviews per proposal.","KWATA3RR":"Australian trial compared traditional grant peer review with a \"reader system\" where senior academics read all proposals in a subdiscipline, finding the reader system produced substantially higher reliability.","YYCAX8PN":"An empirical analysis of 4,000 ESRC grant proposals and 15,000 reviews, finding low agreement between reviewers (correlation 0.2), overly generous applicant-nominated reviewers who nonetheless influence funding, and a single negative review halving a proposal's chance of success.","WC3AX6UH":"Empirical study of 443 reviews from an interdisciplinary conference, finding poor inter-rater reliability and low construct validity across rating dimensions, with same-discipline reviewers' relevance ratings positively, and novelty ratings negatively, predicting later citations.","CM2NSUHH":"A multi-year study of two New Zealand postgraduate scholarship competitions found good to excellent agreement among assessors, with agreement strongest for academic-merit ratings and weaker on other criteria.","GEWF5FVV":"A study of manuscripts at a medical journal found readers, peer reviewers, and methods experts generally agreed on overall favorable ratings but showed poor agreement beyond chance, and experts were more critical than readers or reviewers.","SUR92HZC":"Text analysis of NIH R01 grant critiques from one university finding funded applications received more positive language, and identifying differences in critique wording associated with applicant gender.","KULA287G":"Reports moderate correlations (roughly 0.5 to 0.6) between reviewers rating manuscript quality and publishability at a psychology journal over several years, comparable to reliability seen in other journals but not high enough to trust individual paper judgments confidently.","625FMAPA":"Comparing 1990 and 1995 conference abstract reviews, this study found that increasing the number and expertise of reviewers did not improve interrater agreement, which stayed low in both years.","STUQWYY7":"An editorial note describing an unusual case where a journal's own editorial board became both subjects and evaluators of a study on manuscript review reliability, complicating the normal blind review process.","RJR9QVGF":"German-language study of reviewer agreement across five years of communication-studies conference submissions, finding disagreement on both overall judgments and individual criteria, though averaging across criteria improved consistency.","EJFLDVJU":"Retrospective comparison finding author-suggested reviewers produced reports of similar quality to other reviewers but were substantially more likely to recommend acceptance, regardless of the journal's peer-review model.","MM5CTRWM":"Analysis of thousands of reviews at a general medical journal finding reviewers' recommendations to accept or reject agreed only slightly above chance, yet editors' decisions still tracked reviewer consensus closely.","DRPY6FB5":"Analysis of 120 manuscripts at a behavioural science journal found low agreement between reviewers on specific quality criteria, though reviewer recommendations still predicted a large share of the editor's final decision.","9QWVUD3K":"Compares scoring of 160 conference abstracts by three large language models against 14 human reviewers, finding the models agreed strongly with each other but only moderately with humans, especially on subjective criteria.","LB5K9AI7":"Canadian study of a transdisciplinary funding initiative finding that pre-application networking activities were linked to funding success, while agreement among research, practice, and policy peer reviewers was low.","DK7Z8PE6":"Bibliometric analysis of 100 review requests at a plastic surgery journal examining benefits to editors, authors, and readers, and testing for publication delay or bias against non-Anglo-American submissions.","GIY7LUIX":"Pilot study develops and validates a holistic rubric for editors to score peer review quality, finding it produced a better score distribution and predicted editorial decisions better than a prior rating scale.","2H7BLLRW":"A pre-post study of 55 mentees across five cohorts found that a structured, mentor-guided peer-review training program raised review-quality scores and boosted participants' confidence and comfort with reviewing.","WH32N544":"An experimental study sending 75 journal reviewers identical manuscripts differing only in reported results, finding poor inter-rater agreement and strong bias against papers whose findings contradicted reviewers' theoretical perspectives.","T5FYTZ8E":"A pilot study of structured peer review across 220 Elsevier journals analyzed how reviewers answered nine standardized questions on 107 manuscripts, finding patterns of agreement that the authors used to refine the question set.","93G8L2PU":"Pilot of a nine-question structured review form across 23 Elsevier journals found reviewers engaged with most prompts, with agreement between reviewers highest on manuscript structure and lowest on statistical and interpretive rigor.","F68GYWKS":"Summarizing a research program with the Australian Research Council, this paper reports that grant peer reviews showed poor reliability and little systematic bias except for inflated scores from reviewers nominated by applicants, proposing an alternative \"reader system.\"","TIAHXM4B":"A study of manuscript reviews for an educational psychology journal found only moderate reliability in reviewers' overall recommendations, with agreement on specific rating subscales even lower than for overall judgments.","XHLJRX3Y":"An empirical study of 278 manuscripts submitted to the Journal of Educational Psychology, finding single-reviewer reliability of about .30 for overall recommendations and showing that specific rating dimensions, singly or combined, agreed no better because of idiosyncratic reviewer response biases.","QTGZWS97":"Commentary offering reflections on the peer review process, with no further detail available beyond the title.","F6AMXHF7":"Analysing over 10,000 reviews of grant proposals across disciplines using multilevel statistical models, the study finds no meaningful effect of principal investigator gender on review scores, a result that held across reviewer gender, discipline and country.","XFRJ7EYE":"A case study analysing interrater reliability among referees at The International Journal of Social Work Values and Ethics, finding moderate agreement that exceeds published baselines and offering eight recommendations for improving anonymous journal review.","SAJ2V3PH":"Analysis of NIH grant scoring found panel discussion produced a meaningful shift in outcome for more than one in ten applications, contrary to the assumption that preliminary reviewer scores alone determine the final result.","SZ554632":"Proposes a Bayesian statistical method for estimating inter-rater reliability while accounting for contextual factors like rater gender or experience, illustrated with grant peer review data.","85LYQU4R":"Comparing two grant-ranking methods on the same pilot-project applications, the study finds only modest agreement between them, with top-ranked projects sometimes failing to be funded depending on which reviewer pair was assigned, underscoring a large role for chance.","UFFKJYI4":"An early study measuring how consistently members of an American Psychological Association conference programme committee rated the same submitted papers.","T6YQ33G6":"A study of two-reviewer scoring of primary care conference abstracts finds moderate agreement on objective study-design criteria but much weaker agreement on more subjective elements like perceived importance.","LFVVACBY":"Studying 263 manuscripts and 207 paired reviews at a counseling psychology journal, the authors find that ratings of a paper's overall importance and methodological quality were most linked to reviewer and editorial decisions, with reviewer agreement described as modest.","7EKMP546":"Analysis of over 8,000 Austrian Science Fund proposals found reviewer agreement varied by research field, with humanities showing higher consistency than most other disciplines and biosciences showing comparatively low agreement.","B68EMUVK":"An analysis of nearly 8,500 Austrian Science Fund grant applications (1999-2009) finds no gender effect on funding decisions overall, but reports lower approval odds when reviewer panels have gender parity or a female majority.","74M956VR":"Using matched comparisons, the study finds that chemistry papers designated \"very important\" by a leading journal accumulate substantially more citations, and reach their citation peak faster, than similar non-designated papers.","H9B5WQZN":"Using data from over 8,000 proposals to the Austrian Science Fund, this study compares simulated dichotomous (fund/reject) and continuous funding-allocation systems, showing that a continuous approach based on peer review scores would approve substantially more proposals than a strict cutoff system.","J2WJBPIT":"Analyzes reviewer ratings under eLife's newer \"publish after review\" model versus its earlier gatekeeping model, finding similarly low interrater agreement and reliability on manuscript significance and rigor across both approaches.","NZSRHRPT":"An opinion piece argues that peer review needs strengthening and broader application to better ensure the quality and relevance of policy-oriented applied social research amid shifts in federal research funding priorities.","8IUUWWR8":"German-language article examines the double-blind review process of a sociology journal, questioning whether its editorial and review procedures fairly represent the discipline's full range of subfields and theoretical approaches.","ABE5N6J3":"A study of nine judges rating the quality of 36 review articles found generally consistent agreement across judges of varying research expertise on most quality criteria, based on intraclass correlation coefficients.","I358ZRGB":"A statistical analysis of about 15,000 European Southern Observatory telescope-time proposals over eight years finds that reviewer rankings agree only marginally above chance in the middle quartiles, quantifying limited reproducibility in pre-meeting review.","IV59GNII":"An experimental study with reconstructed NIH-style review panels finds that discussion increases agreement within a panel but decreases agreement between different panels reviewing the same applications, with reviewers' comparative commentary during discussion playing a key role.","85QBBTET":"An experiment replicating the NIH grant review process with independent reviewers rating the same 25 applications, finding essentially no agreement between reviewers on quality, either in scores or in how they translated critiques into numbers.","YXJVI9WJ":"Analysis of nearly 25,000 Marie Curie grant proposals under the EU's Seventh Framework Programme, finding high agreement between remote individual assessments and consensus scores, with disagreement more common for social sciences and humanities panels and lower-scored proposals.","2G4INPDL":"Retrospective study of over 75,000 Marie Curie grant applications (2007-2018) finding that fewer evaluation criteria and virtual panel meetings had little effect on funding outcomes compared with grant type or field.","IX2FS8HU":"Field experiment comparing single- and double-blind review of 530 conference submissions found moderate reliability in both, some demographic effects differing by blinding condition, and no clear reliability or validity advantage for double-blind review.","EK4WFSX9":"A cross-sectional study of 12 Elsevier journals examines whether measured review quality aligns with authors' and editors' satisfaction, finding authors were most satisfied with reviews recommending acceptance, while revision-recommending reviews scored highest on a quality instrument.","MHK2CREE":"Comparison of ChatGPT and human expert reviews of 18 single-case experimental design manuscripts found substantial agreement on quality assessment but weaker agreement on final publication recommendations.","U4JKK6XI":"Survey of authors at a major computer science conference finding they substantially overestimated their papers' acceptance chances and often disagreed with peer-review outcomes and even with their own co-authors' rankings.","MBCQSMRJ":"A mixed-methods evaluation of a Volkswagen Foundation trial in which grant applicants also served as reviewers for other proposals, comparing this distributed model with conventional panel review on efficiency, fairness, and risk of gaming.","UYD3UVJU":"A single-arm pre-post study of a semester-long peer-review training course for doctoral students finds participants' review quality, as judged by journal editors, and self-assessed reviewing skills both improved after completing the course.","YBW8PCPY":"A randomised trial at a journal finding that asking reviewers to reveal their identity to authors had no effect on review quality, publication recommendations or turnaround time, but made reviewers more likely to decline to review.","A7ERVCRX":"Analysis of manuscript and abstract review data from clinical neuroscience journals and conferences finding that agreement between independent reviewers on accept/reject decisions was little better than chance.","Q9DUXILF":"Analyzing reviews of medical conference abstracts, this study found substantial reviewer disagreement, with most score variance attributable to idiosyncratic reviewer judgment rather than the abstract itself, and recommends structured review criteria.","G6K4RHJ6":"A training study with 75 public health professors found that a short training video improved scoring accuracy and interrater reliability for grant review, for both novice and experienced reviewers.","I4ZU3XFU":"Analysing 87 manuscript decisions at a psychology journal, the authors report exact reviewer agreement in about two-thirds of paired ratings, a more favourable reliability picture than an earlier study of journal reviews had suggested.","DVDHBSTH":"Experiment with 605 experienced peer reviewers scoring simulated NIH-style grant summaries, finding generally consistent scoring but noting gender-associated differences in how proposals and reviewers rated them.","Y26RHK69":"A randomised controlled trial of training for reviewers at a general medical journal, finding that short workshop or self-taught packages produced only small, short-lived improvements in review quality and error detection.","52D3XAIF":"Comparing reviewers suggested by authors versus chosen by editors across ten biomedical journals, the study finds similar review quality between the two groups, but author-suggested reviewers were more likely to recommend accepting the manuscript.","UKWAWEDT":"An older paper on manuscript submissions to a psychology journal discusses proposals for standardized rating forms to help discriminate among the many high-quality manuscripts competing for limited publication space.","JMK4CLUJ":"Analysis of nearly 2,000 European COST funding proposals evaluated under a panel-free system found no meaningful funding disadvantage for more interdisciplinary proposals, regardless of scientific field.","QR68P67D":"Studying Stiftelsen Dam's 2020 switch to a two-stage review process, short proposal first, then a longer one for finalists, the study found it cut applicant and reviewer time substantially while producing more reliable scores than the earlier one-stage system.","HPGFXDLZ":"A machine-learning-assisted text analysis of 10,000 peer review reports across journals with varying impact factors, finding that reviews for higher-impact journals tend to be longer and more focused on methodology, with somewhat less attention to formatting or suggested fixes.","NB4IK4KH":"An analysis of the 2016 NeurIPS conference's large-scale review process examines submission and reviewer growth alongside an experiment on collecting ordinal reviewer rankings, aiming to assess review quality and inform future conference design.","M72NE68B":"A field experiment at an academic conference compared human and AI (GPT-4) reviewers, finding human reviewers were somewhat better at spotting AI-written abstracts, and that AI and human quality ratings agreed only loosely overall but converged more on identifying the top-rated abstracts.","3KXX595H":"Bootstrap analysis of fellowship competition data argues that five reviewers per application is a practical optimum, balancing reliability gains against the diminishing returns and cost of adding more reviewers.","QPIKNW58":"An observational study of over 5,000 biomedical grant evaluations compares blinded versus unblinded assessments by the same reviewers, finding scores changed in about a fifth of cases after unblinding, often correlating with applicant reputation.","WNJ55J6C":"Retrospective analysis of 1561 external peer reviewer scores for UK NIHR grant applications, finding that mean reviewer scores predict funding board decisions moderately well, with no detectable gain from using more than four reviewers and equal influence across reviewer expertise types.","ZGLDA6ZW":"An evaluation of structured abstract selection for a European plastic surgery conference, finding that participants' ratings of presentations were reliable but failing to demonstrate the validity of abstract selection or reviewer scores as indicators of presentation quality.","PEMI66JZ":"This intervention study found that introducing a multi-item rating scale with training manuals for a child psychiatry journal's manuscript reviewers modestly improved the interrater reliability of quality ratings compared to prior methods.","QA4UL4N5":"Analysing hundreds of peer reviews, the study finds reviews written during northern-hemisphere winter were harsher and reviews at open-review journals less harsh than anonymous ones, and identifies three latent reviewer types differing in tone and constructiveness.","5PN6GE4C":"Using Google's Gemini model to automatically score peer review quality, this study finds a positive association between higher review scores and a paper's subsequent citation count, with strong agreement between human and AI-assessed review scores.","KARAZVW9":"Analysis of over 11,000 Canadian Institutes of Health Research grant applications finding lower scores associated with female applicants, applied researchers, and reviewers outside the applicant's field, alongside effects from productivity and resubmission history.","V7CCKM45":"Compares rating versus ranking approaches to scoring thousands of Canadian health-research grant applications, finding ranking more reliable and less affected by reviewer characteristics, though both approaches disadvantaged early-career applicants.","7S4GVWVN":"Develops a four-step framework for checking that a grant-selection process is valid and fair, applied to a 2013 US community-service grant program, finding high reviewer agreement and reliability.","SCBRUUIQ":"Analysis of thousands of neuroscience manuscript reviews at PLOS ONE finding reviewers gave more favorable scores to authors closer to them in the co-authorship network, even though instructed to judge only scientific validity.","4RYFCKV2":"Analysing UK REF2021 data, the authors find that the same interdisciplinary journal articles receive notably different quality scores when assessed by different subject panels than when reassessed within one panel, indicating cross-field evaluation inconsistency.","8LVJQ749":"Machine-learning models trained on bibliometric and metadata features predicted UK REF2021 peer-review quality scores with reasonable accuracy in medical, physical science and economics fields, but performed poorly in social sciences, humanities and mathematics.","C9INWUTB":"Compares reviewer scores for rigour, significance, and originality in theoretical physics papers, finding only moderate agreement between reviewers even in this well-defined subject area.","5ZF6XD7M":"A single-author case study tests whether ChatGPT-4 can grade the quality of 51 of the author's own articles against UK Research Excellence Framework criteria, finding weak correlation with self-assessment and limited reliability for fine distinctions.","QCWPTK8N":"Tests whether several medium-sized and smaller large language models can rate the research quality of journal articles as well as bigger commercial models, finding comparable performance from models with at least a few billion parameters, especially when scores are averaged across repeated queries.","RRPR8N5D":"Describes development of a standardized scoring instrument for conference abstracts, finding it achieved good reviewer agreement and that higher scores predicted acceptance.","A42AQLJA":"A study of ChatGPT-assisted grading of 40 Vietnamese undergraduate research proposals found that a Tree-of-Thought prompting approach improved reliability on complex evaluation criteria but reduced it on simpler ones.","E7836ZR3":"Tests ChatGPT-4o's consistency in scoring undergraduate research proposals against human lecturer ratings, finding moderate agreement on straightforward criteria but weaker performance on more abstract, judgment-heavy evaluation areas.","GPQAWJJ9":"Applies Generalizability Theory and Many-Facet Rasch measurement to journal reviews, finding reviewer severity accounts for a substantial share of score variance and that adding more reviewers would not meaningfully fix this.","85KL54GC":"A Rasch model analysis of 710 internal ratings from 23 senior staff assessing 42 academics' outputs, mimicking the UK's Research Excellence Framework, finds most raters reliable but with meaningful variation in severity or leniency across raters.","E7IAN24K":"Commentary discussing UK funders' difficulties recruiting enough peer reviewers and criticisms of multi-stage application processes intended to reduce reviewer burden.","K5N2TIMX":"Using ChatGPT to analyse language in peer-review reports for published neuroscience papers, the study finds most reviews were favourable and polite, but reviews of papers with female first authors were less polite, and papers with female senior authors received more favourable reviews.","JX8TSFKY":"Discusses chance-corrected measures of reviewer agreement on manuscripts within a broader piece on behavioral assessment methodology (fuller context on peer review specifically not available from the excerpt).","FAWCL6KR":"Describes how a mental-health foundation developed and tested a structured rating form for grant applications, finding good internal consistency and interrater reliability among independent raters compared with the foundation's prior subjective review process.","IXVXHAXS":"A pilot study tested a structured rating form designed to make a mental-health foundation's grant review more objective, but found the tool's psychometric performance was weaker than in the original developer's study.","836GCTFJ":"Analysis of reviewer agreement for one developmental psychology journal reports higher reliability than earlier studies of other journals, and argues a particular statistical method better captures cross-journal agreement differences than the intraclass correlation.","T3A36PFH":"An older methodological paper argues that standard intraclass correlation measures understate reviewer agreement on journal manuscripts when manuscript quality variance is low, proposing an alternative statistic (Finn's r) and reanalyzing past agreement data.","HUD5RNL8":"Develops and validates a 14-item tool for rating the transparency of a journal's peer review process, finding transparency ratings correlated with independently assessed review quality and predicted resistance to a hoax paper.","RP9WIA7I":"Analysis of over 6,000 manuscript reviews at one journal from 2002-2008 found female editorial board members were less likely to recommend acceptance and had longer turnaround times than male members, though overall agreement with editors' final decisions did not differ.","E4ZVSVLA":"An older paper studying criteria used to judge journal manuscripts calls for further research into the reliability of such judgments and comparisons between editors' and authors' views of manuscript quality.","GRDE4DDS":"A multi-center pilot randomized trial of 78 neurology residents tests whether mentored peer review of standardized manuscripts improves research-methods knowledge more than unmentored review, finding no significant difference in knowledge gains between groups.","2F6KAXNL":"Examines reliability of double-blind peer review at two UK information-systems conferences, finding agreement close to or somewhat above chance levels and discussing what this implies for how reviewing is organized.","R4QME8PM":"A comparison of two residency-application file review approaches, trait-based versus element-based scoring, with ten reviewers assessing seven files each, finds the trait-based method achieved substantially higher inter-rater reliability despite requiring more review time.","E5YQG6G8":"Describes development and validation of the Review Quality Instrument, a seven-item scale for rating peer reviews, reporting strong internal consistency and good test-retest and inter-rater reliability.","WZVF5TQJ":"A field experiment with the Dutch Research Council in an early-career grant competition, finding that panellists shown only a CV and one-paragraph summary ranked proposals no differently from those shown full proposal texts."}}