Intercoder Reliability | AI in Academia - Dataspheres AI

Intercoder Reliability Fifty-three verbatim segments were independently recoded by two blind coders who saw only the codebook and the raw quotes - no memos...

Intercoder Reliability Fifty-three verbatim segments were independently recoded by two blind coders who saw only the codebook and the raw quotes - no memos, no existing codes, no findings. Agreement was computed pairwise with Cohen’s kappa. Results Pair Theme agreement Theme κ Stance agreement Stance κ Coder A vs Coder B (both blind) 90.6% (48/53) 0.885 94.3% (50/53) 0.909 Original coder vs Coder A 84.9% (45/53) 0.820 69.8% (37/53) 0.542 Original coder vs Coder B 83.0% (44/53) 0.798 66.0% (35/53) 0.480 By the Landis-Koch benchmarks, blind-to-blind agreement is almost perfect on both dimensions; original-to-blind is substantial on theme and only moderate on stance. That asymmetry is the finding. What the disagreement pattern shows The stance divergences are not noise - they are systematic, and they run one way. In eleven of the eighteen divergent segments, the original coder marked a segment critical where both blind coders independently marked it balanced , and nearly all of those segments are statistical or factual statements: the Princeton survey numbers, the 17%-to-18% cheating trend, the Gartner workforce data. The original coder - who had also written the findings - coded the rhetorical work those statistics do inside a critical argument. The blind coders applied the codebook literally: a neutral factual observation is balanced , whatever argument it serves. The codebook said “code the claim, not the speaker.” The blind coders obeyed it more faithfully than its author. This is confirmation drift, it is ordinary, and catching it is precisely what double coding is for. Theme boundaries that need sharpening Integrity vs assessment (segments 4, 8): “the honor code was already a very expensive fiction” - is that a claim about integrity systems or about assessment design? Blind coders read the honor-code subject as integrity; the original read the argumentative target as assessment. Governance vs pedagogy (segments 17, 34): inevitability claims (“there’s no stopping this”) sit between institutional response and teaching practice. Pedagogy vs equity (segments 41, 46): usage-rate statistics about student populations were read either as learning-behavior or access-distribution claims. Codebook amendment v1.1 Statistics rule: a quoted statistic or factual observation is coded balanced on stance unless the segment itself contains evaluative language. The stance of the surrounding argument belongs to the argument’s own segments. Honor-code rule: claims about honor codes, pledges, and enforcement mechanisms are integrity even when deployed inside an assessment-redesign argument; code assessment only when the segment proposes or critiques evaluation design itself. Inevitability rule: adoption-inevitability claims are governance when aimed at institutional response, pedagogy only when they prescribe classroom practice. Adjudication Final codes were set by majority vote across the three coders (a three-way split leaves the original code standing; none occurred on theme). Six theme codes and fifteen stance codes changed. Adjudicated distribution - themes: pedagogy (15), integrity (12), equity (8), labor (6), governance (4), acceleration (3), publishing (3), assessment (2). Stance: balanced (25), critical (17), pro (11). The original single-coder read overstated critical (22 → 17) - worth remembering when reading any single-coder study, including the first version of this one. Do the findings survive? Yes, with one honest downgrade. The convergence finding (both camps diagnose broken assessment), the enforcement-collapse numbers, the acceleration-is-about-research pattern, and the equity results are quote-anchored and unaffected. What softens is tone: on the adjudicated codes this corpus is better described as predominantly balanced with a strong critical current than as critical-leaning. The finding pages have been updated accordingly.