Method & Audit Trail | AI in Academia - Dataspheres AI

Method & Audit Trail This page exists so any claim in the findings can be walked back to the recording that produced it. That walk-back is the whole po...

Method & Audit Trail This page exists so any claim in the findings can be walked back to the recording that produced it. That walk-back is the whole point of keeping a study in one place. The chain Source discovery → live URL verification → adversarial claim verification → transcription → segment coding → blind double-coding and adjudication → findings. Nothing skips a step, and nothing enters the matrix from memory. Transcribed corpus (in the Library) Recording Stance Length Transcript Accelerating Scientific Discovery with AI - Sir Demis Hassabis (Cambridge) pro ~64 min 11,987 words Debate: AI Will Have a Serious Negative Impact on Academia (Academia Europaea) contrast ~49 min 4,886 words AI Didn't Break Education. It Exposed The Lie (House of El: AI) critical ~19 min 3,287 words Practical AI for Instructors and Students Part 1 (Wharton School) pro ~10 min 1,986 words 20 PhDs In the Time of One (Bloomberg Television) balanced ~10 min 1,972 words AI's Role in Higher Ed - Jennifer Pintar (TEDxYoungstown) balanced ~10 min 1,491 words AI and Academic Integrity - What Students Need to Know (Jessica Bernards) balanced ~5 min 742 words Total: fourteen transcripts, about 78,000 words across roughly 8 hours. (The second wave of seven was sampled deliberately to fill gaps the first pass exposed: student perspectives, equity discourse, teaching-centre voices.) Transcripts are timestamped at roughly 30-second intervals so a coder can jump from any coded segment to the moment in the recording. The three matrices Media Corpus (42 rows) - every verified source with item-level themes and stance. Coded Claims Matrix (22 rows) - claims with an explicit evidence basis, from sources not fully transcribed. Coded Segments (Verbatim) (53 rows) - quote, speaker, timestamp, theme, stance, and an analytic memo, drawn from the fourteen transcribed recordings, blind double-coded (κ = 0.885 theme / 0.909 stance) and adjudicated - see Intercoder Reliability . Verification standards Every URL fetched live before inclusion. Four candidate sources were dropped: one dead video and three pages that could not be verified even in a real browser session. Flagship claims passed three-vote adversarial verification - 24 of 25 confirmed. The one refutation (a misattributed podcast title) was corrected before anything was published. Quotes are verbatim from published captions. Where captions garble a word, the segment was either excluded or the surrounding sentence retained for context. Known weaknesses, stated plainly Intercoder reliability now exists (two blind coders + adjudication). The learner exercise is to run a third pass against codebook v1.1 and compare. Caption-derived text carries ASR error; the debate transcript has positional speaker labels only. Source balance improved in wave two: 53 segments now span 14 sources; the largest single source contributes 9. English-language, North-American-skewed, public discourse rather than peer-reviewed literature. Reuse and licensing Transcripts are held here for research and teaching analysis with source and URL attached to every document. Audio and video remain at their original homes - we link, we do not rehost. Fair use is a case-by-case judgment weighed on four statutory factors, and educational status alone does not authorize redistribution.