Loading…

AI in Academia: What the Record Says | AI in Academia - Dataspheres AI

A multimodal research report from the AI in Academia datasphere. Every count in this report links to the dataset that produced it; every quote carries its ...

A multimodal research report from the AI in Academia datasphere. Every count in this report links to the dataset that produced it; every quote carries its speaker and timestamp; every claim that can be checked against an outside source is checked in the appendix. Abstract We built a single coded corpus from three kinds of source: a US Senate subcommittee hearing on AI in K-12 education (1h30m of diarized audio, ten attributed speakers), five Stanford HAI AI+Education Summit panels and one UNESCO Digital Learning Week talk (linked video, analyzed in place), and the datasphere’s existing document corpus. Thirty-one segments were coded against the datasphere’s eight-theme codebook with a per-segment stance. The pattern is consistent across rooms: governance dominates the conversation (9 of 31 segments) and carries most of the caution, while the optimism concentrates in pedagogy (5) and labor (4) — and almost every cautionary claim arrives attached to a concrete policy ask. How the corpus was built The audio spine is the Senate HELP subcommittee hearing of June 2026 [1] — a US federal work in the public domain, captured from the Senate’s own stream, trimmed of its 15 minutes of pre-gavel dead air, and diarized into ten speakers whose labels were checked against the hearing’s official captions. The video sources [2–7] stayed at their origins: each is registered as a linked URL and summarized with Gemini’s native video analysis, so no content was re-hosted. Coding reuses the eight theme codes this datasphere already applies to its text corpus — integrity, assessment, pedagogy, acceleration, labor, equity, governance, publishing — with one stance per segment: pro, con, neutral, or balanced. The result is two live datasets: Coded Segments — AV & Web Corpus (31 rows: verbatim quote, source, speaker, timestamp, theme, stance, analytic memo) and Theme Sentiment Summary (per-theme stance tallies derived from the coded rows). The tables embedded below are those datasets, read live at render time. Findings 1. Governance is the center of gravity, and it skews cautious Nine of 31 segments code to governance — 1 pro, 4 con, 1 neutral, 3 balanced . The cons are specific, not vibes: a witness testified that “there are currently no high quality causal studies on the long term effects of AI on student learning, equity, or social emotional development” (Erin Mote, InnovateEDU, 40:25 in the hearing ), and the ranking member put numbers on the privacy exposure: “fifty two percent of US school districts experienced a cybersecurity incident” and a single vendor breach that “exposed the data of 62,000,000 students and 9,500,000 teachers” ( 49:31 ). That breach is real and independently documented: the PowerSchool incident of December 2024 [11][12], which ended in a federal guilty plea and a four-year sentence [12]. Even the industry witness closed on caution: “As a father, I am nervous… I would certainly encourage studies sooner rather than later” (Joshua Jones, QuantHub, 80:06 ). 2. The optimism is conditional, and it lives in pedagogy and labor The pro segments cluster where AI touches teacher workload and AI literacy: curriculum that starts “before AI… on basic data literacy” (Jones, 66:06 ), tools that “reduce administrative burden… so that they’re in real time with their students” (Sec. Marten, 77:53 ). But nearly every pro comes fenced: human in the loop, teacher training first, “vision before there’s an answer to a vendor” (Marten, 64:34 ). The Stanford panels argue the same shape from the research side — augmentation over replacement, co-design with teachers [4][5]. 3. The same fault lines appear in every room The hearing worries about surveillance and “cognitive surrender” — a term the witness borrowed from a real Wharton paper on AI and human reasoning [9][10] (Mote, 70:33 ). The panels worry about cheating and the loss of human connection [6]. UNESCO widens equity beyond the US frame to geographic and linguistic exclusion [7]. Different institutions, same three tensions: evidence versus speed, protection versus access, augmentation versus replacement. The data, live Both tables are live datasets — recode the corpus and this page, and every chart in the companion presentation , updates with it. Live cards Theme sentiment summary Coded segments (full corpus) Method and limitations Diarization is Deepgram nova-3 with an English hint; speaker names in parentheses come from self-introductions in the transcript, cross-checked against the hearing’s official captions. Timestamps refer to the trimmed recording and land within a few seconds of the spoken words. Video segments carry no timestamps: they are coded from Gemini’s summary of each full recording, which is weaker evidence than a diarized trans