Research Library

The Evidence on AI in K‑12

Evidence Base & Landscape

Featured Review · Stanford SCALE

The Evidence Base on AI in K‑12: A 2026 Review

Screened 800+ papers on AI in K‑12 to find out what’s actually been proven.

Key Findings

  • Only 20 of 800+ papers met the bar for strong causal evidence
  • The clearest available answer to whether AI tools actually work
Open paper →

Think Tank Report · CRPE

Getting Beyond the Lightbulb Stage

A CRPE brief on why AI hasn’t transformed schools yet.

Key Findings

  • Based on 50+ interviews with funders, developers, and district leaders
  • Market, vision, and infrastructure gaps are all still unresolved
Open paper →

Think Tank Report · Bellwether

Productive Struggle

Bellwether’s framework for when AI helps learning versus when it’s a shortcut.

Key Findings

  • Looks at memory, attention, motivation, and self-regulation
  • Ease isn’t automatically good — sometimes it hides a cost
Open paper →

RAND’s first national look at how AI was actually used in classrooms.

Key Findings

  • 18% of teachers used AI for teaching as of fall 2023
  • Warns AI adoption could deepen inequality without intervention
Open paper →

A meta-analysis of ed-tech’s real effect on math learning.

Key Findings

  • 74 studies, 56,886 students: small but real positive effect
  • Smaller effect size for students from low-income backgrounds
Open paper →

Qualitative Study · Friday Institute

Educators’ Perspectives on Generative AI in K‑12

NC State’s Friday Institute on how educators are actually experiencing AI.

Key Findings

  • Compares educator experience against 12 states’ official guidance
  • Flags equity, workload, and privacy as top concerns
Open paper →

Mixed Methods · Research in Learning Tech.

Perceptions and Preparedness of K‑12 Educators

What’s standing in the way of educators feeling ready for AI.

Key Findings

  • Links AI familiarity directly to perceived readiness
  • Insufficient PD is the most-cited barrier
Open paper →

Report · University of Sydney

Young people, learning, and generative AI: a rapid review for PreK-12

A learning-sciences read on what generative AI actually does to young people’s learning

Key Findings

  • Affective gains are common, but weak indicators of durable learning
  • Effects on higher-order reasoning, transfer, and self-regulated learning remain uneven

No named publisher

National Survey · Walton Family Foundation & Gallup

Voices of Gen Z: Year 4 annual survey

Gallup’s fourth annual national read on how Gen Z students feel about school

Key Findings

  • Just 48% of Gen Z students look forward to school most days
  • 33% do not feel they belong; four in ten feel unprepared

Interviews with 36 school and district leaders about how AI changes educators’ jobs

Key Findings

  • Leaders sorted AI uses into routine, instructional, operational, and decision-making categories
  • Hands-on experience using AI predicted more balanced, better-informed leader views

Interviews conducted in late 2023

Open paper →

Survey of 89 rural Idaho K‑12 teachers on whether they are ready for AI

Key Findings

  • Fewer than half of the 89 teachers used generative AI in practice
  • Users applied it to lesson prep and admin, not live teaching

Single-state sample

Open paper →

Meta‑Analysis · Humanities and Social Sciences Communications

The effect of ChatGPT on learning performance, perception, and higher-order thinking

Pooled 51 experiments to size up what ChatGPT does to learning outcomes

Key Findings

  • 51 studies: large effect on learning performance, Hedges’ g = 0.867
  • Smaller gain for higher-order thinking (g = 0.457); 4-8 weeks most stable
Open paper →

Cognitive Offloading & Cognitive Cost

A high school field experiment on AI tutoring with and without guardrails.

Key Findings

  • AI tutoring boosted practice scores, then hurt exam scores once removed
  • Withholding direct answers erased the harm entirely
Open paper →

A randomized experiment inside real secondary-school classrooms.

Key Findings

  • Tests how note-taking habits interact with LLM use
  • Conducted in actual classrooms, not a lab
Open paper →

Panel Study · Working Paper

The generative AI learning penalty: Evidence from Chinese secondary education

Thirty months of data on 26,811 secondary students before and after they adopted AI tools.

Key Findings

  • Homework scores rose 18% while closed-book exam scores fell 20% within six months
  • Entrance exam scores dropped 18% and 24%, the full penalty taking about two years

Unpublished manuscript, not peer reviewed

A ten-year panel showing study time collapsed on exactly the math problems AI can do.

Key Findings

  • High school study time on AI-susceptible math problems fell 31.3%
  • Proctored retention items showed a 25% cumulative decline in odds of answering correctly

Preprint, not peer reviewed

Open paper →

MIT used EEG to watch what happens in the brain during AI-assisted essay writing.

Key Findings

  • Brain connectivity fell as tool support rose; LLM users showed the weakest coupling
  • LLM users reported low essay ownership and could barely quote their own writing

Preprint, not peer reviewed; 54 participants

Open paper →

The randomised study that named metacognitive laziness, comparing ChatGPT against a human expert.

Key Findings

  • The ChatGPT group improved essay scores most but gained no knowledge or transfer
  • Self-regulated learning behaviours differed sharply while intrinsic motivation did not

Higher ed sample

Open paper →

Randomised trials with 1,222 people testing what is left once the AI help is withdrawn.

Key Findings

  • AI users performed significantly worse without the tool and gave up sooner
  • The effect appeared after roughly ten minutes of AI interaction

Preprint, not peer reviewed

Open paper →

An experiment separating students who had AI explain concepts from those who had it write.

Key Findings

  • AI access raised immediate test scores 0.27 standard deviations, holding a week later
  • Automation users lost their quality advantage as soon as the tool was removed

Preprint, not peer reviewed; undergraduate sample

Open paper →

An experiment on whether AI helps or hurts critical thinking depending on when you get it.

Key Findings

  • Early LLM access helped under time pressure but hurt when time was sufficient
  • Starting the task independently produced the opposite pattern
Open paper →

A 666-person study linking heavy AI use to weaker critical thinking through cognitive offloading.

Key Findings

  • Frequent AI use correlated negatively with critical thinking, mediated by cognitive offloading
  • Younger participants showed the most AI dependence and the lowest critical thinking scores

Correlational, not causal

Open paper →

White Paper · EDSAFE AI Alliance

Future-proofing human flourishing: The case for a learning sciences benchmark for AI

Proposes a learning-sciences benchmark for judging whether an AI tool builds or erodes competence.

Key Findings

  • Names a ’Zone of No Development’ and a ’Crutch Effect’ from AI overreliance
  • Proposes three pillars and argues voluntary industry standards are not enough

Empirical Study · Computers in Human Behavior

AI-mediated teaching in K‑12 classrooms

Interviews with 23 high school teachers about what AI quietly takes over in their work.

Key Findings

  • AI cut workload but displaced diagnostic reasoning, sequencing, and evaluative judgment
  • Teacher agency persisted but became conditional on institutional pressure and algorithmic opacity

Teacher interviews, not student outcomes

Open paper →

Tutoring & Personalized Learning

RCT · Stanford

Access Is Not Enough

Two RCTs on whether students actually use AI tutors when given the chance.

Key Findings

  • Nearly half of students never logged into the AI tutor at all
  • Pairing it with a human tutor helped, but not enough to move reading scores
Open paper →

Systematic Review · npj Science of Learning

AI‑Driven Intelligent Tutoring Systems in K‑12

A Nature-family review focused specifically on K‑12 tutoring systems.

Key Findings

  • Covers intelligent tutoring systems built for K‑12 classrooms
  • Maps out where the evidence is strong versus still thin
Open paper →

Meta‑Analysis · Computers & Education

Does ChatGPT Enhance Student Learning?

A meta-analysis of experimental studies, not just surveys.

Key Findings

  • Positive effects found on performance and higher-order thinking
  • Synthesizes results across many independent experiments
Open paper →

Field Study · Preprint

Learning to Prompt

An adaptive tutoring system tested with real high schoolers.

Key Findings

  • Tested with 359 real students across 656 conversations
  • Adaptive prompting outperformed static tutoring scripts
Open paper →

Bias Study · Stanford

Marked Pedagogies

Tested 4 major LLMs on real 8th-grade essays for bias.

Key Findings

  • Feedback shifted based on a student’s race, gender, and learning needs
  • Shows “personalized” AI feedback isn’t neutral
Open paper →

A design case, not an outcomes study, for slower AI tutors.

Key Findings

  • Argues AI tutors should resist giving the fastest answer
  • Prioritizes durable learning over the feeling of instant progress
Open paper →

RCT · Google LearnLM

Teaching with Gemini: measuring Guided Learning’s impact on math progress in Sierra Leone

Preregistered trial of Gemini’s Guided Learning in 48 junior secondary math classrooms

Key Findings

  • 1,763 students: intent-to-treat effect of +0.258 standard deviations on math
  • Full 12-hour dosage reached +0.380 SD, roughly 1.2-1.7 years of progress

Industry-authored; non-US setting

First randomized trial of an AI coach feeding live guidance to math tutors

Key Findings

  • 900 tutors, 1,800 students: 4 percentage point gain in topic mastery
  • Gains hit 9 points for students of lower-rated tutors, $20 per tutor

Preprint, not peer reviewed

Open paper →

Two-year randomized trial of Khanmigo in 18 Tennessee middle schools’ remedial math

Key Findings

  • Gains of 1.3 national percentile ranks per term, about 0.06-0.08 SD yearly
  • 96% tried Khanmigo but the median student messaged it a third of days

Working paper, not peer reviewed

Open paper →

Tested supervised AI tutoring against human tutors with 165 UK secondary students

Key Findings

  • Tutors approved 76.4% of LearnLM’s drafted messages with zero or minimal edits
  • 66.2% versus 60.7% solved novel problems, a 5.5 point edge over humans

Industry-authored; small sample

Open paper →

Empirical Study · Khan Academy

Cognitive engagement in GenAI tutor conversations: at-scale measurement and impact on learning

Scored 200,000 Khanmigo tutoring chats for how hard students were actually thinking

Key Findings

  • About 9,000 students and 200,000 tutoring threads across six U.S. districts
  • Higher cognitive engagement predicted better performance on the next practice item

Works-in-progress paper

Report · Stanford SCALE & NSSA

Research on AI tutoring: what the evidence says and what it means

Sorts AI tutoring into four models and says which ones the evidence supports

Key Findings

  • One study found 40-47% of students never logged into the AI tutor
  • Strongest results came from AI built for tutors, not for students

Summarizes a separate brief; no named publisher

Assessment, Integrity & Critical Thinking

Framing Piece · Assessment & Eval. in HE

The Wicked Problem of AI and Assessment

Argues AI exposed a problem that already existed in assessment.

Key Findings

  • Traditional assessment was never a clean measure of understanding
  • A foundational, widely-cited framing piece
Open paper →

Framing Piece · Assessment & Eval. in HE

Black Box Assessment

Makes the case for assessing visible, documented thinking.

Key Findings

  • Pushes back on trying to “catch” AI use after the fact
  • Proposes assessing process, not just final products
Open paper →

Foundational · Preprint, 2023

Chatting and Cheating

One of the earliest, most-cited looks at ChatGPT-era integrity.

Key Findings

  • 244+ citations since its 2023 preprint release
  • A foundational reference point for the whole field
Open paper →

A landmark, widely-cited bias finding in AI detection tools.

Key Findings

  • Non-native English writers flagged far more often as “AI-written”
  • Raises real fairness concerns for detection-based policies
Open paper →

Empirical Test · Intl J. Educational Integrity

Testing of Detection Tools for AI-Generated Text

A head-to-head test of popular AI-detection tools.

Key Findings

  • Tested multiple detection tools side by side
  • Found them unreliable across the board
Open paper →

Systematic Review · Soc. Sci. & Humanities Open

Reassessing Academic Integrity in the Age of AI

A 2025 systematic review of where the integrity debate stands now.

Key Findings

  • Synthesizes findings across dozens of studies
  • Maps how the conversation has shifted since 2023
Open paper →

Framework · Educational Researcher

The AI3 Model

A cross-national framework for assessment innovation.

Key Findings

  • Spans K‑12 and higher education contexts
  • Published in AERA’s flagship journal
Open paper →

Synthesis · Preprint

The Effortless Trap

Argues placement, not permission, determines AI’s effect on learning.

Key Findings

  • An unguarded AI helper left students ~17% worse on an unaided exam
  • A well-engineered tutor roughly doubled learning instead
Open paper →

Survey Study (Contested) · Societies

AI Tools in Society

Links frequent AI use to lower critical thinking scores.

Key Findings

  • A widely-debated survey of 666 participants
  • Read alongside its published correction
Open paper →

Asks whether AI writing feedback teaches writing or just acts as a marking machine

Key Findings

  • Automated writing evaluation shows qualified gains: fewer errors, more drafts, greater autonomy
  • Current tools still lack the contextual awareness and higher-order judgment teachers bring

Language-learning focus

Open paper →

Systematic Review · Khazar Journal

AI detection tools: a systematic review of empirical evidence and implications for education

Pulled together 25 studies on whether AI text detectors are reliable enough to use

Key Findings

  • 25 studies, 18 in classrooms: accuracy unstable across tools, genres, languages
  • Unsuitable for high-stakes integrity decisions; authors urge assessment redesign instead

Empirical Study · Int’l Journal of Ed Tech in Higher Education

Simple techniques to bypass GenAI text detectors

Tested whether easy tweaks let AI-written text slip past six major detectors

Key Findings

  • Detector accuracy dropped 17.4% once simple bypass techniques were applied
  • 805 texts tested; authors say detectors cannot justify integrity violations

Higher ed focus

Open paper →

Reports on humanizers and autotypers that help students slip past AI detection software

Key Findings

  • Humanizers rewrite AI text; autotypers fake human typing, typos and revisions
  • Some firms sell both detection tools and the apps that evade them

Journalism, not research

Open paper →

Reviewed 67 studies on what ChatGPT does to critical and creative thinking

Key Findings

  • 67 studies: gains appear only in inquiry-oriented, scaffolded instructional designs
  • Unstructured use favored creativity over critical thinking, sometimes declining in both

Higher ed sample

Open paper →

Built and tested video modules teaching K‑12 students six kinds of academic dishonesty

Key Findings

  • 1,181 students across 10 schools improved on pre- and post-tests
  • Junior high students benefited most; modules cover AI misuse and contract cheating

Non-US setting

Open paper →

Pitted five language models against 37 teachers scoring real student essays on ten criteria

Key Findings

  • o1 correlated .74 with teachers and reached .80 internal consistency
  • All models over-rated essays; open-source models correlated weakly with human raters

20 German essays, small set

Open paper →

Agentic AI, Literacy & Policy

Foundational · Computers & Ed: AI

Conceptualizing AI Literacy

The foundational framework most later AI-literacy work builds on.

Key Findings

  • Defines AI literacy as technical knowledge plus critical evaluation
  • Cited across most later AI-literacy frameworks
Open paper →

Framework · Preprint

Addressing the Reality Gap

Three tensions schools face in adopting more autonomous AI.

Key Findings

  • Names feasibility, adaptation speed, and mission alignment
  • A preprint chapter for Springer’s forthcoming AI-in-education handbook
Open paper →

OpenAI’s own usage data on agentic AI adoption.

Key Findings

  • Agentic AI usage grew more than 5x in six months
  • Adoption is spreading well beyond software developers
Open paper →

Perspective · Frontiers in Education

The Cognitive Mirror

A conceptual framework, not a data study, on AI and metacognition.

Key Findings

  • AI’s value may be reflecting a learner’s thinking back to them
  • Proposes building self-regulation, not just giving answers
Open paper →

The finalized OECD framework defining what AI literacy means for primary and secondary students

Key Findings

  • 22 competences across four domains: Engage, Create, Manage, and Shape AI
  • Final version strengthens metacognition, student agency, and healthy skepticism toward AI
Open paper →

Systematic Review · Computers and Education: AI

Unleashing human potential: an AI competency framework for K‑12 education

Scoping review of 54 studies behind a three-part AI competency model for K‑12

Key Findings

  • 54 studies converge on AI knowledge, practical skills, and ethical awareness
  • Existing frameworks underemphasize human values; adds an ’Unleashing’ developmental component
Open paper →

Position Paper · Khalid, Islam & Ajmal

Reimagining curriculum knowledge through agentic AI systems

Argues agentic AI should reshape curriculum from fixed content into a living process

Key Findings

  • Proposes ’AI-augmented curriculum agency’ pairing human intentionality with machine intelligence
  • Casts AI as a knowledge partner rather than a classroom tool

Conceptual, no empirical data

Systematic Review · Applied Resilience and Sustainability

Agentic AI-driven tutoring: a multi-agent cognitive architecture for personalized adaptive learning

Reviews what multi-agent AI tutoring architectures promise and where they still fall short

Key Findings

  • Multi-agent systems improve behavior modeling, real-time feedback, and metacognitive support
  • Privacy, algorithmic transparency, scalability, and missing evaluation standards remain unresolved

Low-profile venue

No papers match that search — try a different term or clear the filter.

A NoteThis reflects papers I’m actively reading, not a comprehensive literature review. I’ll keep updating this as I read more.