← What we believe Evidence library

Home · C2

C2 alive

When learners use an AI tool that hands them finished answers or solutions while they practise, their scores on the practice work go up, but what they have learned is lower when it is later measured without the tool.

Status
alive
Last change
2026-09-29
Serves
Q2 (How do we achieve optimal learning when AI and screens are everywhere?)
Ages
grundskola-hog, gymnasium, vuxen-livslangt (evidence: upper-secondary maths in Turkey, grades 7–12 in one Chinese county, England Year 10 reading, university lab, HE RCT with delayed closed-book, and HE introductory programming); not yet shown for younger pupils or in Swedish classrooms
Depends on
C2 depends-on C1 ("measured without the tool" is C1's accessibility test applied to AI: learning counts only if the target can be reached later without new presentation, and an answer-giving tool is a form of re-presentation).
Bounds
Evidence in the bank: Bastani et al. 2024/2025 (upper-secondary maths RCT, Turkey: answer-giving GPT Base +48 % correct answers during practice, −17 % on the closed-book exam compared with pupils who never had GenAI; kill-condition test survived 2026-09-27); Strömberg, Lei & Wu 2026 (secondary-school DiD in one Chinese county, conference paper, results may be revised: homework scores +18 %, monthly closed-book tests −20 % within six months, entrance exams −18 to −24 % with the full penalty visible after about two years); Fan et al. 2025 (university lab, second-language essay: ChatGPT raised essay scores but not knowledge gain or transfer); Barcaui 2025 (HE RCT, Brazil: unrestricted ChatGPT as study aid → lower surprise retention at 45 days vs traditional study, d = 0.68; effect held after controlling for study time; kill-condition test survived 2026-09-27). Kreijkes et al. 2026 (England Year 10 RCT: LLM-only study of informational text → lower 3-day literal retention d=0.44, comprehension d=0.38, free recall d=0.21 vs note-taking; kill-condition test survived 2026-09-29). Bergh et al. 2026 (HE CS experiment, arXiv: ChatGPT during coding → higher task scores but lower immediate and 48 h recall without the tool, g≈0.71–0.77; kill-condition test survived 2026-09-29). The loss is concentrated in "homework outsourcing" use (about 80 % of AI users in Strömberg et al.); users who keep study time similar to non-users show small losses. The claim is about learning measured without the tool; performance with the tool goes up. Positive pooled effects in ChatGPT/chatbot meta-analyses (Deng et al. 2025, Doo & Park 2026, Wu & Yu 2024, Fan et al. 2026) are recorded as challenges, because the bank notes that most of their outcomes were not measured without the tool.
Does-not-transfer-to
Tutor modes that withhold direct answers. In Bastani et al. 2024 the GPT Tutor arm (same model, avoids giving answers) substantially mitigated the harm; in Oreopoulos et al. 2026 (NUMI) AI inside a mastery rule gave a modest delayed gain over the same platform without AI.
What would kill this
A replicated school or university study, measuring without the tool, in which typical GenAI use during practice does not lower retention compared with a control group and the "homework outsourcing" pattern shows no learning loss; or a replicated upper-secondary RCT in which the unprotected answer-giving arm does not harm the closed-book result compared with control. Bank wording (Swedish, verbatim): "Replikerad K–12/HE-studie med closed-book där typisk GenAI-användning inte skadar retention relativt kontroll, och outsourcing-mönstret saknar learning loss." and "Replikerad gymnasie-RCT där oskyddad GPT Base inte skadar closed-book relativt kontroll och Tutor-armen inte längre skiljer sig i pedagogik från Base."
Tutor consequence
none yet

Attack record

4 attacks: 4 survived

2026-09-27 survived

Attack
A replicated upper-secondary RCT in which the unprotected answer-giving GenAI arm does not harm the closed-book result compared with control.
Target
kill-condition
Source
Bastani et al. 2025
Steelmans to
Kill-condition prong on unprotected answer-giving arms fails if GPT Base matches or beats control on the closed-book exam.
Why
Nightly-attacks 2026-09-27 04:11 Europe/Stockholm (first 04:05 run). Named result: GPT Base improved practice (+48 %) but students scored −17 % vs control on the subsequent closed-book exam; GPT Tutor mitigated the harm. Kill-condition not met. Clause tested: unprotected answer-giving arm vs control on closed-book. Claim field unchanged.
Rewrite
none

2026-09-27 survived

Attack
Unrestricted ChatGPT as a study aid does not lower delayed retention measured without the tool compared with traditional non-AI study in higher education.
Target
kill-condition
Source
Barcaui 2025
Steelmans to
Kill-condition for typical GenAI/answer-giving practice fails if HE learners using unrestricted ChatGPT match or beat traditional study on a delayed closed-book test.
Why
Nightly-attacks 2026-09-27 04:11 Europe/Stockholm. Named RCT result: after 45 days, ChatGPT-assisted undergraduates scored lower on a surprise retention test than traditional learners (57.5 % vs 68.5 %; d = 0.68); effect held after controlling for study time. Kill-condition not met. Clause tested: typical unrestricted ChatGPT use vs control on delayed retention without the tool. Claim field unchanged; Bounds note Barcaui as additional HE closed-book support.
Rewrite
none

2026-09-29 survived

Attack
In a secondary-school RCT, LLM-only study of continuous informational text does not lower 3-day comprehension and retention measured without the tool compared with traditional note-taking.
Target
kill-condition
Source
Kreijkes et al. 2026 | DOI 10.1016/j.compedu.2025.105514
Steelmans to
Kill-condition for typical GenAI study use fails if LLM-only matches or beats non-AI note-taking on delayed closed-book outcomes.
Why
Nightly-attacks 2026-09-29 04:16 Europe/Stockholm. Named pre-registered RCT (405 Year 10; 344 analysed): Notes beat LLM-only on literal retention (d=0.44), comprehension (d=0.38), and free recall (d=0.21) at 3 days without the tool; LLM+Notes also beat LLM-only on retention and comprehension. Students preferred the LLM and rated it more helpful while investing less effort. Kill-condition not met. Clause tested: typical LLM study aid (explanations/summaries) vs traditional study on delayed learning without the tool. Claim field unchanged; Bounds note Kreijkes kill-test survived.
Rewrite
none

2026-09-29 survived

Attack
When undergraduates use ChatGPT while practising introductory programming tasks, delayed cued recall measured without the tool is at least as high as after conventional non-GenAI web search, so typical answer-capable GenAI practice does not lower learning.
Target
kill-condition
Source
Bergh, Tag, Vassar & Renzella 2026 | DOI 10.48550/arXiv.2609.21194
Steelmans to
Kill-condition fails if ChatGPT-assisted coding practice matches or beats non-GenAI search on immediate and 48-hour recall without the tool.
Why
Nightly-attacks 2026-09-29 04:16 Europe/Stockholm. Named between-subjects experiment (n=55 analysed): ChatGPT raised coding scores (89% vs 69%) but lowered immediate recall (41% vs 53%; g=0.71) and 48-hour recall (39% vs 52%; g=0.77); ownership of submitted code lower (45% vs 81%). Kill-condition not met. Clause tested: answer-capable GenAI during practice vs non-GenAI search on learning without the tool (programming). Claim field unchanged; Bounds note Bergh kill-test survived (arXiv preprint).
Rewrite
none

Linked sources

12 linked sources (7 supports, 3 limits, 4 challenges)

Supports (7)

SourceBasis
AI-010 · ChatGPT as a cognitive crutch: Evidence from a randomized controlled trial on knowledge retention (2025)HE RCT: unrestricted ChatGPT study aid → lower Day-45 surprise retention vs traditional (d = 0.68); learning measured without the tool.
AI-003 · Generative AI without guardrails can harm learning: Evidence from high school mathematics (2024 (SSRN preprint); PNAS 2025 (publicerad 2025-06-25))Upper-secondary maths RCT: answer-giving GPT Base +48 % in practice, −17 % on the closed-book exam.
AI-017 · Your Programming Students' Cognition with ChatGPT: Higher Performance, Lower Retention, and Reduced Ownership (2026)HE experiment: ChatGPT during coding → higher task scores but lower immediate and 48 h recall without the tool; kill-condition test survived 2026-09-29 (arXiv preprint).
AI-008 · Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance (2025 (online Dec 2024))University lab RCT: ChatGPT raised essay scores but not knowledge gain or transfer; fewer metacognitive processes.
AI-016 · Effects of LLM use and note-taking on reading comprehension and memory: A randomised experiment in secondary schools (2026)Secondary RCT: LLM-only study → lower 3-day retention/comprehension vs note-taking; learning measured without the tool; kill-condition test survived 2026-09-29.
AI-001 · OECD Digital Education Outlook 2026: Exploring Effective Uses of Generative AI in Education (2026)OECD report: general GenAI tools can raise task quality without raising knowledge/skill acquisition (performance is not learning). Mixed evidence levels.
AI-009 · The Generative AI Learning Penalty: Evidence from Chinese Secondary Education (2026 (CEPR DP21577 / NBER conference paper, juni 2026))Chinese secondary DiD: homework +18 %, monthly closed-book tests −20 % within six months; loss concentrated in homework outsourcing.

Limits (3)

SourceBasis
AI-003 · Generative AI without guardrails can harm learning: Evidence from high school mathematics (2024 (SSRN preprint); PNAS 2025 (publicerad 2025-06-25))The GPT Tutor arm (withholds direct answers) substantially mitigated the harm.
AI-014 · Making AI Tutoring Productive: Evidence from a Mastery-Based Math Practice Experiment (2026)Guard-railed AI inside a mastery rule gave a modest delayed gain over the same platform without AI; not answer-giving, so outside C2.
AI-009 · The Generative AI Learning Penalty: Evidence from Chinese Secondary Education (2026 (CEPR DP21577 / NBER conference paper, juni 2026))Users who keep study time similar to non-users show small losses.

Challenges (4)

SourceBasis
AI-005 · Does ChatGPT enhance student learning? A systematic review and meta-analysis of experimental studies (2025 (online Dec 2024))Positive pooled effect of ChatGPT; the bank notes that without closed-book outcomes this is performance, not learning. Open test for C2.
AI-004 · A Meta-Analysis of ChatGPT’s Influence on Learning Achievement (2026 (IRRODL Vol. 27, No. 1, februari 2026))Positive pooled effect on achievement; closed-book follow-up absent or mixed. Open test for C2.
AI-002 · Exploring the effect of GenAI on learning outcomes in higher education: a three-level meta-analysis (2026 (publicerad 15 maj 2026))Higher-education meta g ≈ 0.50; not strictly closed-book. Open test for C2.
AI-007 · Do AI chatbots improve students learning outcomes? Evidence from a meta-analysis (2024 (online 2023))Positive chatbot effects, stronger in higher education than K–12; closed-book rare. Open test for C2.

Review state: first-pass draft, pending human review.