Learning, under fire

Home · Claims · C2

C2 Answer-giving AI in practice

alive Attack 1 on C2, 2026-09-27: survived (Bastani et al. 2025)Attack 2 on C2, 2026-09-27: survived (Barcaui 2025)Attack 3 on C2, 2026-09-29: survived (Kreijkes et al. 2026)Attack 4 on C2, 2026-09-29: survived (Bergh, Tag, Vassar & Renzella 2026) 4 attacks: 4 survived · serves Q2 · last change 29 September 2026

When learners use an AI tool that hands them finished answers or solutions while they practise, their scores on the practice work go up, but what they have learned is lower when it is later measured without the tool.

In practice

While practising, don't use an AI tool that hands over finished answers or solutions. Practice scores go up, but less is learned when it is later measured without the tool, so check learning without the tool.

Where it holds

Inside the circle: the core of the claim. Outside: every bound and carve-out in the claim file, numbered, with the attack that forced it.

forced by a narrowing attackin the original wordingwritten at creation

Where C2 holdsThe core of C2 sits inside the circle. 7 bounds and carve-outs sit outside the boundary, numbered; the numbered list after the diagram gives each one in full with the attack that forced it.C2 · coreAnswer-giving AI duringpractice raises practicescores but lowers learningmeasured without the tool.1Evidence in the bank: Bastaniet al. 2024/2025 …at creation2Kreijkes et al. 2026 (EnglandYear 10 RCT: LLM-only study of …at creation3Bergh et al. 2026 (HE CSexperiment, arXiv: ChatGPT …at creation4The loss is concentrated in"homework outsourcing" use …at creation5The claim is about learningmeasured without the tool; …at creation6Positive pooled effects inChatGPT/chatbot meta-analyses …at creation7Not: Tutor modes that withholddirect answers.at creation
  1. Evidence in the bank: Bastani et al. 2024/2025 (upper-secondary maths RCT, Turkey: answer-giving GPT Base +48 % correct answers during practice, −17 % on the closed-book exam compared with pupils who never had GenAI; kill-condition test survived 2026-09-27); Strömberg, Lei & Wu 2026 (secondary-school DiD in one Chinese county, conference paper, results may be revised: homework scores +18 %, monthly closed-book tests −20 % within six months, entrance exams −18 to −24 % with the full penalty visible after about two years); Fan et al. 2025 (university lab, second-language essay: ChatGPT raised essay scores but not knowledge gain or transfer); Barcaui 2025 (HE RCT, Brazil: unrestricted ChatGPT as study aid → lower surprise retention at 45 days vs traditional study, d = 0.68; effect held after controlling for study time; kill-condition test survived 2026-09-27).

    Written when the claim was created (25 September 2026); not forced by an attack

  2. Kreijkes et al. 2026 (England Year 10 RCT: LLM-only study of informational text → lower 3-day literal retention d=0.44, comprehension d=0.38, free recall d=0.21 vs note-taking; kill-condition test survived 2026-09-29).

    Written when the claim was created (25 September 2026); not forced by an attack

  3. Bergh et al. 2026 (HE CS experiment, arXiv: ChatGPT during coding → higher task scores but lower immediate and 48 h recall without the tool, g≈0.71–0.77; kill-condition test survived 2026-09-29).

    Written when the claim was created (25 September 2026); not forced by an attack

  4. The loss is concentrated in "homework outsourcing" use (about 80 % of AI users in Strömberg et al.); users who keep study time similar to non-users show small losses.

    Written when the claim was created (25 September 2026); not forced by an attack

  5. The claim is about learning measured without the tool; performance with the tool goes up.

    Written when the claim was created (25 September 2026); not forced by an attack

  6. Positive pooled effects in ChatGPT/chatbot meta-analyses (Deng et al. 2025, Doo & Park 2026, Wu & Yu 2024, Fan et al. 2026) are recorded as challenges, because the bank notes that most of their outcomes were not measured without the tool.

    Written when the claim was created (25 September 2026); not forced by an attack

  7. Does not transfer to: Tutor modes that withhold direct answers.

    Written when the claim was created (25 September 2026); not forced by an attack

    In Bastani et al. 2024 the GPT Tutor arm (same model, avoids giving answers) substantially mitigated the harm; in Oreopoulos et al. 2026 (NUMI) AI inside a mastery rule gave a modest delayed gain over the same platform without AI.

0 of 7 bounds and carve-outs were forced by a specific attack. 4 attacks hit the boundary and did not move it: #1, #2, #3, #4.

How it changed

The ledger records 4 attacks, but no earlier wording of C2 is on record, so the text cannot be replayed step by step. Only the current wording is shown.

When learners use an AI tool that hands them finished answers or solutions while they practise, their scores on the practice work go up, but what they have learned is lower when it is later measured without the tool.

Attack record

  1. #1 survived 27 September 2026 · Bastani et al. 2025 · target: kill-condition

    A replicated upper-secondary RCT in which the unprotected answer-giving GenAI arm does not harm the closed-book result compared with control.

    Strongest reading and verdict reasoning

    Steelmans to: Kill-condition prong on unprotected answer-giving arms fails if GPT Base matches or beats control on the closed-book exam.

    Why: Nightly-attacks 2026-09-27 04:11 Europe/Stockholm (first 04:05 run). Named result: GPT Base improved practice (+48 %) but students scored −17 % vs control on the subsequent closed-book exam; GPT Tutor mitigated the harm. Kill-condition not met. Clause tested: unprotected answer-giving arm vs control on closed-book. Claim field unchanged.

  2. #2 survived 27 September 2026 · Barcaui 2025 · target: kill-condition

    Unrestricted ChatGPT as a study aid does not lower delayed retention measured without the tool compared with traditional non-AI study in higher education.

    Strongest reading and verdict reasoning

    Steelmans to: Kill-condition for typical GenAI/answer-giving practice fails if HE learners using unrestricted ChatGPT match or beat traditional study on a delayed closed-book test.

    Why: Nightly-attacks 2026-09-27 04:11 Europe/Stockholm. Named RCT result: after 45 days, ChatGPT-assisted undergraduates scored lower on a surprise retention test than traditional learners (57.5 % vs 68.5 %; d = 0.68); effect held after controlling for study time. Kill-condition not met. Clause tested: typical unrestricted ChatGPT use vs control on delayed retention without the tool. Claim field unchanged; Bounds note Barcaui as additional HE closed-book support.

  3. #3 survived 29 September 2026 · Kreijkes et al. 2026 · target: kill-condition

    In a secondary-school RCT, LLM-only study of continuous informational text does not lower 3-day comprehension and retention measured without the tool compared with traditional note-taking.

    Strongest reading and verdict reasoning

    Steelmans to: Kill-condition for typical GenAI study use fails if LLM-only matches or beats non-AI note-taking on delayed closed-book outcomes.

    Why: Nightly-attacks 2026-09-29 04:16 Europe/Stockholm. Named pre-registered RCT (405 Year 10; 344 analysed): Notes beat LLM-only on literal retention (d=0.44), comprehension (d=0.38), and free recall (d=0.21) at 3 days without the tool; LLM+Notes also beat LLM-only on retention and comprehension. Students preferred the LLM and rated it more helpful while investing less effort. Kill-condition not met. Clause tested: typical LLM study aid (explanations/summaries) vs traditional study on delayed learning without the tool. Claim field unchanged; Bounds note Kreijkes kill-test survived.

  4. #4 survived 29 September 2026 · Bergh, Tag, Vassar & Renzella 2026 · target: kill-condition

    When undergraduates use ChatGPT while practising introductory programming tasks, delayed cued recall measured without the tool is at least as high as after conventional non-GenAI web search, so typical answer-capable GenAI practice does not lower learning.

    Strongest reading and verdict reasoning

    Steelmans to: Kill-condition fails if ChatGPT-assisted coding practice matches or beats non-GenAI search on immediate and 48-hour recall without the tool.

    Why: Nightly-attacks 2026-09-29 04:16 Europe/Stockholm. Named between-subjects experiment (n=55 analysed): ChatGPT raised coding scores (89% vs 69%) but lowered immediate recall (41% vs 53%; g=0.71) and 48-hour recall (39% vs 52%; g=0.77); ownership of submitted code lower (45% vs 81%). Kill-condition not met. Clause tested: answer-capable GenAI during practice vs non-GenAI search on learning without the tool (programming). Claim field unchanged; Bounds note Bergh kill-test survived (arXiv preprint).

Evidence

12 linked sources: 7 supports, 3 limits, 4 challenges (first-pass draft, pending human review). supports: Bastani et al. 2024, Fan et al. 2025, OECD 2026, Strömberg et al. 2026, Barcaui 2025, Kreijkes et al. 2026, Bergh et al. 2026 · limits: Bastani et al. 2024, Oreopoulos et al. 2026, Strömberg et al. 2026 · challenges: Deng et al. 2025, Doo & Park 2026, Fan et al. 2026, Wu & Yu 2024

Full claim record in the evidence library →