C2 Answer-giving AI in practice
alive 4 attacks: 4 survived · serves Q2 · last change 29 September 2026
When learners use an AI tool that hands them finished answers or solutions while they practise, their scores on the practice work go up, but what they have learned is lower when it is later measured without the tool.
In practice
While practising, don't use an AI tool that hands over finished answers or solutions. Practice scores go up, but less is learned when it is later measured without the tool, so check learning without the tool.
Where it holds
Inside the circle: the core of the claim. Outside: every bound and carve-out in the claim file, numbered, with the attack that forced it.
forced by a narrowing attackin the original wordingwritten at creation
Evidence in the bank: Bastani et al. 2024/2025 (upper-secondary maths RCT, Turkey: answer-giving GPT Base +48 % correct answers during practice, −17 % on the closed-book exam compared with pupils who never had GenAI; kill-condition test survived 2026-09-27); Strömberg, Lei & Wu 2026 (secondary-school DiD in one Chinese county, conference paper, results may be revised: homework scores +18 %, monthly closed-book tests −20 % within six months, entrance exams −18 to −24 % with the full penalty visible after about two years); Fan et al. 2025 (university lab, second-language essay: ChatGPT raised essay scores but not knowledge gain or transfer); Barcaui 2025 (HE RCT, Brazil: unrestricted ChatGPT as study aid → lower surprise retention at 45 days vs traditional study, d = 0.68; effect held after controlling for study time; kill-condition test survived 2026-09-27).
Written when the claim was created (25 September 2026); not forced by an attack
Kreijkes et al. 2026 (England Year 10 RCT: LLM-only study of informational text → lower 3-day literal retention d=0.44, comprehension d=0.38, free recall d=0.21 vs note-taking; kill-condition test survived 2026-09-29).
Written when the claim was created (25 September 2026); not forced by an attack
Bergh et al. 2026 (HE CS experiment, arXiv: ChatGPT during coding → higher task scores but lower immediate and 48 h recall without the tool, g≈0.71–0.77; kill-condition test survived 2026-09-29).
Written when the claim was created (25 September 2026); not forced by an attack
The loss is concentrated in "homework outsourcing" use (about 80 % of AI users in Strömberg et al.); users who keep study time similar to non-users show small losses.
Written when the claim was created (25 September 2026); not forced by an attack
The claim is about learning measured without the tool; performance with the tool goes up.
Written when the claim was created (25 September 2026); not forced by an attack
Positive pooled effects in ChatGPT/chatbot meta-analyses (Deng et al. 2025, Doo & Park 2026, Wu & Yu 2024, Fan et al. 2026) are recorded as challenges, because the bank notes that most of their outcomes were not measured without the tool.
Written when the claim was created (25 September 2026); not forced by an attack
Does not transfer to: Tutor modes that withhold direct answers.
Written when the claim was created (25 September 2026); not forced by an attack
In Bastani et al. 2024 the GPT Tutor arm (same model, avoids giving answers) substantially mitigated the harm; in Oreopoulos et al. 2026 (NUMI) AI inside a mastery rule gave a modest delayed gain over the same platform without AI.
0 of 7 bounds and carve-outs were forced by a specific attack. 4 attacks hit the boundary and did not move it: #1, #2, #3, #4.
How it changed
The ledger records 4 attacks, but no earlier wording of C2 is on record, so the text cannot be replayed step by step. Only the current wording is shown.
When learners use an AI tool that hands them finished answers or solutions while they practise, their scores on the practice work go up, but what they have learned is lower when it is later measured without the tool.
Attack record
- #1 survived 27 September 2026 · Bastani et al. 2025 · target: kill-condition
A replicated upper-secondary RCT in which the unprotected answer-giving GenAI arm does not harm the closed-book result compared with control.
Strongest reading and verdict reasoning
Steelmans to: Kill-condition prong on unprotected answer-giving arms fails if GPT Base matches or beats control on the closed-book exam.
Why: Nightly-attacks 2026-09-27 04:11 Europe/Stockholm (first 04:05 run). Named result: GPT Base improved practice (+48 %) but students scored −17 % vs control on the subsequent closed-book exam; GPT Tutor mitigated the harm. Kill-condition not met. Clause tested: unprotected answer-giving arm vs control on closed-book. Claim field unchanged.
Unrestricted ChatGPT as a study aid does not lower delayed retention measured without the tool compared with traditional non-AI study in higher education.
Strongest reading and verdict reasoning
Steelmans to: Kill-condition for typical GenAI/answer-giving practice fails if HE learners using unrestricted ChatGPT match or beat traditional study on a delayed closed-book test.
Why: Nightly-attacks 2026-09-27 04:11 Europe/Stockholm. Named RCT result: after 45 days, ChatGPT-assisted undergraduates scored lower on a surprise retention test than traditional learners (57.5 % vs 68.5 %; d = 0.68); effect held after controlling for study time. Kill-condition not met. Clause tested: typical unrestricted ChatGPT use vs control on delayed retention without the tool. Claim field unchanged; Bounds note Barcaui as additional HE closed-book support.
In a secondary-school RCT, LLM-only study of continuous informational text does not lower 3-day comprehension and retention measured without the tool compared with traditional note-taking.
Strongest reading and verdict reasoning
Steelmans to: Kill-condition for typical GenAI study use fails if LLM-only matches or beats non-AI note-taking on delayed closed-book outcomes.
Why: Nightly-attacks 2026-09-29 04:16 Europe/Stockholm. Named pre-registered RCT (405 Year 10; 344 analysed): Notes beat LLM-only on literal retention (d=0.44), comprehension (d=0.38), and free recall (d=0.21) at 3 days without the tool; LLM+Notes also beat LLM-only on retention and comprehension. Students preferred the LLM and rated it more helpful while investing less effort. Kill-condition not met. Clause tested: typical LLM study aid (explanations/summaries) vs traditional study on delayed learning without the tool. Claim field unchanged; Bounds note Kreijkes kill-test survived.
When undergraduates use ChatGPT while practising introductory programming tasks, delayed cued recall measured without the tool is at least as high as after conventional non-GenAI web search, so typical answer-capable GenAI practice does not lower learning.
Strongest reading and verdict reasoning
Steelmans to: Kill-condition fails if ChatGPT-assisted coding practice matches or beats non-GenAI search on immediate and 48-hour recall without the tool.
Why: Nightly-attacks 2026-09-29 04:16 Europe/Stockholm. Named between-subjects experiment (n=55 analysed): ChatGPT raised coding scores (89% vs 69%) but lowered immediate recall (41% vs 53%; g=0.71) and 48-hour recall (39% vs 52%; g=0.77); ownership of submitted code lower (45% vs 81%). Kill-condition not met. Clause tested: answer-capable GenAI during practice vs non-GenAI search on learning without the tool (programming). Claim field unchanged; Bounds note Bergh kill-test survived (arXiv preprint).
Evidence
12 linked sources: 7 supports, 3 limits, 4 challenges (first-pass draft, pending human review). supports: Bastani et al. 2024, Fan et al. 2025, OECD 2026, Strömberg et al. 2026, Barcaui 2025, Kreijkes et al. 2026, Bergh et al. 2026 · limits: Bastani et al. 2024, Oreopoulos et al. 2026, Strömberg et al. 2026 · challenges: Deng et al. 2025, Doo & Park 2026, Fan et al. 2026, Wu & Yu 2024