C2 alive
When learners use an AI tool that hands them finished answers or solutions while they practise, their scores on the practice work go up, but what they have learned is lower when it is later measured without the tool.
Bounds: Evidence in the bank: Bastani et al. 2024/2025 (upper-secondary maths RCT, Turkey: answer-giving GPT Base +48 % correct answers during practice, −17 % on the closed-book exam compared with pupils who never had GenAI; kill-condition test survived 2026-09-27); Strömberg, Lei & Wu 2026 (secondary-school DiD in one Chinese county, conference paper, results may be revised: homework scores +18 %, monthly closed-book tests −20 % within six months, entrance exams −18 to −24 % with the full penalty visible after about two years); Fan et al. 2025 (university lab, second-language essay: ChatGPT raised essay scores but not knowledge gain or transfer); Barcaui 2025 (HE RCT, Brazil: unrestricted ChatGPT as study aid → lower surprise retention at 45 days vs traditional study, d = 0.68; effect held after controlling for study time; kill-condition test survived 2026-09-27). Kreijkes et al. 2026 (England Year 10 RCT: LLM-only study of informational text → lower 3-day literal retention d=0.44, comprehension d=0.38, free recall d=0.21 vs note-taking; kill-condition test survived 2026-09-29). Bergh et al. 2026 (HE CS experiment, arXiv: ChatGPT during coding → higher task scores but lower immediate and 48 h recall without the tool, g≈0.71–0.77; kill-condition test survived 2026-09-29). The loss is concentrated in "homework outsourcing" use (about 80 % of AI users in Strömberg et al.); users who keep study time similar to non-users show small losses. The claim is about learning measured without the tool; performance with the tool goes up. Positive pooled effects in ChatGPT/chatbot meta-analyses (Deng et al. 2025, Doo & Park 2026, Wu & Yu 2024, Fan et al. 2026) are recorded as challenges, because the bank notes that most of their outcomes were not measured without the tool.
Attack record: 4 attacks: 4 survived · Evidence links: 12 linked sources (7 supports, 3 limits, 4 challenges)