← What we believe Evidence library

Evidence library

One project that merges a Swedish evidence library on learning (the Kunskapsbank) with a small set of claims that are deliberately attacked (the learning graph). Everything serves two questions:

  1. Q1. What is optimal learning?
  2. Q2. How do we achieve it when AI and screens are everywhere?

Q1 covers cognition, age and maturity, and teaching, whatever the medium. Q2 covers Swedish school in practice.

Claims serving Q1

C0 weakened

After the material has been encoded, practicing retrieval produces more durable retention than spending the same time restudying when the criterion test is delayed, items are correctly retrieved repeatedly in practice with corrective feedback (a single correct retrieval is not enough), and no-feedback multiple-choice practice is excluded. At an immediate (~5 min) test, equal-time restudy can be superior. Once initial learning is equated, the claim does not assert that this advantage grows as final-test delay grows. The claim does not hold when encoding never succeeded, when the practice test is a poor match to what must later be remembered, when practice accuracy is relatively low and no feedback is given, or when “testing” is high-stakes performance rather than practice. The claim does not transfer to novel, meaningless abstract visual materials that lack preexisting semantic scaffolding.

Serves: Q1 (What is optimal learning?)
Ages: not restricted by the claim text; linked bank evidence is mostly adults/university students plus age-spanning reviews (overgripande)

Bounds: Main contrast is for delayed criterion tests, not immediate (~5 min) tests. Delayed contrast requires repeated successful retrievals with corrective feedback (not one-shot success); no-feedback multiple-choice practice is excluded. Delay-growth of the retrieval–restudy advantage is not claimed once initial learning is equated. High semantic relatedness of cue–target materials reduces relative efficacy of retrieval vs restudy (~26%) but does not erase the advantage (Gupta, Pan & Rickard 2026; kill-condition test survived 2026-09-28). Novel meaningless abstract visual materials lacking preexisting semantic scaffolding: McCarter, Huber & Cowell 2025 (six experiments) found restudy ≥ retrieval (sometimes significantly better) despite feedback and delays — outside the claim (cf. Ferreira & Wimber 2023 on meaningful vs low-semantic visual materials).

Attack record: 12 attacks: 6 survived, 6 narrowed · Evidence links: 12 linked sources (7 supports, 5 limits, 2 challenges)

C1 weakened

Encoding has succeeded when the learner can later access the target under cue conditions that match the practice criterion without new presentation of the material — accessibility under that cue set, not mere study exposure or storage-without-access. Under corrective feedback, that accessibility criterion is stricter than what C0's delayed retrieval-over-restudy contrast requires: the contrast can appear equally when unaided practice success is experimentally low.

Serves: Q1 (What is optimal learning?)
Ages: not restricted by the claim text

Bounds: Canonical retrieval-practice designs that motivate C0 (e.g. Roediger & Karpicke 2006 Exp. 1) typically gate the restudy/test contrast on study exposure rather than item-level verified retrieval; C1 asserts a stricter accessibility criterion than those designs verify. Success here is accessibility under the practice/criterion cue set (Tulving & Pearlstone), not storage availability alone and not “was on the page.” Encoding success is criterion-relative (Morris, Bransford & Franks transfer-appropriate processing): it is defined relative to the practice/criterion task, not as task-free storage strength. Kill-condition prong 2 is met under corrective feedback (Gupta, Pan & Rickard 2022, DOI 10.3758/s13421-021-01236-4, full text verified 2026-10-01): with practice success experimentally ~0.22 versus ~0.55 via prior-study load, retrieval-with-feedback still beat restudy at ~48 h by a similar margin (Exp. 1 finals ~0.40 vs 0.24 at low success and ~0.60 vs 0.41 at high, no Training×Repetition interaction), so C0's delayed contrast with feedback does not require high unaided practice-phase accessibility success. Racsmány, Szőllősi & Marián 2020 (DOI 10.3758/s13421-020-01041-5, full text verified 2026-10-02): no-feedback equality residual tested and failed — practice success experimentally varied via prior-study load with no practice feedback and a restudy arm; first-final 1-week TE present but not equal across levels (Exp 1 practice ~25–29%, TE d=0.43; Exp 2 ~66–75%, d=1.33; Exp 3 ~76–81%, d=1.71 / ~61% vs ~28%), so without feedback the contrast does not appear equally when accessibility fails for most items; survived 2026-10-02.

Attack record: 7 attacks: 3 survived, 4 narrowed · Evidence links: 9 linked sources (4 supports, 5 limits)

Claims serving Q2

C2 alive

When learners use an AI tool that hands them finished answers or solutions while they practise, their scores on the practice work go up, but what they have learned is lower when it is later measured without the tool.

Serves: Q2 (How do we achieve optimal learning when AI and screens are everywhere?)
Ages: grundskola-hog, gymnasium, vuxen-livslangt (evidence: upper-secondary maths in Turkey, grades 7–12 in one Chinese county, England Year 10 reading, university lab, HE RCT with delayed closed-book, and HE introductory programming); not yet shown for younger pupils or in Swedish classrooms

Bounds: Evidence in the bank: Bastani et al. 2024/2025 (upper-secondary maths RCT, Turkey: answer-giving GPT Base +48 % correct answers during practice, −17 % on the closed-book exam compared with pupils who never had GenAI; kill-condition test survived 2026-09-27); Strömberg, Lei & Wu 2026 (secondary-school DiD in one Chinese county, conference paper, results may be revised: homework scores +18 %, monthly closed-book tests −20 % within six months, entrance exams −18 to −24 % with the full penalty visible after about two years); Fan et al. 2025 (university lab, second-language essay: ChatGPT raised essay scores but not knowledge gain or transfer); Barcaui 2025 (HE RCT, Brazil: unrestricted ChatGPT as study aid → lower surprise retention at 45 days vs traditional study, d = 0.68; effect held after controlling for study time; kill-condition test survived 2026-09-27). Kreijkes et al. 2026 (England Year 10 RCT: LLM-only study of informational text → lower 3-day literal retention d=0.44, comprehension d=0.38, free recall d=0.21 vs note-taking; kill-condition test survived 2026-09-29). Bergh et al. 2026 (HE CS experiment, arXiv: ChatGPT during coding → higher task scores but lower immediate and 48 h recall without the tool, g≈0.71–0.77; kill-condition test survived 2026-09-29). The loss is concentrated in "homework outsourcing" use (about 80 % of AI users in Strömberg et al.); users who keep study time similar to non-users show small losses. The claim is about learning measured without the tool; performance with the tool goes up. Positive pooled effects in ChatGPT/chatbot meta-analyses (Deng et al. 2025, Doo & Park 2026, Wu & Yu 2024, Fan et al. 2026) are recorded as challenges, because the bank notes that most of their outcomes were not measured without the tool.

Attack record: 4 attacks: 4 survived · Evidence links: 12 linked sources (7 supports, 3 limits, 4 challenges)

C3 weakened

For understanding continuous (linear) text in depth, reading on paper gives somewhat better comprehension on average than reading the same text on a screen among readers beyond early primary school, and readers on screen tend to overestimate how well they have understood. The average paper advantage has not been shown for first-grade beginner readers of short linear texts.

Serves: Q2 (How do we achieve optimal learning when AI and screens are everywhere?)
Ages: overgripande; strongest evidence in university students (vuxen-livslangt; Li & Yan 2024 finds a significant paper advantage only in its university subgroup); school-age evidence from grade 8 and grade 10 in Norway (Jensen, Roe & Blikstad-Balas 2024; Mangen et al. 2013), from Italian grade 7 lengthy informational texts with null medium main effect (Ronconi, Altoè & Mason 2025; kill-condition test survived 2026-09-30, calibration not tested), and from upper-secondary VET (Foss, Støle & Magyari 2026); first-grade beginner readers are outside the average paper advantage (Florit et al. 2025: no screen-inferiority main effect); see Bounds for schoolchildren in Salmerón et al. 2024 and for children aged 1–8

Bounds: The average effect is small (Delgado et al. 2018 g ≈ −0.21; Clinton 2019 g ≈ −0.25) and is scoped to readers beyond early primary school. Delgado's same-text between-participant subgroups favour paper in grades 1–6 (g = −.19), grades 7–12 (g = −.15) and undergraduates (g = −.28), with educational level not a significant moderator. The newest meta-analysis (Li & Yan 2024; 46 experiments, 2000–2022; I² ≈ 92 %) found no overall difference (g = 0.046, ns), but that null is carried by interactive or enhanced digital texts (g = 0.959, digital better), literary texts (mostly children's storybooks) and kindergarten studies; inside this claim's scope its estimates favour paper (university g = −0.432; time-limited g = −0.468; 1000–2000-word texts g = −0.391; informational g = −0.218, p = .053; non-interactive g = −0.172, ns). Its school-age subgroups (primary, middle school) lean digital but are not significant and are dominated by interactive e-books. It pooled no calibration outcome and did not separate comprehension depth or device. Kill-condition test survived 2026-09-28. Foss, Støle & Magyari 2026 (Norwegian VET, n=106, within-subjects, informational texts on own laptop vs print): whole-sample screen inferiority (OR=1.44; print M=4.15 vs digital M=3.42); significant in the lower-comprehension subgroup, trend only in the higher (p=.073); kill-condition test survived 2026-09-29 (calibration not tested). Ronconi, Altoè & Mason 2025 (Italian grade 7, N=191, mixed design; two ~1000-word informational texts; paper vs scrolling desktop PDF): no significant medium main effect on literal or inferential comprehension; lower perceived cognitive load on screen; calibration not tested; kill-condition test survived 2026-09-30 (prong a met for this design; prong b untested). First-grade beginner readers of short linear texts: Florit et al. 2025 (N=58, within-subjects) found no significant medium main effect on main-point, literal, or inferential comprehension; for descriptive texts tablet beat paper on main-point (d ≈ 0.57) and literal (d ≈ 0.59). The advantage is for informational/expository text; for narrative-only text it is not significant (Delgado) or about zero (Clinton). Time limits make the screen disadvantage larger (Delgado: time-limited g = −.26, self-paced g = −.09, ns; Li & Yan 2024: time-limited g = −0.468, free reading g = 0.192, ns); whether any paper advantage remains with self-paced reading is open question 4. Much of the gap is tied to scrolling: with scrolling paper is clearly better, with paginated (no-scroll) display the differences are small or uncertain (Clinton-Lisell & Litzinger 2026). On handheld devices (tablet, e-reader) the disadvantage is about half as large, and between studies it was larger for undergraduates than for schoolchildren, for whom it was not significant (Salmerón et al. 2024). Display technology alone (LCD vs e-ink) does not explain the gap (Siegenthaler et al. 2012). Main-idea understanding is about equal; key points and other details are better on paper (Singer & Alexander 2017). For children aged 1–8, plain digitisation of a story lowers comprehension, but story-congruent digital enhancements can beat paper (Furenes, Kucirkova & Bus 2021). Much of the evidence for the average paper advantage is from studies up to 2017 and from university students.

Attack record: 4 attacks: 3 survived, 1 narrowed · Evidence links: 15 linked sources (10 supports, 10 limits)

How the claims connect

What else is here

Open questions

  1. Does C0’s “delayed” main contrast hold between ~5 min and ~2 days, or must Bounds name a multi-day minimum as in Roediger & Karpicke 2006?
  2. Does C1 require that encoding verification use the same surface format as C0’s eventual criterion test, or is any successful retrieval without re-presentation enough?
  3. Within Roediger & Karpicke 2006, do idea units that failed initial free recall after study still show a delayed testing-over-restudy advantage at the item level?
  4. Does C3’s average paper advantage hold when reading is self-paced, given that neither Delgado 2018 (self-paced g = −.09, ns) nor Li & Yan 2024 (free reading g = 0.192, ns) found a significant difference without time limits — or is the paper advantage specific to time-limited reading? (Replaces the Li & Yan 2024 question, resolved 2026-09-28: C3 survived that kill-test.)
  5. Does C1’s accessibility criterion collapse to “was studied / received feedback re-presentation,” with no residual role for cue-matched access as a distinct criterion — given that under corrective feedback C0 does not require high unaided practice success (Gupta, Pan & Rickard 2022) and that without practice feedback the retrieval-over-restudy contrast shrinks (does not appear equally) when practice success is experimentally low (Racsmány, Szőllősi & Marián 2020, DOI 10.3758/s13421-020-01041-5, kill-condition test survived 2026-10-02)? (Replaces the residual no-feedback equality question, resolved 2026-10-02.)