What PISA 2025 Actually Shows About AI and Student Capability

The line that now has a number

The debate about AI and learning has mostly lacked evidence at scale. On 8 September 2026 the OECD published some: PISA 2025 asked more than 760,000 students in 91 countries and economies how they use AI chatbots, then set those answers against measured performance. The result is not a simple AI-versus-no-AI story. Purpose matters. Frequency matters. Guidance matters.

PISA 2025 was the largest assessment in the programme’s history. More than 760,000 students took part, representing roughly 33 million 15-year-olds. Science was the main domain. For the first time at this scale, the questionnaire distinguished between different things students do with AI chatbots when they sit down to work.

That distinction carries much of the useful signal. Some coverage reduced the release to a simple claim that students who use AI perform worse. The full OECD analysis shows a more complicated pattern: measured performance varies with the task, how often AI is used and whether students are taught to evaluate what the system returns.

What the OECD found

The headline picture frames everything else. PISA 2025 recorded the lowest OECD-average performance ever observed in science, reading and mathematics. Across the OECD, one in five 15-year-olds is now a low performer in all three subjects at once, up from 16% in 2022. Average reading scores fell by 28 points between 2015 and 2025, which the OECD equates to about one and a half years of learning, and mathematics fell by 22 points. Science, the focus of this cycle, held broadly stable between 2022 and 2025 after trending downward from 2009.

Against that backdrop, the AI findings break into three parts.

Task-specific use is associated with lower performance on average. Students who used AI for specific schoolwork tasks such as summarising assigned reading, drafting text or conducting preliminary research scored lower in science than students who did not use AI for those purposes. Averaged across those three tasks, the gap between non-users and the mean of all user frequencies is 19.6 score points, derived from Table I.B1.4.61 behind Figure I.4.13. The pattern is not perfectly linear, however. For summarising and preliminary research, moderate users, roughly once a month to twice a week, outperformed both limited and frequent users. For drafting, that moderate-use advantage was less pronounced.

Frequent drafting shows a larger gap. For drafting writing assignments specifically, the reported difference between students who used AI daily or almost daily and those who used it never or almost never is 28.7 science points after adjustment for socio-economic status: 509.4 for non-users against 480.7 for daily users. That figure is separate from, and should not be confused with, the 28-point decade-long decline in OECD reading scores noted above.

Learning-oriented use behaves differently. Students who used AI about weekly “to help me learn” scored slightly above non-users after accounting for socio-economic profile: 501.3 against 499.0. Students using it for that purpose only occasionally or almost every day tended to score lower. When students had opportunities at school to assess the quality of AI-generated information, frequent learning-oriented users, once a week or more, tended to score slightly higher than non-users and infrequent users. The OECD also notes that these learning opportunities are reported more often by socio-economically advantaged students, raising the possibility of a new AI-literacy divide.

PISA 2025 · THREE VARIABLES SHAPE THE SIGNAL1 · TASKSpecific schoolworkSummarising, drafting, research19.6 ptsnon-user gap, 3 tasksModerate summarising/researchusers outperform limited andfrequent users.NOT A SIMPLE LINEAR EFFECT2 · FREQUENCYDrafting every dayDaily vs never/almost never28.7 pts509.4 vs 480.7, adjustedWeekly “help me learn” use:501.3 vs 499.0Occasional or near-dailylearning use tends to be lower.3 · GUIDANCELearning + evaluationFrequent learning-oriented useslightly higherwith opportunities to assess AIThose opportunities are morecommon among advantagedstudents.GUIDANCE CHANGES THE CONTEXTPurpose matters. Frequency matters. Guidance matters.OECD Table I.B1.4.61, adjusted for socio-economic status. Association only.

Source: OECD, PISA 2025 Results, Volume I, Figures I.4.13 and I.4.14 and the OECD overview. Scores in the underlying analysis are adjusted for socio-economic status. The 28.7-point drafting figure is the difference between never/almost-never use (509.4) and daily/almost-daily use (480.7), taken from Table I.B1.4.61. The OECD states that these relationships do not establish a causal effect of AI on science performance.

What this data cannot tell you

This section exists because the alternative is a headline that will not survive contact with an informed reader.

  • It is not causation. The OECD states directly that these relationships do not necessarily imply a negative impact of AI on science performance. Everything above is an association.
  • The decline predates generative AI. Reading fell 28 points and mathematics 22 points across OECD countries between 2015 and 2025. ChatGPT became publicly available in late 2022. A decade-long trend cannot be assigned to a three-year-old product.
  • Reverse causation remains possible. Students who are struggling may have more reason to reach for a chatbot. That can produce an association without AI having caused the underlying difficulty.
  • Selection also complicates the guided-use finding. Opportunities to assess AI-generated information are more common among socio-economically advantaged students. Some of the observed advantage may reflect differences in schools, students or learning environments rather than the instruction itself.
  • The surrounding causes are broader. The OECD discusses teacher shortages, distraction, screen use, attention and declining reading habits among the wider conditions shaping performance. PISA does not assign a causal share to each factor, and neither should anyone reading it.
  • “Students are doing worse” is not uniformly true. In the United States, NAEP’s 2025 long-term-trend results showed 9-year-olds improving in reading and mathematics relative to 2022, while 13-year-olds showed no significant change from 2023. Within PISA itself, some education systems improved in science even as the OECD average remained weak.

One limitation this article cannot resolve on its own is whether the OECD-average pattern holds system by system. Adrian Mizzi’s exploratory secondary analysis of the underlying PISA 2025 data, supplied by the OECD, addresses exactly that. Across 73 education systems, modelling reading performance against four reported purposes of AI use while adjusting for curiosity, perseverance, disciplinary climate, family support and socio-economic status, he reports learning support positively associated with reading in 71 of 73 systems, drafting negatively associated in 68, and summarisation negatively associated in 64. He describes the work as exploratory, notes that reverse causality cannot be excluded, and cautions that multiple comparisons inflate the chance of spurious findings. His outcome measure is reading; the figures above are science, the main domain in this cycle. Two outcome measures, two levels of aggregation, the same direction. The OECD has since made the related point in its own commentary, that AI use appears associated with better learning outcomes when students are explicitly taught how to use it.

None of that makes the AI result unimportant. It makes the useful question narrower. The evidence is not asking whether a generation uses AI. It is asking what cognitive work is being delegated, how often, and under what guidance.

Evidence consistent with the Performance–Capability Gap

BBGK uses a specific name for the distance between what a person can produce with AI assistance and what they can reliably do on their own. We call it the Performance–Capability Gap, and it was defined before this dataset existed.

PISA did not measure that gap. It did not collect the assignments students produced with AI, compare their quality with unaided work, or run a controlled before-and-after experiment. What it measured was self-reported AI-use patterns on one side and standardised science performance on the other. That is observational evidence and should be described as such.

But it is population-scale evidence consistent with the gap, and the shape of the result matters.

BBGK reading

When the purpose of an activity is to build a capability, delegating the cognitive production can raise visible output without assuring that the underlying capability develops with it. Using AI as a scaffold can preserve more of the reasoning work. If the Performance–Capability Gap is real, heavier delegation should coincide with weaker independent performance more often than moderate, guided support. PISA cannot observe the assisted artefact side of that gap. It can observe the independent-assessment side, and the pattern it reports is consistent with the prediction.

That is BBGK’s inference, not an OECD finding. The honest conclusion is not that PISA proved the framework. It is that a framework built from smaller studies and first principles now has a very large observational dataset pointing in the same direction.

Why output is the wrong measure of capability

Andreas Schleicher, the OECD’s Director for Education and Skills, puts the mechanism plainly in the PISA 2025 foreword. We do not become fit by watching sport. Learning comes through productive cognitive struggle with new material. His implication for AI is equally direct: the technology should act as a scaffold rather than a crutch, enabling thought rather than short-circuiting it.

The struggle is not simply a cost of learning that the tool can remove. Often, the struggle is the learning. Remove too much of it and visible performance can improve while capability remains unchanged, or potentially weakens.

Where the boundary sits

If the Gap describes the risk, the AI Delegation Boundary Framework describes where to place the tool. But the boundary cannot be determined by the surface label of a task. It depends on what the task is for.

  • Automate when the artefact is the objective. If the underlying cognitive skill is not what is being learned, tested or relied upon, AI can take more of the production burden. Formatting, routine transformation, boilerplate and low-risk first-pass synthesis can belong here, with appropriate verification.
  • Assist when the thinking still matters. Use AI to explain, question, challenge, generate alternatives, provide feedback or test comprehension while the person retains the reasoning work. This is closest to the moderate, learning-oriented use that PISA finds compatible with stronger performance.
  • Own directly when performing the cognition is the point. Comprehension, synthesis, judgment, core reasoning and any skill currently being learned or independently evaluated should not be delegated away merely because AI can produce a plausible artefact.

The same surface task can therefore sit in different tiers. Summarising a meeting for distribution may be automatable. Summarising a difficult chapter to build reading comprehension is a different task, even if both outputs are called a summary.

PISA does not prove this boundary. It does make one principle harder to ignore: delegation decisions should be based on the capability you are trying to preserve, not on whether the software can generate the output.

This is not a schools problem

Here is the part that should concern anyone who is not a teacher.

PISA had a rare advantage. It could set reported AI-use patterns against an assessment that measured what students could do in the test environment. Few organisations using AI routinely run an equivalent comparison. Companies measure delivery speed, output volume, throughput and turnaround time because those are operationally visible. Much less often do they measure what people can still do independently when the tool is unavailable, unreliable or wrong.

For many knowledge workers whose AI-assisted output has improved, neither the individual nor the employer has a direct measure of whether independent capability improved with it, stayed flat, or declined. The dependency can form quietly because it can look like productivity while it is forming. The weakness becomes visible when the system fails or when judgment is required to catch a plausible error. BBGK has written separately about what happens at that moment.

The 15-year-olds in this dataset were measured. Most workplaces are not.

Singapore is at the top, and still investing in reading

In PISA 2025 the highest-performing systems across science, mathematics and reading were the Chinese jurisdictions of Beijing, Shanghai, Jiangsu and Zhejiang, and Singapore. The comparison that follows is not evidence that one policy caused that performance. The timing, populations and measures are different.

Two days before the PISA release, Singapore’s National Library Board launched ReadSG, a five-year national movement intended to build lifelong reading habits. Its starting target is deliberately small: at least 15 minutes of reading a day. As one component of the movement, the ReadSG Challenge uses CrowdTask SG to add goal-setting, streaks, experience points and redeemable virtual coins to those daily reading sessions.

The chronology matters. ReadSG was announced before the PISA results and was not created in response to them. Nor does it target the same population as PISA. The comparison here is about policy orientation, not causality.

The design rewards an input rather than an artefact. The challenge does not ask participants to produce a summary, a score or a deliverable. It rewards time spent reading. That is not a capability metric because capability is not being measured. It is a behavioural input intended to support capability-building.

It also targets a problem the OECD is worried about. PISA’s broader analysis repeatedly returns to concentration, deep reading, critical interpretation and the risks of digital distraction. NLB makes the same strategic bet in different language: in a high-stimulation digital environment, sustained long-form reading is a habit worth deliberately protecting.

Whether a reward system builds durable intrinsic motivation is an open question. A widely cited meta-analysis by Deci, Koestner and Ryan found that some expected tangible rewards can reduce later intrinsic motivation, although the effects depend on reward type, context and design. ReadSG will therefore be more interesting as a behavioural experiment than as proof that paying people to read works.

The useful point is narrower: one of the world’s highest-performing education systems is not treating sustained attention as a background condition it can take for granted. It is treating reading habit as something that must be designed for.

What to do with this

Six moves, in order of how quickly they can be made.

  • Name the purpose before you delegate. Before opening a chat window, ask whether the objective is to produce an artefact or to build understanding. The same task label can justify different levels of AI assistance.
  • Watch frequency, not just permission. PISA’s pattern is non-linear. Moderate use can look different from near-daily use. “Allowed” versus “banned” is too crude a control variable.
  • Build one unassisted check. Pick a capability you now exercise with AI and test it periodically without assistance. Not as a moral exercise. As instrumentation.
  • Teach evaluation, not mere access. PISA does not test whether bans work. What it does show is that frequent learning-oriented users with opportunities to evaluate AI-generated information tend to perform better than comparable users without that context.
  • Measure capability as well as output. Delivery speed and throughput are performance metrics. Add at least one measure that tells you whether people can still reason, diagnose or decide when the tool is absent or wrong.
  • Protect the inputs that build judgment. Reading, analysis, problem framing, verification and reflection are slower than generating an artefact. If organisations reward only visible output, they create an incentive to remove exactly the cognitive work they still need people to own.

PISA 2025 does not settle the argument over AI and learning. It does something more useful. It replaces a binary question with an operational one.

What are we delegating, how often, under what guidance, and which human capability must remain when the tool is gone?

The 15-year-olds were measured. Most workplaces are not. That is the one to take home.

Frequently asked

Did PISA 2025 find that AI causes lower test scores?

No. PISA 2025 found associations between patterns of AI use and science performance. The OECD states that these relationships do not establish that AI caused the score differences. Reading and mathematics performance had also been declining across OECD countries well before generative AI became widely available.

What did PISA 2025 find about AI use for schoolwork?

Students using AI for specific schoolwork tasks such as summarising, drafting and preliminary research scored 19.6 points lower in science on average than non-users of AI for those purposes. But frequency matters: moderate summarising and research users outperformed limited and frequent users, and the pattern for learning-oriented use was different again.

How large was the drafting gap in PISA 2025?

For drafting writing assignments, the difference between students using AI daily or almost daily and those using it never or almost never is 28.7 science points after adjusting for socio-economic status: 509.4 against 480.7. This is distinct from the separate 28-point decline in OECD reading scores between 2015 and 2025.

What happened when students used AI to help them learn?

Students using AI about weekly “to help me learn” scored slightly above non-users after accounting for socio-economic profile, 501.3 against 499.0. Occasional and almost-daily learning-oriented users tended to score lower. Frequent learning-oriented users who also had opportunities at school to assess AI-generated information tended to score slightly higher than non-users and infrequent users.

What is the Performance–Capability Gap?

The Performance–Capability Gap is a BBGK framework describing the distance between what a person can produce with AI assistance and what they can reliably do independently. PISA 2025 did not measure that gap directly, but its observational pattern is consistent with the framework’s central prediction.

What is Singapore’s ReadSG programme?

ReadSG is a five-year national reading movement launched by Singapore’s National Library Board on 6 September 2026. It encourages at least 15 minutes of reading a day. The separate ReadSG Challenge adds goal-setting, streaks, experience points and redeemable virtual coins through CrowdTask SG. It launched before the PISA 2025 results and should not be treated as a response to, or explanation for, Singapore’s PISA performance.

Primary sources. OECD, PISA 2025 Results (Volume I): Future-Ready Students, 2026; OECD, Student school life and beyond, including Figures I.4.13 and I.4.14; OECD, Foreword, Andreas Schleicher; OECD, PISA 2025 press release, 8 September 2026; U.S. Department of Education / NCES, NAEP 2025 Long-Term Trend Assessment; National Library Board Singapore, Launch of ReadSG, 6 September 2026. OECD, PISA findings on artificial intelligence use, reading skills and learning, 2026. Cross-system analysis: Adrian Mizzi, exploratory secondary analysis of PISA 2025 underlying data across 73 education systems, September 2026.

Supporting research. Deci, E. L., Koestner, R. and Ryan, R. M., A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation, Psychological Bulletin, 1999. All score figures in this article are taken from OECD, PISA 2025 Database, Table I.B1.4.61 (underlying data for Figure I.4.13), and Table I.B1.4.70 (Figure I.4.14), via the OECD StatLink at stat.link/o1j4k0. Scores are the socio-economically adjusted series.

Published 19 September 2026 by AHS Shohel Ahmed, Founder and Principal Analyst, BBGK.

Comments

comments

AHS Shohel Ahmed
About the Author
AHS Shohel Ahmed is founder and principal analyst of BBGK (Beyond Boundaries Global Knowledge). More than a decade of documented client work across AI strategy, search and discoverability, digital authority, and executive communication — with Top Rated Plus standing on Upwork and 30,000+ hours of verified client delivery. He currently serves as Director of AI Strategy & Digital Authority for Griffin Capital Funding, OHA HVAC Plumbing Chimney and Fireplaces, and GiveTaxFree.