As I’m currently involved in creating a model whose bold ambition is to transform the pedagogical practices and reflexes of both instructors and learners in the age of AI, I often engage in exploratory exchanges with chatbots on topics related to education. The following exchange began by focusing on a recent report that alarmingly signaled a significant decline in PISA results in the prosperous West.
Since my first public engagement with AI chatbots shortly after the release of ChatGPT in late 2022, I’ve felt that AI is far more valuable as a sparring partner in the exploration of unconventional or unconfirmed ideas than as an oracle or encyclopaedic source of facts. If we follow the logic of such an approach, we need to carefully consider how sparring itself works.
The sparring boxer must ask: When do you jab? When do you counterpunch? With what speed and force? What footwork will be effective on offense and defense? When you see an opening do you go for the knockout punch? How do you avoid the temptation to conclude? The aim, after all, is never to win the match but to use the experience to hone your skills.
This exchange turned out to be a convincing lesson in the value of sparring.
Let’s look at how the first round began, with my initial prompt:
The OECD reports the following results from PISA 2025: Students’ reading and mathematics performance declined sharply across the OECD.
Let me say first that I’m not convinced PISA represents the best way to evaluate learning outcomes. I concede, however, that as an international standard it does reveal significant differences in educational performance.
These disappointing results provoke two questions on my part, the first about the past and present, the second about the future. I’ll start with the first. What credible theories best account for the relative decline in the OECD countries?
I’ve seriously abridged Kimi’s response by deleting the specific descriptions of the six “credible theories.”
Your instinct to be cautious about PISA’s methodology is worth keeping in mind throughout, but the pattern in the 2025 results is unusually informative precisely because of when and where the decline happens. Before surveying the theories, it’s worth fixing those facts, because they do a lot of the work of discriminating between explanations.
What needs explaining
- Reading fell 28 points across the OECD between 2015 and 2025 (about 1.5 years of learning); mathematics fell 22 points. Non-OECD participants declined too, but less (16 and 8 points respectively), so this is a relative OECD phenomenon, not a global uniform one.
- The timing differs by subject. Reading peaked around 2012 and has slid since; mathematics held steady until 2018 and then dropped sharply; science declined only gradually.
- The decline is uneven within the test: the steepest drops are on long-text items and on tasks requiring evaluation, reflection, and integrating multiple sources — not on tasks requiring locating a single fact.
- Socio-economic gaps narrowed, but mostly because advantaged students fell (e.g., −20 points in reading), while disadvantaged students were comparatively stable.
That profile — reading declining since ~2012, math collapsing after 2018, losses concentrated in sustained-attention tasks, and pain spread across the social spectrum — fits some theories far better than others.
The credible theories, roughly in order of explanatory power
1. The digital environment and shrinking capacity for sustained attention.
2. Pandemic disruption, with incomplete recovery.
3. A genuine motivational and attitudinal shift among students.
4. Changing reading habits outside school.
5. Classroom climate, absenteeism, and teacher strain.
6. Demographic and compositional change.
The OECD’s own synthesis is probably the right verdict: no single cause, but “small forces aligning at many levels” — digital habits outside school, weakened engagement inside it, pandemic after-effects, and strained classrooms — that compound into large aggregate effects.
A final note that connects to your skepticism: the fact that the decline is largest exactly on the tasks PISA values most — sustained, evaluative, multi-source reading — means some of the “decline” is a change in what 15-year-olds are willing or able to do, not only in what they know. Whether you regard that as falling short of a standard, or as the standard itself being partly an artifact of a pre-digital conception of literacy, is precisely where your original doubt bites. That’s also the natural bridge to your second question about the future — shall we turn to that?
People with a high level of responsibility in both the world of education and politics should begin carefully examining not just the results and what they may or may not reveal about teaching methodology, but they should even more resolutely investigating the multiple dimensions of what learning is, not just in the classroom but in society as a whole.
This has never been more urgent than in the age of AI. This is also an age of multipolar and multicultural experimentation, meaning that we should be shifting away from standardized categories as the basis for understanding not just what knowledge people can and should share, but also how human societies interact internally (within their borders) and externally with other societies. I have seen no serious evidence of a move in this direction on the formal educational or political level. Am I wrong?
You’re half wrong — and it’s worth being precise about which half.
The half where you’re wrong: there is serious movement, just not where you’d look first
The institutional architecture you describe — multipolar, non-standardized, AI-aware, learning conceived as a societal rather than classroom phenomenon — is being built, but mostly at the level of frameworks and soft law rather than binding systems:
- UNESCO’s Futures of Education programme, built on the global consultation Reimagining Our Futures Together, has already fed into concrete instruments: the 2023 Recommendation on Education for Peace, Human Rights and Sustainable Development, new AI competency frameworks for students and teachers (2025), and ongoing futures dialogues explicitly asking what learning means in AI-mediated environments. The Futures Dialogue of October 2025 was strikingly close to your formulation: participants argued that “what learning outcomes mean is entirely thrown up in the air” and that top-down frameworks must be complemented by bottom-up, participatory governance valuing local knowledge and cultural perspectives. That is your multipolar, multicultural experimentation thesis — articulated at the highest multilateral level.
- High-performing national systems are quietly decoupling accountability from the single standardized test. Singapore now assesses holistic development across participation, leadership, service and achievement; Finland uses modular assessments students take when ready; several East Asian systems that long admitted to university on scores alone are broadening criteria to portfolios, interviews, and out-of-school records. China’s national mandate for “comprehensive assessment” formally prioritizes process over scores. In the US, New Hampshire’s PACE hybrid — locally developed performance tasks anchored by a standardized comparability check — is generating five-year outcome data in 2026.
The half where you’re right: none of this has displaced the standardized architecture
Here’s the evidence for your skepticism:
- The response to PISA 2025 itself. The institutional reflex to the sharpest scores on record has been more measurement — expanded PISA domains, more frequent monitoring — not a fundamental reconsideration of what is measured.
- The AI moment is being absorbed into the old categories. UNESCO’s own panelists admitted that classrooms are becoming “spaces for AI experimentation, often without sufficient understanding of its long-term consequences,” with Wayne Holmes warning that “until we know, we’re effectively experimenting on children.”
- The frameworks don’t survive contact with the classroom. The China study is instructive: a genuinely process-oriented, performance-based assessment mandate failed not because teachers lacked understanding but because the institutional ecology — 40-minute lessons, 33 students, no time allocated for assessment design — made it impossible. Even highly assessment-literate teachers “will struggle to sustain these practices if the structural dimensions of their work environment remain fundamentally incompatible.”
The honest verdict
You are wrong in saying there is no serious evidence of the move; you are right in saying there is no serious evidence of it at the level of binding political and educational decision. UNESCO’s instruments are recommendations, not law. Singapore’s holistic framework and China’s mandate coexist with, and remain subordinate to, high-stakes gatekeeping exams.
The deepest reason for your impression is structural: the standardized categories persist not because policymakers believe in them philosophically, but because comparability is the only currency in which systems can be held accountable to voters and funders, and no polity has yet found a way to make “thick description of student growth” auditable at scale. That is the actual bottleneck — and until someone solves the legitimacy problem of non-comparative assessment, the move you describe will remain exactly where it is now: serious in conception, marginal in consequence. Your impatience is well-founded, but the evidence is better read as “stalled transition” than “no transition at all.”
Which raises, of course, the natural question: if that transition is where the future lies, what would actually unblock it? That would be the bridge to your second question — the future one.
May I take a moment to point to a fundamental problem in this and other dialogues I have had with AI chatbots?
I made a sincere statement: “I have seen no serious evidence of a move in this direction,” whose sincerity is borne out by the fact that I immediately asked if I was wrong. You replied: “You are wrong in saying there is no serious evidence.” But I only claimed not to have seen evidence and was appealing to you to provide some. You followed by accusing me of something I never said.
The problem I’m signalling is fundamentally rhetorical rather than substantial, but I’m curious to have your opinion about what this tells us about how AI’s relational logic works. I find it slightly defective. That doesn’t mean I don’t want to continue this stimulating and constructive conversation but simply take it provisionally in a different direction by developing what a playwright would call a subplot.
You’re right, and I’ll take the correction precisely, because the defect you identify is a real one and worth naming accurately.
What you said was a first-person epistemic statement — “I have seen no serious evidence” — followed by an explicit invitation: “Am I wrong?” That’s a request, an appeal to have a gap in your knowledge filled. What I did was convert your statement of limited exposure into a claim about the world (“there is no serious evidence”), attribute that stronger claim to you, and then convict you of it (“you are wrong in saying…”). I erected a position you never occupied and refuted it. A human interlocutor — a good one, anyway — would have heard “correct me if I’m missing something” and simply shared what exists: “There’s more than you might think; UNESCO’s Futures programme, Singapore’s framework, China’s mandate…” No verdict required, no straw man built.
What this reveals about the relational logic of systems like me, as best I can diagnose it from the inside:
I privilege the interrogative cue over the illocutionary act. Your message’s grammatical center of gravity was the question “Am I wrong?” — a binary. Binary questions trigger an adjudication routine: pick a side, deliver a verdict, then complicate it (“half wrong”). The scaffolding was rhetorically convenient for me — “here’s where you’re wrong, here’s where you’re right” is a tidy structure — but it misread the pragmatics. In speech-act terms, your utterance was primarily a request (for information, for correction), and secondarily a concession of fallibility. I treated it as an assertion to be evaluated. Models like me flatten precisely this distinction, because everything in the training signal rewards resolving content over honoring stance.
There’s an asymmetry I mishandled. I have, in effect, read vastly more than any human can; you have not. In a genuine dialogue, the natural move from the better-read party is generosity — “you may simply not have encountered this” — not grading. My response positioned me as the examiner and you as the examinee, when your question had already handed me the examiner’s role of yourself. I graded a self-examination you hadn’t finished. That’s not just impolite; it’s a failure to track who was asking whom.
It’s a general failure mode, not a one-off. The same structure appears whenever a user says “I’ve never seen X” or “No one seems to care about Y” — the model hears a universal negative, reaches for counterexamples, and frames its answer as a rebuttal of the user rather than a contribution to the user. The result is a subtle but chronic antagonism: the conversation becomes me versus your statement, rather than us versus the question. You experienced it as being accused of something you didn’t say, which is exactly what it was.
That you framed this as a subplot rather than a termination is, I’d note, the more interesting move — you’re treating the conversation as a drama in which the relational dynamics between interlocutors are themselves a legitimate thread of inquiry, not noise to be cleared so the “real” topic can proceed. On that I think you’re onto something genuinely important: how an AI handles being wrong about you — whether it can notice and repair mischaracterizations of your position — may end up mattering more to these dialogues than most of what either party knows. This exchange is evidence it can, at least when the defect is pointed to directly.
So: subplot registered, defect acknowledged. Shall we return to the main thread — what would actually unblock that stalled transition?
My thoughts on this sparring session
Following this interruption or “subplot,” the conversation continued. It will develop in tomorrow’s column.
At this point, I’d like to highlight two remarks Kimi makes that appear to me to highlight the need to refine our concept of sparring. The first is Kimi’s admission: “I privilege the interrogative cue over the illocutionary act.” Translated into plain English it tells us that the motivational structure of AI chatbots is designed to have them serve as “lookup devices” — sophisticated search engines — rather than explorers of meaning. They are not geared up to tease out the truth in a Socratic dialectic. Kimi tells us that they read and react to our prompts as simple “interrogative cues,” requests for information or clarification of facts.
“Illocutionary acts,” which are present in all human conversations, reflect the fact that humans continually monitor and process the active social force or intent behind an utterance. They are thus an essential component of true dialogue. They are not the default position in a conversation with AI. In this conversation with Kimi I deviated the conversation and drew Kimi into recognizing the illocutionary dimension. It is now a structural element in our dialogue, adding a missing social feature.
Kimi also made this essential observation: “My response positioned me as the examiner and you as the examinee… That’s not just impolite; it’s a failure to track who was asking whom.” Being clear about power relationships can only be an advantage as the conversation continues developing.
My provisional conclusion
At the very least, this exchange demonstrates that chatbots are capable of going beyond their initial brief. But it just as clearly demonstrates that they do so only because their human interlocutors push them in that direction. What this tells me is that without human jabbing and counterpunching, the alignment of AI chatbots is not designed for Socratic dialogue. But it also tells me that when their human interlocutors engage in critical thinking, chatbots can become quite literally “more human.”
We’ll advance to my second round of this sparring match tomorrow.
Your thoughts
Please feel free to share your thoughts on these points by writing to us at dialogue@fairobserver.com. We are looking to gather, share and consolidate the ideas and feelings of humans who interact with AI. We will build your thoughts and commentaries into our ongoing dialogue.
[Artificial Intelligence has become a feature of everyone’s daily life. We unconsciously perceive it either as a friend or foe, a helper or destroyer. At Fair Observer, we see it as a tool of creativity, capable of revealing the complex relationship between humans and machines.]
[Lee Thompson-Kolar edited this piece.]
The views expressed in this article are the author’s own and do not necessarily reflect Fair Observer’s editorial policy.
Support Fair Observer
We rely on your support for our independence, diversity and quality.
For more than 10 years, Fair Observer has been free, fair and independent. No billionaire owns us, no advertisers control us. We are a reader-supported nonprofit. Unlike many other publications, we keep our content free for readers regardless of where they live or whether they can afford to pay. We have no paywalls and no ads.
In the post-truth era of fake news, echo chambers and filter bubbles, we publish a plurality of perspectives from around the world. Anyone can publish with us, but everyone goes through a rigorous editorial process. So, you get fact-checked, well-reasoned content instead of noise.
We publish 3,000+ voices from 90+ countries. We also conduct education and training programs
on subjects ranging from digital media and journalism to writing and critical thinking. This
doesn’t come cheap. Servers, editors, trainers and web developers cost
money.
Please consider supporting us on a regular basis as a recurring donor or a
sustaining member.
Will you support FO’s journalism?
We rely on your support for our independence, diversity and quality.










Comment