Large Language Models Pass the History Exam But Miss the «History»: A Polish High School Exit Exam Matura Benchmark
summary
The gist
This study introduces the first LLM history benchmark grounded in Polish national curriculum, evaluating eight leading LLMs on the Polish high school exit exams (Matura) in history—three official
In short
A study benchmarked eight large language models against the Polish Matura history exam, finding they scored significantly higher than human averages. However, hosts discuss how these models fail to grasp cultural specifics and temporal reasoning. They pass based on knowledge retrieval but not on deep historical inquiry or contextual understanding.
Key concepts
- Matura Benchmark
- This is a real, nationally grounded high school exit exam used as a test for AI models. It provides 'ecological validity' because it uses actual grading criteria and real student questions, unlike synthetic tests that are made up by researchers.
- Source Conflation
- This failure mode occurs when the model treats historical documents as factual sources about the world, rather than analyzing them as artifacts of a specific object or time. It is a fundamental issue with how models process context.
- Temporal Disorientation
- This is when the model correctly identifies actors in history but places them in the wrong chronological period. It shows they struggle with sequencing events and understanding the correct timeline of historical inquiry.
- Cultural Specificity
- The models consistently score lower on content specific to Polish history compared to global history. This penalty suggests a lack of robust knowledge regarding non-Anglophone or Western European topics.
Terminology used across episodes
This episode discusses
- Large Language Models Pass the History Exam But Miss the <<History>>: A Polish High School Exit Exam Matura Benchmark · Paper Radio
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Time Awareness in Large Language Models: Benchmarking Fact Recall Across Time
- Measuring Massive Multitask Language Understanding
- LLMzSz: a comprehensive LLM benchmark for Polish
- LLMs' Reading Comprehension Is Affected by Parametric Knowledge and Struggles with Hypothetical Statements
- GPT-4o System Card
The paper
Large Language Models Pass the History Exam But Miss the <<History>>: A Polish High School Exit Exam Matura Benchmark · Read on arXiv
Adrian Trzoss, Kacper Dudzic, Wiktor Werner, Marcin Moskalewicz
Adam Mickiewicz University · IDEAS Research Institute · AMU Center for Artificial Intelligence · Poznan University of Medical Sciences · Maria Curie-Sklodowska University · WSB Merito University
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Large Language Models Pass the History Exam But Miss the «History»: A Polish High School Exit Exam Matura Benchmark".
Jane: The paper was written by Adrian Trzoss, Kacper Dudzic, Wiktor Werner and Marcin Moskalewicz from Adam Mickiewicz University and IDEAS Research Institute and AMU Center for Artificial Intelligence and Poznan University of Medical Sciences and Maria Curie-Sklodowska University and WSB Merito University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper with a title that really made me stop and think: "Large Language Models Pass the History Exam But Miss the «History»: A Polish High School Exit Exam Matura Benchmark." Jane, that title is doing a lot of work.
Jane: It really is, Tom. And honestly, it's a bit of a gut punch, isn't it? The researchers took eight leading AI models and had them sit the actual Polish high school history exit exam, the Matura, from two thousand twenty-three to two thousand twenty-five. And every single model absolutely crushed the human average. We're talking scores in the high 80s and 90s, while the human examinees averaged around forty-four percent.
Tom: So they pass with flying colors, but the title says they miss the history. What's the catch?
Jane: The catch is in the details. The aggregate scores look amazing, but when you break them down, the rankings are all over the place. A model that's third overall might be first on one type of question and sixth on another. And there's a consistent penalty on Polish history versus global history. The models just aren't as good when the content is culturally specific to Poland.
Tom: Right, so it's not just about knowing facts. It's about the kind of reasoning that's being tested. Lu, you've been nodding along—what's your take on that title?
Lu: I think the title is actually a very precise diagnosis. The paper identifies two specific failure modes that explain the "missing the history" part. One is what they call source conflation, where the model treats a historical document as a source of facts about the world, rather than as an object to be analyzed in its own right. The other is temporal disorientation, where the model gets the right actors but places them in the wrong period. So they're not just missing a few facts—they're missing the fundamental stance of historical inquiry.
Jane: And that's the part that scares me a little. Because a student who scores forty-four percent on this exam clearly doesn't know the material. But a model that scores ninety-six percent might be confidently wrong in ways that are much harder to spot.
Tom: Meng, from a practical standpoint, what does this mean for the student who's using one of these chatbots to study for their history final?
Meng: It means they should be careful. The paper shows that on the hardest questions, the models often fail completely, scoring zero across all three runs. And those hardest questions aren't the ones you'd expect. For humans, the hardest questions involve photo and text combinations. For the models, it's the temporal reasoning questions that trip them up. So a student asking for help on a question about sequencing events might get a very confident, very wrong answer.
Tom: So the models are like that student who memorized the textbook but doesn't really get the subject. They can ace the multiple choice but stumble on the essay that asks them to actually think.
Jane: Exactly. And that's why this benchmark is so valuable. It's not a synthetic test designed by researchers. It's a real exam, with real grading criteria, taken by real students. And it's testing a skill that most benchmarks ignore: interpretative reasoning about the past. That's a big deal.
Lu: And it's not just about history. This kind of benchmark, grounded in a national curriculum, can tell us a lot about how these models handle culturally embedded knowledge. If they struggle with Polish history, they're probably struggling with other non-Anglophone content too. That's a significant blind spot.
Tom: So we've got models that ace the test but miss the point. What exactly did the researchers find when they dug into the scores? That's coming up next.
Summary: Tom: So we're back with "Large Language Models Pass the History Exam But Miss the «History»: A Polish High School Exit Exam Matura Benchmark." And Jane, we've established that the models score incredibly high overall. But I want to dig into that score breakdown, because that's where the real story is.
Jane: Absolutely, Tom. The paper creates a three-tier ranking. At the top, you've got Claude Sonnet four point six and Gemini three point one Pro, and their scores are so close that statistically you can't separate them. Then there's a second tier with Grok four and GPT-five point four. And then everyone else clusters together. But here's the thing—that ranking only holds for short-answer questions. When you look at the essays, almost every model scores near perfect. The essays don't discriminate between models at all.
Tom: So the short-answer questions are where the models separate themselves. And what about that Polish versus global split we mentioned?
Jane: That's the most striking finding. Almost every model scores higher on global history than on Polish history. For the top two models, the gap is negligible. But for the other six, it's a real penalty, ranging from about two to almost eleven percentage points. And it causes actual rank reordering. Grok four is third overall, but it drops below GPT-five point four when you only look at Polish content.
Lu: That's a really important result, Jane. It suggests that the models' knowledge is not uniform. They have a deep, robust understanding of, say, the French Revolution, but a much shallower grasp of the specifics of Polish political history. And that's not surprising—the training data is overwhelmingly in English and focused on Anglophone or Western European topics. But it's a real limitation for anyone using these models in a Polish educational context.
Meng: And the source type matters too. The paper shows that photo and text tasks are the easiest for the models, probably because there's redundant information. But text-only tasks show the widest variance between models. Gemini two point five Pro, which is a strong model overall, completely collapses on text-only tasks, scoring the lowest of any model on any source type.
Tom: That's wild. So a model can be great at one thing and terrible at another, and you'd never know it from the aggregate score. What about the hardest questions? What did those look like?
Jane: They were almost all about temporal reasoning or culturally specific Polish content from the twentieth century. There's one question that really stands out. It asks the models to determine which of three documents was created first, based on the content and their own knowledge. The models could recognize the documents, but they couldn't order them chronologically. Gemini and Grok four got it right every time. GPT-five point four got it right once. And the other five models scored zero on all three attempts.
Tom: So they know what the documents are, but they don't know when they happened. That's the temporal disorientation the paper talks about.
Lu: Exactly. And there's a fascinating inversion here. The hardest questions for human examinees, according to the official reports, are the photo and text combinations. But that's precisely the question type where the models perform most consistently. So the models and the humans are struggling with completely different things. That tells you the models aren't just doing the same task faster—they're doing a different task altogether.
Meng: And that's a problem if you're using the exam to evaluate the models. You might think they're ready for the test, but they're really just good at a different kind of test that happens to look similar.
Tom: So the models are acing the exam, but for the wrong reasons. What does the paper suggest we should do about it? That's next.
Improvements: Tom: We're back with "Large Language Models Pass the History Exam But Miss the «History»: A Polish High School Exit Exam Matura Benchmark." And we've seen the models crush the test but stumble on the reasoning. So what does this paper suggest we actually do with these findings?
Jane: Well, Tom, the paper is careful not to over-prescribe. It's primarily a benchmark, a way of measuring. But the failure modes they identify point toward what needs to improve. The source conflation problem—where models treat a document as a window onto the past rather than as an artifact of the past—that's a fundamental issue with how they process context. They need to be trained or prompted to maintain a critical distance from the source material.
Lu: And I think that's the key insight. The models lack what you might call historical stance. They can retrieve facts about the past, but they can't reason within a period. They can't put themselves in the mindset of someone living in 1950s Poland and understand why a particular propaganda poster would have been effective. They're always looking at history from the outside, as a collection of facts, rather than from the inside, as a set of lived experiences and constraints.
Meng: From a practical engineering standpoint, the paper also highlights the instability of the rankings. A model that's best overall might be worst on a specific task type. That means you can't just pick the top model and assume it's the right tool for every job. You need to match the model to the task. And the paper's data gives you a way to do that.
Tom: So instead of one model to rule them all, you might have a portfolio of models, each specialized for a different kind of question?
Meng: Exactly. And that's actually feasible with the API-based approach they used. You can route each question to the model that's most likely to get it right.
Jane: There's also a broader point about benchmarks themselves. The paper argues that nationally grounded, human-referenced benchmarks are necessary. Synthetic tests, where researchers write their own questions, just don't capture the same thing. The Matura is a real exam, with real stakes, and it's been refined over decades. Using it as a benchmark gives you a level of ecological validity that you just can't get from a made-up question set.
Lu: And that's why I think this paper is going to be influential beyond just the Polish context. It provides a template for how to build similar benchmarks in other countries and other languages. We need more of these, not just for history, but for any subject where interpretative reasoning matters. Literature, philosophy, law—these are all domains where the models might be passing the test while missing the point.
Tom: So the improvement isn't just about making the models better at history. It's about building better benchmarks to find out where the models are actually weak.
Jane: Right. And the paper also notes some limitations. They only used three years of exams, which limits the statistical power. They only tested closed-source models, so we can't see what's happening under the hood. And they didn't include a Polish monolingual model, because the best one, Bielik, doesn't support image inputs. So there's a clear path for future work.
Tom: So we've got a benchmark, we've got failure modes, we've got a path forward. What's the big picture here? That's our final segment.
Conclusion: Tom: So we've spent the show on "Large Language Models Pass the History Exam But Miss the «History»: A Polish High School Exit Exam Matura Benchmark." And I think the title really does capture the whole thing. These models are incredibly capable, but they're not thinking about history the way a historian does.
Jane: That's the core message, Tom. The models dramatically outperform human examinees, but the aggregate scores hide a lot of instability. Rankings change depending on the task type, the source material, and whether the content is Polish or global. And the two failure modes—source conflation and temporal disorientation—show that the models are missing the fundamental stance of historical inquiry.
Lu: And that's the part that matters most, I think. Passing an exam is not the same as understanding the subject. The models have learned to produce answers that look correct, but they haven't learned to reason about the past in a way that's historically situated. They're not placing themselves within the period they're analyzing.
Meng: From my perspective, the practical takeaway is that we need to be careful about how we use these models in education. They're great tools, but they're not reliable for every kind of question. The paper gives us a map of where they're strong and where they're weak, and we should use that map to guide our expectations.
Tom: And for the students listening, what should they take away?
Jane: I'd say this: use these tools to help you organize facts and get a broad overview, but don't trust them for the kind of deep, contextual reasoning that a history exam really tests. They might get the score, but they might not get the history. And if you're studying for a test, you want to understand the subject, not just mimic the answers.
Tom: That's a perfect way to put it. So we're saying goodbye to this paper, but we're taking its lessons with us. The benchmark is a reminder that we need to look beyond the scores and ask what the models are actually doing.
Jane: And that's a question worth asking for every subject, not just history. Thanks for joining us, everyone. We'll be back soon with another paper, and we'll see if we can't find the history in that one too.
Tom: Take care, and keep questioning.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization