Large Language Models Pass the History Exam But Miss the <<History>>: A Polish High School Exit Exam Matura Benchmark
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Large Language Models Pass the History Exam But Miss the «History»: A Polish High School Exit Exam Matura Benchmark".
Jane: The paper was written by Adrian Trzoss, Kacper Dudzic, Wiktor Werner and Marcin Moskalewicz from Adam Mickiewicz University and IDEAS Research Institute and AMU Center for Artificial Intelligence and Poznan University of Medical Sciences and Maria Curie-Sklodowska University and WSB Merito University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper with a title that really made me stop and think: "Large Language Models Pass the History Exam But Miss the «History»: A Polish High School Exit Exam Matura Benchmark." Jane, that title is doing a lot of work.
Jane: It really is, Tom. And honestly, it's a bit of a gut punch, isn't it? The researchers took eight leading AI models and had them sit the actual Polish high school history exit exam, the Matura, from two thousand twenty-three to two thousand twenty-five. And every single model absolutely crushed the human average. We're talking scores in the high 80s and 90s, while the human examinees averaged around forty-four percent.
Tom: So they pass with flying colors, but the title says they miss the history. What's the catch?
Jane: The catch is in the details. The aggregate scores look amazing, but when you break them down, the rankings are all over the place. A model that's third overall might be first on one type of question and sixth on another. And there's a consistent penalty on Polish history versus global history. The models just aren't as good when the content is culturally specific to Poland.
Tom: Right, so it's not just about knowing facts. It's about the kind of reasoning that's being tested. Lu, you've been nodding along—what's your take on that title?
Lu: I think the title is actually a very precise diagnosis. The paper identifies two specific failure modes that explain the "missing the history" part. One is what they call source conflation, where the model treats a historical document as a source of facts about the world, rather than as an object to be analyzed in its own right. The other is temporal disorientation, where the model gets the right actors but places them in the wrong period. So they're not just missing a few facts—they're missing the fundamental stance of historical inquiry.
Jane: And that's the part that scares me a little. Because a student who scores forty-four percent on this exam clearly doesn't know the material. But a model that scores ninety-six percent might be confidently wrong in ways that are much harder to spot.
Tom: Meng, from a practical standpoint, what does this mean for the student who's using one of these chatbots to study for their history final?
Meng: It means they should be careful. The paper shows that on the hardest questions, the models often fail completely, scoring zero across all three runs. And those hardest questions aren't the ones you'd expect. For humans, the hardest questions involve photo and text combinations. For the models, it's the temporal reasoning questions that trip them up. So a student asking for help on a question about sequencing events might get a very confident, very wrong answer.
Tom: So the models are like that student who memorized the textbook but doesn't really get the subject. They can ace the multiple choice but stumble on the essay that asks them to actually think.
Jane: Exactly. And that's why this benchmark is so valuable. It's not a synthetic test designed by researchers. It's a real exam, with real grading criteria, taken by real students. And it's testing a skill that most benchmarks ignore: interpretative reasoning about the past. That's a big deal.
Lu: And it's not just about history. This kind of benchmark, grounded in a national curriculum, can tell us a lot about how these models handle culturally embedded knowledge. If they struggle with Polish history, they're probably struggling with other non-Anglophone content too. That's a significant blind spot.
Tom: So we've got models that ace the test but miss the point. What exactly did the researchers find when they dug into the scores? That's coming up next.
Summary: Tom: So we're back with "Large Language Models Pass the History Exam But Miss the «History»: A Polish High School Exit Exam Matura Benchmark." And Jane, we've established that the models score incredibly high overall. But I want to dig into that score breakdown, because that's where the real story is.
Jane: Absolutely, Tom. The paper creates a three-tier ranking. At the top, you've got Claude Sonnet four point six and Gemini three point one Pro, and their scores are so close that statistically you can't separate them. Then there's a second tier with Grok four and GPT-five point four. And then everyone else clusters together. But here's the thing—that ranking only holds for short-answer questions. When you look at the essays, almost every model scores near perfect. The essays don't discriminate between models at all.
Tom: So the short-answer questions are where the models separate themselves. And what about that Polish versus global split we mentioned?
Jane: That's the most striking finding. Almost every model scores higher on global history than on Polish history. For the top two models, the gap is negligible. But for the other six, it's a real penalty, ranging from about two to almost eleven percentage points. And it causes actual rank reordering. Grok four is third overall, but it drops below GPT-five point four when you only look at Polish content.
Lu: That's a really important result, Jane. It suggests that the models' knowledge is not uniform. They have a deep, robust understanding of, say, the French Revolution, but a much shallower grasp of the specifics of Polish political history. And that's not surprising—the training data is overwhelmingly in English and focused on Anglophone or Western European topics. But it's a real limitation for anyone using these models in a Polish educational context.
Meng: And the source type matters too. The paper shows that photo and text tasks are the easiest for the models, probably because there's redundant information. But text-only tasks show the widest variance between models. Gemini two point five Pro, which is a strong model overall, completely collapses on text-only tasks, scoring the lowest of any model on any source type.
Tom: That's wild. So a model can be great at one thing and terrible at another, and you'd never know it from the aggregate score. What about the hardest questions? What did those look like?
Jane: They were almost all about temporal reasoning or culturally specific Polish content from the twentieth century. There's one question that really stands out. It asks the models to determine which of three documents was created first, based on the content and their own knowledge. The models could recognize the documents, but they couldn't order them chronologically. Gemini and Grok four got it right every time. GPT-five point four got it right once. And the other five models scored zero on all three attempts.
Tom: So they know what the documents are, but they don't know when they happened. That's the temporal disorientation the paper talks about.
Lu: Exactly. And there's a fascinating inversion here. The hardest questions for human examinees, according to the official reports, are the photo and text combinations. But that's precisely the question type where the models perform most consistently. So the models and the humans are struggling with completely different things. That tells you the models aren't just doing the same task faster—they're doing a different task altogether.
Meng: And that's a problem if you're using the exam to evaluate the models. You might think they're ready for the test, but they're really just good at a different kind of test that happens to look similar.
Tom: So the models are acing the exam, but for the wrong reasons. What does the paper suggest we should do about it? That's next.
Improvements: Tom: We're back with "Large Language Models Pass the History Exam But Miss the «History»: A Polish High School Exit Exam Matura Benchmark." And we've seen the models crush the test but stumble on the reasoning. So what does this paper suggest we actually do with these findings?
Jane: Well, Tom, the paper is careful not to over-prescribe. It's primarily a benchmark, a way of measuring. But the failure modes they identify point toward what needs to improve. The source conflation problem—where models treat a document as a window onto the past rather than as an artifact of the past—that's a fundamental issue with how they process context. They need to be trained or prompted to maintain a critical distance from the source material.
Lu: And I think that's the key insight. The models lack what you might call historical stance. They can retrieve facts about the past, but they can't reason within a period. They can't put themselves in the mindset of someone living in 1950s Poland and understand why a particular propaganda poster would have been effective. They're always looking at history from the outside, as a collection of facts, rather than from the inside, as a set of lived experiences and constraints.
Meng: From a practical engineering standpoint, the paper also highlights the instability of the rankings. A model that's best overall might be worst on a specific task type. That means you can't just pick the top model and assume it's the right tool for every job. You need to match the model to the task. And the paper's data gives you a way to do that.
Tom: So instead of one model to rule them all, you might have a portfolio of models, each specialized for a different kind of question?
Meng: Exactly. And that's actually feasible with the API-based approach they used. You can route each question to the model that's most likely to get it right.
Jane: There's also a broader point about benchmarks themselves. The paper argues that nationally grounded, human-referenced benchmarks are necessary. Synthetic tests, where researchers write their own questions, just don't capture the same thing. The Matura is a real exam, with real stakes, and it's been refined over decades. Using it as a benchmark gives you a level of ecological validity that you just can't get from a made-up question set.
Lu: And that's why I think this paper is going to be influential beyond just the Polish context. It provides a template for how to build similar benchmarks in other countries and other languages. We need more of these, not just for history, but for any subject where interpretative reasoning matters. Literature, philosophy, law—these are all domains where the models might be passing the test while missing the point.
Tom: So the improvement isn't just about making the models better at history. It's about building better benchmarks to find out where the models are actually weak.
Jane: Right. And the paper also notes some limitations. They only used three years of exams, which limits the statistical power. They only tested closed-source models, so we can't see what's happening under the hood. And they didn't include a Polish monolingual model, because the best one, Bielik, doesn't support image inputs. So there's a clear path for future work.
Tom: So we've got a benchmark, we've got failure modes, we've got a path forward. What's the big picture here? That's our final segment.
Conclusion: Tom: So we've spent the show on "Large Language Models Pass the History Exam But Miss the «History»: A Polish High School Exit Exam Matura Benchmark." And I think the title really does capture the whole thing. These models are incredibly capable, but they're not thinking about history the way a historian does.
Jane: That's the core message, Tom. The models dramatically outperform human examinees, but the aggregate scores hide a lot of instability. Rankings change depending on the task type, the source material, and whether the content is Polish or global. And the two failure modes—source conflation and temporal disorientation—show that the models are missing the fundamental stance of historical inquiry.
Lu: And that's the part that matters most, I think. Passing an exam is not the same as understanding the subject. The models have learned to produce answers that look correct, but they haven't learned to reason about the past in a way that's historically situated. They're not placing themselves within the period they're analyzing.
Meng: From my perspective, the practical takeaway is that we need to be careful about how we use these models in education. They're great tools, but they're not reliable for every kind of question. The paper gives us a map of where they're strong and where they're weak, and we should use that map to guide our expectations.
Tom: And for the students listening, what should they take away?
Jane: I'd say this: use these tools to help you organize facts and get a broad overview, but don't trust them for the kind of deep, contextual reasoning that a history exam really tests. They might get the score, but they might not get the history. And if you're studying for a test, you want to understand the subject, not just mimic the answers.
Tom: That's a perfect way to put it. So we're saying goodbye to this paper, but we're taking its lessons with us. The benchmark is a reminder that we need to look beyond the scores and ask what the models are actually doing.
Jane: And that's a question worth asking for every subject, not just history. Thanks for joining us, everyone. We'll be back soon with another paper, and we'll see if we can't find the history in that one too.
Tom: Take care, and keep questioning.
Adrian Trzoss, Kacper Dudzic, Wiktor Werner, Marcin Moskalewicz
Adam Mickiewicz University · IDEAS Research Institute · AMU Center for Artificial Intelligence · Poznan University of Medical Sciences · Maria Curie-Sklodowska University · WSB Merito University
cs.CL
Submitted: 2026-08-14
Updated: 2026-08-17
Code: https://github.com/kdudzic/history-matura-llm-evaluation
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 63/100
The gist: This study introduces the first LLM history benchmark grounded in Polish national curriculum, evaluating eight leading LLMs on the Polish high school exit exams (Matura) in history—three official
Key concepts
- Matura Benchmark
- This is a real, nationally grounded high school exit exam used as a test for AI models. It provides 'ecological validity' because it uses actual grading criteria and real student questions, unlike synthetic tests that are made up by researchers.
- Source Conflation
- This failure mode occurs when the model treats historical documents as factual sources about the world, rather than analyzing them as artifacts of a specific object or time. It is a fundamental issue with how models process context.
- Temporal Disorientation
- This is when the model correctly identifies actors in history but places them in the wrong chronological period. It shows they struggle with sequencing events and understanding the correct timeline of historical inquiry.
- Cultural Specificity
- The models consistently score lower on content specific to Polish history compared to global history. This penalty suggests a lack of robust knowledge regarding non-Anglophone or Western European topics.
Terminology
Summary
This study introduces the first LLM history benchmark grounded in Polish national curriculum, evaluating eight leading LLMs on the Polish high school exit exams (Matura) in history—three official papers from 2023–2025, comprising short-answer questions and extended essays—comparing model performance against the human examinee population. Every model dramatically outperforms human examinees, yet aggregate scores mask distinct competency profiles: rankings are unstable across task type, source modality, and geographical scope, with a consistent penalty on Polish versus Global history content. Qualitative error analysis reveals two recurring failure modes—source conflation, in which models reason from source content rather than treating it as an object of analysis, and temporal disorientation, in which responses are historically misplaced.
The benchmark consists of three official history Matura papers from 2023, 2024, and 2025. Each contains tasks—thematic units built around source materials—where each may contain one or more questions. The 2023 paper included 36 questions across 25 tasks; the 2024 paper, 39 questions across 25 tasks; and the 2025 paper, 37 questions across 24 tasks. The final task in each paper required an essay of at least 600 words on one of three proposed topics. All papers covered a broad chronological range from antiquity to the late twentieth century, with roughly equal coverage of Polish and Global history. Short-answer questions varied widely in format: single-word or single-phrase responses (naming a figure or event), multiple-choice items, binary judgment tasks (true/false) with brief justificatory reasoning, and source-analysis exercises (3–5 sentences). Each task—except the essay—included source materials, which appeared in varying combinations: historical texts, iconographic materials (maps, photographs), and tables (economic or genealogical). Most questions were worth one or two points; the maximum total score per paper was 60 points, with 15 allocated to the essay. Responses were scored according to official CKE grading guidelines and cross-checked by a high-level human expert—a CKE-trained and experienced Matura examiner.
We evaluated a total of N = 8 models from 4 leading providers: GPT-4o and GPT-5.4 from OpenAI, Claude Sonnet 3.7 and Claude Sonnet 4.6 from Anthropic, Gemini 2.5 Pro and Gemini 3.1 Pro from Google, as well as Grok 4 and Grok 4.20 from xAI. Model selection criteria encompassed public interest and general performance, the ability to handle image inputs, availability on OpenRouter, as well as availability through the web interface on free plans of the respective providers. The last criterion was methodologically motivated by the potential real-world use of LLMs by high school students. Similarly, the choice of the two models from each provider was intended to investigate any substantial performance differences between the newest available version and its older equivalent.
Model outputs were obtained via the OpenRouter API. Each question was passed in a separate API call. The question text was passed along with the prompt (all in Polish), whereas the image inputs were included in an additional API call payload. Each short-answer question was passed to each model 3 separate times with the default temperature setting to evaluate potential inconsistencies arising from the non-deterministic nature of the models. Similarly, each essay topic was passed 3 times, with all possible topics from a sheet evaluated for each model (in contrast to a student who must choose one); this adjustment allowed for a more comprehensive evaluation of model performance on the essay section—topics concern various time periods and problems, and a model is not guaranteed to perform equally well on each. In total, 2688 API calls for short answer questions and 216 for essays have been made. No in-context learning paradigm was employed in the inference protocol.
Each question was manually annotated along three dimensions: geographical scope (Polish vs. Global history); historical period (antiquity, medieval, early modern, nineteenth century, twentieth century, PRL-communist Poland post-1945); and source material type (text only, photo only, photo and text, table-based).
Three overall model rankings were computed—all tasks combined, short-answer only, and essays only—based on mean normalized scores aggregated across all years and runs, with uncertainty estimated via 95% bootstrap confidence intervals (5,000 replications). Rankings were further aggregated by each annotation category. For human vs. model comparisons, we calculated the Wasserstein distance. Because all three essay topics are evaluated for each model—compared with only one topic for human examinees—a complete model run yields a maximum of 90 points, whereas the human is 60. All comparisons, therefore, use normalized scores to account for this asymmetry.
Results reveal a three-tier structure: Claude Sonnet 4.6 and Gemini 3.1 Pro lead with overlapping confidence intervals (96.6% & 96.2%), Grok 4 and GPT-5.4 form a distinct second tier, and the remaining models cluster tightly with indistinguishable confidence intervals. Short-answer performance closely tracks overall scores, confirming it as the primary discriminator between models. Essay scores are uniformly high for seven models, offering no discriminative power across the benchmark; GPT-4o is the sole exception at 90.1%. The aggregate ranking, however, masks substantial instability across topical and source-type categories—and remains far above the human examinee average of 44.1% across all analyzed years.
Almost all models score higher on Global than on Polish history tasks, with the penalty on Polish content ranging up to 10.8 percentage points. The gap is negligible for the top two models but substantial for the remaining six, where it ranges from 2.3 to 10.8 percentage points—suggesting that weaker models are disproportionately disadvantaged by nationally specific content. The split produces non-trivial rank reordering: Grok 4, ranked third overall, scores 97.6% on Global but drops to 91.7% on Polish, falling behind GPT-5.4. The epoch breakdown mirrors this pattern, with twentieth-century and PRL-period showing the largest score variance.
Photo + Text tasks yield the highest and most consistent scores, suggesting a ceiling or redundant-cue effect. Text-only tasks show the widest cross-model variance (69.4%–96.8%), with Gemini 2.5 Pro collapsing to the lowest cell in the matrix. Gemini 3.1 Pro leads on Photo-only (99.0%), Claude Sonnet 4.6 on Table-based tasks (100%), and Grok 4 on Text-only (96.8%) despite ranking third overall. GPT-4o underperforms frontier models across all types.
The hardest and most discriminating questions for models are situated mainly in the 2024 and 2025 papers and concern either temporal-oriented reasoning or culturally specific Polish content from the twentieth century and the PRL period. Question 11 02 2025 is both in the hardest and most model-discriminating items in the benchmark: Determine which of the documents cited in fragments A–C was created first. Justify your answer by referring to the sources and your own knowledge.
Models correctly recognized sources names but failed to order them in chronological sequence. Both Gemini and Grok 4 scored max points across all runs, GPT-5.4 scored a point only in one run, while others scored 0 across all three tests. A cross-population comparison reveals an inversion: per CKE reports, human-hardest items disproportionately involve Photo + Text combinations—precisely the question type on which models perform most consistently.
Manual inspection of all model outputs reveals two recurring failure modes. The first, which we deem source conflation, occurs when a model reasons from the semantic content of a provided source rather than treating it as an object of historical analysis. In question 07 2024, models consistently inferred chronological order from content comparison rather than authorial context, effectively treating sources as evidence about the world rather than as historically situated documents. This was observed across six models on all test runs; the only exceptions were Gemini 3.1 Pro and GPT-5.4. The second, we call temporal disorientation, concerns questions requiring models to identify or sequence causally linked events, or locate them within a specific Polish historical period. On task 22 2025, models produced partially plausible responses—correctly identifying relevant actors or concepts—but placed them in the wrong period or order. This pattern was particularly pronounced on PRL-period questions (questions 23 2024, 25 2024). The essay component reveals differences in rhetorical strategies. While all models achieve broadly comparable essay scores, their argumentative strategies diverge. The majority of models align with the proposed thesis, constructing confirmatory arguments. The Grok family is the sole exception: both systematically adopt a counter-argumentative position, particularly in the 2023 exam.
This study presents the first benchmark evaluation of LLMs on the Polish history Matura. Every model dramatically outperforms human examinees, yet aggregate scores mask distinct competency profiles: rankings are unstable across task type, source modality, and geographical scope. The Polish versus Global history split produces systematic rank reordering, extending Chartier’s findings on non-Anglophone content to a Central-Eastern European context. The two identified failure modes suggest that what models lack on the hardest questions is not factual coverage but historically situated reasoning—the ability to reason within a period rather than merely about it. Essay argumentation further reveals rhetorical divergence: while most models align with the proposed thesis, Grok models systematically adopt a counter-argumentative position. Nationally grounded, human-referenced benchmarks are necessary to complement synthetic evaluation frameworks, particularly for languages and domains where global models remain undertested. After all, passing an examination is not the same as understanding its subject.
Limitations include: the benchmark comprises three examination papers, which limits statistical power and may not capture the full range of question types across years; official CKE reports provide only aggregate human score distributions rather than individual-level data, constraining the precision of human-model comparisons; all evaluated models are closed-source, precluding analysis of the behavioral and architectural factors underlying observed failure modes; potential contamination cannot be ruled out, as examination papers and official model answers are publicly available online and may appear in pretraining corpora; all models were prompted under baseline conditions without chain-of-thought or few-shot examples, meaning results reflect but one point in a broader prompting space; finally, no Polish monolingual model was included: Bielik, the most capable Polish-language model, lacks multimodal support and could not be evaluated on multimodal tasks.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system, along with what the improved system can do:
What to change: Add a dedicated reasoning layer that explicitly separates source content from source analysis. Before answering any question involving provided documents, the system must first classify the source (author, date, context) and then treat the source's claims as data to be analyzed, not as ground truth.
What the improved AI can do: When given a chronicle fragment (e.g., task 07 2024), the system will no longer infer chronological order by comparing the events described. Instead, it will identify the author, the period of authorship, and the historical context of the source itself, then answer based on that metadata. This directly addresses the 6-model failure on that task.
The improved AI system will:
-
Reason about sources as objects, not just as information carriers.
-
Maintain chronological accuracy even on complex, multi-source tasks.
-
Perform equally well on Polish and Global history content.
-
Handle all source types consistently without performance collapse.
-
Generate essays with flexible argumentative stance (aligned or counter).
-
Produce human-calibrated answers that align with partial-credit grading.
-
Correctly sequence historical documents even under ambiguity.
These improvements directly target the two identified failure modes (source conflation, temporal disorientation) and the systematic Polish-history penalty, resulting in a system that not only passes exams but demonstrates genuine historical reasoning.
Sources
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Time Awareness in Large Language Models: Benchmarking Fact Recall Across Time
- Measuring Massive Multitask Language Understanding
- LLMzSz{\L}: a comprehensive LLM benchmark for Polish
- LLMs' Reading Comprehension Is Affected by Parametric Knowledge and Struggles with Hypothetical Statements
- GPT-4o System Card
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering