From Evaluated Models to Evaluation Aids: A Multi-Evidence Study of LLM-Based Difficulty Calibration for Programming Examinations
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "From Evaluated Models to Evaluation Aids: A Multi-Evidence Study of LLM-Based Difficulty Calibration for Programming Examinations".
Jane: The paper was written by Hongfei Yan, Jiangkai Xiong, Yiqing Li and Chong Chen from School of Computer Science, Peking University and Yuanpei College, Peking University and School of Government, Beijing Normal University and National Key Laboratory for Multimedia Information Processing and Beijing Key Laboratory of AI Systems.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title and Authors: Tom: Welcome back to the arXiv channel, everyone. I'm Tom, and with me is Jane. Today we're looking at a paper with a title that really makes you stop and think: "From Evaluated Models to Evaluation Aids: A Multi-Evidence Study of LLM-Based Difficulty Calibration for Programming Examinations."
Jane: That title is doing a lot of work, Tom. It's basically saying we've spent years testing whether AI can write code, and now this paper flips the whole thing around and asks whether AI can help us figure out how hard exam questions are for students.
Tom: Exactly. And it comes out of Peking University, with the first author Hongfei Yan and colleagues. The core idea is that when you have multiple classes of the same programming course, each with a different teacher writing a different final exam, how do you know if one exam was harder than another? That's a fairness problem.
Jane: And the old way of doing this is just to compare average scores, but that's flawed because different cohorts of students have different abilities. So the paper proposes using large language models as a kind of reference point — a consistent judge that can look at all the exam problems and rate their difficulty.
Tom: Right. And what's clever is that they don't just ask the AI to guess difficulty. They actually had ten different AI models sit the same exam as one hundred twenty students, in real time, submitting code to the same online judge system. That's the first stage of their evidence.
Jane: And the results were striking. The AI pass rate correlated with student pass rate at a Spearman rho of zero point eight six six. That's a strong rank correlation, meaning the problems that AI found hard, students also found hard.
Tom: But here's the thing I love about this paper — they don't stop there. They know that running ten models on every exam is expensive, so they test a cheaper method: just giving the AI the problem statement and reference solution, and asking it to rate difficulty on a scale. And that also works remarkably well across seventy-nine problems from eleven different classes.
Jane: So the title really captures the shift. AI goes from being the thing we evaluate, to being a tool that helps us evaluate. That's a fundamental repositioning of what these models are for.
Tom: And it has real consequences for students. If one class gets a brutally hard exam and another gets an easy one, comparing grades across classes is meaningless. This gives teachers a way to discuss that fairly.
Jane: But we should be careful. The paper is very explicit that this is not about automatically adjusting grades. It's about giving teachers better information to have better conversations.
Tom: Right. And that's where we're headed next — the actual findings from the first stage experiment. Stay with us.
Paper Summary: Jane: So Tom, we've set the stage with the title. Now let's talk about what the paper actually found in its first big experiment. This is the synchronous exam where ten AI models and one hundred twenty students tackled the same eight programming problems.
Tom: And the headline number is that Spearman correlation of zero point eight six six between AI pass rate and student pass rate. But what's even more interesting is the composite difficulty index they built. It's not just pass rate — it also factors in how many attempts the AI needed and how close the running time came to the time limit.
Jane: Right. And that composite index correlated with student pass rate at-zero point nine zero five. So the higher the AI difficulty score, the lower the student pass rate. That's a very tight relationship.
Tom: But here's where it gets nuanced. There's one problem, I30547, that only one model could solve — ChatGPT. And the student pass rate on that problem was zero percent. So the AI scale correctly identified it as extremely hard.
Jane: But there were also problems where AI did much better than students. Like T30913, where AI pass rate was one hundred percent but only twenty-five percent of students passed. That tells us AI is not a perfect predictor of absolute difficulty — it's better at ranking problems relative to each other.
Tom: Exactly. And that's why the paper is so careful about language. They say the AI scale is good for relative difficulty ordering, not for predicting exact pass rates.
Jane: And then they take this and extend it. They use a cheaper review-based method where the AI just reads the problem and reference solution and scores it on six dimensions: concept difficulty, implementation difficulty, debugging difficulty, complexity risk, reading difficulty, and overall difficulty.
Tom: And across seventy-nine problems from eleven parallel classes, the overall difficulty score correlated with student pass rate at-zero point eight seven one. That's almost as strong as the full solving-based experiment, but at a fraction of the cost.
Jane: Which is a big deal for practical use. You can't run ten models on every exam, but you can run one review batch on every problem in a course.
Tom: And they also looked at non-attempt rate — the proportion of students who didn't even try a problem. That correlated at zero point eight zero zero with AI difficulty. So harder problems don't just have lower pass rates; students actively give up on them.
Jane: That's a really important process measure. It tells you something about time pressure and student confidence, not just raw ability.
Tom: And this is where the paper gets really interesting for me — they also found that the AI difficulty levels correspond to different error patterns. Problems rated high on complexity risk tend to have more time limit exceeded submissions.
Jane: So the AI isn't just giving a single number; it's providing a diagnostic profile. That's much more useful for teachers who want to understand why students struggle.
Tom: Exactly. And that's what we'll dig into next — the specific improvements and methods the paper proposes. Don't go anywhere.
Improvements Suggested: Tom: Welcome back. Jane and I have been talking about the core findings, but now I want to focus on what this paper actually proposes as improvements to how we handle programming exams.
Jane: And the biggest one is the idea of exposure adjustment. See, some exam problems are taken directly from the practice item bank. Students may have seen them before. So the AI might rate them as moderately hard, but students find them easier because they've practiced them.
Tom: Right. And the paper introduces an exposure discount — they reduce the effective difficulty of those exposed problems by twenty-five percent in their main analysis. But here's the clever part: they ran a sensitivity analysis testing discounts from zero to forty percent.
Jane: And what did they find? That the correlation direction doesn't change regardless of the discount. That's important because it means their conclusions aren't just an artifact of picking the right number.
Tom: But there's a surprise in there too. The correlation was actually strongest when the discount was zero — meaning no adjustment at all. That's counterintuitive. You'd think accounting for exposure would improve the fit.
Jane: And the paper explains that honestly. It says the exposure discount is more about documenting context than improving prediction. In their sample, the exposed problems weren't actually easier for students — possibly because they were concentrated in one class with other factors at play.
Tom: That's a really honest finding. A lot of papers would have just reported the adjusted numbers and moved on. This one shows you the sensitivity analysis and admits the adjustment doesn't help prediction.
Jane: And then there's the longitudinal improvement. They took the same review pipeline and applied it to twenty-six problems from four semesters of the same teacher's Data Structures and Algorithms B course. The correlations held up — -zero point eight two nine with pass rate and zero point eight eight three with non-attempt rate.
Tom: So the same ruler works across time, not just across classes in the same semester. That means teachers can track whether their exams are getting harder or easier over the years.
Jane: But then they hit a boundary. They tried the same thing on Introduction to Computing B — one hundred six problems across sixteen exams — and the problem-level correlation dropped to-zero point five five two, and the exam-level correlation nearly vanished.
Tom: And that's a really valuable negative result. It tells you the AI difficulty scale works best in algorithm-heavy courses, but in introductory courses, cohort differences dominate. The same exam can produce wildly different results depending on who's taking it.
Jane: So the improvement isn't just "use AI to rate difficulty." It's "use AI to rate difficulty, but know when it works and when it doesn't."
Tom: And that's the mark of mature research. They're not overselling their tool. They're mapping its boundaries.
Jane: Which brings us to the first page of the paper and the framing of the whole study. Let's get into that.
First Page Discussion: Jane: So Tom, let's go back to the very beginning of the paper — the abstract and introduction — because that's where they lay out the research questions and the philosophical stance.
Tom: And the philosophical stance is really important. They're drawing on Messick's validity theory and Kane's argument-based approach to assessment. That's heavy educational measurement theory, but it translates to a simple idea: a measurement tool is only valid if you can defend what you're using it for.
Jane: Right. And the paper is very clear that they're not using AI to grade students. They're using AI to provide evidence about problems and exams. That's a crucial distinction.
Tom: And they frame it as four research questions. RQ1 asks whether AI solving performance can rank problems consistently with student pass rates. RQ2 asks whether cheaper AI review can scale to more exams. RQ3 asks what explains the boundaries of the AI ruler. And RQ4 asks whether it works longitudinally.
Jane: And the first page also introduces the two forms of the AI ruler — solving-based and review-based. The solving-based one actually submits code and gets judged. The review-based one just reads the problem and gives structured scores.
Tom: And the paper argues that the solving-based evidence is the precondition for trusting the review-based extension. If AI can't rank problems by actually solving them, why would we trust its reading-based judgments?
Jane: That's a logical chain. And it's why they ran the synchronous exam first — to establish that baseline validity before scaling up.
Tom: The first page also mentions the policy context — Chinese educational reform documents calling for AI and big data in evaluation. So this isn't just academic curiosity; it's responding to a real policy push.
Jane: And that gives the paper practical weight. It's not just "here's a cool thing AI can do." It's "here's how AI can help us make fairer exams, which is a stated policy goal."
Tom: But the paper is also careful about ethics. They say the AI ruler must not be used for individual student evaluation or automatic grade adjustment. It's a discussion tool for teachers, not a decision machine.
Jane: And they're honest about the limitations. The reviewer in their main analysis runs through a third-party endpoint with a model label that can't be authenticated. So they call it an "identity-bounded exploratory ruler."
Tom: That level of transparency is rare. Most papers would just say "we used GPT-five point six" and move on. This one says "we can't actually prove what model this is, so we're telling you."
Jane: And that builds trust. When a paper tells you its weaknesses upfront, you can believe its strengths.
Tom: Absolutely. And that honesty carries through to the conclusion, where they lay out exactly what the AI ruler can and cannot do.
Conclusion: Jane: So Tom, we've covered a lot of ground on "From Evaluated Models to Evaluation Aids." Let's pull it together for our listeners.
Tom: The core message is that AI can serve as a reliable reference for ranking programming exam difficulty, but only within clearly defined boundaries. The solving-based experiment showed strong correlation with student performance, and the cheaper review-based method scaled that up to seventy-nine problems with almost the same strength.
Jane: And the longitudinal data showed it works across semesters for the same teacher, but it breaks down in introductory courses where cohort differences dominate. That's a critical boundary to know about.
Tom: The paper also introduced the exposure discount concept — accounting for whether students have seen problems before — and showed through sensitivity analysis that their conclusions don't depend on the exact discount value.
Jane: And throughout, they emphasized that this is a tool for teachers, not a replacement for them. AI provides evidence; teachers make judgments.
Tom: The ethical boundaries are clear: no individual student evaluation, no automatic grade adjustment, and full provenance tracking of AI outputs so you know exactly what you're looking at.
Jane: And that provenance requirement is actually a contribution in itself. They built a pipeline that archives prompts, schemas, raw responses, and hashes — so the AI evidence is auditable.
Tom: Which matters because AI outputs aren't deterministic. Even at temperature zero, they found jitter in repeated reviews. But the jitter didn't change the main correlations, which gives us confidence in the findings.
Jane: So what's the takeaway for our listeners? If you're teaching programming, this paper gives you a practical way to compare exam difficulty across classes and semesters. If you're doing research, it gives you a framework for validating AI-based measurement tools.
Tom: And if you're a student, it means your exam grades might be interpreted more fairly — because teachers will have better evidence about whether your exam was genuinely harder.
Jane: It's a good note to end on. Thanks for joining us, and we'll see you for the next paper.
Tom: Take care, everyone.
Hongfei Yan, Jiangkai Xiong, Yiqing Li, Chong Chen
School of Computer Science, Peking University · Yuanpei College, Peking University · School of Government, Beijing Normal University · National Key Laboratory for Multimedia Information Processing · Beijing Key Laboratory of AI Systems
cs.CY, cs.AI, cs.PL
Submitted: 2026-07-13
Updated: 2026-08-11
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 63/100
The gist: This paper investigates whether large language models (LLMs) can serve as auxiliary evidence sources for calibrating the difficulty of programming examinations in university courses, rather than
Key concepts
- AI Difficulty Calibration
- Using Large Language Models to help determine how difficult programming exam questions are. This involves having AI models assess problems based on various criteria, such as solving performance or reading problem statements.
- Exposure Discount
- A concept where the effective difficulty of an exam problem is reduced if students have previously seen it in practice materials. The paper tested this adjustment to see how it affects the correlation between AI difficulty scores and student pass rates.
- Solving-Based vs. Review-Based Ruler
- Two methods for using AI to measure difficulty. The solving-based ruler requires the AI to actually solve code, while the review-based ruler has the AI read a problem and provide structured scores based on dimensions like concept difficulty.
Terminology
Summary
This paper investigates whether large language models (LLMs) can serve as auxiliary evidence sources for calibrating the difficulty of programming examinations in university courses, rather than merely being evaluated as code-generation targets. The study is motivated by the problem that "different teachers make differing judgments regarding knowledge-point coverage, algorithmic depth, implementation complexity, boundary data, problem statements, and time limits, which leaves student scores without a stable reference point in cross-class comparison, course-quality analysis, and instructional improvement."
The paper repositions LLMs from evaluation targets in code-generation benchmarks to auxiliary evidence sources for interpreting exam difficulty
and builds a multi-evidence framework combining AI evidence, aggregated student performance, item exposure, online-judge process data, and teacher interpretation.
The research is organized around four research questions: (RQ1) whether synchronous AI answering data can form a relative difficulty ordering consistent with student pass rates in a real online judge environment; (RQ2) whether auditable API structured review can reflect student performance at exam and problem levels in parallel-class scenarios; (RQ3) how item-bank exposure, item-count pressure, non-attempt rate, and error-type distribution explain the applicability boundary of the AI difficulty ruler; and (RQ4) whether the auditable API review ruler supports longitudinal comparison of exam difficulty and problem risk across semesters and course levels.
The study uses four types of data. The first is a synchronous computer-based examination experiment in which 10 large language models and students answered the same exam problem set, with all code judged via OJ feedback.
The exam contains 8 programming problems covering array processing, sorting, simulation, graph theory, dynamic programming, data-structure maintenance, and combinatorial counting, with 120 students as statistical objects. The second dataset comes from the final computer-based examinations of 11 parallel classes in the same semester: a total of 79 programming problems,
with each exam having problem statements, reference answers, and class-wide whole-exam score distributions. The third dataset comprises 26 problems across 4 final computer-based examinations of the same teacher's Data Structures and Algorithms B
(Spring 2024, Spring 2025, Fall 2025, and Spring 2026). The fourth dataset consists of 106 problems across 16 past final computer-based examinations of the same teacher's Introduction to Computing B
(Fall 2008 through Fall 2025).
The paper distinguishes two forms of the AI ruler. The first is the solving-based AI ruler,
where the AI reads the problem statement, generates code, submits it to the OJ, and performs a limited number of revision rounds based on judging feedback,
producing process data such as pass rate, number of attempts, running time, and error type. The second is the review-based AI ruler,
where "the AI does not submit code but, based on the problem statement, constraints, and reference answer, provides structured scores for concept difficulty, implementation difficulty, debugging difficulty, complexity risk, statement-reading difficulty, and overall difficulty" on a 1-to-5 scale.
For the solving-based experiment, the paper constructs a composite difficulty index: dp = α(1 − ΣAC/M) + β·min(1, Attempts/5) + γ·min(1, Time/TL), with α=0.6, β=0.2, γ=0.2. The weights are justified by the pass rate being the most direct proxy indicator of difficulty
while the number of attempts reflects the process cost of solving, and the running-time ratio reflects computational-efficiency pressure.
For the review-based calibration, the paper uses a single-reviewer design
whose historical selection basis comes from the first-stage solving experiment: among the 10 participating models, ChatGPT was the only model that passed all 8 problems (AK) and cracked the extremely difficult problem I30547.
However, the paper explicitly discloses that "the review batch actually used in this paper is invoked through a third-party OpenAI-compatible endpoint, and the model identifier in both the request and the endpoint's return is gpt-5.6-sol; this label cannot certify that the upstream is an official OpenAI model. The three review batches (79 problems, 26 problems, 106 problems) were run on July 10, 2026
under a unified protocol: temperature=0, reasoning effort=high, and JSON-schema-constrained output, with complete provenance metadata including
the model identifier, invocation date, endpoint, temperature, the SHA-256 of the prompt and schema, the raw API response, and run metadata."
The paper also introduces an item-bank exposure variable, distinguishing intrinsic difficulty
from effective exam difficulty.
For the 12 practice item-bank original problems, the paper adopts a conservative discount discount p=0.25,
with sensitivity analysis across discount levels 0, 0.10, 0.25, and 0.40.
The first-stage results show that "the Spearman rank correlation between AI pass rate and student problem pass rate is 0.866, with an exact two-sided permutation p=0.0119; the Spearman rank correlation between the composite difficulty index dp and student problem pass rate is-0.905, with an exact two-sided permutation p=0.0046. The paper notes that
the composite index exhibits stronger ranking consistency than pass rate alone, indicating that the number of attempts and time cost during solving carry additional explanatory value. The AI solving performance formed
a clear gradient: basic problems were passed by all or nearly all models, while
the extremely hard problem was passed by only 1 model (ChatGPT, which solved all 8 problems). The paper also notes absolute-level differences:
on medium-high-difficulty problems such as T30913, M30947, and U30919, the AI pass rate is markedly higher than the student pass rate, indicating that
the AI scale is suitable for relative difficulty ranking and problem-setting review, but not for directly predicting the proportion of students who pass."
The cross-sectional review-based calibration results show that the 11 parallel-class exams differ in item count, high-difficulty ratio, and item-bank exposure.
The overall difficulty of the 79 problems covers levels 1 to 5 (13/26/25/10/5 problems). At the problem level, "the Spearman correlation between raw AI overall difficulty and student pass rate is-0.871 (p<0.001, Fisher-approximate 95% CI [-0.916, -0.805]), and-0.857 after exposure adjustment. After removing the 12 item-bank original problems,
the Spearman correlation between raw AI difficulty and pass rate is still-0.866." The correlation between raw AI difficulty and non-attempt rate is 0.800. At the exam level, the correlations are weaker: across all 11 exams, the Spearman correlation between adjusted AI difficulty and normalized student mean is-0.461 (Pearson r=-0.598), which can only be interpreted as exploratory evidence.
The sensitivity analysis of the exposure discount shows that the direction and strength of the problem-level correlation are generally stable: across the four discount levels, the Spearman ρ between adjusted AI difficulty and pass rate ranges from-0.780 to-0.871.
Notably, "the correlation is strongest when the discount is 0 (i.e., no adjustment), and it weakens monotonically as the discount increases; this is contrary to the assumption that 'the exposure discount improves explanatory power.' The paper concludes that
the exposure discount can therefore only serve as a contextualizing explanatory parameter, not a causal estimate of student performance."
The error-type analysis shows that among the 79 problems, the dominant error type is Wrong Answer for 65 problems, Runtime Error for 8 problems, Time Limit Exceeded for 4 problems, and Compile Error for 2 problems.
The Spearman correlation between the AI complexity-risk score and the student TLE-submission share is 0.350, the strongest relationship with TLE among the six review dimensions,
though its strength is only moderate.
The longitudinal Data Structures and Algorithms B results show that the mean AI difficulty rises semester by semester from 2.000 in Spring 2024 to 3.375 in Spring 2026, and the high-difficulty ratio rises from 0 to 0.500.
On the item-level data of 26 items, overall AI difficulty is strongly negatively correlated with student pass rate (Pearson r=-0.795, Spearman ρ=-0.829) and strongly positively correlated with the non-attempt rate (Pearson r=0.848, Spearman ρ=0.883).
The Introduction to Computing B sample marks the boundary of the scale: the overall difficulty of the 106 items covers levels 1 to 5 (8/41/31/24/2 items), and 26 items are rated as high-difficulty,
but "at the item level the Pearson/Spearman between overall AI difficulty and student pass rate is-0.550/-0.552, and with the non-attempt rate is 0.541/0.564—the correlation direction is correct and the strength is moderate, clearly weaker than the-0.86/-0.83 magnitude of the two Data Structures and Algorithms B samples. At the exam-paper level across the 16 exams,
the Spearman correlation between mean AI difficulty and mean student pass rate is only-0.063 (Pearson r=-0.168), close to zero. The paper explains that
cohort differences rather than item difficulty dominate exam-paper-level performance," citing the contrast between domestic and international student papers of the same year.
The paper presents four illustrative cases: Class 16 shows high AI difficulty and low student performance
with a normalized mean student performance of 0.191 and low-performance rate of 0.636; Class 15 shows that a paper with the highest AI difficulty (3.375) can still maintain a relatively high student completion level
(mean 3.817 problems passed, low-performance rate 0.067) due to structural division of labor between 'basic coverage' and 'discriminating top students'
; Class 9 with 13 items shows that item count, the switching cost between items, the reading burden, and the pressure of time allocation themselves constitute sources of difficulty at the exam-paper level
; and the 12 item-bank original items across three papers illustrate that whether an item has been seen before and whether it comes from an in-course practice system will change the effective difficulty and the evaluative meaning of the paper.
The paper's conclusions emphasize that "the AI difficulty ruler is only suitable for serving item-setting review, exam-structure analysis, problem-level risk diagnosis, parallel-class fairness discussion, longitudinal quality tracking, and benchmark construction; it cannot directly predict individual student performance, cannot replace teachers' professional judgment, and still less can automatically trigger score adjustment. The validity of the evidence is bounded by three source limits:
the review is completed via a third-party OpenAI-compatible endpoint, and the model label cannot authenticate the upstream as an official OpenAI model; the reviewer is a single reviewer, and there is as yet no source-complete multi-reviewer reliability; the repeated-problem test shows that temperature=0 does not constitute a guarantee of output determinism, but the scoring jitter does not change the main correlations."
The paper also discusses ethical boundaries and governance requirements, including that AI results can only be used for auxiliary analysis at the problem and exam levels, and cannot be directly used to evaluate individual student ability,
that student data should be limited to anonymized group statistics,
that the model version, prompt version, review date, endpoint, temperature, prompt hash, and scoring schema must be recorded together with the results,
that the AI ruler should not automatically trigger score adjustment,
and that students should enjoy the right to be informed regarding AI-assisted evaluation.
Limitations acknowledged include: the solving-based experiment has only 8 problems and 10 AI models; the review-based evidence comes from a single batch run of a single reviewer via a third-party endpoint; the first author's insider expert knowledge as the instructor of Class 15 may bring item-setter perspective bias; the nested structure of 79 problems within 11 exams limits statistical inference; the analysis lacks finer process indicators such as student ability stratification and complete submission sequences; the whole-exam student performance mixes two statistical calibers; the exposure discount is a conservative empirical setting with exposed problems concentrated in Class 8; there are slight differences between ranking-page statistics and source-document aggregated buckets; and there is a potential training-data contamination problem, though the paper notes mitigating factors including that "the 10 models in the first stage answered in the same environment, and if ChatGPT's memory advantage were caused by contamination, it would be hard to explain why the other models did not reach the same level on all problems."
Future research directions include expanding problem-level student process data, introducing multiple AI reviewers and teacher-expert scoring, expanding to multi-semester multi-course multi-teacher samples with hierarchical or IRT models, and building an AI-calibration benchmark for educational-scenario programming exams
that incorporates parallel-class exam problems, longitudinal exam problems, problem-exposure annotations, AI answering-process data, AI review data, student-group statistics, online-judge process data, and external code-benchmark anchors.
Improvements for AI systems
Based on the paper, here are specific improvements that can be made to AI systems, along with what the improved systems can do:
-
Improvement: Train or fine-tune models to output structured, multi-dimensional difficulty scores (concept, implementation, debugging, complexity risk, reading difficulty, overall) for programming problems, not just solve them.
-
What the improved system can do: Given a problem statement, constraints, and reference solution, it can produce a 1–5 scale difficulty profile that correlates with real student pass rates (Spearman ρ = −0.871) and non-attempt rates (ρ = 0.800) at the problem level, enabling teachers to pre-screen exam items.
-
Improvement: Extend the evaluation loop so that after generating code, the system records: pass/fail per problem, number of repair attempts, and runtime-to-time-limit ratio. Then compute a composite difficulty index (e.g., α=0.6 pass rate, β=0.2 attempts, γ=0.2 time ratio).
-
What the improved system can do: Produce a relative difficulty ordering of exam problems that matches student performance (Spearman ρ = −0.905), even when pass rates alone are not discriminative (e.g., problems with high pass rate but near-time-limit runtime).
-
Improvement: Integrate an item-bank exposure flag into the review input, and apply a configurable discount (0.00–0.40) to the raw difficulty score to estimate effective exam difficulty.
-
What the improved system can do: Distinguish intrinsic difficulty from effective difficulty in contexts where students have seen the problem before, and flag cases where exposure does not reduce difficulty (as observed in this study), preventing over-adjustment.
-
Improvement: Wrap the review call so that every request stores: endpoint, temperature, prompt/schema SHA-256 hashes, raw response, run metadata, and model identifier. Add a validation step that cross-checks normalized output against raw responses and hashes.
-
What the improved system can do: Provide fully auditable, reproducible difficulty reviews that can be verified by third parties, and detect when a script or non-AI process is mislabeled as AI output (a failure mode identified in the paper).
-
Improvement: Train the reviewer to predict, for each problem, the dominant student error type (WA, TLE, RE, CE) and the expected proportion of erroneous submissions, based on the problem’s complexity risk and implementation difficulty dimensions.
-
What the improved system can do: Alert teachers to specific risk mechanisms (e.g., high TLE risk for naive algorithms, high RE risk for boundary handling) before the exam, enabling targeted teaching interventions.
-
Improvement: Implement a warning system that detects when the AI difficulty scale is being used at the exam level in courses where cohort composition dominates performance (e.g., introductory courses with international vs. domestic students).
-
What the improved system can do: Automatically flag when exam-level difficulty comparisons are invalid (e.g., near-zero correlation with student performance), preventing automatic grade adjustments or unfair cross-class comparisons.
-
Improvement: When the same problem statement and reference code are submitted multiple times, the system should report the variance in overall difficulty scores (e.g., ±1 level on 5 of 13 duplicate groups in this study).
-
What the improved system can do: Quantify its own output instability, allowing users to treat single-reviewer scores as exploratory rather than deterministic, and to compute confidence intervals for difficulty estimates.
-
Improvement: Extend the review pipeline to accept student stratification data (e.g., by prior OJ experience, course-selection background) and output difficulty scores per subgroup.
-
What the improved system can do: Detect when the same problem presents different difficulty to different student groups, enabling discussion of within-class fairness—a capability the paper explicitly identifies as missing.
The improved system can serve as a source-auditable, problem-level difficulty calibration tool that:
-
Pre-screens exam items with multi-dimensional difficulty scores that correlate with real student performance.
-
Provides a solving-based difficulty index that captures pass rate, attempt cost, and time pressure.
-
Adjusts for item-bank exposure with a configurable discount.
-
Maintains full provenance for every review call, preventing misattribution.
-
Predicts error-type mechanisms and flags complexity risks.
-
Detects when its own scale is invalid at the exam level (e.g., in introductory courses).
-
Reports its own output instability on duplicate inputs.
-
Can be extended to subgroup-level analysis for within-class fairness.
It cannot, and should not, be used to evaluate individual students, automatically adjust grades, or replace teacher judgment.
Sources
Related papers
- Reasoning Enhances Robustness to Prompt Injection in LLM-Based Consensus
- Generative AI Purpose-built for Social and Mental Health: A Real-World Pilot
- PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks
- What is an intelligent system?
- AI University: An LLM-Powered Learning Assistant for Engineering---A Finite Element Method Case Study
- Generative AI Use in Entrepreneurship: An Integrative Review and an Empowerment-Entrapment Framework