From Evaluated Models to Evaluation Aids: A Multi-Evidence Study of LLM-Based Difficulty Calibration for Programming Examinations

summary

Video file (mp4)

The gist

This paper investigates whether large language models (LLMs) can serve as auxiliary evidence sources for calibrating the difficulty of programming examinations in university courses, rather than

In short

The episode discusses a paper using Large Language Models to calibrate programming exam difficulty. The study showed AI can rank problem difficulty, correlating with student pass rates, using both solving-based and cheaper review-based methods. Key findings include the need for context like exposure adjustment and identifying course boundaries where the AI tool is most effective.

Key concepts

AI Difficulty Calibration
Using Large Language Models to help determine how difficult programming exam questions are. This involves having AI models assess problems based on various criteria, such as solving performance or reading problem statements.
Exposure Discount
A concept where the effective difficulty of an exam problem is reduced if students have previously seen it in practice materials. The paper tested this adjustment to see how it affects the correlation between AI difficulty scores and student pass rates.
Solving-Based vs. Review-Based Ruler
Two methods for using AI to measure difficulty. The solving-based ruler requires the AI to actually solve code, while the review-based ruler has the AI read a problem and provide structured scores based on dimensions like concept difficulty.

Terminology used across episodes

This episode discusses

The paper

From Evaluated Models to Evaluation Aids: A Multi-Evidence Study of LLM-Based Difficulty Calibration for Programming Examinations · Read on arXiv

Hongfei Yan, Jiangkai Xiong, Yiqing Li, Chong Chen

School of Computer Science, Peking University · Yuanpei College, Peking University · School of Government, Beijing Normal University · National Key Laboratory for Multimedia Information Processing · Beijing Key Laboratory of AI Systems

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "From Evaluated Models to Evaluation Aids: A Multi-Evidence Study of LLM-Based Difficulty Calibration for Programming Examinations".

Jane: The paper was written by Hongfei Yan, Jiangkai Xiong, Yiqing Li and Chong Chen from School of Computer Science, Peking University and Yuanpei College, Peking University and School of Government, Beijing Normal University and National Key Laboratory for Multimedia Information Processing and Beijing Key Laboratory of AI Systems.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title and Authors: Tom: Welcome back to the arXiv channel, everyone. I'm Tom, and with me is Jane. Today we're looking at a paper with a title that really makes you stop and think: "From Evaluated Models to Evaluation Aids: A Multi-Evidence Study of LLM-Based Difficulty Calibration for Programming Examinations."

Jane: That title is doing a lot of work, Tom. It's basically saying we've spent years testing whether AI can write code, and now this paper flips the whole thing around and asks whether AI can help us figure out how hard exam questions are for students.

Tom: Exactly. And it comes out of Peking University, with the first author Hongfei Yan and colleagues. The core idea is that when you have multiple classes of the same programming course, each with a different teacher writing a different final exam, how do you know if one exam was harder than another? That's a fairness problem.

Jane: And the old way of doing this is just to compare average scores, but that's flawed because different cohorts of students have different abilities. So the paper proposes using large language models as a kind of reference point — a consistent judge that can look at all the exam problems and rate their difficulty.

Tom: Right. And what's clever is that they don't just ask the AI to guess difficulty. They actually had ten different AI models sit the same exam as one hundred twenty students, in real time, submitting code to the same online judge system. That's the first stage of their evidence.

Jane: And the results were striking. The AI pass rate correlated with student pass rate at a Spearman rho of zero point eight six six. That's a strong rank correlation, meaning the problems that AI found hard, students also found hard.

Tom: But here's the thing I love about this paper — they don't stop there. They know that running ten models on every exam is expensive, so they test a cheaper method: just giving the AI the problem statement and reference solution, and asking it to rate difficulty on a scale. And that also works remarkably well across seventy-nine problems from eleven different classes.

Jane: So the title really captures the shift. AI goes from being the thing we evaluate, to being a tool that helps us evaluate. That's a fundamental repositioning of what these models are for.

Tom: And it has real consequences for students. If one class gets a brutally hard exam and another gets an easy one, comparing grades across classes is meaningless. This gives teachers a way to discuss that fairly.

Jane: But we should be careful. The paper is very explicit that this is not about automatically adjusting grades. It's about giving teachers better information to have better conversations.

Tom: Right. And that's where we're headed next — the actual findings from the first stage experiment. Stay with us.

Paper Summary: Jane: So Tom, we've set the stage with the title. Now let's talk about what the paper actually found in its first big experiment. This is the synchronous exam where ten AI models and one hundred twenty students tackled the same eight programming problems.

Tom: And the headline number is that Spearman correlation of zero point eight six six between AI pass rate and student pass rate. But what's even more interesting is the composite difficulty index they built. It's not just pass rate — it also factors in how many attempts the AI needed and how close the running time came to the time limit.

Jane: Right. And that composite index correlated with student pass rate at-zero point nine zero five. So the higher the AI difficulty score, the lower the student pass rate. That's a very tight relationship.

Tom: But here's where it gets nuanced. There's one problem, I30547, that only one model could solve — ChatGPT. And the student pass rate on that problem was zero percent. So the AI scale correctly identified it as extremely hard.

Jane: But there were also problems where AI did much better than students. Like T30913, where AI pass rate was one hundred percent but only twenty-five percent of students passed. That tells us AI is not a perfect predictor of absolute difficulty — it's better at ranking problems relative to each other.

Tom: Exactly. And that's why the paper is so careful about language. They say the AI scale is good for relative difficulty ordering, not for predicting exact pass rates.

Jane: And then they take this and extend it. They use a cheaper review-based method where the AI just reads the problem and reference solution and scores it on six dimensions: concept difficulty, implementation difficulty, debugging difficulty, complexity risk, reading difficulty, and overall difficulty.

Tom: And across seventy-nine problems from eleven parallel classes, the overall difficulty score correlated with student pass rate at-zero point eight seven one. That's almost as strong as the full solving-based experiment, but at a fraction of the cost.

Jane: Which is a big deal for practical use. You can't run ten models on every exam, but you can run one review batch on every problem in a course.

Tom: And they also looked at non-attempt rate — the proportion of students who didn't even try a problem. That correlated at zero point eight zero zero with AI difficulty. So harder problems don't just have lower pass rates; students actively give up on them.

Jane: That's a really important process measure. It tells you something about time pressure and student confidence, not just raw ability.

Tom: And this is where the paper gets really interesting for me — they also found that the AI difficulty levels correspond to different error patterns. Problems rated high on complexity risk tend to have more time limit exceeded submissions.

Jane: So the AI isn't just giving a single number; it's providing a diagnostic profile. That's much more useful for teachers who want to understand why students struggle.

Tom: Exactly. And that's what we'll dig into next — the specific improvements and methods the paper proposes. Don't go anywhere.

Improvements Suggested: Tom: Welcome back. Jane and I have been talking about the core findings, but now I want to focus on what this paper actually proposes as improvements to how we handle programming exams.

Jane: And the biggest one is the idea of exposure adjustment. See, some exam problems are taken directly from the practice item bank. Students may have seen them before. So the AI might rate them as moderately hard, but students find them easier because they've practiced them.

Tom: Right. And the paper introduces an exposure discount — they reduce the effective difficulty of those exposed problems by twenty-five percent in their main analysis. But here's the clever part: they ran a sensitivity analysis testing discounts from zero to forty percent.

Jane: And what did they find? That the correlation direction doesn't change regardless of the discount. That's important because it means their conclusions aren't just an artifact of picking the right number.

Tom: But there's a surprise in there too. The correlation was actually strongest when the discount was zero — meaning no adjustment at all. That's counterintuitive. You'd think accounting for exposure would improve the fit.

Jane: And the paper explains that honestly. It says the exposure discount is more about documenting context than improving prediction. In their sample, the exposed problems weren't actually easier for students — possibly because they were concentrated in one class with other factors at play.

Tom: That's a really honest finding. A lot of papers would have just reported the adjusted numbers and moved on. This one shows you the sensitivity analysis and admits the adjustment doesn't help prediction.

Jane: And then there's the longitudinal improvement. They took the same review pipeline and applied it to twenty-six problems from four semesters of the same teacher's Data Structures and Algorithms B course. The correlations held up — -zero point eight two nine with pass rate and zero point eight eight three with non-attempt rate.

Tom: So the same ruler works across time, not just across classes in the same semester. That means teachers can track whether their exams are getting harder or easier over the years.

Jane: But then they hit a boundary. They tried the same thing on Introduction to Computing B — one hundred six problems across sixteen exams — and the problem-level correlation dropped to-zero point five five two, and the exam-level correlation nearly vanished.

Tom: And that's a really valuable negative result. It tells you the AI difficulty scale works best in algorithm-heavy courses, but in introductory courses, cohort differences dominate. The same exam can produce wildly different results depending on who's taking it.

Jane: So the improvement isn't just "use AI to rate difficulty." It's "use AI to rate difficulty, but know when it works and when it doesn't."

Tom: And that's the mark of mature research. They're not overselling their tool. They're mapping its boundaries.

Jane: Which brings us to the first page of the paper and the framing of the whole study. Let's get into that.

First Page Discussion: Jane: So Tom, let's go back to the very beginning of the paper — the abstract and introduction — because that's where they lay out the research questions and the philosophical stance.

Tom: And the philosophical stance is really important. They're drawing on Messick's validity theory and Kane's argument-based approach to assessment. That's heavy educational measurement theory, but it translates to a simple idea: a measurement tool is only valid if you can defend what you're using it for.

Jane: Right. And the paper is very clear that they're not using AI to grade students. They're using AI to provide evidence about problems and exams. That's a crucial distinction.

Tom: And they frame it as four research questions. RQ1 asks whether AI solving performance can rank problems consistently with student pass rates. RQ2 asks whether cheaper AI review can scale to more exams. RQ3 asks what explains the boundaries of the AI ruler. And RQ4 asks whether it works longitudinally.

Jane: And the first page also introduces the two forms of the AI ruler — solving-based and review-based. The solving-based one actually submits code and gets judged. The review-based one just reads the problem and gives structured scores.

Tom: And the paper argues that the solving-based evidence is the precondition for trusting the review-based extension. If AI can't rank problems by actually solving them, why would we trust its reading-based judgments?

Jane: That's a logical chain. And it's why they ran the synchronous exam first — to establish that baseline validity before scaling up.

Tom: The first page also mentions the policy context — Chinese educational reform documents calling for AI and big data in evaluation. So this isn't just academic curiosity; it's responding to a real policy push.

Jane: And that gives the paper practical weight. It's not just "here's a cool thing AI can do." It's "here's how AI can help us make fairer exams, which is a stated policy goal."

Tom: But the paper is also careful about ethics. They say the AI ruler must not be used for individual student evaluation or automatic grade adjustment. It's a discussion tool for teachers, not a decision machine.

Jane: And they're honest about the limitations. The reviewer in their main analysis runs through a third-party endpoint with a model label that can't be authenticated. So they call it an "identity-bounded exploratory ruler."

Tom: That level of transparency is rare. Most papers would just say "we used GPT-five point six" and move on. This one says "we can't actually prove what model this is, so we're telling you."

Jane: And that builds trust. When a paper tells you its weaknesses upfront, you can believe its strengths.

Tom: Absolutely. And that honesty carries through to the conclusion, where they lay out exactly what the AI ruler can and cannot do.

Conclusion: Jane: So Tom, we've covered a lot of ground on "From Evaluated Models to Evaluation Aids." Let's pull it together for our listeners.

Tom: The core message is that AI can serve as a reliable reference for ranking programming exam difficulty, but only within clearly defined boundaries. The solving-based experiment showed strong correlation with student performance, and the cheaper review-based method scaled that up to seventy-nine problems with almost the same strength.

Jane: And the longitudinal data showed it works across semesters for the same teacher, but it breaks down in introductory courses where cohort differences dominate. That's a critical boundary to know about.

Tom: The paper also introduced the exposure discount concept — accounting for whether students have seen problems before — and showed through sensitivity analysis that their conclusions don't depend on the exact discount value.

Jane: And throughout, they emphasized that this is a tool for teachers, not a replacement for them. AI provides evidence; teachers make judgments.

Tom: The ethical boundaries are clear: no individual student evaluation, no automatic grade adjustment, and full provenance tracking of AI outputs so you know exactly what you're looking at.

Jane: And that provenance requirement is actually a contribution in itself. They built a pipeline that archives prompts, schemas, raw responses, and hashes — so the AI evidence is auditable.

Tom: Which matters because AI outputs aren't deterministic. Even at temperature zero, they found jitter in repeated reviews. But the jitter didn't change the main correlations, which gives us confidence in the findings.

Jane: So what's the takeaway for our listeners? If you're teaching programming, this paper gives you a practical way to compare exam difficulty across classes and semesters. If you're doing research, it gives you a framework for validating AI-based measurement tools.

Tom: And if you're a student, it means your exam grades might be interpreted more fairly — because teachers will have better evidence about whether your exam was genuinely harder.

Jane: It's a good note to end on. Thanks for joining us, and we'll see you for the next paper.

Tom: Take care, everyone.

More episodes

← Home