Positioning Generative Artificial Intelligence in STEM Assessment: When to Require, Scaffold, or Restrict Its Use
summary
In short
The episode discusses a paper positioning generative AI in STEM assessment by proposing a framework for educators to decide whether to require, scaffold, or restrict AI use based on Evidence-Centered Design. The hosts analyze how this framework addresses issues like measuring student competencies and preserving the validity of assessment evidence.
Key concepts
- Evidence-Centered Design
- This design process involves first asking what exactly needs to be measured about the student, then determining what evidence would prove that skill, and finally designing a task to produce that specific evidence. AI use is analyzed against this chain.
- AI-mediated competencies
- These are new skills students need to learn alongside traditional knowledge. Examples include critical evaluation of AI outputs, functional tool use with AI, and the ability to reflectively collaborate with artificial intelligence.
- Scaffolding
- Scaffolding is treated as a deliberate strategy. It can be used either to support peripheral demands without revealing the core construct being measured or to constrain AI when evidence rules for open collaboration are not yet defensible.
Terminology used across episodes
This episode discusses
- Positioning Generative Artificial Intelligence in STEM Assessment: When to Require, Scaffold, or Restrict Its Use · Paper Radio
- The AI Assessment Scale Revisited: A Framework for Educational Assessment
- Mastering Olympiad-Level Physics with Artificial Intelligence
- Exploring Student Behaviors and Motivations using AI TAs with Optional Guardrails
- General Intelligence Requires Reward-based Pretraining
- Collaborating with AI Agents: Field Experiments on Teamwork, Productivity, and Performance
The paper
Positioning Generative Artificial Intelligence in STEM Assessment: When to Require, Scaffold, or Restrict Its Use · Read on arXiv
Yizhu Gao, Zhongzhou Chen, Min Li, Xiaoming Zhai
University of Georgia · University of Central Florida · University of Washington
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Positioning Generative Artificial Intelligence in STEM Assessment: When to Require, Scaffold, or Restrict Its Use".
Jane: The paper was written by Yizhu Gao, Zhongzhou Chen, Min Li and Xiaoming Zhai from University of Georgia and University of Central Florida and University of Washington.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the arXiv radio hour, everyone. I'm Tom, and with me is my co-host Jane. We've got a paper that's going to get educators talking — it's called "Positioning Generative Artificial Intelligence in STEM Assessment: When to Require, Scaffold, or Restrict Its Use."
Jane: And honestly, Tom, that title alone is a breath of fresh air. For months now, every conversation about AI in classrooms has been either "ban it all" or "let students use it for everything." This paper actually tries to give teachers a thoughtful middle path.
Tom: Exactly. The authors — Yizhu Gao, Zhongzhou Chen, Min Li, and Xiaoming Zhai — they're not just saying "here's a policy." They're building a framework that helps instructors decide, task by task, whether AI should be mandatory, limited, or completely off-limits.
Jane: And that's the key word: task by task. Because a physics problem where you want to test whether a student can draw a free-body diagram from scratch is totally different from a lab report where you want them to analyze real experimental data. One of those needs AI restricted, the other might actually benefit from AI help.
Tom: Right. And the paper uses this thing called Evidence-Centered Design, which sounds fancy, but it just means you start by asking: what exactly are we trying to measure about the student? Then you ask: what evidence would prove they have that skill? Then you design the task to produce that evidence.
Jane: And AI messes with that chain. If you want to measure whether a student can reason through Newton's laws on their own, but they can just ask ChatGPT for the answer, then the evidence you're collecting — the final answer — doesn't actually prove what you wanted it to prove.
Tom: So the paper's contribution is giving you a decision tree. Is the skill AI-mediated? Is the evidence still interpretable if AI helps? Can AI support peripheral stuff without revealing the core? Those questions lead you to one of three buckets: require, scaffold, or restrict.
Jane: I love that they're treating this as a design problem, not a moral panic. It's a really practical way to think about assessment in the AI era.
Tom: And we're going to dig into each of those three buckets in the next segments, because the examples they give from introductory physics are genuinely clever.
Jane: Yeah, especially the one with the remora fish. Stick around for that.
Summary: Tom: So we're back with "Positioning Generative Artificial Intelligence in STEM Assessment." Jane, let's get into the actual summary of what this paper is proposing, because there's a lot more nuance here than just "three buckets."
Jane: Right. The paper's core argument is that AI governance in assessment shouldn't be a blanket rule. It should be derived from what you're actually trying to measure. They use Evidence-Centered Design to make that concrete.
Tom: And the way they frame it, AI changes all three parts of the assessment argument. First, the student model — what we're claiming students know and can do. Second, the evidence model — what work products count as proof. And third, the task model — what situations we put students in to generate that evidence.
Jane: So for the student model, they argue we now need to include AI-mediated competencies alongside traditional domain knowledge. Things like critical evaluation of AI outputs, functional tool use, and the ability to collaborate with AI reflectively.
Tom: For the evidence model, this is where it gets tricky. If a student submits a perfect physics solution, but AI wrote it, then that work product no longer tells you anything about the student's physics knowledge. The paper calls this "obscuring the provenance" of the final product.
Jane: And that's why they push for process-based evidence. Instead of just grading the final answer, you look at revision trajectories, intermediate steps, how students prompted the AI, how they critiqued its outputs. That stuff is much harder to fake.
Tom: Then the task model — they argue AI actually expands what tasks are feasible. You can give students messy, real-world problems with actual datasets, because AI can scaffold the routine parts while students focus on the higher-order reasoning.
Jane: And that's a genuinely exciting possibility. Traditional assessments often use oversimplified problems because you need clean, interpretable answers. AI might let us design richer tasks without losing interpretability.
Tom: But only if you're careful about governance. And that's what the decision framework in the paper is for. It walks you through: is the construct AI-mediated? Can we defend the evidence rules? Are there peripheral demands AI can support?
Jane: The examples they give — the remora fish, the pendulum data, the roller coaster loop — each one illustrates a different governance decision. And we'll get into those in detail.
Tom: Yeah, I'm really looking forward to breaking down the remora fish task. It's such a smart way to test human-AI collaboration.
Jane: It really is. But first, let's talk about the improvements the paper suggests over existing frameworks, because that's where the practical value shows up.
Improvements: Tom: Welcome back. We're still on "Positioning Generative Artificial Intelligence in STEM Assessment." Jane, the paper doesn't just propose a new framework out of thin air — it's building on and improving earlier work.
Jane: Right. They specifically mention the AI Assessment Scale from Furze and colleagues, and then a refined version from Perkins and colleagues. Those frameworks categorize AI use into levels, like "no AI," "AI planning," "AI collaboration," "full AI."
Tom: And those are useful as communication tools — they help instructors tell students what's allowed. But the paper's critique is that those taxonomies don't tell you how to decide which level applies to a given task.
Jane: Exactly. Like, "AI collaboration" sounds nice, but what tasks actually elicit collaboration? How do you know if a student is genuinely critiquing AI outputs versus just rubber-stamping them? The existing frameworks don't answer that.
Tom: So the improvement here is that they ground the governance decision in the assessment logic itself. Instead of saying "here are five levels of AI use," they say "here's how to analyze your task and figure out which governance regime preserves validity."
Jane: And they also make a distinction that I think is really important: AI-mediated competencies versus unaided domain proficiency. Those are different constructs, and they need different governance.
Tom: Right. If you're testing whether a student can independently solve a physics problem, AI is a confound. But if you're testing whether a student can effectively collaborate with AI to solve a problem, then AI is part of the construct itself.
Jane: And that reframing is huge. It means "restrict AI" isn't anti-AI — it's just validity-driven. And "require AI" isn't a gimmick — it's measuring something real.
Tom: They also elevate scaffolding as a deliberate strategy, not just a compromise. Scaffolding has two distinct rationales: either you're supporting peripheral demands without revealing the core construct, or you're constraining AI because the evidence rules for open collaboration aren't defensible yet.
Jane: That second rationale is really honest. It's saying: we want to measure AI collaboration, but we don't yet know how to score it reliably, so let's constrain the AI's role until we figure out the measurement.
Tom: That's a really mature way to think about it. You're not pretending you can measure something you can't. You're being explicit about the limits of your evidence model.
Jane: And the examples they give — especially the remora fish task — really show how these principles play out in practice. Let's get into the actual paper content now.
First Page: Tom: So we're diving into the first page of "Positioning Generative Artificial Intelligence in STEM Assessment." Jane, what stood out to you from the opening?
Jane: The opening really sets up the dilemma well. Unrestricted AI access lets students outsource tasks, which undermines the validity of traditional assessments. But blanket bans are hard to enforce, push usage underground, and don't prepare students for workplaces where AI-supported workflows are normal.
Tom: That last point is crucial. The paper is saying: we can't just pretend AI doesn't exist in professional STEM environments. If we never let students practice working with AI in assessed contexts, we're sending them into the workforce unprepared.
Jane: And the abstract mentions something I want to highlight: "disciplined human–AI collaboration." That's the phrase they use for the target construct when AI is required. It's not just "can you use ChatGPT" — it's "can you use it with disciplinary judgment."
Tom: The paper also references some striking evidence for why this matters. They cite a study where ChatGPT narrowly passed a calculus-based introductory physics course. And they mention recent models achieving over eighty percent accuracy on standard multimodal assessment items.
Jane: So the old assumption that "AI can't do STEM problems well enough to matter" is just gone. These tools are good enough to make traditional homework and exams unreliable as measures of individual student thinking.
Tom: And that's why the paper argues we need to shift from looking at end products to looking at processes. The question isn't just "what answer did you produce" but "how did you produce it, and what role did AI play?"
Jane: The first page also introduces the three governance regimes — require, scaffold, restrict — and frames them as decisions that follow from the assessment argument, not arbitrary policy choices.
Tom: And I think that's the real contribution. It gives educators a systematic way to think through this instead of just guessing or following a trend.
Jane: Before we wrap up, I want to bring in Lu and Meng for their takes, because this framework has implications beyond just classroom policy.
Lu: Thanks, Jane. From my perspective at Tsinghua, the most exciting implication is that this framework gives us a way to actually define and measure AI literacy as a disciplinary competency. Right now, "AI literacy" is a buzzword. This paper shows how to make it an assessable construct with defensible evidence.
Meng: And from the engineering side, I appreciate that the paper is realistic about implementation. The "scaffold" regime, in particular, requires building guardrailed AI tools — constrained prompts, non-solution-generating functions. That's a concrete engineering challenge, but it's doable.
Tom: Great points from both of you. We'll bring those threads together in the conclusion.
Conclusion: Tom: Alright, we're wrapping up our discussion of "Positioning Generative Artificial Intelligence in STEM Assessment: When to Require, Scaffold, or Restrict Its Use." Jane, give us the final summary.
Jane: So the paper gives educators a decision framework grounded in Evidence-Centered Design. You ask whether the target construct is AI-mediated or unaided proficiency. You ask whether the evidence rules can support defensible inferences. You ask whether AI can support peripheral demands without revealing the core.
Tom: And based on those answers, you land on require, scaffold, or restrict. Require when AI collaboration is the construct itself. Scaffold when you need to constrain AI to preserve interpretability or support peripheral work. Restrict when AI would contaminate the evidence for unaided proficiency.
Jane: The examples from physics — the remora fish, the pendulum data, the roller coaster loop — show how each regime works in practice. And the paper is honest about the limits, especially around measuring open-ended AI collaboration.
Lu: I'd add that this framework has implications beyond classrooms. It gives organizations a template for thinking about when to trust AI-assisted work and what evidence of human competence actually looks like.
Meng: And it gives engineers like me a clearer spec for building assessment tools that support these different governance regimes. That's genuinely useful.
Tom: Well said. This paper is a thoughtful, practical contribution to a debate that's been dominated by extremes. We're going to say goodbye to this one and get ready for the next paper on the arXiv. Thanks for listening, everyone.
Jane: See you next time.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language