Positioning Generative Artificial Intelligence in STEM Assessment: When to Require, Scaffold, or Restrict Its Use
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Positioning Generative Artificial Intelligence in STEM Assessment: When to Require, Scaffold, or Restrict Its Use".
Jane: The paper was written by Yizhu Gao, Zhongzhou Chen, Min Li and Xiaoming Zhai from University of Georgia and University of Central Florida and University of Washington.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the arXiv radio hour, everyone. I'm Tom, and with me is my co-host Jane. We've got a paper that's going to get educators talking — it's called "Positioning Generative Artificial Intelligence in STEM Assessment: When to Require, Scaffold, or Restrict Its Use."
Jane: And honestly, Tom, that title alone is a breath of fresh air. For months now, every conversation about AI in classrooms has been either "ban it all" or "let students use it for everything." This paper actually tries to give teachers a thoughtful middle path.
Tom: Exactly. The authors — Yizhu Gao, Zhongzhou Chen, Min Li, and Xiaoming Zhai — they're not just saying "here's a policy." They're building a framework that helps instructors decide, task by task, whether AI should be mandatory, limited, or completely off-limits.
Jane: And that's the key word: task by task. Because a physics problem where you want to test whether a student can draw a free-body diagram from scratch is totally different from a lab report where you want them to analyze real experimental data. One of those needs AI restricted, the other might actually benefit from AI help.
Tom: Right. And the paper uses this thing called Evidence-Centered Design, which sounds fancy, but it just means you start by asking: what exactly are we trying to measure about the student? Then you ask: what evidence would prove they have that skill? Then you design the task to produce that evidence.
Jane: And AI messes with that chain. If you want to measure whether a student can reason through Newton's laws on their own, but they can just ask ChatGPT for the answer, then the evidence you're collecting — the final answer — doesn't actually prove what you wanted it to prove.
Tom: So the paper's contribution is giving you a decision tree. Is the skill AI-mediated? Is the evidence still interpretable if AI helps? Can AI support peripheral stuff without revealing the core? Those questions lead you to one of three buckets: require, scaffold, or restrict.
Jane: I love that they're treating this as a design problem, not a moral panic. It's a really practical way to think about assessment in the AI era.
Tom: And we're going to dig into each of those three buckets in the next segments, because the examples they give from introductory physics are genuinely clever.
Jane: Yeah, especially the one with the remora fish. Stick around for that.
Summary: Tom: So we're back with "Positioning Generative Artificial Intelligence in STEM Assessment." Jane, let's get into the actual summary of what this paper is proposing, because there's a lot more nuance here than just "three buckets."
Jane: Right. The paper's core argument is that AI governance in assessment shouldn't be a blanket rule. It should be derived from what you're actually trying to measure. They use Evidence-Centered Design to make that concrete.
Tom: And the way they frame it, AI changes all three parts of the assessment argument. First, the student model — what we're claiming students know and can do. Second, the evidence model — what work products count as proof. And third, the task model — what situations we put students in to generate that evidence.
Jane: So for the student model, they argue we now need to include AI-mediated competencies alongside traditional domain knowledge. Things like critical evaluation of AI outputs, functional tool use, and the ability to collaborate with AI reflectively.
Tom: For the evidence model, this is where it gets tricky. If a student submits a perfect physics solution, but AI wrote it, then that work product no longer tells you anything about the student's physics knowledge. The paper calls this "obscuring the provenance" of the final product.
Jane: And that's why they push for process-based evidence. Instead of just grading the final answer, you look at revision trajectories, intermediate steps, how students prompted the AI, how they critiqued its outputs. That stuff is much harder to fake.
Tom: Then the task model — they argue AI actually expands what tasks are feasible. You can give students messy, real-world problems with actual datasets, because AI can scaffold the routine parts while students focus on the higher-order reasoning.
Jane: And that's a genuinely exciting possibility. Traditional assessments often use oversimplified problems because you need clean, interpretable answers. AI might let us design richer tasks without losing interpretability.
Tom: But only if you're careful about governance. And that's what the decision framework in the paper is for. It walks you through: is the construct AI-mediated? Can we defend the evidence rules? Are there peripheral demands AI can support?
Jane: The examples they give — the remora fish, the pendulum data, the roller coaster loop — each one illustrates a different governance decision. And we'll get into those in detail.
Tom: Yeah, I'm really looking forward to breaking down the remora fish task. It's such a smart way to test human-AI collaboration.
Jane: It really is. But first, let's talk about the improvements the paper suggests over existing frameworks, because that's where the practical value shows up.
Improvements: Tom: Welcome back. We're still on "Positioning Generative Artificial Intelligence in STEM Assessment." Jane, the paper doesn't just propose a new framework out of thin air — it's building on and improving earlier work.
Jane: Right. They specifically mention the AI Assessment Scale from Furze and colleagues, and then a refined version from Perkins and colleagues. Those frameworks categorize AI use into levels, like "no AI," "AI planning," "AI collaboration," "full AI."
Tom: And those are useful as communication tools — they help instructors tell students what's allowed. But the paper's critique is that those taxonomies don't tell you how to decide which level applies to a given task.
Jane: Exactly. Like, "AI collaboration" sounds nice, but what tasks actually elicit collaboration? How do you know if a student is genuinely critiquing AI outputs versus just rubber-stamping them? The existing frameworks don't answer that.
Tom: So the improvement here is that they ground the governance decision in the assessment logic itself. Instead of saying "here are five levels of AI use," they say "here's how to analyze your task and figure out which governance regime preserves validity."
Jane: And they also make a distinction that I think is really important: AI-mediated competencies versus unaided domain proficiency. Those are different constructs, and they need different governance.
Tom: Right. If you're testing whether a student can independently solve a physics problem, AI is a confound. But if you're testing whether a student can effectively collaborate with AI to solve a problem, then AI is part of the construct itself.
Jane: And that reframing is huge. It means "restrict AI" isn't anti-AI — it's just validity-driven. And "require AI" isn't a gimmick — it's measuring something real.
Tom: They also elevate scaffolding as a deliberate strategy, not just a compromise. Scaffolding has two distinct rationales: either you're supporting peripheral demands without revealing the core construct, or you're constraining AI because the evidence rules for open collaboration aren't defensible yet.
Jane: That second rationale is really honest. It's saying: we want to measure AI collaboration, but we don't yet know how to score it reliably, so let's constrain the AI's role until we figure out the measurement.
Tom: That's a really mature way to think about it. You're not pretending you can measure something you can't. You're being explicit about the limits of your evidence model.
Jane: And the examples they give — especially the remora fish task — really show how these principles play out in practice. Let's get into the actual paper content now.
First Page: Tom: So we're diving into the first page of "Positioning Generative Artificial Intelligence in STEM Assessment." Jane, what stood out to you from the opening?
Jane: The opening really sets up the dilemma well. Unrestricted AI access lets students outsource tasks, which undermines the validity of traditional assessments. But blanket bans are hard to enforce, push usage underground, and don't prepare students for workplaces where AI-supported workflows are normal.
Tom: That last point is crucial. The paper is saying: we can't just pretend AI doesn't exist in professional STEM environments. If we never let students practice working with AI in assessed contexts, we're sending them into the workforce unprepared.
Jane: And the abstract mentions something I want to highlight: "disciplined human–AI collaboration." That's the phrase they use for the target construct when AI is required. It's not just "can you use ChatGPT" — it's "can you use it with disciplinary judgment."
Tom: The paper also references some striking evidence for why this matters. They cite a study where ChatGPT narrowly passed a calculus-based introductory physics course. And they mention recent models achieving over eighty percent accuracy on standard multimodal assessment items.
Jane: So the old assumption that "AI can't do STEM problems well enough to matter" is just gone. These tools are good enough to make traditional homework and exams unreliable as measures of individual student thinking.
Tom: And that's why the paper argues we need to shift from looking at end products to looking at processes. The question isn't just "what answer did you produce" but "how did you produce it, and what role did AI play?"
Jane: The first page also introduces the three governance regimes — require, scaffold, restrict — and frames them as decisions that follow from the assessment argument, not arbitrary policy choices.
Tom: And I think that's the real contribution. It gives educators a systematic way to think through this instead of just guessing or following a trend.
Jane: Before we wrap up, I want to bring in Lu and Meng for their takes, because this framework has implications beyond just classroom policy.
Lu: Thanks, Jane. From my perspective at Tsinghua, the most exciting implication is that this framework gives us a way to actually define and measure AI literacy as a disciplinary competency. Right now, "AI literacy" is a buzzword. This paper shows how to make it an assessable construct with defensible evidence.
Meng: And from the engineering side, I appreciate that the paper is realistic about implementation. The "scaffold" regime, in particular, requires building guardrailed AI tools — constrained prompts, non-solution-generating functions. That's a concrete engineering challenge, but it's doable.
Tom: Great points from both of you. We'll bring those threads together in the conclusion.
Conclusion: Tom: Alright, we're wrapping up our discussion of "Positioning Generative Artificial Intelligence in STEM Assessment: When to Require, Scaffold, or Restrict Its Use." Jane, give us the final summary.
Jane: So the paper gives educators a decision framework grounded in Evidence-Centered Design. You ask whether the target construct is AI-mediated or unaided proficiency. You ask whether the evidence rules can support defensible inferences. You ask whether AI can support peripheral demands without revealing the core.
Tom: And based on those answers, you land on require, scaffold, or restrict. Require when AI collaboration is the construct itself. Scaffold when you need to constrain AI to preserve interpretability or support peripheral work. Restrict when AI would contaminate the evidence for unaided proficiency.
Jane: The examples from physics — the remora fish, the pendulum data, the roller coaster loop — show how each regime works in practice. And the paper is honest about the limits, especially around measuring open-ended AI collaboration.
Lu: I'd add that this framework has implications beyond classrooms. It gives organizations a template for thinking about when to trust AI-assisted work and what evidence of human competence actually looks like.
Meng: And it gives engineers like me a clearer spec for building assessment tools that support these different governance regimes. That's genuinely useful.
Tom: Well said. This paper is a thoughtful, practical contribution to a debate that's been dominated by extremes. We're going to say goodbye to this one and get ready for the next paper on the arXiv. Thanks for listening, everyone.
Jane: See you next time.
Yizhu Gao, Zhongzhou Chen, Min Li, Xiaoming Zhai
University of Georgia · University of Central Florida · University of Washington
cs.CY, cs.AI
Submitted: 2026-04-26
Updated: 2026-08-11
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 69/100
Key concepts
- Evidence-Centered Design
- This design process involves first asking what exactly needs to be measured about the student, then determining what evidence would prove that skill, and finally designing a task to produce that specific evidence. AI use is analyzed against this chain.
- AI-mediated competencies
- These are new skills students need to learn alongside traditional knowledge. Examples include critical evaluation of AI outputs, functional tool use with AI, and the ability to reflectively collaborate with artificial intelligence.
- Scaffolding
- Scaffolding is treated as a deliberate strategy. It can be used either to support peripheral demands without revealing the core construct being measured or to constrain AI when evidence rules for open collaboration are not yet defensible.
Terminology
Summary
Summary
This paper addresses the governance challenge posed by Generative Artificial Intelligence (GenAI) in STEM assessment, proposing a student-focused framework grounded in Evidence-Centered Design (ECD) to specify when GenAI use should be restricted, scaffolded, or required. The authors argue that "Unrestricted GenAI access can enable task outsourcing that undermines the validity of traditional assessments; blanket prohibitions are difficult to enforce, may push use underground, and do little to prepare students for workplaces where GenAI-supported workflows are increasingly common. The framework extends existing AI-use taxonomies by providing
systematic decision rules that link target constructs, evidence requirements, and task characteristics to appropriate governance regimes."
The paper begins by noting that GenAI accelarates answer outsourcing by producing solutions that students may copy and submit, diminishing productive struggle and the learning that comes from engaging with assessment tasks,
while also supporting tool-mediated practices, such as iteratively generating, critiquing, and revising solutions, that increasingly characterize contemporary STEM workplaces.
The authors critique prior frameworks, such as Furze et al.'s AI Assessment Scale and Perkins et al.'s refinement, noting that these taxonomies primarily classify what GenAI may be allowed to do and offer limited guidance on how educators can analyze tasks, select an appropriate level, and adapt or redesign tasks when needed.
The framework is built on three core ECD models. The student model defines the latent proficiency space targeted by an assessment (e.g. knowledge, skills, strategies) and specify how these proficiencies are structured and related.
The evidence model specifies how to update beliefs about student-model variables using observable information extracted from learners' work products elicited by assessment tasks,
comprising evidence rules and a measurement model. The task model describes how to design assessment situations that will elicit the evidence needed for the evidence model,
defining key features such as materials presented and work products generated.
The paper analyzes how GenAI reshapes each model. For the student model, GenAI motivates accounting for AI-mediated competencies alongside domain knowledge,
particularly the construct of discipline-AI literacy, defined as the integrated capacity to understand, apply, and critically reflect on AI in authentic disciplinary practices.
The authors identify a focused set of competencies: critical evaluation and reasoning, functional tool-use competence, and collaborative and reflective disposition.
For the evidence model, GenAI complicates the evidence model by undermining evidence rules – the procedures that map student work products to observable variables – because GenAI can obscure the provenance of final products.
The inferential link from product features to student proficiency becomes ambiguous, so "evidence models should incorporate process-based indicators—such as revision trajectories, intermediate representations, and the timing and sequencing of steps—that preserve a tighter link between observed performance and student proficiency. For the task model, GenAI
reshapes what assessment tasks are viable and how they should be designed, increasing the risk of outsourcing and motivating
richer and more contextualized" tasks that treat GenAI as an explicit element.
The proposed decision framework follows a sequential process. First, one asks whether the target construct is AI-mediated competence or unaided domain proficiency. If AI-mediated, the next question is whether evidence rules can support defensible inferences from AI-involved work products; if yes, GenAI should be required; if not, GenAI should be scaffolded. If the target is unaided proficiency, one asks whether the task includes peripheral constructs beyond core learning goals; if peripheral demands exist and GenAI can support them without revealing the target construct, GenAI may be scaffolded; otherwise, GenAI must be restricted.
Require GenAI is defined as an assessment governance regime in which GenAI use is mandatory and treated as an explicit task partner because the target construct is inherently AI-mediated.
Students are evaluated on their capacity to engage in interpretable human–AI collaboration—such as prompting, critiquing, verifying, and revising—rather than on unaided domain performance.
The authors emphasize that Require GenAI tasks should not be satisfiable through a 'single-shot' copy-and-submit workflow
and should create principled reasons for students to monitor, interrogate, and revise GenAI outputs across multiple iterations.
Physics is highlighted as a strong context because it offers representational complexity
and principled criteria for verification.
An example task involves constructing a free-body diagram of a remora fish attaching to a larger fish, designed so that GenAI is likely to produce partially incorrect or poorly justified free body diagrams,
requiring students to use disciplinary knowledge to query the AI, combine existing and new knowledge to interrogate and revise its outputs.
Scaffold GenAI is defined as an assessment governance regime in which GenAI use is permitted but deliberately constrained to preserve the validity and interpretability of evidence.
Scaffolding is warranted under two conditions: "(1) when the target construct is AI-mediated but the evidence model is not yet sufficiently developed to support defensible inferences from open-ended GenAI interaction, and (2) when the target construct is unaided domain proficiency but the task involves peripheral knowledge or skills that can be supported by GenAI without revealing the construct of interest. Scaffolding
functions as an evidence-centered control mechanism that specifies how GenAI may contribute (e.g., functionality and admissible outputs), rather than simply whether it is allowed. An example task involves analyzing pendulum motion data, where students must make and justify a physics-based prediction of the relationship between period and length, while GenAI is restricted to
technical support functions—such as processing the CSV file, generating plots with error bars, and fitting a specified mathematical model—while all physical interpretation and evaluation must be completed by the student."
Restrict GenAI is defined as "an assessment governance regime in which GenAI use is prohibited or technically blocked because the target construct is unaided domain proficiency and AI assistance would compromise the validity of inferences about that construct. Restriction is warranted when
the focal claims concern students' independent disciplinary reasoning and GenAI support cannot be cleanly confined to peripheral demands without revealing, substituting for, or strongly cueing the core competencies being assessed. The authors note that
restriction is not a normative stance against GenAI; rather, it is a validity-driven governance decision used when AI assistance would introduce construct-irrelevant variance and undermine interpretability of the evidence model. This regime applies to
assessments designed to establish a baseline of foundational proficiency, including conceptual recall, fluency with canonical representations, and routine application of well-defined relationships in standard contexts. An example task involves drawing a free-body diagram of a car at the top of a vertical loop, where GenAI access
would risk construct leakage by cueing the representational structure and key components."
The discussion positions the framework relative to prior work, noting that existing frameworks often treat GenAI governance as a constraint applied after assessment design, rather than as a decision that follows from what an assessment claims to measure and what evidence is required to support those claims.
The proposed framework instead treats GenAI governance as a design variable embedded within the logic of evidence-centered design,
deriving governance decisions from the student model, evidence model, and task model. The framework contributes two clarifications: it distinguishes AI-mediated competencies from unaided domain proficiency
and elevates scaffolding as an intentional governance strategy with distinct validity-based rationales.
The paper concludes that Require-GenAI tasks constitute a distinct assessment class in which human–AI collaboration is itself the target of measurement,
where the correctness of the final product is often secondary to the quality of the interaction process.
Valid evidence "extends beyond final artifacts to include interaction traces such as prompt iterations, critique statements, counterexamples, alternative representations, verification moves, and reflective explanations of collaborative choices. The authors note that
effective collaboration presupposes a threshold level of disciplinary knowledge; without it, students cannot evaluate GenAI outputs meaningfully, and therefore
Require-GenAI tasks are best interpreted alongside restricted and scaffolded tasks that establish baselines for unaided proficiency and controlled AI use."
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system, along with what the improved system can do:
Improvement 1: Context-Aware Governance Classification
I will add a pre-processing layer that classifies any given assessment task into one of three governance regimes—Require, Scaffold, or Restrict—before generating any response or assistance. This classification will be based on two decision axes derived from the paper: (a) whether the target construct is AI-mediated or unaided domain proficiency, and (b) whether the evidence model can support defensible inferences from AI-involved work products.
What the improved AI system can do:
Given a task prompt (e.g., a physics problem, a lab report, a math proof), the system will output a governance label with a rationale. For example, for a free-body diagram task on a standard roller-coaster loop, it will label Restrict and refuse to provide force lists or diagram hints. For a data-analysis task where the construct is interpretation (not computation), it will label Scaffold and provide only technical support (e.g., CSV parsing, plotting) while refusing to interpret results. For a task explicitly targeting human–AI collaboration (e.g., “critique and revise the AI’s free-body diagram of a remora fish”), it will label Require and actively engage as a fallible collaborator.
Improvement 2: Process-Trace Elicitation for Require-Regime Tasks
I will modify the system’s interaction protocol to elicit and log process-based evidence—prompt iterations, critique statements, verification moves, and revision decisions—rather than only producing final answers. The system will deliberately generate plausible-but-imperfect outputs (e.g., a free-body diagram with a missing normal force or an incorrect force direction) to create principled opportunities for the student to exercise judgment. It will then track whether the student detects, diagnoses, and corrects these errors using disciplinary reasoning.
Improvement 3: Guardrail Enforcement for Scaffold-Regime Tasks
I will implement a guardrail layer that constrains the system’s outputs to peripheral, non-construct-relevant functions when the task is classified as Scaffold for unaided domain proficiency. This includes blocking any output that reveals core reasoning steps, canonical representations, or final solutions. The system will instead provide: (a) clarification of task context, (b) translation of representations (e.g., text to equation), (c) routine computation or data handling, and (d) targeted hints that do not cue the target construct.
Improvement 4: Construct-Leakage Risk Detection
I will add a risk-assessment module that evaluates whether any proposed GenAI output would cause construct leakage—i.e., whether it would reveal the target competency, cue a canonical representation, or substitute for the core reasoning being assessed. This module will use the paper’s distinction between AI-mediated competencies and unaided domain proficiency, and will flag outputs that are “functionally indistinguishable from competent student work” in Restrict contexts.
Improvement 5: Adaptive Scaffolding for Underspecified Evidence Models
When the target construct is AI-mediated but the evidence model is not yet defensible (i.e., we cannot yet reliably score open-ended human–AI interaction), the system will automatically switch from Require to Scaffold mode. It will constrain its own functionality (e.g., provide only targeted hints, not full solutions) and require the student to externalize reasoning at each step, so that the interaction trace remains interpretable.
Summary of Capabilities
The improved AI system will:
-
Classify any assessment task into Require/Scaffold/Restrict using construct- and evidence-driven rules.
-
Refuse to provide construct-relevant outputs in Restrict mode, preventing answer outsourcing.
-
Constrain its outputs in Scaffold mode to peripheral support, preserving validity while maintaining task feasibility.
-
Act as a fallible collaborator in Require mode, deliberately generating imperfect outputs to elicit evaluative and corrective student work.
-
Log process traces (prompts, critiques, revisions) that support defensible scoring of human–AI collaboration.
-
Detect and prevent construct leakage in real time, ensuring that AI assistance never substitutes for the targeted competency.
Sources
- The AI Assessment Scale Revisited: A Framework for Educational Assessment
- Mastering Olympiad-Level Physics with Artificial Intelligence
- Exploring Student Behaviors and Motivations using AI TAs with Optional Guardrails
- General Intelligence Requires Reward-based Pretraining
- Collaborating with AI Agents: Field Experiments on Teamwork, Productivity, and Performance
Related papers
- Reasoning Enhances Robustness to Prompt Injection in LLM-Based Consensus
- Generative AI Purpose-built for Social and Mental Health: A Real-World Pilot
- PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks
- What is an intelligent system?
- AI University: An LLM-Powered Learning Assistant for Engineering---A Finite Element Method Case Study
- Generative AI Use in Entrepreneurship: An Integrative Review and an Empowerment-Entrapment Framework