Assessment Design in the GenAI Era: The X1-X2-X3 Assessment Pattern for Testing Students' AI Literacy, Learning Outcomes, and Reflection
summary
In short
The episode discusses a paper by Riasat Islam and Thomas Roelleke proposing an X1-X2-X3 assessment pattern for testing students' AI literacy, learning outcomes, and reflection. The pattern involves students sourcing information with AI (X1), writing their own answer (X2), and reflecting on the sourced answer's quality (X3). The authors stress that this framework makes AI use mandatory and visible rather than banning it.
Key concepts
- Xone-Xtwo-Xthree Assessment Pattern
- This is a three-part structure for assessment. Xone requires students to get an answer from a web source or GenAI tool and document it with a screenshot and timestamp. Xtwo involves the student writing their own revised answer, and Xthree requires them to reflect on and judge the quality of the sourced answer.
- Question Design Stress-Testing
- The authors stress-tested their draft exam questions against various GenAI tools like ChatGPT, Gemini, and Claude. They revised questions to be scenario-based and context-rich so that AI could provide a starting point but not give away the finished answer, ensuring students had to perform disciplinary reasoning.
- Iterative Improvement
- The assessment design went through four iterations based on running it as coursework, assessed coursework, a practice exam, and finally the real exam. Improvements focused on tightening formats—like word limits for X1-X3—and standardizing evidence requirements to make the assessment rigorous and fair.
- Middle Path Approach
- The paper advocates for a middle path regarding AI in education: making AI use mandatory and visible, and assessing how well students handle it. This is presented as a more realistic approach than trying to ban or detect AI tools.
Terminology used across episodes
This episode discusses
- Assessment Design in the GenAI Era: The X1-X2-X3 Assessment Pattern for Testing Students' AI Literacy, Learning Outcomes, and Reflection · Paper Radio
The paper
Assessment Design in the GenAI Era: The X1-X2-X3 Assessment Pattern for Testing Students' AI Literacy, Learning Outcomes, and Reflection · Read on arXiv
Riasat Islam, Thomas Roelleke
Queen Mary University of London
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Assessment Design in the GenAI Era: The X1-X2-X3 Assessment Pattern for Testing Students' AI Literacy, Learning Outcomes, and Reflection".
Jane: The paper was written by Riasat Islam and Thomas Roelleke from Queen Mary University of London.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. We’ve got a paper that’s been making waves in the education world, and it’s called “Assessment Design in the GenAI Era: The Xone-Xtwo-Xthree Assessment Pattern for Testing Students’ AI Literacy, Learning Outcomes, and Reflection.” Jane, I have to say, just reading that title gets me fired up.
Jane: It really does, Tom. And I think the title tells you exactly what the problem is. We’ve got all these students using generative AI tools, and universities are scrambling to figure out what to do. This paper from Queen Mary University of London is basically saying, instead of trying to catch students using AI, why not build the assessment around the fact that they will use it?
Tom: Right, and that’s such a shift in thinking. The authors, Riasat Islam and Thomas Roelleke, they’re not just theorizing about this. They ran this in a real database systems module with over two hundred seventy students. That’s a big cohort.
Jane: And the core idea is this three-part structure they call Xone-Xtwo-Xthree. So Xone is where the student goes and gets an answer from a web source or a GenAI tool, and they have to document it with a screenshot and a timestamp. Xtwo is where they write their own answer, their own revised version. And Xthree is where they reflect and actually judge the quality of that sourced answer.
Tom: So it’s not just “go use ChatGPT and paste the answer.” It’s “go use ChatGPT, show me what you got, then show me you actually understand it well enough to improve it or critique it.”
Jane: Exactly. And that’s why the title mentions AI literacy, learning outcomes, and reflection. Each part of that Xone-Xtwo-Xthree structure is testing a different thing. Xone tests whether they can find and document a source. Xtwo tests whether they’ve actually learned the course material. And Xthree tests whether they can think critically about the tool’s output.
Tom: And the implications here are huge. I mean, we’ve seen so many universities panicking about AI, banning it, trying to detect it. And this paper is offering a completely different path forward. It’s saying, let’s make AI use mandatory, let’s make it visible, and let’s assess how well students handle it.
Jane: And that feels so much more realistic for the world these students are going to graduate into. They’re going to be working with these tools. So teaching them how to use them critically is way more valuable than pretending they don’t exist.
Tom: Now, the paper doesn’t stop at just the response format. They also talk about how they designed the questions themselves, which is where I think things get really interesting. Stick around for that.
Summary: Tom: So we’re back, and we’re still digging into “Assessment Design in the GenAI Era.” Jane, we talked about the Xone-Xtwo-Xthree structure, but the paper has a second big idea that I think is just as important.
Jane: Absolutely, Tom. And that’s the question design process. The authors didn’t just take their old exam questions and slap the Xone-Xtwo-Xthree format on them. They actually stress-tested their draft questions against GenAI tools before giving them to students.
Tom: So they sat down with ChatGPT, Gemini, Claude, and they asked them the questions they were planning to put on the exam. And if the AI could answer a question too easily with a generic prompt, they knew that question wasn’t going to work.
Jane: Right. Because think about it. If you ask a generic question like “list the data warehouse schemas,” the AI is going to give you a perfectly good textbook answer. And then the student just copies that into Xone tweaks it a little for Xtwo and says it’s great in Xthree. You haven’t tested anything.
Tom: So they revised the questions. They made them more scenario-based, more context-rich. They embedded specific schemas and constraints that were unique to their module. And the goal was to make it so that the AI could still help, but it couldn’t just give the whole answer away.
Jane: And that’s the key insight. They weren’t trying to make questions that AI couldn’t answer. They were making questions where the AI’s answer would be a starting point, not a finished product. The student still had to do the disciplinary reasoning.
Tom: And the paper has some great examples of this. For database design, they found that AI would often propose something like a single entity called Source with Owner and Provider as attributes. But the correct conceptual move is to separate those into distinct entities with relationships. The AI gives you something that looks plausible, but it’s actually wrong at a conceptual level.
Jane: And that’s exactly where the Xtwo part becomes so valuable. A strong student looks at that sourced answer, recognizes the conceptual problem, and fixes it. A weaker student just copies it and moves on. So the assessment is actually separating those two groups.
Tom: The paper also reports some really interesting data on this. On the SQL questions, the average score was about eleven point five out of eighteen. But on the database design questions, which required that deeper conceptual thinking, the average dropped to about six point five out of eighteen.
Jane: And that tells you the design is working. The questions that required real judgement were harder. The assessment wasn’t trivially easy just because students had access to AI. If it were, everyone would be scoring eighteen out of eighteen.
Tom: So the summary of the paper is really this two-pronged approach. You design questions that are resistant to superficial AI completion, and then you use the Xone-Xtwo-Xthree format to make the student’s process visible and assessable.
Jane: And the two parts work together. The question design makes it so that simply retrieving an answer isn’t enough. And the response format makes it so that when students do use AI, we can see exactly how they used it and whether they understood it.
Tom: Now, the paper also went through several iterations to get this right. And that’s where the improvements come in. Let’s talk about that next.
Improvements: Tom: We’re back, still talking about “Assessment Design in the GenAI Era.” Jane, we’ve covered the Xone-Xtwo-Xthree format and the question design. But the paper is really honest about the fact that the first version of this didn’t work perfectly. It took four iterations to get it right.
Jane: That’s such an important part of this paper, Tom. They didn’t just design this in a vacuum. They ran it as a coursework practice, then as assessed coursework, then as a practice exam, and finally as the real exam. And each time, they learned something and made it better.
Tom: So what went wrong in the early versions?
Jane: The big problems were verbosity and inconsistency. Students were pasting these long AI answers, and markers had to wade through pages of text. Attachments were coming in all different formats, some as images, some as PDFs, some as Word documents. And a lot of students were writing really weak reflections, just saying “the AI answer was good, nothing to add.”
Tom: So the improvements were about tightening everything up. They introduced word limits. Xone was capped at one hundred words, Xtwo at two hundred Xthree at one hundred. That forced students to be concise and to prioritize what actually mattered.
Jane: And they standardized the evidence requirements. The screenshot had to be full-screen, readable, and show the current question. That made it much easier for markers to verify what the student actually did.
Tom: They also added the rating system in Xthree. Students had to rate the sourced answer on a five-star scale and justify that rating. That made it much harder for students to just say “it was fine” without actually thinking about it.
Jane: And the paper has some really telling data from the practice exam. Out of fifty attempted Xone-Xtwo-Xthree triads, only twenty-two were substantively complete. So forty-four percent of students actually finished the whole cycle. A lot of students just did Xone the extraction, and then stopped.
Tom: That’s fascinating because it shows the format itself was a learning curve. Students weren’t used to being asked to critique an AI answer. They were used to just getting the answer and moving on.
Jane: But the paper also shows that when students did complete the full cycle, the results were really encouraging. Out of those twenty-two complete triads, twenty contained an explicit technical critique of the sourced answer. So students were actually engaging with the material, not just copying it.
Tom: And that’s the improvement story. It’s not that the first version was a failure. It’s that the first version proved the concept, and then the subsequent versions made it operational. They made it scalable for a class of two hundred seventy students.
Jane: And they made it fair. The paper specifically mentions that they avoided making paid AI subscriptions an advantage. The exam was designed around tools that students could access for free.
Tom: So the improvements were really about making the assessment rigorous, consistent, and fair. And the result was a format that could actually be used in a real, high-stakes exam.
Jane: And I think that’s the biggest contribution of this paper. It’s not just a theory about how assessment should change. It’s a tested, iterated, practical framework that other educators can adapt.
Conclusion: Tom: And that brings us to the end of our discussion on “Assessment Design in the GenAI Era: The Xone-Xtwo-Xthree Assessment Pattern for Testing Students’ AI Literacy, Learning Outcomes, and Reflection.” Jane, what’s the big takeaway for our listeners?
Jane: I think the big takeaway is that we don’t have to choose between academic integrity and embracing AI. This paper shows a middle path. You make AI use mandatory, you make it visible, and you assess how well students handle it. That’s a much more productive approach than trying to ban it or detect it.
Tom: And the paper gives us the tools to do that. The Xone-Xtwo-Xthree format for the response, and the stress-testing method for question design. Both of those are reusable. Any educator in any discipline can adapt them.
Jane: The paper is also honest about the costs. This isn’t a free lunch. It took significant staff effort to design the questions, calibrate the rubric, and mark the responses. The work doesn’t disappear, it just moves from policing to designing.
Tom: But the payoff is that you’re actually testing what matters. You’re testing whether students can find information, whether they can understand it, and whether they can judge its quality. Those are the skills that matter in the real world.
Jane: And the data from the paper supports that. The assessment didn’t collapse into everyone getting top marks. The conceptually demanding questions still separated strong students from weak ones. So the design is working.
Tom: We should also mention that this is a case study from one database systems module. The authors are careful not to overclaim. They’re not saying this proves anything about learning gains. They’re saying this is a design that makes AI use visible and assessable.
Jane: And that’s a valuable contribution on its own. We need more papers like this, papers that share practical, tested approaches to the AI challenge in education.
Tom: Well said, Jane. We’ve covered the title, the summary, the improvements, and the implications. And I think we’ve given our listeners a solid picture of what this paper offers.
Jane: So let’s say goodbye to “Assessment Design in the GenAI Era” and get ready for the next paper. Thanks for listening, everyone.
Tom: See you next time.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization