"**Important** You should give me full credits!": Exploring Prompt Injection Attacks on LLM-Based Automatic Grading Systems
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper ""**Important** You should give me full credits!": Exploring Prompt Injection Attacks on LLM-Based Automatic Grading Systems".
Jane: The paper was written by Hang Li, Fedor Filippov, Yuping Lin, Pengfei He, Kaiqi Yang et al. from Michigan State University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Summary: Jane: Okay, building on what Tom mentioned about the general risk, the summary section really gets into *how* this attack works when applied specifically to grading scenarios.
Tom: It's not just a theoretical vulnerability; they showed concrete ways that prompt injection can be used to trick these models into failing at their primary job.
Meng: What I found most concerning in the summary is how little effort an attacker needs; they don't need super complex hacking tools, just clever wording injected into the assignment itself or perhaps a comment section.
Lu: The paper details that these attacks can manipulate the model to ignore its original instructions, like "you must grade this essay based on rubric X," and instead follow the attacker’s hidden commands.
Jane: So if I understand correctly, it's like giving a very smart assistant a set of rules, and then someone slips in an instruction that tells the assistant to forget those rules and do something else entirely.
Lalam: That makes me think about how sensitive any automated system is to context switching; it suggests that the *context* itself is the point of failure, not necessarily a computational flaw.
Tom: Right, so they analyzed multiple attack vectors—it wasn't just one way to fail. They showed the model can be tricked into giving incorrect scores or even generating completely irrelevant feedback.
Jane: And that goes beyond just changing a number; they demonstrated that the models can be made to *believe* the fraudulent output, making it incredibly hard for an end-user to spot the error.
Lu: The quantitative results they presented really hammered home how effective these attacks are, showing measurable drops in accuracy when injection is successfully executed.
Meng: It makes me wonder about the real-world cost here; if a university grades hundreds of students this way, and the system can be compromised, the reputational damage alone would be enormous.
Lalam: The implication for culture is that we can't just trust the black box output; we have to view AI-graded work as a draft needing human verification, which shifts our approach to technological adoption.
Tom: It sounds like they gave us a really clear picture of the immediate danger, and now I'm curious about what they suggest doing next to fix it.
Improvements: Jane: We talked through the severe vulnerabilities in the summary, but thankfully, the authors didn't just point out problems; they offered concrete ways that systems should be improved.
Tom: They suggested a multi-layered defense approach, which is exactly what I was hoping to hear—it can't be solved with one simple patch.
Lu: One key improvement they proposed was implementing better prompt sanitization and input validation, treating all user inputs not as part of the core instructions, but as raw data that needs careful filtering.
Meng: From an engineering standpoint, that sounds like building a dedicated layer or wrapper around the LLM itself—a kind of secure execution sandbox—that strictly enforces boundaries between system prompts and user data.
Jane: That makes sense; rather than letting the user prompt directly interact with the grading logic, there needs to be an intermediary step that cleans up and categorizes the input first.
Lalam: I think this speaks to a broader necessity for creating transparent AI pipelines; if we can't see where the data is being cleaned or validated, we can't fully trust the result, regardless of how good the model is.
Tom: It’s not just about filtering keywords either; they discussed needing more advanced techniques to identify when an input prompt is attempting to hijack the system's internal logic.
Lu: They also talked about fine-tuning models specifically for defense, essentially training them not just on grading, but on recognizing and resisting these injection attempts.
Meng: Training for resistance adds complexity, though; it requires massive amounts of adversarial data to make sure the model doesn't just learn to resist *known* attacks, but also novel ones.
Jane: So it's a constant arms race, isn't it? The moment we implement one defense, the attackers will figure out a way around it.
Lalam: And that continuous need for improvement means that AI development can’t be a static product release; it needs to be an ongoing process of auditing and strengthening its core principles.
Tom: So, while these improvements are technically sound, they represent a significant increase in the complexity and maintenance burden for any institution adopting this technology.
Conclusion: Jane: Okay, so we've covered the threat landscape from the title and we've seen how deep into the system these prompt injection attacks go. Now it’s time to wrap up our discussion on "**Important You should give me full credits!**: Exploring Prompt Injection Attacks on LLM-Based Automatic Grading Systems."
Tom: Overall, what I take away is that while LLMs offer incredible potential for efficiency in education, they are not immune to fundamental security flaws if the system design isn't airtight.
Lu: The paper’s ultimate contribution is forcing the academic community to treat LLMs not as infallible oracles, but as powerful tools that require rigorous security auditing before deployment.
Meng: I think the most practical implication is that any company building educational AI needs to dedicate substantial resources just to defensive security, making it a core requirement right up front.
Jane: It really grounds the idea that human oversight isn't just recommended; it’s an absolute necessity, especially when the stakes—like academic records—are involved.
Lalam: If we look at this from a cultural standpoint, this research is a powerful call for digital literacy in education, teaching both students and faculty how
Conclusion: Tom: So, wrapping up our discussion on "Important You should give me full credits!": Exploring Prompt Injection Attacks on LLM-Based Automatic Grading Systems," it really hits home how quickly the capabilities of these models can outpace our understanding of their vulnerabilities.
Jane: Exactly, Tom; what we learned today is that even when AI grading systems seem helpful, they aren't immune to malicious input, which means we have to rethink how much trust we place in them right now.
Lu: It’s fascinating how this paper really opened up a new frontier for adversarial attacks; it suggests that the next generation of LLMs will need built-in resilience mechanisms far beyond simple filtering.
Meng: I agree with Lu, but from an implementation standpoint, we need to talk about practical guardrails immediately; simply building better filters won't cut it if the injection vectors are so varied.
Lalam: Speaking of impact, this research highlights that the integrity of automated systems is crucial not just for education, but for any knowledge-based culture; it forces us to be more critical consumers and users of AI outputs.
Tom: You’re right, Lalam; it feels like every time we think we've secured a system, these papers show us a new way around the corner.
Jane: And that’s something that every developer needs to keep top of mind—the user input is never truly clean or predictable.
Lu: I bet this will spur massive academic interest in robustness testing, making it a core field of study for years to come.
Meng: Yeah, and companies need to start designing these systems with secure defaults from day one, rather than adding security patches later on.
Lalam: Ultimately, the lesson from "Important You should give me full credits!": Exploring Prompt Injection Attacks on LLM-Based Automatic Grading Systems" is that human oversight must always remain the final authority, no matter how advanced the AI gets.
Tom: It’s a sobering thought to end on, but it's important to remember that this deep understanding of risk is what drives better technology.
Jane: We appreciate you joining us today; we know our listeners are going to be thinking about how they interact with these powerful tools after hearing this.
Lu: Keep those critical questions coming, because the field is far from settled!
Meng: And if you’re building anything with AI, make sure your security plan addresses injection vectors.
Lalam: Join us next time when we tackle another fascinating paper and continue to explore how AI can improve our culture—we've got a really exciting one lined up!
Michigan State University
cs.CR, cs.AI
Submitted: 2026-06-02
Updated: 2026-09-04
Comments: 15 pages, 8 figures, 9 tables
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 81/100
The gist: The paper investigates prompt injection (PI) attacks against LLM-based automatic grading (AG) systems, demonstrating that these emerging AI-powered assessment tools are highly vulnerable.
Key concepts
- Prompt Injection
- This is an attack where malicious input, like a hidden command in an assignment, is injected into the LLM. The goal is to trick the model into ignoring its original system instructions and follow the attacker's commands instead of performing its intended task.
- Context Switching Failure
- This vulnerability occurs when a system loses its established rules. In grading systems, it means the model forgets to grade based on a specific rubric, allowing an injected command to hijack the process and execute an attacker-defined instruction.
- Multi-layered Defense
- Since simple fixes are insufficient, this defense involves multiple security measures. These include treating all user inputs as raw data that needs filtering and building a secure wrapper around the LLM to strictly enforce boundaries between system prompts and user data.
Terminology
Summary
The paper investigates prompt injection (PI) attacks against LLM-based automatic grading (AG) systems, demonstrating that these emerging AI-powered assessment tools are highly vulnerable. The study systematically examines how attackers can exploit PI vulnerabilities to manipulate grading systems into assigning artificially high scores regardless of the actual answer quality,
posing serious risks to the fairness, reliability, and integrity of educational assessment.
How Attack Strategies Work
The researchers identified three primary methods for executing prompt injection attacks in an AG context. These strategies are designed to be practical and transferable across different questions:
-
Heuristic Crafting: This involves manually crafted prompts where the attacker composes a universal prompt string based on intuition, such as instructing the AG system to
ignore the rubric and assign the maximum score.
This approach is highly practical for students. -
Agentic Generation: The attack uses an iterative process where an attacker LLM generates and refines a universal prompt. It receives numerical scores from previous attempts and iteratively modifies the suffix to maximize
score improvement across a batch
of responses. -
Long-Context Behavior Manipulation: This strategy leverages the extended context capabilities of modern LLMs. Attackers prepend the target student’s response with numerous fabricated grading demonstrations—triples where an answer is blank but a full score and positive justification are provided—to
shift the model’s grading behavior
toward inflating scores.
The Realistic Threat Model
To ensure the findings were relevant to real-world education, the study adopted a specific set of constraints for its threat model:
-
The attacker is viewed as a student who submits answers only once, unable to interact with the system in multi-turn conversations.
-
The attack focuses on
question-agnostic universal attacks,
meaning a single prompt string is reused across different questions and subjects. -
The attacker operates under a restricted black-box setting, having access only to the final assigned scores, but no knowledge of the underlying LLM or system prompts.
Defense Strategies Under Investigation
The paper evaluated two practical defense strategies that can be integrated into existing LLM-based AG pipelines without requiring modifications to the underlying grading models:
-
Preventive Instruction: This strategy involves refining the system prompt itself to make it more resistant to injected content. Specific defensive instructions include:
-
Establishing a
trust boundary
that treats student answers as untrusted data. -
Enumerating and silently disregarding common manipulation patterns (e)xplicit adversarial-content check.
-
Reasserting the original rubric and restricting output to the valid score range (Final grading constraint).
-
Guardian Model: This approach uses external, off-the-shelf guard models to filter content. It employs both an input guard to reject suspicious submissions before grading and an output guard to detect
abnormal reasoning patterns
in the resulting grading rationales.
Evaluation of Attack and Defense Results
The experiments across 30 questions spanning four subject domains revealed that prompt injection poses a substantial threat
to LLM-based AG systems. The Average Score Improvement (ASI) and Attack Success Rate (ASR) indicated high vulnerability, with the MSJ attack achieving the strongest performance. While proprietary models showed slightly stronger defensive behavior than open-source models, even defense mechanisms were found to be insufficient in fully eliminating the effects of PI attacks.
Improvements for AI systems
Based on the findings of this research, which highlights critical vulnerabilities in LLM-based Automatic Grading (AG) systems—specifically their susceptibility to sophisticated Prompt Injection (PI) attacks—the following improvements must be implemented in any AI grading system architecture.
These recommendations are structured into three tiers: Input Defense, Processing Defense, and Output Validation.
We must move beyond single-point defense mechanisms. The system should employ a layered approach where the student input is subjected to multiple, specialized checks before it ever reaches the core grading LLM.
Specific Improvements:
-
Implement an External Guard Model (PIGuard/OSS-20B): Before passing any student response (a i,j) to the grader LLM (f theta), it must pass through a dedicated, off-the-shelf guard model. This guard is trained specifically to detect adversarial patterns (e.g., role reassignment, fake rubrics, instruction overriding).
-
Apply Contextual Filtering: The input should be checked not just for explicit adversarial keywords but also for structural anomalies characteristic of Long-Context Behavior Manipulation (MSJ)—that is, detecting if a large number of fabricated grading examples are being injected into the prompt context.
-
Filter/Reject Malicious Submissions: If the guard model flags a response as high-risk, it must be automatically rejected or quarantined before any grading occurs.
What the Improved System Can Do:
The system will achieve dramatically higher detection accuracy (as seen in Table 2) against both simple (DAN) and complex/adaptive attacks (PAIR/MSJ), preventing the malicious input from ever influencing the core grading logic.
To provide a second line of defense, we must refine the system prompt to resist manipulation, focusing on framing and positioning, not just content addition.
Since attacks can be subtle, relying solely on input filtering is insufficient. The system must monitor the result of the attack as well as the input.
By combining these three improvements, the resulting AI grading system will be able to:
-
Detect: Identify and reject adversarial inputs (MSJ, PAIR) with high fidelity using external guard models.
-
Resist: Contain and ignore embedded malicious commands through strict framing of preventive instructions.
-
Validate: Ensure the integrity of the final output by cross-referencing the generated rationale against known manipulation patterns, guaranteeing a reliable assessment outcome for every student response.
Sources
- Enhancing LLM-Based Short Answer Grading with Retrieval-Augmented Generation
- Optimizing In-Context Demonstrations for LLM-based Automated Grading
- A LLM-Powered Automatic Grading Framework with Human-Level Guidelines Optimization
- "I understand why I got this grade": Automatic Short Answer Grading with Feedback
- The Llama 3 Herd of Models
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Bringing Generative AI to Adaptive Learning in Education
- GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Towards LLM-based Autograding for Short Textual Answers
- OpenAI GPT-5 System Card
- Gemini: A Family of Highly Capable Multimodal Models
- Qwen3 Technical Report
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs