Making AI-Generated Feedback Matter: From Provision to Student Enactment

arXiv:2608.11625 · cs.AI · Submitted 2026-08-13 · Read on arXiv

Omar Alsaiari, Nilufar Baghaei, Jason M. Lodge, Dragan Gaševi'c, Naomi Winstone, Hassan Khosravi

The University of Queensland · University of Surrey · The University of Hong Kong

cs.AI

Submitted: 2026-08-13

Updated: 2026-08-17

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: This study examined how different AI-mediated feedback workflows were associated with students' behavioural engagement, self-assessment confidence, and submitted-work quality.

Terminology

Summary

This study examined how different AI-mediated feedback workflows were associated with students' behavioural engagement, self-assessment confidence, and submitted-work quality. Rather than treating AI-generated feedback comments as static products, the study compared three ways of organising students' encounters with AI-generated feedback comments or optional AI assistance: Directed Feedback, Self-Directed Feedback, and Enacted Feedback. Overall, the findings suggest that the educational value of AI-generated feedback comments depends not only on their availability or quality, but also on how students are guided to interpret, prioritise, discuss, and act on them. This supports contemporary feedback scholarship, which views feedback as an active process of learner sense-making and action rather than the passive transmission of information.

The study used a quasi-experimental sequential cohort design in which each implementation period corresponded to one feedback workflow condition. The Directed Feedback condition was implemented in Semester 1, 2025, the Self-Directed Feedback condition in Semester 2, 2025, and the Enacted Feedback condition in Semester 1, 2026. The dataset comprised 13,037 students, 51,296 student-authored resources, and 70 course offerings across the three feedback workflow conditions. Specifically, the Directed Feedback condition included 3,723 students, 14,425 resources, and 21 course offerings; the Self-Directed Feedback condition included 3,951 students, 15,548 resources, and 25 course offerings; and the Enacted Feedback condition included 5,363 students, 21,323 resources, and 24 course offerings.

For RQ1 (Uptake, revision counts, and event-flow transitions), a likelihood-ratio test showed a statistically significant overall effect of condition on uptake, χ2(2) = 5115.97, p <.001. Observed uptake was highest in Enacted Feedback, where 6,165 of 21,323 resources showed a downstream edit (28.9%), followed by Directed Feedback, where 2,716 of 14,425 resources showed an immediate edit (18.8%), and Self-Directed Feedback, where 32 of 15,548 resources showed an edit after requesting AI assistance (0.2%). Model-estimated probabilities followed the same pattern: Enacted Feedback, 26.2% (95% CI [25.3%, 27.2%]); Directed Feedback, 14.1% (95% CI [13.3%, 15.0%]); and Self-Directed Feedback, 0.1% (95% CI [0.1%, 0.2%]). Tukey-adjusted pairwise comparisons indicated that the odds of uptake were higher in Enacted Feedback than in Directed Feedback, OR = 2.16, simultaneous 95% CI [1.96, 2.38], corresponding to a model-estimated probability difference of 12.1 percentage points, simultaneous 95% CI [10.7, 13.5], p <.001. Directed Feedback also showed higher odds of uptake than Self-Directed Feedback, OR = 134.36, simultaneous 95% CI [87.84, 205.53], corresponding to a model-estimated probability difference of 14.0 percentage points, simultaneous 95% CI [13.0, 15.0], p <.001. Enacted Feedback showed higher odds of uptake than Self-Directed Feedback, OR = 290.18, simultaneous 95% CI [190.05, 443.07], corresponding to a model-estimated probability difference of 26.1 percentage points, simultaneous 95% CI [25.0, 27.3], p <.001.

For revision counts, a likelihood-ratio test showed a statistically significant overall effect of condition, χ2(2) = 5766.41, p <.001. Mean revision counts after capping were highest in Enacted Feedback (M = 0.87, SD = 1.82), followed by Directed Feedback (M = 0.46, SD = 1.26), and Self-Directed Feedback (M = 0.01, SD = 0.24). Model-estimated revision counts followed the same ordering: Enacted Feedback, 0.602 revisions per resource (95% CI [0.572, 0.635]); Directed Feedback, 0.239 (95% CI [0.222, 0.257]); and Self-Directed Feedback, 0.0034 (95% CI [0.0027, 0.0042]). Tukey-adjusted pairwise comparisons showed that Enacted Feedback exceeded Directed Feedback by 0.363 revisions per resource, simultaneous 95% CI [0.325, 0.402], IRR = 2.52, simultaneous 95% CI [2.29, 2.77], p <.001. Enacted Feedback exceeded Self-Directed Feedback by 0.599 revisions per resource, simultaneous 95% CI [0.561, 0.637], IRR = 177.67, simultaneous 95% CI [136.99, 230.41], p <.001. Directed Feedback exceeded Self-Directed Feedback by 0.236 revisions per resource, simultaneous 95% CI [0.215, 0.256], IRR = 70.46, simultaneous 95% CI [54.27, 91.49], p <.001.

Event-flow transitions were analysed using first-order Markov models. In Directed Feedback, 78.6% of transitions from AI Feedback moved directly to Self-Assessment, while a smaller proportion moved immediately to editing states (8.2% to Question and 10.0% to Options). In Self-Directed Feedback, only 1.0% of drafted resources moved from Draft Resource to AI Assistance after support was requested, while 99.0% did not enter the optional AI assistance pathway. In Enacted Feedback, 63.0% of resources moved from AI Feedback to selected suggestions and 37.0% proceeded without selecting suggestions, with 28.9% showing downstream editing after feedback comments.

For RQ2 (Self-assessment confidence), a likelihood-ratio test showed a statistically significant overall effect of condition, χ2(2) = 139.06, p <.001. Observed confidence was high in all conditions: Enacted Feedback (M = 4.19, SD = 0.71), Directed Feedback (M = 4.13, SD = 0.72), and Self-Directed Feedback (M = 4.02, SD = 0.76). EMMs on the expected rating scale indicated the same descriptive ordering: Enacted Feedback, 4.20 (95% CI [4.18, 4.21]); Directed Feedback, 4.13 (95% CI [4.11, 4.15]); and Self-Directed Feedback, 4.03 (95% CI [4.00, 4.05]). Tukey-adjusted pairwise comparisons showed that Enacted Feedback was associated with higher cumulative odds of reporting higher confidence than Directed Feedback, OR = 1.41, simultaneous 95% CI [1.21, 1.65], with an expected-rating difference of 0.070, simultaneous 95% CI [0.039, 0.101], p <.001. Enacted Feedback was also associated with higher cumulative odds of reporting higher confidence than Self-Directed Feedback, OR = 2.44, simultaneous 95% CI [2.04, 2.91], with an expected-rating difference of 0.170, simultaneous 95% CI [0.136, 0.204], p <.001. Confidence was also higher in Directed Feedback than in Self-Directed Feedback, OR = 1.73, simultaneous 95% CI [1.43, 2.08], with an expected-rating difference of 0.100, simultaneous 95% CI [0.066, 0.134], p <.001.

For RQ3 (Submitted-work quality), a likelihood-ratio test showed a statistically significant overall effect of condition, χ2(2) = 251.30, p <.001. Observed moderation outcomes were high across conditions, with Enacted Feedback showing the highest mean score (M = 4.22, SD = 0.51), followed by Self-Directed Feedback (M = 4.18, SD = 0.47), and Directed Feedback (M = 4.13, SD = 0.55). Back-transformed EMMs on the original 0–5 scale indicated that Enacted Feedback had the highest estimated moderation score, 4.328 (95% CI [4.317, 4.338]), followed by Self-Directed Feedback, 4.244 (95% CI [4.231, 4.257]), and Directed Feedback, 4.191 (95% CI [4.177, 4.205]). Tukey-adjusted pairwise comparisons showed that Enacted Feedback exceeded Directed Feedback by 0.137 points, simultaneous 95% CI [0.116, 0.158], with an expected-proportion odds ratio of 1.24, simultaneous 95% CI [1.20, 1.28], p <.001. Enacted Feedback exceeded Self-Directed Feedback by 0.083 points, simultaneous 95% CI [0.063, 0.103], with an expected-proportion odds ratio of 1.15, simultaneous 95% CI [1.11, 1.18], p <.001. Self-Directed Feedback exceeded Directed Feedback by 0.054 points, simultaneous 95% CI [0.031, 0.076], with an expected-proportion odds ratio of 1.08, simultaneous 95% CI [1.05, 1.12], p <.001.

The study concludes that the educational value of AI-generated feedback comments appears to depend on how students are supported to use them, not simply on whether comments are provided. The Enacted Feedback workflow was associated with higher workflow-specific uptake, more revisions, higher self-assessment confidence, and higher submitted-work quality than either comparison workflow. The central distinction drawn in this study is between access and enactment. Making AI-generated feedback comments available, or giving students optional access to AI assistance, was not accompanied by the level of feedback use observed when enactment was scaffolded. The optional pathway showed very limited uptake, indicating that availability alone did not lead most students to engage with AI support. The strongest engagement and outcomes were observed in the workflow that prompted students to select feedback suggestions, consider their relevance, engage in targeted AI-supported dialogue, and revise their work.

Improvements for AI systems

Improvements to AI Systems:

  1. Add an Enactment Scaffold Mode – Instead of only generating feedback comments or offering optional AI assistance, the AI should actively prompt students to select specific suggestions, justify their relevance, and then guide them through targeted revision steps. This turns passive feedback into a structured action loop.

  2. Implement a Suggestion Selection Interface – The AI should present feedback as discrete, selectable suggestions (e.g., Improve thesis clarity, Add evidence for claim X) rather than as a block of text. Students pick which to apply, and the AI tracks their choices to drive follow-up dialogue.

  3. Build a Guided Revision Dialogue – After a student selects a suggestion, the AI should engage in a short, focused conversation (e.g., Why is this suggestion relevant? What change would you make?) before allowing them to edit. This mirrors the Enacted Feedback workflow’s high engagement.

  4. Add a Post-Feedback Edit Tracker – The AI should automatically detect whether a student made a downstream edit after receiving feedback, and if not, send a nudge or offer a mini-prompt (e.g., You selected two suggestions—would you like to revise now?). This increases uptake from 0.2% (optional) toward 29% (enacted).

  5. Incorporate Confidence Calibration Prompts – Before and after feedback use, the AI should ask students to rate their confidence on a 1–5 scale. This not only measures self-assessment but also primes metacognitive reflection, which was associated with higher confidence in the enacted condition.

  6. Enable Feedback-to-Revision Mapping – The AI should link each feedback comment to specific revision actions (e.g., Add a counterargument → Insert one sentence addressing an opposing view). This makes the path from comment to edit explicit, increasing revision counts (0.60 vs. 0.24 per resource).

  7. Add a Workflow-Specific Uptake Dashboard – For instructors, the AI should report not just whether feedback was viewed, but whether it was selected, discussed, and acted upon. This allows real-time intervention when students are not enacting feedback.

What the Improved AI System Can Do:

  • Triple the rate of feedback-driven revisions by guiding students through selection, justification, and revision steps (from 19% in directed mode to 29% in enacted mode).

  • Raise submitted-work quality scores by 0.14 points on a 0–5 scale compared to simple feedback delivery, and by 0.08 points compared to optional AI assistance.

  • Increase students’ self-assessment confidence by 0.17 points (on a 1–5 scale) over optional AI help, by embedding confidence checks into the feedback loop.

  • Reduce passive feedback consumption – students no longer just read comments; they actively choose, discuss, and apply them, leading to 2.5× more revisions per resource.

  • Provide instructors with actionable analytics on which feedback suggestions are selected, discussed, and turned into edits, enabling targeted support for under-engaged students.

  • Automatically detect and recover from “feedback drop-off” – if a student views feedback but doesn’t edit, the AI can re-engage them with a targeted prompt, closing the gap between access and enactment.

Sources

Related papers