Generative Tutorial: Towards Live Contextualized Visual Instructions for Physical Tasks
cs.HC, cs.AI
Submitted: 2026-09-21
Updated: 2026-09-21
License: http://creativecommons.org/licenses/by/4.0/
The gist: Visual instructions for physical tasks are typically authored in one context and followed in another, requiring users to translate demonstrated tools, materials, and spatial relationships into their
Terminology
Abstract
Visual instructions for physical tasks are typically authored in one context and followed in another, requiring users to translate demonstrated tools, materials, and spatial relationships into their own environment. We introduce Generative Tutorial, a conceptual framework for live visual instruction that depicts intended outcomes and actions within the user's environment and task flow. A formative evaluation of state-of-the-art image and video generation identifies failures and potential benefits across 15 physical tasks. Drawing on these findings, we build an augmented-reality prototype system that proactively generates goal images and demonstration videos using observed workspace context and predicted visual outcomes of preceding actions. A 24-participant lab study found higher task performance quality, greater perceived workspace correspondence, and shorter step-confirmation intervals with the system than with pre-authored guidance. Qualitative findings highlighted how contextual resemblance shapes trust, how generation errors affect interpretation, and how guidance delivery should adapt to users' needs, informing future designs.
Sources
- WorldSmith: Iterative and Expressive Prompting for World Building with a Generative AI
- Vid2Coach: Transforming How-To Videos into Task Assistants
- Generative Lecture: Making Lecture Videos Interactive with LLMs and AI Clone Instructors
- Substantial, Decomposable, and Invisible: Visual Context Misalignment in Instructional Videos for Physical Tasks
- AROMA: Mixed-Initiative AI Assistance for Non-Visual Cooking by Grounding Multi-modal Information Between Reality and Videos
- eXplainMR: Generating Real-time Textual and Visual eXplanations to Facilitate UltraSonography Learning in MR
- Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
- Guided Reality: Generating Visually-Enriched AR Task Guidance with LLMs and Vision Models
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support