Empowering Children to Create AI-Enabled Augmented Reality Experiences

arXiv:2508.08467 · cs.HC, cs.AI, cs.GR, cs.PL · Submitted 2026-08-12 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Empowering Children to Create AI-Enabled Augmented Reality Experiences".

Jane: The paper was written by Lei Zhang, Shuyao Zhou, Amna Liaqat, Tinney Mak, Brian Berengard et al. from New Jersey Institute of Technology and Princeton University and The Clubhouse Network.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: Welcome back to the arXiv audio channel, everyone. I'm Tom, and today we're cracking open a paper that honestly made me smile the moment I read the title. It's called "Empowering Children to Create AI-Enabled Augmented Reality Experiences."

Jane: And I'm Jane. Tom, this is such a refreshing topic. We talk a lot about AI for adults, but this team from Princeton and the New Jersey Institute of Technology is asking a much more interesting question: what happens when you hand the keys to a seven-year-old?

Tom: Exactly. The authors are Lei Zhang, Shuyao Zhou, Amna Liaqat, Tinney Mak, Brian Berengard, Emily Qian, and Andrés Monroy-Hernández. And they built this app called Capybara.

Jane: Capybara. I love that. The cute, chill animal that's all over the internet right now. It's a block-based programming environment that runs in augmented reality on an iPad.

Tom: Right. So instead of just playing a game where a character does what the developer coded, kids get to build the whole thing themselves. They drag and drop code blocks, just like Scratch, but the output is a three dee character standing in their living room.

Jane: And here's the kicker, Tom. The kids can generate that character using their own voice. They speak a prompt, the app uses generative AI to create a three dee model, and then they can program it to react to the physical world around them.

Tom: I mean, when I was a kid, the most advanced thing I did was make a sprite bounce off the edge of a screen. These kids are making a panda wear a wizard hat and balance on a banana in their kitchen.

Jane: And that's the core of what the paper is arguing for. We've had constructionist tools like Scratch for decades, but the ceiling for what kids can make has been stuck in 2D. This team is trying to raise that ceiling.

Tom: It's a big swing. And the fact that they tested it with twenty children across the United States and Argentina makes it feel real, not just a lab demo.

Jane: Absolutely. And I think the cultural piece is important too. They didn't just test in one place. They wanted to see if kids from different backgrounds would engage with it differently.

Tom: So, Jane, what's the one thing you want listeners to remember about this paper before we go deeper?

Jane: That the goal isn't just to make a cool app. It's to shift kids from being consumers of AI-powered AR to being creators of it. That's a mindset shift with huge implications.

Tom: And we're just getting started. Next up, we're going to dig into the actual system and how it works under the hood.

Summary and Core System: Tom: Alright, we're back with "Empowering Children to Create AI-Enabled Augmented Reality Experiences." Jane, we talked about the big idea. Now let's get into the meat of the system. What does Capybara actually do?

Jane: So there are three big features that set it apart from anything else out there for kids. First, you can create three dee characters and accessories using speech-to-three dee generation. You hold a button, say "a cute panda with a wizard hat," and the app generates it.

Tom: And that's not just a filter. That's a full three dee mesh that gets placed into the AR scene. But here's the thing that blew my mind. The app auto-rigs that character. It figures out where the joints are so the character can move.

Lu: That's the part I want to jump in on, Tom. Auto-rigging is a notoriously hard problem in computer graphics. The team used an algorithm based on Baran and Popović's work from two thousand seven which is a classic. It's computationally light enough to run on a tablet, which is impressive.

Jane: And Lu, that's what enables the second feature. Because the character has a skeleton, kids can puppeteer it. They can record their own body movements, and the character mirrors them. So if you do a jumping jack, the panda does a jumping jack.

Meng: But that's where I start asking engineering questions. Running an auto-rigging algorithm on-device, plus a YOLOv11 object detection model, plus ARKit tracking, all at the same time on an iPad? That's a lot of compute.

Tom: Meng, that's a fair point. The paper says the auto-rigging takes less than five minutes on an iPad Pro with the M4 chip. So it's not real-time, but it's acceptable for a creative workflow.

Meng: Okay, that's reasonable. And they're using YOLOv11s, which is the small version, so it's optimized for edge devices. I'm glad they didn't try to run a massive model.

Jane: And the third feature is what really bridges the gap between the virtual and physical worlds. Kids can program the character to react to physical objects. They use a "touches object" block, pick a banana from a list, and then the character can detect when it's near a real banana in the camera feed.

Lu: The clever part is the interaction design. They didn't just make it a technical demo. They made it a programming concept. The "touches object" block is an event trigger, just like in Scratch. So kids are learning computational thinking while they're playing.

Tom: And they also have "touches zone" blocks, where kids can define a region on the floor. So you could program the character to do a dance when it steps into the "dance zone."

Jane: Right. And the whole thing is driven by a "when character is touched" event block. The kid sticks their hand out in front of the camera to pat the character, and that starts the program.

Meng: I like that. It's a physical trigger for a digital action. It makes the programming feel tangible.

Tom: So we've got the three pillars: generative AI for assets, puppeteering for animation, and computer vision for physical interaction. And they tested it with kids. We'll get into those results next.

Improvements and User Study Insights: Tom: Welcome back. We're still on "Empowering Children to Create AI-Enabled Augmented Reality Experiences." We've covered what the system does. Now, Jane, what did the kids actually do with it?

Jane: The study involved twenty children aged seven to sixteen split across the US and Argentina. And the results are genuinely encouraging. Every single participant was able to complete the warm-up exercise, which required them to use all the core features.

Lu: That's a strong signal. The learning curve is shallow enough that even a seven-year-old can get from zero to a working AR program in under twenty minutes of instruction.

Tom: And the open-ended tasks were where it got fun. One girl in the US created a character that looked like herself, using a long prompt describing her hair and clothes. She was literally putting herself into the AR world.

Jane: And in Argentina, a group created a clown wig accessory and programmed the character to perform in front of a set of physical toys. They were building narratives that blended their physical space with the virtual character.

Meng: So the customization features were the big draw. The kids felt ownership over the characters because they made them.

Jane: Exactly. One participant said it made the experience "more like yourself, more expressing yourself." That's the creative freedom they were going for.

Tom: But it wasn't all smooth sailing. The paper is honest about the challenges. The biggest one was AI alignment. The kids would ask for something, and the AI would give them something slightly off.

Lu: Right. One kid asked for a capybara and got one with a tail and "really, really big teeth." Another noticed that when the character did jumping jacks, the stomach folded "like a tortilla." The AI isn't perfect, and kids notice.

Meng: And that's a double-edged sword. On one hand, it teaches them that AI has limitations. On the other hand, it can frustrate them and distract them from their creative goals. The paper notes that some kids spent too much time refining prompts instead of building their story.

Jane: That's a real design tension. How do you keep the magic of AI without letting it derail the creative process? The authors suggest future work could give kids more agency, like letting them manually adjust the rigging joints.

Tom: And there's another insight I loved. The kids were actually reasoning about how the AI works. One participant noted that object detection might fail if the physical object looks different from the training data, like a different brand of the same item.

Lu: That's AI literacy in action. They're not just using the tool; they're building mental models of how it works. That's a huge pedagogical win.

Jane: The kids also said they wanted more. They wanted sound effects, the ability to generate virtual environments like a school or a hospital, and more sophisticated interactions like the character picking up a physical orange.

Meng: So the ceiling is still not high enough for them. That's a good problem to have.

Tom: It is. And it points to a future where these tools are even more expressive. But we've got to wrap up soon. Let's get to the big picture in our final segment.

Conclusion: Tom: We're at the end of our discussion on "Empowering Children to Create AI-Enabled Augmented Reality Experiences." Jane, give it to me straight. What's the legacy of this paper?

Jane: I think the legacy is proving that kids can be the authors of complex AI-powered AR experiences, not just the audience. The Capybara system shows that with the right design, a seven-year-old can generate three dee assets, rig them, animate them, and program them to interact with the physical world.

Tom: And that's a fundamentally different relationship with technology. Instead of consuming a game, they're building a world.

Lu: And the implications go beyond just fun. The paper shows that kids engage with computational thinking concepts like loops and conditionals, and they start to build AI literacy by reasoning about the models' limitations. That's a foundation for critical thinking in a world full of AI.

Meng: From an engineering standpoint, I'm impressed they got this running on a tablet. The on-device auto-rigging and real-time object detection are no small feats. It makes me think about how we can push more of this to edge devices.

Jane: And Lalam, you've been quiet. What's your take on the cultural impact?

Lalam: I think the cultural impact is about democratizing creation. When you give children the tools to express themselves with AI and AR, you're not just teaching them to code. You're giving them a new language for storytelling. The fact that the study was run in both the US and Argentina shows that this desire to create is universal. The kids in both countries made personal, playful, and meaningful experiences.

Tom: That's a beautiful way to put it. So, as we say goodbye to this paper, what's the one thing you want our listeners to remember?

Jane: That the future of AI and AR isn't just about smarter apps. It's about empowering the next generation to build with these tools. Capybara is a stepping stone, and I can't wait to see what these kids build next.

Tom: And on that note, we're signing off on "Empowering Children to Create AI-Enabled Augmented Reality Experiences." Thanks for listening, and we'll see you on the next one.

Lei Zhang, Shuyao Zhou, Amna Liaqat, Tinney Mak, Brian Berengard, Emily Qian, Andrés Monroy-Hernández

New Jersey Institute of Technology · Princeton University · The Clubhouse Network

cs.HC, cs.AI, cs.GR, cs.PL

Submitted: 2026-08-12

Comments: Accepted to ACM UIST 2025

DOI: 10.1145/3746059.3747662

Code: https://github.com/ultralytics/ultralytics

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 67/100

The gist: This paper introduces Capybara, an AR-based and AI-powered visual programming environment that empowers children to create, customize, and program 3D characters overlaid onto the physical world.

Key concepts

Capybara
A block-based programming environment running in augmented reality on an iPad. It allows children to build experiences by dragging and dropping code blocks.
Speech-to-three dee generation
A feature where children speak a prompt, such as "a cute panda with a wizard hat," and generative AI creates the three-dimensional model for the character.

Terminology

Summary

This paper introduces Capybara, an AR-based and AI-powered visual programming environment that empowers children to create, customize, and program 3D characters overlaid onto the physical world. Capybara enables children to create virtual characters and accessories using text-to-3D generative AI models, and to animate these characters through auto-rigging and body tracking. In addition, our system employs vision-based AI models to recognize physical objects, allowing children to program interactive behaviors between virtual characters and their physical surroundings. We demonstrate the expressiveness of Capybara through a set of novel AR experiences. We conducted user studies with 20 children in the United States and Argentina. Our findings suggest that Capybara can empower children to harness AI in authoring personalized and engaging AR experiences that seamlessly bridge the virtual and physical worlds.

The paper's contribution is threefold: 1) Capybara, an authoring tool for children to create AI-enabled AR experiences, which seamlessly combines AI techniques with children’s content creation process; 2) a set of example experiences demonstrating the expressiveness of the tool; and 3) empirical insights gained from our user studies with children, highlighting its benefits and challenges of using Capybara.

The system has three primary design goals: D1 - Lowering the barrier of entry for programming AI-enabled AR behaviors, D2 - Enabling higher customizability, and D3 - Facilitating interactions between the virtual and physical worlds.

Capybara is based on a block-based programming environment in AR. The code blocks feature an interlocking, puzzle-like design that visually guides users in constructing functional programs. These blocks are categorized into five distinct groups, each with a unique color for easy identification: (1) Event Blocks (Yellow), (2) Motion Blocks (Dark Blue), (3) Looks Blocks (Purple), (4) Control Blocks (Orange), and (5) Sensing Blocks (Light Blue).

The system integrates Generative AI (GenAI) functionalities to allow children to create personalized AR experiences through customized characters and accessories. The character customization process is initiated through the Settings menu, where users can select Custom Character to begin. Accessory customization, on the other hand, is triggered directly within the inventory interface using the start wear code block. Both character and accessory generation workflows follow a similar UI pattern. Upon initiating the generation process, users see a dedicated creation page asking them to describe their desired character or accessory using speech. The system uses the Google Gemini 1.5 Flash LLM to review the prompt and ensure the appropriateness of the request for a young audience. Once the prompt passes the filter, the system calls an external text-to-3D-object generation service.

Capybara also enables users to customize the character’s animation via puppeteering. The system refines and modifies the open-sourced auto-rigging algorithm proposed by Baran and Popović. In Capybara, the animation feature is integrated into the block-based programming environment by providing the code block named start animation. Users can drag the codeblock and place it on the canvas, which then guides them to define the animation clip that can be programmatically replayed.

A key limitation of existing authoring tools for children is that the outcome experiences remain in the digital world. One of the design goals of Capybara is to facilitate children to create interactions between the virtual and the physical world. The system leverages off-the-shelf vision-based AI models to detect objects in the user’s physical surroundings. Users can program interactions with physical objects using the touches object block. The system adapted the YOLOv11s model, a deep neural network for real-time object detection, and curated a set of 24 recognizable objects. The system also introduced the touches zone block as part of the Sensing Blocks. Zones are user-defined regions on a physical surface that also support collision detection, enabling interactions with parts of the physical surroundings that are not covered by object detection.

To demonstrate the complete authoring workflow of combining the above functionalities and showcase the expressiveness of Capybara, the paper offers a set of novel experiences, including: Getting ready for the day, Alphabet learning, and Ping-pong gameplay.

To understand the benefits and challenges of Capybara, the researchers conducted a series of user studies with children in the United States and Argentina. They recruited 20 participants (12 female and 8 male, age 7–16), including nine groups. Each study session lasted approximately 90 minutes. The study session began with an introduction and a walkthrough of the system that lasted approximately 20 minutes. After the walkthrough, participants were asked to complete a warm-up exercise, where they were asked to create the example experience depicted in Figure 7.a from scratch. Then, participants completed a 30-minute open-ended authoring task. After finishing all activities, participants filled out a system usability survey and concluded with an open-ended discussion.

The results show that all participants were able to complete the given warm-up exercise using Capybara. Participants generally felt that Capybara was easy to use (M=3.80, SD=1.01). They also thought that they enjoyed playing Capybara (M=4.35, SD=0.99) and that they felt confident when playing Capybara (M=4.40, SD=0.82).

The findings highlight several creative usage patterns: creating animations of physical activities for storytelling, generating 3D characters for both realistic and fantastical experiences, and building interactions between the virtual and the physical world.

The qualitative feedback revealed several benefits and challenges. Participants thought that Capybara feels unique due to its 3D coding environment and AI-enabled customization and interaction. They also found that Capybara offers creative freedom due to AI-enabled customization. However, the integration of AI also introduced challenges of AI alignments and distractions. Participants reported that Capybara facilitates programming interactions with the physical world. Finally, the findings suggest pedagogical potential for fostering Computational Thinking and AI literacy.

The discussion section highlights opportunities and challenges of the design of Capybara and discusses implications for future research at the intersection of AR, AI, and child-centered programming environments. The paper discusses programming intricate interactions between the virtual and physical worlds, balancing AI-enabled customization and children’s agency, fostering Computational Thinking and AI literacy, and safety, privacy, and equity towards large-scale deployment.

The paper also acknowledges limitations. System limitations include the auto-rigging algorithm's assumptions (closed and connected mesh, proportions roughly matching the predefined rigging skeleton), the speech-to-3D module's generation time, and the object detection's limited set of recognizable objects. Evaluation limitations include the small number of participants (N = 20) and the short interaction time, which may not generalize to all future users.

Improvements for AI systems

Based on the paper, here are specific improvements I can make to AI systems and what the improved system can do:


  • Improvement: Implement a feedback loop that compares the generated 3D mesh against the user's prompt using a vision-language model (e.g., CLIP-based scoring). If the semantic similarity is low, automatically re-prompt the text-to-3D service with refined keywords or request a regeneration.

  • What it can do: Reduces the frustration children experienced when generated characters didn't match their intent (e.g., a capybara that looked like a rat with a tail). The system would proactively detect mismatches and offer alternatives, preserving creative flow.

  • Improvement: Before rigging, use a lightweight neural network to estimate the generated mesh's limb proportions relative to the predefined skeleton. If proportions deviate beyond a threshold, automatically apply a non-uniform scaling or a retargeting step to better fit the skeleton, or suggest a different prompt to the user.

  • What it can do: Prevents the tortilla stomach artifact during animation (where the character's body folds unnaturally). The system would ensure that even non-standard characters (e.g., long-limbed or squat) animate smoothly, expanding the range of creatable characters.

  • Improvement: Monitor user behavior during the authoring process. If a user spends more than a threshold time (e.g., 3 minutes) on prompt refinement without making progress on the overall program, trigger a gentle, non-intrusive suggestion (e.g., You can use the character now and refine it later) or offer a set of pre-validated prompt templates that are more likely to succeed.

  • What it can do: Mitigates the observed distraction where children became absorbed in prompting and lost sight of their creative goals. The system would help maintain a balance between exploration and completion, fostering a more structured creative process.

  • Improvement: Integrate a monocular depth estimation model (e.g., MiDaS) alongside YOLOv11. When an object is detected, estimate its 3D bounding box using depth cues, rather than relying solely on 2D bounding boxes and raycasting.

  • What it can do: Enables more accurate collision detection between the virtual character and physical objects, especially when objects are partially occluded or at varying distances. This supports more intricate interactions, such as the character stepping onto a physical book or avoiding a wall, as requested by participants.

  • Improvement: Extend the zone detection to allow users to label zones with semantic meanings (e.g., dance floor, bed, danger zone) using speech. The system then uses a language model to interpret these labels and automatically adjusts the character's behavior (e.g., playing a dance animation when entering a dance floor zone).

  • What it can do: Reduces the programming burden for children by allowing them to express intent in natural language rather than explicit conditional logic. For example, a child could say this zone is a pool and the character would automatically perform a swimming animation when touching it, without needing to manually program each behavior.

  • Improvement: When replaying a recorded animation (from puppeteering), apply a lightweight physics-based correction (e.g., using inverse kinematics) to prevent self-intersection and unnatural folding, especially for characters with non-standard proportions.

  • What it can do: Eliminates the unsettling folding artifacts observed during jumping jacks. The character's animations would look more natural, increasing immersion and reducing the cognitive dissonance that could distract from the storytelling experience.

  • Improvement: Based on the user's interaction patterns (e.g., repeated failures with object detection), dynamically inject short, age-appropriate explanations about how the AI works (e.g., The AI looks for patterns it learned from many pictures. If your object looks different, it might not recognize it.). Use a large language model to generate these explanations in real-time.

  • What it can do: Addresses the participants' misconceptions about AI (e.g., expecting perfect recognition regardless of brand or appearance). The system would turn errors into teachable moments, fostering AI literacy without requiring external instruction.

  • Improvement: Extend the speech-to-3D pipeline to also generate sound effects and ambient environments. For example, a child could say add a rain sound or make a forest background, and the system would use a text-to-audio model (e.g., AudioLDM) and a text-to-scene model to generate these assets.

  • What it can do: Directly addresses participants' desires for sound effects, animal noises, and virtual environments (e.g., a TV, a mountain). This would allow children to create richer, more immersive AR experiences without needing separate tools or advanced skills.

  • Improvement: Implement a two-tier moderation system: first, use the on-device LLM for immediate screening; if the prompt is flagged as borderline, escalate to a cloud-based LLM with a more comprehensive policy. Additionally, store all generated assets locally by default, with an explicit opt-in for cloud processing.

  • What it can do: Balances safety and privacy. The system would reduce the risk of inappropriate content slipping through while minimizing data exposure, making it more suitable for large-scale deployment in schools and homes.

  • Improvement: Allow multiple devices to share the same AR space (using ARKit's collaborative sessions). Children can program different characters or zones that interact with each other's creations in real-time.

  • What it can do: Supports social and collaborative learning, as seen in the pair-based study sessions. It would enable children to build joint stories or games, fostering teamwork and communication skills, and making the tool more engaging for classroom use.

In summary, the improved AI system would:

  • Generate 3D content that more reliably matches children's intent.

  • Animate characters more naturally, even with non-standard shapes.

  • Guide children to stay focused on their creative goals.

  • Understand the physical world more accurately for richer interactions.

  • Support natural language for defining complex behaviors.

  • Educate children about AI in real-time.

  • Expand creative possibilities to sound and environments.

  • Ensure safety and privacy at scale.

  • Enable collaborative, social AR experiences.

Abstract

Despite their potential to enhance children's learning experiences, AI-enabled AR technologies are predominantly used in ways that position children as consumers rather than creators. We introduce Capybara, an AR-based and AI-powered visual programming environment that empowers children to create, customize, and program 3D characters overlaid onto the physical world. Capybara enables children to create virtual characters and accessories using text-to-3D generative AI models, and to animate these characters through auto-rigging and body tracking. In addition, our system employs vision-based AI models to recognize physical objects, allowing children to program interactive behaviors between virtual characters and their physical surroundings. We demonstrate the expressiveness of Capybara through a set of novel AR experiences. We conducted user studies with 20 children in the United States and Argentina. Our findings suggest that Capybara can empower children to harness AI in authoring personalized and engaging AR experiences that seamlessly bridge the virtual and physical worlds.

Sources

Related papers