iAm.md: Robot Skill Self-Assessment through Agentic Introspection for Unknown Open-Vocabulary Domains
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "iAm.md: Robot Skill Self-Assessment through Agentic Introspection for Unknown Open-Vocabulary Domains".
Rosa: The gist The paper presents iAm.md, a Markdown standard and generation framework, that allows anchoring agentic introspection in robot behavior generation through open-vocabulary semantic mapping and persistent object records.
Dev: First, who's behind it and why it matters.
Paper summary: Rosa: To wrap up this paper, we're talking about "iAm.md: Robot Skill Self-Assessment through Agentic Introspection for Unknown Open-Vocabulary Domains" by Guarino, Musumeci, Suriani and Nardi. They laid out a framework where an embodied foundation model can use persistent scene memory and the robot's own description to assess tasks in open-vocabulary settings.
Dev: The core contribution here is that they established iAm.md as a Markdown standard for describing the robot's persistent embodiment, including its URDF structure, sensing capabilities, and explicit unknowns.
Taro: This standardized description allows the agent to introspect by combining queries from the semantic map with this robot description when trying to generate code for a natural language request.
Rosa: They proved that this combination supports skill self-assessment and executable task generalization during their work on simulated navigation and manipulation tasks, showing how it helps the agent revise its plan based on execution outcomes.
Dev: In terms of what this means for us, it suggests that by providing a foundation model with both scene memory and a structured representation of the robot's configuration via iAm.md, we give it a much better way to handle tasks where it doesn't have explicit training data.
Taro: It shifts the focus from just generating plausible-sounding plans to having an agent that actively checks its assumptions against physical reality before making a move in an unknown environment.
Rosa: The paper shows that using iAm.md, in conjunction with semantic maps and live ROS two observations, leads to lower construction costs and verified task completion compared to other methods they compared <ref:2610.10962#pg2>.
Dev: They also showed that this approach allows for few-shot generalization from just one simulated run when extending the tool to a new object geometry without needing to re-read the iAm.md evidence.
Taro: The authors of iAm.md are essentially showing how to create a system where an embodied agent can use persistent memory and self-description as tools for more reliable, introspective task execution in complex, open-vocabulary domains.
Conclusion: Rosa: So, we’re wrapping up on iAm.md, this paper by Guarino and his team about using an AI agent to figure out what a robot can actually do in a situation it hasn't seen before.
Dev: Exactly, it’s about this standard format for describing the robot itself so that the agent can check its own work against the real hardware.
Taro: The authors show how combining that persistent memory with live observations lets the agent self-assess its skill level without making a huge mistake on a single try.
Rosa: It seems like they’re moving away from just letting an AI generate code and trusting it blindly, to having it actually inspect its own plans against the robot's known capabilities.
Dev: That inspection loop, where the agent checks the ROS two graph while trying to execute a command, that sounds like a crucial layer for dealing with those context window problems we’ve been seeing in generation.
Taro: It addresses the "grounding failure" directly by forcing the agent to verify if its plan actually matches what's possible given the physical setup and software interfaces.
Rosa: For someone just listening, it means an AI robot stops being a black box that just spits out code, and starts being something that has a built-in way to ask itself, "Can I really do this right now?"
Dev: And the results they got were pretty solid; they saw lower construction costs when using this iAm.md evidence compared to other methods in their test.
Taro: Plus, they showed this system can even handle new objects or different shapes after only one simulated run without needing all that setup info again.
Rosa: So, the big picture here is moving toward AI agents that aren't just smart predictors but are actually grounded in the reality of their physical world and their own robot description.
Dev: It’s about building a more reliable execution loop for these foundation model agents when they step outside the perfect training data they were built on.
Taro: And what we haven't seen yet is how robust this becomes when you take it out of the simulation and onto a physical robot in a truly messy, open environment.
Vincenzo Guarino, Emanuele Musumeci, Vincenzo Suriani, Daniele Nardi
Department of Computer, Control and Management Engineering “Antonio Ruberti”, Sapienza University of Rome
cs.RO, cs.AI
Submitted: 2026-10-07
Updated: 2026-10-07
Project page: https://yurimachine.github.io/iAm.md
The gist: The gist The paper presents iAm.md, a Markdown standard and generation framework, that allows anchoring agentic introspection in robot behavior generation through open-vocabulary semantic mapping and
Key concepts
- Agentic Introspection
- This is when an AI agent inspects the information available inside its own planning process before generating code. It allows the agent to judge how well it can actually perform a task given the robot's capabilities and the current environment, addressing potential 'grounding failures' where generated behaviors might be unsupported.
- Open-Vocabulary Semantic Mapping
- This technique combines local vision-language detections with object segmentation to create persistent object records. These records store spatial data like centroids and bounding boxes, which downstream programs use to resolve natural language references by comparing new observations against stored supports.
- iAm.md
- This is a standardized text format for describing a robot's persistent embodiment, including its physical structure, software functions, and known unknowns. It allows foundation models to read this context when planning actions, anchoring their generated plans in concrete robot reality.
Terminology
Summary
The gist The paper presents iAm.md, a Markdown standard and generation framework, that allows anchoring agentic introspection in robot behavior generation through open-vocabulary semantic mapping and persistent object records.
Agentic Introspection and Grounding Failure
Agentic AI based on Large Language Model generalization capabilities offers a wide range of potential applications, including planning for embodied tasksDue to limited context windows or hallucinatory phenomena in the next-token prediction formulation, behaviors may be generated without establishing whether the deployed robot and the observed environment actually support the requested operation, in what we call a grounding failure
. Robot self-assessment frames the missing judgment as the ability to predict, estimate, or measure how well it can perform a task in a given context and environment
. Here agentic introspection means the agent is capable of inspecting the representations available inside its own planning process before committing to generated code.
Open-Vocabulary Semantic Mapping
The methodology involves combining local vision-language detections and object segmentation and refer them to persistent object records in this intermediate standardized representationA locally served vision-language model (Qwen3.6-35B via llama.cpp) proposes whole-object detections with open-vocabulary labels and bounding boxes. After non-maximum suppression, SAM 2 produces a binary mask for each surviving proposal. Each validated mask is reprojected through depth into a map-frame point cloud, cleaned by morphological closing, foreground filtering, and DBSCAN clustering. These stored voxels yield centroids, bounds, and oriented bounding boxes exported as a JSON snapshot and RViz markers that downstream programs query to resolve NL references. Association is decided by a voxeloverlap score between the new observation A and each stored support B, s(A, B) = max(A ∩ B / A, A ∩ B / B) (1).
iAm.md: Embodiment Evidence
iAm.md is a standard for describing the persistent embodiment of a configured robot deployment in plain text that a foundation-model agent can read in contextA conforming document follows a fixed schema from identity and embodiment summary, through the URDF-derived physical structure, actuation, and sensing, to the higher-level software functions exposed by the robot stack and the unknowns that remain after generation. Every block carries explicit cardinality, a value that cannot be defensibly established is recorded as unknown in place, and software-function records separate the acceptance criteria of a command interface from the physical outcome of the commanded operation. An instance is generated once per deployment from the URDF, augmented with the live ROS 2 graph and configuration artifacts for otherwise-unobservable interfaces.
Agentic Introspection and Execution Loop
The system places a tool-calling foundation-model agent between a task request and the deployed robotPersistent robot facts enter the agent context through iAm.md, whereas the current pose and camera observation describe the state in which the request must be executed. The semantic map supplies object identities and spatial support through ROS 2 topics and a serialized snapshot, so generated programs resolve task targets without placing transient observations in the persistent robot description. Each NL request is one execution episodeThe agent inspects the ROS 2 graph and interface schemas through the same command tool it uses to act, so interface availability is established from the deployed system during the episode. Skill self-assessment runs inside the agent loop: the agent relates the request to observed objects, then checks the robot description and live interfaces needed to realize the operation/navigation needs object location and navigation interface, whereas manipulation additionally requires manipulator and gripper contracts. When evidence supports the operation the agent proceeds to observation, program construction, or a physical command; otherwise a structured outcome records a blocked result.
Experimental Findings on iAm.md Contribution
The two evaluations identify complementary roles for persistent and live evidence during skill selfassessmentThe semantic map supplies the object referents and spatial estimates needed to interpret a request; iAm.md relates those estimates to the configured robot; and the resulting assessment is expressed through executable checks that change when execution exposes an unsupported assumption. In a controlled comparison, access to iAm.md was associated with lower construction cost and verified completionThe documented agent produced a verified tool in 50.4 minutes, versus the 110-minute budget without a verified grasp for the undocumented agent, expending roughly twice the calls and tokens. Retaining that tool, the agent extended it to a second task with different object geometry without rereading iAm.md-evidence of few-shot generalization, from a single simulated run.
Conclusion
We showed that an embodied foundation-model agent can combine persistent scene memory, a robot self-description, and live ROS 2 observations to assess and realize NL tasks in an open-vocabulary domain. The semantic map recovered all twelve annotated objects with one duplicate record. In a controlled comparison, access to iAm.md was associated with lower construction cost and verified completion. Retaining that tool, the agent extended it to a second task with different object geometry without rereading iAm.md-evidence of few-shot generalization, from a single simulated run. Future work will evaluate iAm.md generation across embodiments and add an independent safety boundary validating trajectories before actuation on physical robots.
--- Page 1 ---
iAm.md: Robot Skill Self-Assessment through Agentic Introspection for Unknown Open-Vocabulary DomainsVincenzo Guarino, Emanuele Musumeci, Vincenzo Suriani and Daniele NardiDepartment of Computer, Control and Management Engineering “Antonio Ruberti”, Sapienza University of Rome, Via Ariosto 25, 00185 Rome, ItalyAbstract Agentic AI based on Large Language Model generalization capabilities offers a wide range of potential applications, including planning for embodied tasks. For example, embodied agents based on Foundation models can generate plausible plans in autonomous robotics scenarios. Due to limited context windows or hallucinatory phenomena in the next-token prediction formulation, behaviors may be generated without establishing whether the deployed robot and the observed environment actually support the requested operation, in what we call a grounding failure
. Thanks to the recent improvements in reasoning capabilities of foundation models, autonomous robot behavior generation problem can be formulated as a code generation problem. We present iAm.md, a Markdown standard and generation framework, that allows anchoring this process in complementary forms of deployment evidence. Through open-vocabulary semantic mapping, we combine local vision-language detections and object segmentation and refer them to persistent object records in this intermediate standardized representation, allowing agentic introspection. We then study this new technique on a simulated TIAGo, on navigation-and-manipulation tasks, showing how this standardized representation jointly supports skill selfassessment and executable task generalization. Keywords Robot skill self-assessment, agentic introspection, open-vocabulary semantic mapping, foundation-model agents, robot code generation, ROS 2
--- Page 2 ---
semantic memory that uses local vision-language inference and segmentation to provide persistent object identities and locations; (ii) iAm.mda structured Markdown standard and generation framework representing specific robot information, resolved interfaces, and explicit unknowns. and (iii) an introspective embodied agent that combines semantic-map queries, introspection, task code generation, and robot execution for NL requests
--- Page 3 ---
Figure 1: System architecture of the robot-description generation and embodied-agent execution stages.The generation agent extracts resolved robot evidence from the ROS 2 environment and records it in iAm.md. During task execution, the embodied agent combines this description with the semantic map and NL request, produces code for successive subgoals, and uses execution outcomes for introspective revision.
--- Page 4 ---
4.1. Semantic-Map EvaluationTwelve objects in a simulated scene were independently annotated as labelled spheres in the map frame through an interactive RViz interface.
Improvements for AI systems
-
The system can perform
skill self-assessment
by relating a task request toobserved objects
and checking against a structured robot description, allowing it to determineif evidence supports the operation
or record astructured outcome records a blocked result.
-
The agent gains capability in generalization through the process where plans are updated based on execution outcomes, as shown by the ability to
extend it to a second task with different object geometry without rereading iAm.md-evidence of few-shot generalization, from a single simulated run.
-
The system can overcome
grounding failure
by using open-vocabulary semantic mapping combined with persistent object records to anchor planning processes incomplementary forms of deployment evidence,
ensuring generated actions are supported by the actual deployed environment.
Sources
- Evaluating Large Language Models Trained on Code
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- SAM 2: Segment Anything in Images and Videos
- Open3D: A Modern Library for 3D Data Processing
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving