DexHoldem: An Agentic Robotics Benchmark for Dexterous Manipulation in Texas Hold'em

arXiv:2605.18727 · cs.RO, cs.AI · Submitted 2026-05-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "DexHoldem: An Agentic Robotics Benchmark for Dexterous Manipulation in Texas Hold'em".

Rosa: DexHoldem introduces a novel system-level benchmark for evaluating embodied agents that couples dexterous manipulation skills with agentic perception within a real-world Texas Hold'em tabletop setting.

Dev: First, who's behind it and why it matters.

Title and authors: Rosa: So, we're diving into the DexHoldem paper today, which looks at how to really test embodied systems when they have to handle complex physical tasks in a real environment. It seems the focus is moving beyond just testing isolated skills and seeing if an AI can actually manage a whole situation.

Dev: Exactly what I thought. The title itself, "DexHoldem: An Agentic Robotics Benchmark for Dexterous Manipulation in Texas Hold'em," makes it pretty clear that this isn't just about a simple pick-and-place task; it involves perception, policy execution, and a whole game state.

Taro: I'm interested in how they set up the benchmark. It sounds like they’re trying to catch systems that can handle the whole loop, not just one tiny part of it.

Rosa: That's right, and what's compelling is that they use a real-world Texas Hold'em tabletop setting with a ShadowHand platform for this research, which adds a lot of realism compared to pure simulation.

Dev: From an engineering standpoint, the system overview shows how they close the loop by parsing observations into game states and routing instructions before executing policies, so we can actually see where those potential latency or failure modes might creep in.

Taro: And that’s where I want to push: what happens when the world misbehaves? The paper mentions testing agents on whether they can recover from perceptual errors during closed-loop deployment, which is crucial for real-world autonomy.

Rosa: Right, and that recovery aspect is a big deal because it means we're looking at agents that can handle unexpected physical disturbances while still trying to maintain the game context.

Dev: I saw some numbers regarding the policy benchmark, where π0 point 5 managed a task completion rate of sixty-one point two percent, but then they noted that when counting disruptive completions too, π0 point 5 and another model tied on the scene-preserving success rate at forty-seven point five percent.

Taro: That tie on the scene-preserving success rate is interesting because it suggests that even if an agent can finish a task, it might not be doing so in a way that keeps the table usable for whatever comes next.

Title and authors: Rosa: Precisely, and this leads us into what they call instruction-conditioned dexterous manipulation, which is key—it’s not just about getting the card in your hand, but getting it there without making a mess of the rest of the game.

Dev: The agentic perception benchmark also shows a clear bottleneck where isolated sub-capabilities are strong, but routing-critical chip-state fields, like opponent chip inventory accuracy peaking at forty-three point eight percent, remain especially unreliable.

Taro: That low accuracy on those critical fields suggests that even if the perception module can see the scene well, getting it to make the right decision for long-horizon planning is where the real difficulty lies.

Rosa: So, what they did in terms of improvements to their approach was introducing these three distinct benchmarks: a standardized physical policy benchmark, an agentic perception benchmark, and a system-level evaluation of both working together.

Dev: The authors suggest that this combined approach is necessary because existing benchmarks usually only test one or the other—either isolated motor skills or simulation-based planning—and DexHoldem evaluates the whole integrated system.

Taro: I think what they are improving is the methodology itself by forcing these components to interact under a shared observation-action interface, which tests how errors accumulate across that chain.

Rosa: And the system-level evaluation specifically probes this "compounding closed-loop reliability gap," showing that agents and policies can solve parts of the benchmark but their errors just pile up over many states and primitive dispatches.

Dev: That accumulation is a major concern for me as a controls engineer; it means we need to focus heavily on how the system handles those repeated waiting, verification, and recovery events without timing out or escalating to an unnecessary human request.

Taro: That ties into the idea of embodied decision routing, where the agent needs to correctly map high-level game states directly into the correct sequence of low-level dexterous actions with minimal misrouting errors.

Title and authors: Rosa: Right, and that brings us to the real-world implications: this research suggests that for agents to be truly useful in complex physical tasks, they need robust methods for both fine motor control and accurate, structured game state tracking simultaneously.

Dev: If we can tackle those chip-state perception bottlenecks mentioned earlier, it opens the door for much more reliable strategic decision-making in embodied AI systems operating in dynamic physical spaces.

Taro: The impact here is that it moves the goalposts from just "can the hand do this?" to "can the agent use its perception to make a coherent, long-term plan based on what it sees and knows about the game state?"

Rosa: It really shows how important it is for these agents to understand not just where objects are, but their exact relationship within the context of a specific game strategy.

Dev: I wonder how this relates to other work we've seen, like FlowDPG or TCBiRRT, because those focus on policy and planning in different domains; DexHoldem tests if that same philosophy holds up when perception and complex physical contact are involved.

Taro: It seems the real value is in proving that combining those elements—dexterous manipulation with structured state awareness—is where the current limitations of models like GPT five point five and π0 point 5 become most apparent.

Rosa: So, to wrap up this discussion on DexHoldem, it’s a comprehensive setup that highlights the need for integrated evaluation when building agents for real-world physical interaction in structured environments like Texas Hold'em.

Dev: We see that while individual components can perform well, the system-level view reveals these compounding reliability gaps that need addressing through better state tracking and recovery logic.

Taro: The implication is that future work needs to focus heavily on building perception modules specifically tuned for the highly structured, but sometimes noisy, visual data found in tabletop games.

Rosa: It’s a solid piece of research because it sets a new standard for what it means to evaluate an agent that needs to be both physically dexterous and strategically aware in a shared physical setting.

The paper's summary: Rosa: So, we just got through the abstract for DexHoldem, which basically sets up this new way to test if an AI can handle real physical manipulation combined with smart game strategy in a table game setting.

Dev: Yeah, it lays out that they're not just looking at one skill or one planning method; they're testing the whole loop—perception feeding policy execution, all while dealing with the mess of a real-world environment.

Taro: From my perspective as an autonomy researcher, this is exciting because it moves away from isolated skills and forces the AI to make decisions based on a structured game state that changes constantly.

Rosa: Exactly, and what really hits me is how they frame it—they’re testing if an agent can actually perceive a changing physical scene, pick the right action for that moment, and keep track of the game context over a long sequence.

Dev: And from my control engineering standpoint, it’s interesting because they explicitly focus on closing that loop, which means they have to deal with latency and failure modes when the perception doesn't match what the policy expects.

Taro: I’m really focused on the agentic perception side here; it seems like they’re trying to see if an AI can correctly parse all those different game challenges—like who owns a turn or what chips are where—to route its next move.

Rosa: And that's where the benchmark gets interesting because they found some real bottlenecks, showing that even when individual parts are strong, getting the chip inventory right is proving tough.

Dev: That's a critical point for me; if the perception module can’t get those specific numerical states right, then the entire decision-making chain falls apart regardless of how good the physical policy is.

Taro: It suggests that we need to develop perception modules specifically tuned for that kind of structured data found in tabletop games, not just general vision models.

Rosa: And when we look at the system-level evaluation, they’re showing how these component errors actually compound across many captured states and primitive dispatches, which is a serious reliability concern.

Dev: That compounding gap is what worries me most; it means that even if an agent gets the first few steps right, subsequent failures can lead to total breakdown without a clean recovery mechanism.

Taro: It really hammers home the need for robust recovery logic that can handle those repeated waiting and verification events gracefully instead of just crashing.

Rosa: Overall, what this paper contributes is providing a standardized framework that evaluates dexterous execution and agentic perception as two interdependent parts of an embodied system.

Dev: The real implication here is that we’re moving toward evaluating integrated systems rather than just looking at individual models in isolation, which should help us build more reliable robotics.

Taro: It pushes the research toward creating agents that aren't just good at one thing, but can handle the messy reality of a physical task within a dynamic context.

Rosa: So, this work really sets a new bar for what we expect from AI systems intended for real-world physical interaction in complex environments.

Dev: It’s definitely an important step toward creating embodied agents that can operate reliably outside of just controlled simulation settings.

The paper's improvements: Rosa: So, we just went through how DexHoldem suggests improving things by focusing on those three distinct benchmarks: policy, perception, and system-level evaluation working together.

Dev: That’s right; it’s about moving away from just testing isolated skills or single planning methods and instead forcing them to interact under a shared interface.

Taro: I think the real improvement is in how they structure the agentic perception benchmark, making sure it tests parsing structured game states like loop stage and turn ownership, not just general scene understanding.

Rosa: And that directly addresses my question about whether this works outside the lab; if an AI can handle that kind of structured perception, it opens up possibilities for real-world applications where the environment isn't perfectly rendered in a simulation.

Dev: From an engineering standpoint, I’m looking at how they propose improving state tracking and recovery logic to handle those accumulated errors across multiple steps without needing constant human intervention.

Taro: Exactly, and that leads into the idea of better embodied decision routing, where the agent learns to map high-level game needs directly into a specific sequence of low-level physical actions with minimal misrouting.

Rosa: That’s a huge step because if an agent can correctly route its high-level strategy into precise physical movements, it starts looking like something that could actually function in a complex physical workspace.

Dev: And I’m interested in the data efficiency aspect they mentioned; it seems that once you have enough real-world dexterous data, initialization from a pre-trained policy can help speed up fitting the target action distribution under real constraints.

Taro: That’s interesting because it suggests that we don't always need to start training from scratch for these complex embodied tasks if we can leverage some existing skill knowledge.

Rosa: Ultimately, the implication is that future research needs to focus on building perception modules specifically tuned for the kind of structured, but often noisy, visual data found in these kinds of physical interactions.

Dev: And that brings up a huge question for us: how reliable are these systems when they encounter unexpected physical contact or friction disturbances during those long-horizon tasks?

Taro: That’s the core challenge; we need to see if they can reliably interpret and replan actions based on how physical forces affect thin objects, even when operating in simulated environments that use reconstructed visual data.

Rosa: So, this research is essentially laying out a path for building agents that aren't just skilled manipulators but are also strategically aware decision-makers capable of handling the real-world physics and uncertainty of a tabletop game.

Conclusion: Rosa: So we’ve wrapped up our deep dive into DexHoldem: An Agentic Robotics Benchmark for Dexterous Manipulation in Texas Hold'em, summarizing how this paper sets a new standard for evaluating integrated AI systems in physical tasks.

Dev: It really shows that the gap between isolated skill mastery and robust, closed-loop decision-making is where the real engineering work is needed right now.

Taro: I think what stands out most is how it forces us to look at perception not just as a vision module, but as a component vital for high-level strategic routing in dynamic environments.

Rosa: And from a field robotics standpoint, the fact that they test this in a real Texas Hold'em setting gives us confidence that these concepts could translate beyond the lab and into actual physical interaction scenarios.

Dev: I’m still thinking about those failure modes; if we can get the loop rate tight enough to minimize latency during those repeated verification steps, it makes the entire system much more viable for deployment.

Taro: For autonomy research, it suggests that future agents need better ways to handle perceptual uncertainty when making long-horizon decisions in a physical world.

Rosa: It certainly does; we need systems that can maintain context and make smart choices even when things get messy or unexpected, like the chip inventory tracking issues they pointed out.

Dev: So, the main implication is that we have to design for reliability across these stacked components—policy, perception, and routing—rather than just focusing on one part in isolation.

Taro: That’s the big picture here; it validates the idea that a successful embodied agent must be a cohesive unit where all parts work together seamlessly under pressure.

Rosa: It really is exciting because seeing this level of structured evaluation for complex physical tasks gives us a much clearer target for what we need to build next.

Dev: I agree; focusing on reducing those compounding reliability gaps across the entire execution chain is the most practical engineering hurdle ahead.

Taro: Moving forward, I think we should be looking at how this framework can be adapted for other complex physical environments that require both fine motor skills and real-time strategic reasoning.

cs.RO, cs.AI

Submitted: 2026-05-18

Updated: 2026-10-01

Comments: 35 Pages

Project page: https://dexholdem.github.io/Dexholdem

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 73/100

The gist: DexHoldem introduces a novel system-level benchmark for evaluating embodied agents that couples dexterous manipulation skills with agentic perception within a real-world Texas Hold'em tabletop

Key concepts

DexHoldem
A novel system-level benchmark for evaluating embodied agents. It couples dexterous manipulation skills with agentic perception within a real-world Texas Hold'em tabletop setting to test how AI manages a whole situation, not just isolated skills.
Agentic Perception Benchmark
A part of DexHoldem that tests an agent's ability to correctly parse structured game states, such as turn ownership and chip inventory accuracy. The discussion notes that this component shows bottlenecks in accurately tracking critical numerical states.
Compounding Closed-Loop Reliability Gap
A major concern where errors from individual components—like perception or policy execution—accumulate across many captured states and primitive dispatches. This gap means agents can fail over long sequences without a clean recovery mechanism.
Embodied Decision Routing
The idea that an agent must correctly map high-level game needs directly into the correct sequence of low-level dexterous actions with minimal misrouting errors. This is crucial for moving strategy into precise physical movements.

Terminology

Summary

DexHoldem introduces a novel system-level benchmark for evaluating embodied agents that couples dexterous manipulation skills with agentic perception within a real-world Texas Hold'em tabletop setting. This research matters because existing benchmarks often evaluate isolated motor skills or simulation-based planning, failing to test whether an agent can perceive a changing physical scene, choose context-appropriate actions, and maintain structured game state awareness for long-horizon decision making. DexHoldem addresses this gap by requiring agents to execute precise, contact-rich manipulation of thin cards and chips while simultaneously recovering from perceptual errors in a closed-loop deployment.

DexHoldem System Overview

DexHoldem is built around a real-world ShadowHand platform interacting with a Texas Hold'em tabletop environment. It provides 1,470 teleoperated demonstrations across 14 atomic card and chip primitives, including card pickup and placement together with chip pushing and pulling across multiple denominations. The system closes the loop by parsing observations into a game state, routing instructions, executing policies, and allowing for recovery from failures. This setup allows for the evaluation of instruction-conditioned dexterous tabletop manipulation in a setting where objects like thin cards require contact-rich manipulation under friction and disturbance uncertainty.

Policy Benchmarking

The policy benchmark isolates atomic dexterous execution from game-level decision making, consisting of a standardized suite of 14 language-instructed primitives. For each primitive, DexHoldem provides 105 teleoperated demonstrations, totaling 1,470 demonstrations. Policies are evaluated under a shared observation-action interface that maps three camera views and proprioception to a 30-dimensional joint-position space. The benchmark uses a four-level outcome rubric to score rollouts:

  1. Scene-preserving success (SPSR): the requested primitive is completed and the table remains usable for subsequent actions.

  2. Disruptive completion (DC): the goal is achieved but the execution disturbs the scene enough to prevent normal continuation.

  3. Task failure (TF): the primitive is not completed, but the scene remains stable enough for retry.

  4. Disruptive failure (DF): the primitive fails and the environment must be reset before continuing.

Agentic Perception Benchmarking

DexHoldem includes an agentic perception benchmark that tests whether agents can visually parse structured tabletop game state for downstream decision routing. Each problem corresponds to a sampled tabletop state presented with its predecessor-state context, requiring the perceiver to decompose the state into eight challenges: loop stage (LS), turn ownership (TO), blind information (BI), community cards (CC), current bet chips (CB), robot chip inventory (RCI), opponent chip inventory (OCI), and showdown outcome (SO). Success is defined as exact match over the challenges applicable to that problem. The benchmark reveals a bottleneck, where while isolated sub-capabilities can be strong, routing-critical chip-state fields remain especially unreliable, such as opponent chip inventory accuracy peaking at 43.8%.

System-Level Evaluation

The system-level evaluation composes the dexterous policy and agentic perception interfaces in real two-player Texas Hold'em tabletop rollouts. The agent captures an agent-view image, parses it into the structured state, routes through deterministic workflow gates (waiting, verification, recovery), and dispatches a dexterous primitive when physical motion is required. This evaluation probes how component errors accumulate across a closed loop. Operational counters track events such as request human primitives, wait-branch events, and recovery dispatches. The study identifies a compounding closed-loop reliability gap: current agents and dexterous policies can each solve parts of the benchmark, but their errors and delays accumulate across many captured states and primitive dispatches.

Key Findings

The results demonstrate that while policy models like π0.5 obtain the highest task completion rate (61.2%), they also tie on the stricter scene-preserving success rate (47.5%). The agentic perception results show that GPT 5.5 obtains the best average field-wise accuracy (66.8%) but struggles with strict problem-level accuracy (34.3%). System-level case studies reveal that closed-loop execution is dominated by repeated waiting, verification, continuation, and occasional recovery rather than by a single high-level decision. The benchmark ultimately evaluates dexterous tabletop execution, agentic perception, and embodied decision routing in a shared physical setting.

Contributions

  1. Collection of a real-world Texas Hold'em dexterous manipulation dataset with 1,470 teleoperated demonstrations across 14 primitives.

  2. Introduction of a real-world dexterous hand policy benchmark under a shared multi-view observation–action interface and scene-preservation rubric.

  3. Introduction of an agentic perception benchmark to evaluate structured tabletop game state parsing for downstream routing.

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed DexHoldem: Playing Texas Hold’em with Dexterous Embodied System. This paper introduces a novel benchmark, DexHoldem, designed to evaluate the complex interplay between agentic perception, policy execution (dexterous manipulation), and closed-loop decision-making in a real-world setting.

Here are the specific improvements that can be made to AI systems based on this research:


)Improvements for AI Systems Based on DexHoldem Research


The core improvement lies in moving beyond isolated skill mastery or general vision toward developing agents capable of closed-loop, instruction-conditioned dexterous manipulation within a semantically structured environment. The following specific advancements can be targeted:

  1. textbf Enhanced State Tracking and Recovery Logic (Agentic Perception Improvement):

  2. An AI system can achieve robust recovery from accumulated errors during long-horizon tasks by explicitly modeling the structured game state (loop stage, turn ownership, chip inventory) rather than relying solely on raw visual features.

  3. The improved system will be able to detect and correct chip-state perception bottlenecks (e.g., missing opponent's bet chips or tracking small denomination chips like 5s or 10s), allowing it to correctly route to the appropriate decision branch (e.g., waiting for a turn change vs. calling a bet).

  4. textbf Precision in Fine-Grained Dexterous Control (Policy Improvement):

  5. The AI system can execute complex, instruction-conditioned primitives with high precision, specifically regarding contact and placement on delicate objects like thin cards and chips, by training policies against the scene-preserving success rate rubric rather than just task completion.

  6. The improved policy will exhibit superior interaction precision, meaning it won't just achieve a goal (e.g., placing a card), but it will execute the action in a way that leaves non-target objects (cards or chips) exactly where they need to be for subsequent, long-horizon actions, thus preventing scene disturbance.

  7. textbf Robust System-Level Closed-Loop Execution (End-to-End Improvement):

  8. The complete embodied agent system will exhibit higher reliability during sustained interaction because it learns the accumulation of errors across multiple steps—specifically managing the compounding closed-loop reliability gap. This means the agent can better manage sequences involving repeated waiting, recovery dispatches, and primitive retries without escalating to a human intervention request prematurely.

  9. A system will demonstrate superior embodied decision routing, meaning it can correctly map high-level game states (like needing to 'Call' or 'Wait') directly into the correct sequence of low-level dexterous actions (e.g., selecting the appropriate card pickup and chip pushing primitives) with minimal misrouting errors.

  10. A system will show improved data efficiency when learning in the real world, as evidenced by the RDT fine-tuning study, suggesting that once sufficient dexterous data is available, initialization from a pre-trained policy provides a modest optimization advantage for fitting the target action distribution under real-world constraints.

  11. A system will be capable of real-to-simulation reconstruction fidelity, where the agent can reliably interpret and replan actions based on its understanding of how physical contact dynamics (friction, compliance) affect thin objects, even when operating in a simulated environment that uses reconstructed visual data.

Sources

Related papers