N 0-Foundation: Towards the Age of Tactile Intelligence
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "N 0-Foundation: Towards the Age of Tactile Intelligence".
Dev: The paper introduces a comprehensive benchmark suite designed to advance tactile intelligence by testing robotic manipulation skills across diverse, contact-rich tasks.
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: Well, so we're starting with the paper titled "N0-Foundation: Towards the Age of Tactile Intelligence," and I want to talk about what that title actually means for us in the field. It sounds like they're aiming for something much deeper than just looking at objects visually.
Dev: I agree, Rosa, it suggests a focus on building a foundation where touch is central to how robots understand their environment rather than just seeing it through cameras. The authors are NeoteAI Team and Fudan TEAI Team, and they’re tackling this with a lot of data collection from various embodiments.
Taro: From an autonomy standpoint, I'm curious if this foundation means the AI can handle situations where visual input is completely missing or misleading because the robot needs to rely purely on physical feedback.
Rosa: Exactly, Taro, that’s the core idea; they are pushing for a system that integrates tactile sensing hardware with large-scale multimodal data to achieve this understanding. It moves beyond simple vision inputs by focusing on what happens during contact.
Dev: And the scale of their data collection is substantial; they've put together NeoData, which includes more than thirty thousand hours of synchronized visual and tactile demonstrations across six different robot embodiments. That's a lot to process for any system trying to learn generalized skills.
Taro: Having that kind of diverse dataset across multiple robot designs is impressive because it suggests the underlying representation learning should be quite robust against changes in hardware or physical setup. I wonder how transferable those learned representations truly are when we move outside the lab.
Rosa: That's exactly what we need to figure out; they mention releasing OpenNeoData, a five thousand-hour subset of that data, which is really important for letting other researchers test this approach openly.
Dev: Having that open-source subset means the community can start building on this infrastructure immediately without having to wait for the full dataset to be fully processed by everyone. It’s a good step toward practical deployment, I think.
Taro: That accessibility is key for accelerating the pace of development in autonomous systems; if we can see and test these concepts outside their controlled environment, it helps us understand the real challenges of deployment.
The paper's summary: Rosa: So, let's talk about what "N0-Foundation: Towards the Age of Tactile Intelligence" actually delivers in terms of its methodology. Essentially, they are presenting a unified framework that integrates tactile sensing hardware with large-scale multimodal data to create a new way for robots to learn manipulation skills.
Dev: The core of it is engineering the underlying tactile infrastructure, including something called NeoReal and NeoSim, which provides both real-world tasks and simulated tasks for policy testing. This infrastructure is designed to support scalable data collection from various robot embodiments using a Tactile Universal Manipulation Interface or N0-TacUMI.
Taro: I'm interested in the specific mechanisms they use for this integration; how does the system actually combine those different sensor inputs—vision, touch, and joint states—into a single representation?
Rosa: They construct NeoData with over thirty thousand hours of synchronized visual and tactile demonstrations spanning four hundred fifty tasks. Crucially, they introduce a unified formulation for tactile representation learning that uses dense three-axis force fields as a common physical supervision space across different tactile sensor designs.
Dev: That force field approach sounds like it’s the key to achieving hardware-agnostic representations, which is what they are aiming for when they release NeoForce, their visuo-tactile representation model. This means the representation learned should not be tied to a specific sensor type.
Taro: So, when we look at the results reported in terms of testing policies on this infrastructure, what's the main finding regarding performance on these tasks? Are they achieving high success rates compared to previous methods?
Rosa: The paper shows that policies trained under this new formulation perform well across a wide range of contact-rich manipulation tasks. They demonstrate the ability to learn transferable skills from the large-scale and heterogeneous tactile data available in NeoData.
Dev: It’s interesting because they explicitly state that vision alone is often insufficient for contact-rich manipulation, which highlights why this multimodal approach is necessary for success in these specific scenarios.
The paper's improvements: Rosa: Now that we've looked at the core setup, I want to focus on the specific suggested improvements they propose to make this foundation even stronger and more useful for real-world deployment. These aren't just incremental tweaks; they are about pushing the boundaries of what this system can achieve.
Dev: They suggest a few things, including developing a system that incorporates stochastic dynamics modeling and characterization for sensor drift, which means accounting for the noise in motors and environmental physics, not just assuming perfect physics.
Taro: That’s crucial because if we don't account for unmodeled dynamics, the AI might fail catastrophically when deployed in a real setting where friction or unexpected resistance is higher than simulated. How does this change the robustness of the policy?
Rosa: The improved AI system will be trained to be stochastically robust; instead of just succeeding under ideal randomized conditions, it will learn policies that maintain performance margins even when the underlying physical model deviates slightly from the simulation's assumptions.
Dev: And on top of that, they propose goal-conditioned inverse reinforcement learning to infer the intent or cost function behind demonstrations rather than just mimicking trajectories. That shifts the focus from following expert moves to understanding *why* those moves are successful in a more causal sense.
Taro: Inferring the underlying cost function sounds like it gives us a way to adapt when the environment changes, because if we know what constraint is critical—like maintaining a specific seating force during insertion—the AI can reason about corrective actions instead of just blindly following an expert's path.
Conclusion: Rosa: So, wrapping up our discussion on "N0-Foundation: Towards the Age of Tactile Intelligence," it really shows how we are moving toward a new era where robots understand physical nature through contact. We’ve seen how this unified framework brings together data, hardware, and learning to tackle complex manipulation challenges.
Dev: It’s exciting because it suggests that future embodied AI won't rely solely on vision but will need to incorporate touch for true physical understanding. The transition from visual observation to sensing subtle resistance during tasks like nesting cups together is a significant step in that direction.
Taro: I just want to add that the implications are huge because this work opens up new avenues for how we can design autonomy systems that are inherently grounded in physical reality, not just computation.
Rosa: Absolutely, Taro; this research provides a solid path forward for building systems that interact with the world in a much more intuitive way.
Dev: To sum up the paper "N0-Foundation: Towards the Age of Tactile Intelligence," it’s a comprehensive approach to creating a foundation for tactile intelligence.
Taro: It really sets a high bar for what embodied AI can achieve in terms of physical interaction fidelity.
Fudan University · NeoteAI Team (Project)
cs.RO, cs.CV, cs.LG
Submitted: 2026-08-30
Updated: 2026-09-20
Comments: 13 figures, 5 tables
Project page: http://research.neoteai.com/n0-foundation
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 89/100
The gist: The paper introduces a comprehensive benchmark suite designed to advance tactile intelligence by testing robotic manipulation skills across diverse, contact-rich tasks.
Key concepts
- N 0-Foundation
- "N 0-Foundation" is a comprehensive framework designed to build a foundation for tactile intelligence. It integrates tactile sensing hardware with large-scale multimodal data to help robots understand their environment through physical contact rather than just visual input.
- NeoData
- NeoData is the large dataset created by the authors, containing over thirty thousand hours of synchronized visual and tactile demonstrations across six different robot embodiments. This data is used to train policies for complex manipulation tasks.
- Tactile Universal Manipulation Interface (N0-TacUMI)
- This interface is part of the infrastructure engineered by the authors. It supports scalable data collection from various robot embodiments, allowing researchers to test their approach across different physical setups.
Terminology
Summary
The paper introduces a comprehensive benchmark suite designed to advance tactile intelligence by testing robotic manipulation skills across diverse, contact-rich tasks. This framework moves beyond purely visual inputs by integrating detailed simulated tactile feedback, enabling the rigorous evaluation of policies that must sense subtle physical interactions—such as edge contact or seating force—to achieve success.
Simulation Platform and Tactile Ground Truth
The NeoSim benchmark is built upon UniVTAC, a visuo-tactile simulation platform leveraging Isaac Sim and Isaac Lab. The core innovation lies in simulating the sensor gel as a finite element soft body
which outputs tactile RGB images, marker motion, and gel depth maps. Furthermore, the system simulates a markerless particle gel surface,
advecting speckle coating based on both marker motion and the gel depth map. Crucially for training fidelity, per-vertex contact forces are derived from the contact solver and interpolated into a dense force field that serves as the tactile ground truth in NeoSim.
Each task is instantiated using a modular configuration file that dictates asset specifications, initial object layout, and randomization ranges.
Scope of Contact-Rich Tasks
The benchmark encompasses 12 distinct tasks, divided into single-arm and dual-arm categories, designed to stress specific physical capabilities. The four single-arm tasks include:
-
Pour Ball: Requiring the policy to
maintain a stable grasp force and modulate the pouring angle
due to compliant contents. -
Unplug and Plug Charger: A canonical
force-guided mating task
that depends on sensing seating force and small resistance during connection. -
Plug USB: Testing alignment where success hinges on
tactile detection of edge contact.
-
Grasp Chip: Probing
fine-grained force limiting,
as the policy must operate within a narrow force window to avoid crushing or slipping.
The eight dual-arm tasks test bimanual coordination and stable contact, such as:
-
Insert Screw: Requiring
coordinated force sensing on both arms
for alignment and proper seating. -
Place Gears: Where tactile feedback confirms that each gear is seated flush against the peg, a state that is
visually ambiguous.
-
Cup Handover: The moment requires both arms to sense the shared contact so that the cup is neither dropped nor crushed during transfer.
Training and Policy Generalization Settings
Policies undergo a structured pretrain then post train recipe. Initially, each policy is pretrained on over 400,000 episodes from NeoData to acquire broad contact-rich manipulation priors.
In the specialization stage, the policy is trained using 300 demonstrations collected specifically for that task. Demonstrations are automatically collected by scripted expert policies built on the cuRobo motion planner, ensuring that all evaluated policies are trained on the same 100 successful episodes per task. Every raw episode stores synchronized external and wrist RGB views, two finger tactile streams, and both joint and end-effector states to facilitate conversion into the native input format of various baselines.
Rigorous Evaluation Protocols
Evaluation is designed to test generalization across multiple dimensions. Each task is evaluated over 20 trials (in NeoReal) or 100 rollouts (in NeoSim). To ensure controlled randomization, the robot arm is reset to an "
Improvements for AI systems
Given the high stakes implied by the context—where errors can cost millions of dollars—the current benchmarks, while excellent in defining complex manipulation tasks, require significant enhancements in robustness, causal understanding, and verifiable safety guarantees before they can reliably underpin real-world industrial AI systems.
My improvements focus on hardening the system's ability to handle uncertainty beyond simple randomization and ensuring its actions are physically justifiable.
The current simulation (UniVTAC/Isaac Sim) is strong, but real-world variability often involves unmodeled physics or sensor noise that goes beyond simple state randomization.
Improvement: Integration of Stochastic Dynamics Modeling and Sensor Drift Characterization.
-
Specific Change: The simulation must incorporate parameterized, non-Gaussian stochastic noise models for actuation (motor friction variation, backlash) and environmental dynamics (e.g., fluid resistance modeling beyond simple linear drag). Furthermore, the tactile sensor model needs to simulate drift and temperature dependency, not just perfect readings.
-
What the Improved AI System Can Do: The resulting policy will be trained to be stochastically robust. Instead of merely succeeding under ideal randomized conditions, it will learn policies that maintain performance margins even when the underlying physical model deviates slightly from the simulation's assumptions (i.e., it learns to compensate for unmodeled dynamics). This dramatically reduces the risk of catastrophic failure upon deployment due to minor hardware imperfections.
The current training relies heavily on scripted expert demonstrations (cuRobo) and pretraining on large, general datasets (NeoData). This creates a strong dependence on the quality and scope of the initial data.
The current scoring system reports success rate and progressive score, which are excellent metrics of competence but lack formal guarantees of safety or long-term stability in complex environments.
Area Current Limitation Proposed Improvement Key Capability Gained by AI System
:---:---:---:---
Dynamics (Sim) Assumes perfect physics; lacks noise modeling. Stochastic Dynamics Modeling & Sensor Drift Simulation. Stochastic Robustness: Performance maintained despite unmodeled physical variations. (Reduces catastrophic failure risk.)
Learning (Policy) Relies on trajectory imitation from demonstrations (tau). Goal-Conditioned IRL with Causal Priors. Intent Understanding & Adaptability: Can reason about why a task succeeds and adapt to novel conditions. (Moves beyond mere mimicry.)
Safety (Eval) Success/Failure is binary or score-based; lacks formal guarantees. Formal Verification: Reachability Analysis & Violation Cost Functions. Provable Safety: Guarantees that all actions remain within a mathematically defined, safe operational envelope. (Essential for high-stakes deployment.)
Abstract
We present N 0-Foundation, a paradigm for tactile-enabled embodied manipulation, which integrates tactile sensing hardware, large-scale multimodal data, tactile representation learning, and standardized evaluation. First, we engineer the infrastructure for scalable data collection, including a vision-based tactile sensor, a tactile Universal Manipulation Interface (UMI), and a synchronized visuo-tactile data collection system supporting both robot embodiments and UMI-based demonstrations. Leveraging this infrastructure, we construct NeoData, which contains more than 30000 hours of synchronized visual and tactile demonstrations, spanning six embodiments, 450 tasks, and billions of paired RGB and tactile frames collected through a mixture of real-robot teleoperation and UMI-based demonstrations. To facilitate open research, we further release OpenNeoData, a 5000-hour open-source subset of NeoData. The dataset addresses a central limitation of existing manipulation corpora, critical for deformable-object manipulation, precise assembly, delicate force control, and sustained surface interaction. Capitalizing on the large-scale, heterogeneous tactile measurements, we propose NeoForce, a visuo-tactile representation model that learn transferable tactile representations across different sensor designs. To enable systematic evaluation of tactile embodied models built upon our infrastructure, datasets and tactile representations, we further propose a comprehensive benchmark, which combines the real-world NeoReal suite and the simulated NeoSim suite for standardized evaluation. Experiments across both suites show that policies benefit from the physical contact state rather than from the device-specific appearance of the tactile signal. We release the dataset, the representation, and the benchmark, aiming at supporting future work on tactile-enabled embodied manipulation.
Sources
- RT-1: Robotics Transformer for Real-World Control at Scale
- InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation
- Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution
- UniVTAC: A Unified Simulation Platform for Visuo-Tactile Manipulation Data Generation, Learning, and Benchmarking
- RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies
- TacUMI: A Multi-Modal Universal Manipulation Interface for Contact-Rich Tasks
- Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots
- RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot
- LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
- Vision Pretraining for Dense Spatial Perception
- EBench: Elemental Diagnosis of Generalist Mobile Manipulation Policies
- Infinite Worlds with Versatile Interactions
- Sensor-Invariant Tactile Representation
- UMI on Legs: Making Manipulation Policies Mobile with Manipulation-Centric Whole-body Controllers
- SoftVTBench: A Safety-Aware Visuo-Tactile Benchmark for Physically Constrained Robotic Manipulation of Deformable Objects
- DuoBench: A Reproducible Benchmark for Bimanual Manipulation in Simulation and the Real World
- DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
- ViTaMIn-B: A Reliable and Efficient Visuo-Tactile Bimanual Manipulation Interface
- Causal World Modeling for Robot Control
- Unified Video Action Model
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving