TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation
cs.AI
Submitted: 2026-06-10
Updated: 2026-08-27
Comments: 18 pages, 11 figures
Code: https://github.com/lvkailin0118/TouchThinker
License: http://creativecommons.org/licenses/by/4.0/
The gist: Touch is a key modality for embodied agents to understand the physical world.
Terminology
Abstract
Touch is a key modality for embodied agents to understand the physical world. Although recent work has incorporated tactile signals into language systems for tactile commonsense reasoning, scaling such systems to realistic open-world settings remains challenging due to two key bottlenecks: (1) current tactile reasoning datasets remain limited in format and scale, providing insufficient supervision for reasoning from tactile observations to physical commonsense and hindering the learning of transferable tactile commonsense; (2) tactile signals are inherently redundant and action-specific, yet existing methods often overlook these properties, resulting in inefficient representations with limited semantic expressiveness. To address these limitations, we propose TouchThinker, a tactile-language framework that scales tactile commonsense reasoning to the open world from both data and representation perspectives. First, we construct TouchThinker-1M, a million-scale, multi-source tactile reasoning dataset covering 415 objects, 8 scenarios, and 7 sensor types, providing a solid data foundation for open-world generalization. We further introduce TouchThinker-Bench, an open-world benchmark with more realistic and diverse tasks. Then, we propose action-aware modeling mechanism to improve tactile representation efficiency and enable efficient reasoning. Experimental results demonstrate that TouchThinker achieves competitive performance against state-of-the-art models across multiple datasets. Our code and dataset will be made available at: https://github.com/lvkailin0118/TouchThinker.
Sources
- A Survey of Vision-Language Pre-Trained Models
- AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-tactile Sensors
- Sparsh: Self-supervised touch representations for vision-based tactile sensing
- OpenAI GPT-5 System Card
- Surveying the MLLM Landscape: A Meta-Review of Current Surveys
- Qwen-Image Technical Report
- Touch and Go: Learning from Human-Collected Vision and Touch
- Demonstrating the Octopi-1.5 Visual-Tactile-Language Model
- Octopi: Object Property Reasoning with Large Tactile-Language Models
- Transferable Tactile Transformers for Representation Learning Across Diverse Sensors and Tasks
- OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection