Continual Field-Adaptive Models (CFAMs) for Post-Deployment Physical AI
cs.RO, cs.AI
Submitted: 2026-09-03
Updated: 2026-09-28
License: http://creativecommons.org/licenses/by/4.0/
The gist: Unattended interactive autonomy - machines that step into danger in place of humans and complete tasks with human tools - remains a missing capability in mission-critical operations.
Terminology
Abstract
Unattended interactive autonomy - machines that step into danger in place of humans and complete tasks with human tools - remains a missing capability in mission-critical operations. These domains offer scarce training data and only onboard compute, yet deployed systems must face novelty without erasing prior competence. We introduce Continual Field-Adaptive Models (CFAMs), which learn efficiently in the lab and continue learning after deployment through autonomous, gradient-free, on-device updates. CFAM uses a complementary learning architecture with a frozen slow-learning component and a fast-learning Capsule Field. The slow component contains three cortices: Sensor, which maps multimodal input into 3D-grounded geometry; Reasoning, which decomposes tasks into skills and evaluates outcomes; and Action, which executes geometric skills. The Capsule Field stores field learning one-shot and gradient-free as Competence Capsules. Skill installation is few-shot in the lab and continual in the field; open-world novelty is outside scope. We evaluate CFAM across five embodiments: manipulator, quadruped, humanoid, quadrotor, and off-road vehicle. Baselines (pi0, CogACT, SpatialVLA) use the same in-house multi-embodiment dataset for physical-platform comparisons. CFAM reaches the operating point of a standard policy trained on the full prior-training dataset using 40% of the data, or 2.5x fewer trajectories. At test time, autonomous capture of verified near-OOD cases improves action success by 13.9 percentage points. In sequential simulation, backward transfer is-0.5 percentage points versus-11.4 for LoRA. CFAM therefore provides a bounded form of post-deployment physical intelligence: few-shot skill learning, autonomous field growth from verified near-OOD experience, and retention of prior competence.
Sources
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- REFLECT: Summarizing Robot Experiences for Failure Explanation and Correction
- Open-Ended Instructable Embodied Agents with Memory-Augmented Large Language Models
- Statler: State-Maintaining Language Models for Embodied Reasoning
- MoS-VLA: A Vision-Language-Action Model with One-Shot Skill Adaptation
- MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
- On-the-Fly VLA Adaptation via Test-Time Reinforcement Learning
- EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
- Residual Policy Learning
- ENPIRE: Agentic Robot Policy Self-Improvement in the Real World
- Autonomous Integration and Improvement of Robotic Assembly using Skill Graph Representations
- MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
- RoboMemory: A Brain-inspired Multi-memory Agentic Framework for Interactive Environmental Learning in Physical Embodied Systems
- EchoVLA: Robotic Vision-Language-Action Model with Synergistic Declarative Memory for Mobile Manipulation
- 3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model
- ViReSkill: Vision-Grounded Replanning with Skill Memory for LLM-Based Planning in Lifelong Robot Learning
- Uni-Skill: Building Self-Evolving Skill Repository for Generalizable Robotic Manipulation
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving