CoLT-Drive: Counterfactual Long-Tail Benchmarking and Knowledge-Preserving Adaptation for Driving Affordance Prediction
cs.CV, cs.AI, cs.CL, cs.RO
Submitted: 2026-08-31
Updated: 2026-08-31
Comments: Accepted by EMNLP 2026
Code: https://github.com/tangzhengxu/CoLT-Drive
License: http://creativecommons.org/licenses/by/4.0/
The gist: Long-tail autonomous driving failures are often framed as rare-object recognition errors.
Terminology
Abstract
Long-tail autonomous driving failures are often framed as rare-object recognition errors. We argue that this view is incomplete: the decision-critical question is not only whether a model recognizes an unusual object, but whether it infers how that object changes the ego vehicle's feasible high-level actions. We formalize this problem as decision-level driving affordance prediction, where a model maps a front-view image, ego-motion history, and navigation command to a structured longitudinal--lateral meta-action. To evaluate this capability, we introduce CoLT-Drive, a 3,536-sample counterfactual long-tail benchmark that inserts rare objects into otherwise fixed driving scenes and measures whether models predict acceptable action pairs. To improve deployable small VLMs, we propose KPA, a knowledge-preserving adaptation framework that combines structured perception-to-decision prompting, SLERP-based expert merging, and RegMoE, a regime-aware LoRA mixture-of-experts module. KPA preserves the pretrained model's open-world knowledge while allocating lightweight adaptation capacity to different driving decision regimes. Experiments on an in-domain driving split and CoLT-Drive show that KPA achieves 60.8% pair accuracy on CoLT-Drive, outperforming the pretrained Qwen3-VL-2B baseline (50.3%) and LoRA SFT (32.4%) while maintaining competitive in-domain accuracy. Our benchmark and code are available at https://huggingface.co/datasets/tangzx2024/CoLT-Drive and https://github.com/tangzhengxu/CoLT-Drive.
Sources
- Qwen3-VL Technical Report
- Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
- LoRA Learns Less and Forgets Less
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- NuPlan: A closed-loop ML-based planning benchmark for autonomous vehicles
- Devil is in Narrow Policy: Unleashing Exploration in Driving VLA Models
- Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases
- Lifelong Language Pretraining with Distribution-Specialized Experts
- Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models
- LoRAMoE: Alleviate World Knowledge Forgetting in Large Language Models via MoE-Style Plugin
- Deconfounded Lifelong Learning for Autonomous Driving via Dynamic Knowledge Spaces
- Can Vehicle Motion Planning Generalize to Realistic Long-tail Scenarios?
- DriveAction: A Benchmark for Exploring Human-like Driving Decisions in VLA Models
- EMMA: End-to-End Multimodal Model for Autonomous Driving
- DriveLMM-o1: A Step-by-Step Reasoning Dataset and Large Multimodal Model for Driving Scenario Understanding
- Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving
- AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning
- MixLoRA: Enhancing Large Language Models Fine-Tuning with LoRA-based Mixture of Experts
- Drive-R1: Bridging Reasoning and Planning in VLMs for Autonomous Driving with Reinforcement Learning
- Robustness Is a Function, Not a Number: A Factorized Comprehensive Study of OOD Robustness in Vision-Based Driving
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models