UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations
Peng Li, Qianqian Xu, Shilong Bao, Yangbangyan Jiang, Qingming Huang
University of Chinese Academy of Sciences · Institute of Computing Technology, Chinese Academy of Sciences · Beijing Academy of Artificial Intelligence · Hunan University
cs.CV, cs.AI
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: This paper has been accepted to ECCV 2026 AI City Challenge Workshop
Code: https://github.com/Roclp/UniTraffic-Agent
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: UniTraffic-Agent is introduced as the MR-CAS solution for Track 3 of the 10th AI City Challenge, which includes the main Traffic Anomaly Reasoning (TAR) task and two out-of-domain evaluations: FETV
Terminology
Summary
UniTraffic-Agent is introduced as the MR-CAS solution for Track 3 of the 10th AI City Challenge, which includes the main Traffic Anomaly Reasoning (TAR) task and two out-of-domain evaluations: FETV for fisheye traffic events and PSI-VQA for pedestrian intention reasoning. The system follows an observe–reason–act–verify workflow that samples timestamped visual evidence, reasons over all questions from the same clip in one request, and converts responses through task-specific action adapters.
During observation, the agent constructs a compact frame set that combines global video coverage with samples near question-specific timestamps.
During reasoning, all questions and output fields associated with a clip are processed jointly to establish a shared event interpretation.
Task-specific adapters then convert this interpretation into the official format, and a verifier checks identifiers and retries unresolved cases using cached visual evidence.
The paper's contributions are summarized as: Unified traffic-video agent framework
supporting heterogeneous tasks across surveillance, fisheye, and dashcam videos; Timestamp-aware observation and reasoning
combining global frame coverage with question-specific temporal evidence and clip-level joint reasoning; and Task-specific action adapters
with validation procedures, leading to MR-CAS ranking 2nd on FETV and 4th on PSI-VQA on the official Public leaderboards.
On the official Public leaderboards, MR-CAS ranks 16th on TAR with a score of 0.5780, 2nd on FETV with 0.4884, and 4th on PSI-VQA with 64.4161.
For TAR, the system performs well on constrained questions, matching the leader on MCQ accuracy and reaching 0.9604 on MCQ-OE F1,
but the main gap comes from long-form tasks, including scene description, temporal description, and summarization.
For FETV, MR-CAS achieves 0.4884 on FETV, only 0.0007 below the leading Public score,
being competitive on description, violator type, image-position fields, and final lane,
with the main weakness is road-topology reasoning, especially intersection type, where fisheye distortion makes geometric interpretation difficult.
For PSI-VQA, MR-CAS scores higher than the leading entry on Open-QA Cue-F1,
but BCQ and temporal localization scores remain lower than the leader.
Implementation details state that each video is processed in a single request with at most 32 sampled frames, including 16 globally distributed frames,
frames are cached as JPEG images with quality 100 and a maximum side length of 768 pixels,
and gpt-5.5
is used as the primary model with gpt-5.4
for recovery. The conclusion notes that remaining errors mainly arise from reference-aligned long-form generation, fisheye road geometry, pedestrian intent prediction, and temporal boundary estimation,
suggesting explicit actor tracking and geometry-aware temporal reasoning are promising directions for future traffic-video agents.
Improvements for AI systems
Improvements to AI Systems:
- Unified Multi-Task Video Agent Architecture
-
Implement a single agent framework that dynamically switches between surveillance, fisheye, and dashcam inputs without task-specific retraining.
-
Use a shared “observe–reason–act–verify” loop with modular adapters for output formatting, enabling seamless deployment across heterogeneous traffic-video benchmarks.
- Timestamp-Aware Sparse Sampling with Global Coverage
-
Replace uniform frame sampling with a hybrid strategy: 16 globally distributed frames for context + question-specific frames near relevant timestamps.
-
Cache frames as high-quality JPEGs (quality 100, max side 768px) to reduce latency and enable retries without re-encoding.
- Joint Clip-Level Reasoning over All Questions
-
Process all questions and output fields for a single video clip in one LLM request, forcing the model to build a coherent event interpretation rather than answering in isolation.
-
This reduces contradictory outputs and improves consistency for multi-part tasks (e.g., scene description + temporal localization + summarization).
- Task-Specific Action Adapters with Verification
-
Develop adapters that convert free-form LLM responses into strict official formats (e.g., JSON schemas for TAR, FETV, PSI-VQA).
-
Add a verifier that checks identifier validity (e.g., object IDs, lane numbers) and automatically retries unresolved cases using cached visual evidence, improving robustness on out-of-domain data.
- Geometry-Aware Temporal Reasoning for Fisheye Videos
-
Integrate a geometric distortion correction module (e.g., fisheye-to-perspective mapping) before feeding frames to the LLM, specifically for road-topology tasks like intersection type classification.
-
Use explicit lane and road-boundary extraction from corrected frames to aid reasoning about spatial relationships.
- Explicit Actor Tracking for Long-Form Generation
-
Add a lightweight object tracker (e.g., ByteTrack or SORT) that assigns consistent IDs across frames, then inject these IDs into the LLM prompt as structured context.
-
This improves temporal description and summarization by grounding references (e.g., “the red car” → “vehicle 3”) and reducing hallucinated entities.
- Two-Model Cascade with Fallback Recovery
-
Use a primary LLM (e.g., gpt-5.5) for initial reasoning and a secondary model (gpt-5.4) for recovery when the verifier flags errors.
-
This improves reliability on ambiguous tasks (e.g., pedestrian intention) without doubling inference cost for all cases.
- Confidence-Weighted Multi-Model Ensemble for Temporal Localization
-
For PSI-VQA’s temporal boundary estimation, run the reasoning with both models and combine their predicted start/end times using a confidence-weighted average.
-
This mitigates single-model bias and improves BCQ scores.
What the Improved AI System Can Do:
-
Achieve top-1 performance on TAR by closing the gap on long-form tasks (scene description, temporal description, summarization) through explicit actor tracking and joint clip reasoning.
-
Surpass the leading FETV score by correcting fisheye distortion before reasoning, enabling accurate intersection-type classification and road-topology inference.
-
Outperform the PSI-VQA leader by combining improved temporal localization (via ensemble) with already superior Open-QA Cue-F1, using explicit pedestrian trajectory tracking to predict intention more accurately.
-
Deploy in real-time traffic monitoring with low latency (≤32 frames per clip, cached JPEGs) and high reliability (verifier + fallback model), handling mixed camera types (surveillance, fisheye, dashcam) in a single unified pipeline.
-
Generalize to unseen traffic-anomaly tasks without retraining, thanks to the modular adapter design and shared reasoning core.
Abstract
Traffic video understanding has become an important problem in intelligent transportation, as road videos provide direct evidence for accidents, violations, and interactions between vehicles and vulnerable road users. A useful system should explain how a traffic event develops, why it happens, and when the relevant interaction occurs, yet this remains difficult for multimodal large language models (MLLMs) because traffic videos contain sparse events and varied viewpoints. We introduce UniTraffic-Agent, the MR-CAS solution for Track 3 of the 10th AI City Challenge, which includes Traffic Anomaly Reasoning (TAR) and two out-of-domain evaluations: FETV for fisheye traffic events and PSI-VQA for pedestrian intention reasoning. UniTraffic-Agent follows an observe--reason--act--verify workflow that samples timestamped visual evidence, reasons over all questions from the same clip in one request, and converts responses through task-specific action adapters. On the official Public leaderboards, MR-CAS ranks 16th on TAR with a score of 0.5780, 2nd on FETV with 0.4884, and 4th on PSI-VQA with 64.4161. The code is available at https://github.com/Roclp/UniTraffic-Agent.
Sources
- AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond
- ReconBoost: Boosting Can Achieve Modality Reconcilement
- OpenworldAUC: Towards Unified Evaluation and Optimization for Open-world Prompt Tuning
- Warehouse Spatial Question Answering with LLM Agent
- Gemini: A Family of Highly Capable Multimodal Models
- VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners
- BERTScore: Evaluating Text Generation with BERT
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models