UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations

arXiv:2608.13031 · cs.CV, cs.AI · Submitted 2026-08-13 · Read on arXiv

Peng Li, Qianqian Xu, Shilong Bao, Yangbangyan Jiang, Qingming Huang

University of Chinese Academy of Sciences · Institute of Computing Technology, Chinese Academy of Sciences · Beijing Academy of Artificial Intelligence · Hunan University

cs.CV, cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: This paper has been accepted to ECCV 2026 AI City Challenge Workshop

Code: https://github.com/Roclp/UniTraffic-Agent

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: UniTraffic-Agent is introduced as the MR-CAS solution for Track 3 of the 10th AI City Challenge, which includes the main Traffic Anomaly Reasoning (TAR) task and two out-of-domain evaluations: FETV

Terminology

Summary

UniTraffic-Agent is introduced as the MR-CAS solution for Track 3 of the 10th AI City Challenge, which includes the main Traffic Anomaly Reasoning (TAR) task and two out-of-domain evaluations: FETV for fisheye traffic events and PSI-VQA for pedestrian intention reasoning. The system follows an observe–reason–act–verify workflow that samples timestamped visual evidence, reasons over all questions from the same clip in one request, and converts responses through task-specific action adapters. During observation, the agent constructs a compact frame set that combines global video coverage with samples near question-specific timestamps. During reasoning, all questions and output fields associated with a clip are processed jointly to establish a shared event interpretation. Task-specific adapters then convert this interpretation into the official format, and a verifier checks identifiers and retries unresolved cases using cached visual evidence.

The paper's contributions are summarized as: Unified traffic-video agent framework supporting heterogeneous tasks across surveillance, fisheye, and dashcam videos; Timestamp-aware observation and reasoning combining global frame coverage with question-specific temporal evidence and clip-level joint reasoning; and Task-specific action adapters with validation procedures, leading to MR-CAS ranking 2nd on FETV and 4th on PSI-VQA on the official Public leaderboards.

On the official Public leaderboards, MR-CAS ranks 16th on TAR with a score of 0.5780, 2nd on FETV with 0.4884, and 4th on PSI-VQA with 64.4161. For TAR, the system performs well on constrained questions, matching the leader on MCQ accuracy and reaching 0.9604 on MCQ-OE F1, but the main gap comes from long-form tasks, including scene description, temporal description, and summarization. For FETV, MR-CAS achieves 0.4884 on FETV, only 0.0007 below the leading Public score, being competitive on description, violator type, image-position fields, and final lane, with the main weakness is road-topology reasoning, especially intersection type, where fisheye distortion makes geometric interpretation difficult. For PSI-VQA, MR-CAS scores higher than the leading entry on Open-QA Cue-F1, but BCQ and temporal localization scores remain lower than the leader.

Implementation details state that each video is processed in a single request with at most 32 sampled frames, including 16 globally distributed frames, frames are cached as JPEG images with quality 100 and a maximum side length of 768 pixels, and gpt-5.5 is used as the primary model with gpt-5.4 for recovery. The conclusion notes that remaining errors mainly arise from reference-aligned long-form generation, fisheye road geometry, pedestrian intent prediction, and temporal boundary estimation, suggesting explicit actor tracking and geometry-aware temporal reasoning are promising directions for future traffic-video agents.

Improvements for AI systems

Improvements to AI Systems:

  1. Unified Multi-Task Video Agent Architecture
  • Implement a single agent framework that dynamically switches between surveillance, fisheye, and dashcam inputs without task-specific retraining.

  • Use a shared “observe–reason–act–verify” loop with modular adapters for output formatting, enabling seamless deployment across heterogeneous traffic-video benchmarks.

  1. Timestamp-Aware Sparse Sampling with Global Coverage
  • Replace uniform frame sampling with a hybrid strategy: 16 globally distributed frames for context + question-specific frames near relevant timestamps.

  • Cache frames as high-quality JPEGs (quality 100, max side 768px) to reduce latency and enable retries without re-encoding.

  1. Joint Clip-Level Reasoning over All Questions
  • Process all questions and output fields for a single video clip in one LLM request, forcing the model to build a coherent event interpretation rather than answering in isolation.

  • This reduces contradictory outputs and improves consistency for multi-part tasks (e.g., scene description + temporal localization + summarization).

  1. Task-Specific Action Adapters with Verification
  • Develop adapters that convert free-form LLM responses into strict official formats (e.g., JSON schemas for TAR, FETV, PSI-VQA).

  • Add a verifier that checks identifier validity (e.g., object IDs, lane numbers) and automatically retries unresolved cases using cached visual evidence, improving robustness on out-of-domain data.

  1. Geometry-Aware Temporal Reasoning for Fisheye Videos
  • Integrate a geometric distortion correction module (e.g., fisheye-to-perspective mapping) before feeding frames to the LLM, specifically for road-topology tasks like intersection type classification.

  • Use explicit lane and road-boundary extraction from corrected frames to aid reasoning about spatial relationships.

  1. Explicit Actor Tracking for Long-Form Generation
  • Add a lightweight object tracker (e.g., ByteTrack or SORT) that assigns consistent IDs across frames, then inject these IDs into the LLM prompt as structured context.

  • This improves temporal description and summarization by grounding references (e.g., “the red car” → “vehicle 3”) and reducing hallucinated entities.

  1. Two-Model Cascade with Fallback Recovery
  • Use a primary LLM (e.g., gpt-5.5) for initial reasoning and a secondary model (gpt-5.4) for recovery when the verifier flags errors.

  • This improves reliability on ambiguous tasks (e.g., pedestrian intention) without doubling inference cost for all cases.

  1. Confidence-Weighted Multi-Model Ensemble for Temporal Localization
  • For PSI-VQA’s temporal boundary estimation, run the reasoning with both models and combine their predicted start/end times using a confidence-weighted average.

  • This mitigates single-model bias and improves BCQ scores.


What the Improved AI System Can Do:

  • Achieve top-1 performance on TAR by closing the gap on long-form tasks (scene description, temporal description, summarization) through explicit actor tracking and joint clip reasoning.

  • Surpass the leading FETV score by correcting fisheye distortion before reasoning, enabling accurate intersection-type classification and road-topology inference.

  • Outperform the PSI-VQA leader by combining improved temporal localization (via ensemble) with already superior Open-QA Cue-F1, using explicit pedestrian trajectory tracking to predict intention more accurately.

  • Deploy in real-time traffic monitoring with low latency (≤32 frames per clip, cached JPEGs) and high reliability (verifier + fallback model), handling mixed camera types (surveillance, fisheye, dashcam) in a single unified pipeline.

  • Generalize to unseen traffic-anomaly tasks without retraining, thanks to the modular adapter design and shared reasoning core.

Abstract

Traffic video understanding has become an important problem in intelligent transportation, as road videos provide direct evidence for accidents, violations, and interactions between vehicles and vulnerable road users. A useful system should explain how a traffic event develops, why it happens, and when the relevant interaction occurs, yet this remains difficult for multimodal large language models (MLLMs) because traffic videos contain sparse events and varied viewpoints. We introduce UniTraffic-Agent, the MR-CAS solution for Track 3 of the 10th AI City Challenge, which includes Traffic Anomaly Reasoning (TAR) and two out-of-domain evaluations: FETV for fisheye traffic events and PSI-VQA for pedestrian intention reasoning. UniTraffic-Agent follows an observe--reason--act--verify workflow that samples timestamped visual evidence, reasons over all questions from the same clip in one request, and converts responses through task-specific action adapters. On the official Public leaderboards, MR-CAS ranks 16th on TAR with a score of 0.5780, 2nd on FETV with 0.4884, and 4th on PSI-VQA with 64.4161. The code is available at https://github.com/Roclp/UniTraffic-Agent.

Sources

Related papers