Semantic Edge-Cloud Communication for Real-Time Urban Traffic Surveillance with ViT and LLMs over Mobile Networks
cs.NI, cs.AI
Submitted: 2025-09-25
Updated: 2025-09-25
Comments: 17 pages, 12 figures
DOI: 10.1109/TNSE.2025.3617881
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Real-time urban traffic surveillance is vital for Intelligent Transportation Systems (ITS) to ensure road safety, optimize traffic flow, track vehicle trajectories, and prevent collisions in smart
Terminology
Abstract
Real-time urban traffic surveillance is vital for Intelligent Transportation Systems (ITS) to ensure road safety, optimize traffic flow, track vehicle trajectories, and prevent collisions in smart cities. Deploying edge cameras across urban environments is a standard practice for monitoring road conditions. However, integrating these with intelligent models requires a robust understanding of dynamic traffic scenarios and a responsive interface for user interaction. Although multimodal Large Language Models (LLMs) can interpret traffic images and generate informative responses, their deployment on edge devices is infeasible due to high computational demands. Therefore, LLM inference must occur on the cloud, necessitating visual data transmission from edge to cloud, a process hindered by limited bandwidth, leading to potential delays that compromise real-time performance. To address this challenge, we propose a semantic communication framework that significantly reduces transmission overhead. Our method involves detecting Regions of Interest (RoIs) using YOLOv11, cropping relevant image segments, and converting them into compact embedding vectors using a Vision Transformer (ViT). These embeddings are then transmitted to the cloud, where an image decoder reconstructs the cropped images. The reconstructed images are processed by a multimodal LLM to generate traffic condition descriptions. This approach achieves a 99.9% reduction in data transmission size while maintaining an LLM response accuracy of 89% for reconstructed cropped images, compared to 93% accuracy with original cropped images. Our results demonstrate the efficiency and practicality of ViT and LLM-assisted edge-cloud semantic communication for real-time traffic surveillance.
Sources
- DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving
- Semantic Communications: Principles and Challenges
- Transformer-Aided Semantic Communications
- Leveraging Multimodal-LLMs Assisted by Instance Segmentation for Intelligent Traffic Monitoring
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- YOLOv11: An Overview of the Key Architectural Enhancements
- UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding
- Importance-Aware Image Segmentation-based Semantic Communication for Autonomous Driving
- Generative AI-driven Semantic Communication Framework for NextG Wireless Network
- Generative AI-Enhanced Multi-Modal Semantic Communication in Internet of Vehicles: System Design and Methodologies
- Context-Aware Semantic Communication for the Wireless Networks
- Rethinking Generative Semantic Communication for Multi-User Systems with Large Language Models
- Vision Transformer Based Semantic Communications for Next Generation Wireless Networks
- Efficient Semantic Communication Through Transformer-Aided Compression
- Advancing Autonomous Vehicle Intelligence: Deep Learning and Multimodal LLM for Traffic Sign Recognition and Robust Lane Detection
- Large Language Models (LLMs) for Semantic Communication in Edge-based IoT Networks
- Large Language Model-Based Semantic Communication System for Image Transmission
- Large Language Model-Driven Distributed Integrated Multimodal Sensing and Semantic Communications
- Multimodal LLM Integrated Semantic Communications for 6G Immersive Experiences
- Leveraging Edge Intelligence and LLMs to Advance 6G-Enabled Internet of Automated Defense Vehicles
Related papers
- HiFiNet: Hierarchical Fault Identification in Wireless Sensor Networks via Edge-Based Classification and Graph Aggregation
- Embodied AI in 6G Networks: From Intelligent Connectivity to Physical Intelligence
- Lightweight GenAI for Network Traffic Generation: Fidelity, Augmentation, and Classification
- EdgePoW: Adaptive Ingress-Aware Defense with Non-Interactive PoW Against Volumetric SYN Floods
- SoK: Where Do Flow Labels Come From? Auditing Label Provenance in Encrypted Traffic Benchmarks
- What is Normal? A Big Data Observational Science Model of Anonymized Internet Traffic