Traffic Co-Simulation Framework Empowered by Infrastructure Camera Sensing and Reinforcement Learning

arXiv:2412.03925 · eess.SY, cs.LG, cs.SY, eess.IV · Submitted 2025-09-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Traffic Co-Simulation Framework Empowered by Infrastructure Camera Sensing and Reinforcement Learning".

Jane: The paper was written by Talha Azfar, Kaicong Huang, Andrew Tracy, Sandra Misiewicz, Chenxi Liu et al. from Rensselaer Polytechnic Institute and Capital Region Transportation Council and University of Utah.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Alright, welcome back to the show, everybody. We've got a paper that's going to make you look at every red light a little differently. It's called "Traffic Co-Simulation Framework Empowered by Infrastructure Camera Sensing and Reinforcement Learning."

Jane: And Tom, I have to say, just reading that title got me excited, because it's pulling together three things that usually live in separate worlds. You've got the high-fidelity three dee world of CARLA, the large-scale traffic flow simulation of SUMO, and then you've got reinforcement learning agents making the decisions.

Tom: Exactly. And the key phrase in that title for me is "infrastructure camera sensing." Because for years, we've been training these smart traffic light agents on perfect data, like the simulator just hands them the exact number of cars waiting. But that's not how the real world works, right?

Jane: Not at all. In the real world, you've got a camera mounted on a pole, maybe it's raining, maybe a truck is blocking the view, maybe the sun is glaring right into the lens. The detection is going to be messy. And this paper is saying, let's embrace that mess and see if our agents can still do a good job.

Lu: And that's the part that really caught my attention. I'm Lu, by the way. The fact that they're not just assuming perfect detection is a huge step. They're actually testing the robustness of the multi-agent reinforcement learning under faulty or sparse sensing conditions. That's the difference between a lab experiment and something that could actually be deployed.

Tom: So Lu, you're saying this is closing the gap between the simulation and the real street corner?

Lu: Precisely. The "simulation-reality gap" is the biggest hurdle in this field. You can build a perfect agent in a perfect world, but the moment you put it on a real intersection, it falls apart because the inputs are noisy. This paper is trying to bridge that.

Meng: And from an engineering standpoint, I'm Meng, the fact that they're using YOLO for detection is smart. It's fast, it's well-known, and it runs in real-time. But I'm curious about the computational load. Running forty-three cameras through a neural network while keeping two simulators in sync sounds like a recipe for a slow simulation.

Jane: That's a great point, Meng, and we'll get into the numbers later. But the fact that they even attempted it, and got it working, is a testament to the framework they built. It's not just a theoretical idea; they actually ran the experiments.

Tom: So we've got the title, we've got the promise of robustness, and we've got a real test-bed. What I want to know is, how did they actually pull this off? How do you get CARLA and SUMO to talk to each other and have the cameras feed the agents? That's our next topic.

Summary: Jane: So we've established that this paper, "Traffic Co-Simulation Framework Empowered by Infrastructure Camera Sensing and Reinforcement Learning," is trying to make traffic signal control more realistic. Now let's talk about what they actually did, the summary of the whole thing.

Tom: Right. And the summary is basically a loop. You have SUMO, which is the traffic simulator, handling the flow of thousands of vehicles. Then you have CARLA, which is the high-fidelity three dee world, rendering those same vehicles in a realistic cityscape.

Lu: And the bridge between them is the cameras. They mounted virtual cameras on the traffic light poles in CARLA. These cameras take pictures, and a YOLO-based computer vision system counts the vehicles. That count becomes the "state" that the reinforcement learning agents in SUMO use to decide what to do with the traffic lights.

Meng: So the agents aren't cheating by reading the simulator's internal data. They're only looking at what a real camera would see. That's the crucial part. They're using the visual feed as the sole input for the adaptive signal control.

Jane: Exactly, Meng. And they trained these agents using multi-agent reinforcement learning, where each traffic light is its own agent. They tested four different reward functions to see which one taught the agents the best behavior.

Tom: And the results were pretty striking. In their heavy traffic scenario, the static traffic lights gave an average speed of four point five seven meters per second. But the best MARL agent, the one trained on average speed reward, boosted that to six point zero two meters per second. That's a massive improvement in flow.

Lu: And they didn't just test with perfect information. They ran the whole thing with the camera detections, which included errors. The detection wasn't perfect, but the agents still performed well. That's the robustness we were talking about.

Meng: But I'm guessing not all reward functions were created equal. Did some of them fall apart when the sensing got noisy?

Jane: Oh, absolutely. The queue length based reward was a disaster. It performed terribly. And the reason is fascinating. During heavy congestion, the queue would spill back beyond what the camera could see. So the agent always saw a "full queue," and it never learned that giving a green light actually helped.

Tom: That's a great example of why this kind of testing is so important. You think you've trained a great agent, and then you realize it was relying on information that isn't available in the real world. So we've got the setup and the results. But what's the actual improvement here? What did they build that's new? Let's dig into that.

Improvements: Tom: So we're back with "Traffic Co-Simulation Framework Empowered by Infrastructure Camera Sensing and Reinforcement Learning." We've talked about the setup and the results. Now, what's the actual improvement? What did this team build that nobody else has?

Jane: For me, the biggest improvement is the automation of the camera setup. In the past, you'd have to manually place cameras in a simulation and hope they were pointing at the right thing. This paper programmatically accesses the traffic light pole locations in CARLA, spawns the cameras, and then uses coordinate matching to figure out the yaw angle so each camera is aimed perfectly at the stopping line.

Lu: And that's a huge deal for scalability. If you want to test a network with a hundred intersections, you can't manually configure a hundred cameras. This framework does it automatically. It's a step towards a real digital twin, where you can just plug in a new city map and have the whole sensing infrastructure set up for you.

Meng: I agree, but I think the more important improvement is the robustness testing. They didn't just build the framework and say "look, it works." They actively degraded the sensing. They used YOLOv5, which had a mean absolute error of two point two one cars per network snapshot, and then they swapped in YOLOv8, which was more accurate at one point seven three.

Tom: And the performance followed the accuracy. With YOLOv8, the best agent's average speed went from six point zero two to six point seven eight meters per second, getting much closer to the ground truth performance of seven point one one. So they showed that better sensing directly translates to better traffic flow.

Jane: But the really interesting part, the part that I think is the biggest improvement, is that they showed the agents can still work even with imperfect detection. The agents trained on average speed reward were relatively robust. They only took a small performance hit compared to the ground truth scenario.

Lu: Exactly. That's the key finding. You don't need perfect sensing to get a benefit. You just need a reward function that's resilient to noise. The wait time and pressure based agents were much more sensitive to the detection errors, which tells us a lot about how we should design these systems.

Meng: So the improvement isn't just the code, it's the insight. It's knowing that if you're deploying this in the real world, you should probably use a speed-based reward, because it's going to be more forgiving when your cameras aren't perfect.

Tom: That's a fantastic point. So we've got the automated setup and the robustness analysis. But I'm still curious about the nitty-gritty of the first page. What's the actual problem statement? Why is this even necessary? Let's look at the introduction.

First Page: Jane: So we're diving into the first page of "Traffic Co-Simulation Framework Empowered by Infrastructure Camera Sensing and Reinforcement Learning." And the first page sets the stage for why this whole thing is necessary.

Tom: Right. It starts with the problem of traffic congestion. It's not just an annoyance; it's a huge economic and environmental cost. And the paper points out that traditional traffic signal control, the fixed-time schedules and simple rule-based systems, just don't adapt well to dynamic conditions.

Lu: And that's where reinforcement learning comes in. It's a way to learn adaptive control strategies. But the paper makes a critical observation: most of these RL studies assume perfect vehicle detection. They assume the agent knows exactly how many cars are waiting, which is just not realistic.

Meng: And that's the gap they're addressing. They mention the Capital Region of New York State, where most signals are pre-timed or actuated, but not coordinated for network-wide optimization. So there's a real-world motivation here, not just an academic exercise.

Jane: Exactly. And they're proposing this co-simulation framework as a virtual test-bed. A place where you can test these AI-driven control strategies with realistic sensor input before you ever deploy them on a real street. It's a safe way to fail.

Tom: So the first page is really about the "why." Why do we need this? Because the current tools don't let us test with realistic sensors, and the current control methods aren't smart enough. And they're saying, "Hey, we've built a bridge between the high-fidelity world and the large-scale simulation world to fix that."

Lu: And I think it's worth pointing out that they're not just talking about traffic lights. The first page hints at the bigger picture. This framework is relevant to connected automated vehicles, or CAVs, which depend on infrastructure-based sensing. If we can prove that traffic control works with noisy camera data, that's a big step for CAV deployment.

Meng: So the first page is the motivation, and it's a strong one. It's saying, "We have a real problem, the current solutions are inadequate, and here's a new way to test better solutions." It sets up the rest of the paper perfectly.

Jane: And it sets up the conclusion, which is where we see the real-world impact. We've got the problem, the solution, and the results. Now we need to talk about what this means for the future of our cities. Let's wrap this up.

Conclusion: Tom: Alright, we've reached the end of our discussion on "Traffic Co-Simulation Framework Empowered by Infrastructure Camera Sensing and Reinforcement Learning." And I think we can all agree this is a significant piece of work.

Jane: Absolutely. To summarize, they built a co-simulation framework that combines CARLA's realistic three dee visuals with SUMO's large-scale traffic flow. They mounted virtual cameras on traffic light poles, used YOLO to detect vehicles, and fed that visual data to multi-agent reinforcement learning agents controlling the signals.

Lu: And the key takeaway for me is the robustness. They showed that even with imperfect detection, the agents trained on average speed reward could significantly improve traffic flow compared to static signals. The detection errors didn't break the system, which is crucial for real-world deployment.

Meng: And from a practical standpoint, they gave us a clear roadmap. They compared YOLOv5 and YOLOv8, showed that better accuracy improves performance, but also showed that you don't need perfect accuracy to get a benefit. That's a green light for engineers looking to deploy this.

Tom: And the implications are huge. This isn't just about traffic lights. This framework lays the groundwork for digital twins of entire urban traffic networks. You could use it to test variable speed limits, parking management, even emergency vehicle routing, all with realistic sensor input.

Jane: And for connected automated vehicles, this is a stepping stone. If we can prove that infrastructure-based sensing and control works reliably, even with noisy data, that builds confidence for CAV integration. It's a bridge between simulation and the real world.

Lu: I think the most exciting part is the future work they mentioned. They want to add multi-object tracking, like DeepSORT, to improve temporal consistency. And they want to explore custom-trained detectors for the CARLA environment. That could push the accuracy even higher.

Meng: And they mentioned edge computing. Running the detection on the camera itself, or on a roadside unit, instead of a central server. That's the kind of distributed architecture that makes this scalable to a whole city.

Tom: Well said, everyone. It's been a fantastic discussion. We've said goodbye to "Traffic Co-Simulation Framework Empowered by Infrastructure Camera Sensing and Reinforcement Learning," and I'm already looking forward to the next paper. Thanks for listening, and we'll see you next time.

Jane: See you soon, everyone.

Talha Azfar, Kaicong Huang, Andrew Tracy, Sandra Misiewicz, Chenxi Liu, Ruimin Ke

Rensselaer Polytechnic Institute · Capital Region Transportation Council · University of Utah

eess.SY, cs.LG, cs.SY, eess.IV

Submitted: 2025-09-18

Updated: 2026-08-12

Journal ref: Journal of Intelligent Transportation Systems, 2025

DOI: 10.1080/15472450.2025.2559410

Code: https://github.com/LucasAlegre/sumo-rl

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 69/100

Terminology

Summary

Summary

This paper proposes and evaluates a co-simulation framework that integrates the CARLA and SUMO simulators to enable multi-agent reinforcement learning (MARL) for adaptive traffic signal control (TSC) using infrastructure camera-based vehicle detection. The study addresses the gap between simulation-based TSC research and real-world deployment, where perfect vehicle detection is often assumed but rarely available.

The framework uses CARLA for high-fidelity 3D urban environments and vehicle dynamics, while SUMO handles large-scale traffic flow simulation and signal control. Virtual cameras are mounted on traffic light poles within CARLA, and a YOLO-based computer vision system detects and counts vehicles in real-time. This visual data is used as input for MARL agents controlling traffic signals in SUMO. The paper states: Cameras mounted on traffic light poles within the CARLA environment use a YOLO-based computer vision system to detect and count vehicles, providing realtime traffic data as input for adaptive signal control in SUMO.

The methodology involves several key components. First, the framework programmatically accesses traffic light pole locations in CARLA, spawns virtual cameras at these locations, and uses coordinate matching between CARLA and SUMO to identify the stopping line of each lane. The camera's yaw angle is automatically calculated to ensure proper aiming. Second, MARL agents are trained using Q-learning with four different reward functions: difference of wait times, average speed, queue length, and pressure. Each traffic light is an independent agent with its own state space (current phase, minimum green time elapsed, lane occupancy, and queue length) and action space (index of the next green phase). The Q-table is updated using the equation: Q(s, a) ← Q(s, a) + α[r + γ maxa' Q(s', a') − Q(s, a)].

The training was performed on the Town04 map from CARLA, which contains a downtown grid-like region with 12 traffic lights. The paper notes: The East-West signal spacing is about 40-50 meters, while the North-South spacing is 50-70 meters. Hyperparameters were tuned with learning rate α = 0.0071, discount factor γ = 0.97, and initial exploration rate ε = 0.05 decaying to 0.005, over 300 training episodes.

Experiments were conducted under two traffic demand scenarios: medium (387 vehicles at 100 to 300 veh/hr) and heavy (591 vehicles at 200 to 400 veh/hr). The most congested signal in the medium demand scenario had a v/C ratio of 0.88 (LOS 'E'), while the heavy demand scenario had a v/C ratio of 1.26 (LOS 'F', oversaturated).

Results showed that MARL-based TSC significantly improved traffic flow compared to static and actuated signal control. The agents trained with the average speed reward function performed best in both test scenarios. In the heavy traffic scenario with camera-based detection, the average speed reward agent achieved a mean speed of 6.02 m/s (compared to 4.57 m/s for static signals), total travel time of 208,443 seconds (compared to 355,699 seconds for static), and mean waiting time of 77.9 seconds (compared to 321 seconds for static). The paper states: Out of the four variants tested, the agents trained with the reward function based on the average speed of vehicles near the intersection performed the best in both test scenarios by every measure.

The study also evaluated robustness under faulty or sparse sensing. Vehicle detection errors were analyzed, showing that YOLOv5 had a mean absolute error of 2.21 cars and root mean square error of 3.48 cars across the network of 43 approaches. YOLOv8, a newer model, achieved higher accuracy with mean absolute error of 1.73 and root mean square error of 2.80, but was slower to execute. The paper notes: while increased accuracy improves performance, even imperfect detection can yield positive results with the right MARL training.

The paper identifies specific failure modes. Queue length based reward agents performed poorly because during congestion, traffic backed up onto the highway outside the camera's observation range, causing the agent to always observe a full queue and fail to learn effective phase changes. The paper explains: the agent always observes that the queue is full, and any action it takes will not change the size of the queue at the next measurement. Interestingly, the camera detection version performed better than ground truth for this reward because detection errors caused different Q-table rows to be accessed, forcing phase changes.

The paper also discusses the interpretability of learned policies through Q-table inspection, providing an example: State: (0, 1, 0, 7, 0, 9, 0, 7, 0, 0) → Q-values: (51.12, 0.01) indicating the signal is in phase 0 with 7 cars from the West queued and 9 cars from the East moving, leading to a high Q-value for remaining in phase zero.

The framework includes error handling mechanisms, such as forcing a phase change if a signal is stuck for over 60 seconds, and using the last known state if a camera fails to capture an image. The paper concludes by discussing extensions to digital twin technologies, where real-world sensor data from cameras and connected automated vehicles (CAVs) could be integrated with the co-simulation environment for continuous learning and adaptive control.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system, along with the resulting capabilities:


  1. Vision-to-State Translation Module for MARL:
  • Improvement: I will build a dedicated module that bridges the gap between raw, high-dimensional image data from infrastructure cameras and the structured state vectors (e.g., lane occupancy, queue length) required by pre-trained MARL agents. This module will use a YOLO-based object detector, filter detections by vehicle type, apply a region-of-interest mask to isolate specific lanes, and aggregate counts across multiple cameras for a single intersection.

  • Capability: The system can now control traffic signals in a realistic 3D environment (CARLA) using only visual input, without relying on ground-truth simulator data. This makes it directly transferable to real-world deployments where only camera feeds are available.

  1. Robust MARL Training with Imperfect Sensing:
  • Improvement: I will implement a training pipeline that explicitly accounts for noisy or faulty detection. The paper shows that agents trained on certain reward functions (e.g., queue length) are highly sensitive to detection errors. I will incorporate a noise-injection layer during training, simulating over-counting and under-counting of vehicles, to force the agents to learn policies that are robust to imperfect observations.

  • Capability: The resulting AI system will be significantly more resilient to real-world sensor failures, occlusions, and misdetections. It will not catastrophically fail (e.g., get stuck in a phase) when the input data is noisy, ensuring stable and safe traffic control even with sub-optimal hardware.

  1. Reward Function Selection and Optimization:
  • Improvement: I will implement and compare the four reward functions from the paper (wait time difference, average speed, queue length, pressure). I will use the paper's findings to prioritize the average speed reward, which demonstrated superior robustness to sensing errors. I will also implement a mechanism to dynamically switch or blend reward functions based on real-time congestion levels.

  • Capability: The system will learn more effective and stable control policies. It will prioritize overall network throughput and minimize individual delays, especially in heavy traffic, without being misled by transient spikes in queue length or pressure caused by detection errors.

  1. Automated Camera Calibration and Spawning:
  • Improvement: I will implement an automated process that programmatically accesses traffic light pole locations in a 3D environment (CARLA), calculates the correct yaw and pitch angles to aim the camera at the lane's stopping line, and spawns the virtual camera. This eliminates manual configuration and allows for scaling to large networks.

  • Capability: The system can be rapidly deployed to any new intersection or city without human intervention for camera setup. This scalability is critical for practical, city-wide deployment of intelligent traffic management.

  1. Asynchronous Control Loop for Real-Time Operation:
  • Improvement: I will implement an asynchronous control loop that decouples the high-frequency simulation (20 Hz) from the computationally expensive object detection (1 Hz). The simulation will pause periodically to perform batched detection across all cameras, then resume, ensuring data consistency and preventing performance bottlenecks.

  • Capability: The system can operate in near real-time, even with many cameras and a computationally heavy vision model. This makes it feasible for live traffic management, where decisions must be made within seconds.

  • Control a network of traffic signals in a high-fidelity 3D simulation using only camera input, mimicking real-world infrastructure.

  • Maintain stable and efficient traffic flow even when vehicle detection is imperfect, demonstrating robustness to sensor noise, occlusions, and weather conditions.

  • Outperform traditional fixed-time and actuated signal control in terms of average speed, total travel time, and waiting time, especially under heavy congestion.

  • Automatically adapt to new intersections by self-calibrating cameras and using pre-trained, transferable MARL agents.

  • Operate in real-time, making control decisions every few seconds based on the latest visual data, making it suitable for live deployment.

  • Provide a testbed for evaluating the impact of different sensor qualities (e.g., YOLOv5 vs. YOLOv8) and failure modes on system performance, allowing for informed infrastructure investment decisions.

Abstract

Traffic simulations are commonly used to optimize urban traffic flow, with reinforcement learning (RL) showing promising potential for automated traffic signal control, particularly in intelligent transportation systems involving connected automated vehicles. Multi-agent reinforcement learning (MARL) is particularly effective for learning control strategies for traffic lights in a network using iterative simulations. However, existing methods often assume perfect vehicle detection, which overlooks real-world limitations related to infrastructure availability and sensor reliability. This study proposes a co-simulation framework integrating CARLA and SUMO, which combines high-fidelity 3D modeling with large-scale traffic flow simulation. Cameras mounted on traffic light poles within the CARLA environment use a YOLO-based computer vision system to detect and count vehicles, providing real-time traffic data as input for adaptive signal control in SUMO. MARL agents trained with four different reward structures leverage this visual feedback to optimize signal timings and improve network-wide traffic flow. Experiments in a multi-intersection test-bed demonstrate the effectiveness of the proposed MARL approach in enhancing traffic conditions using real-time camera based detection. The framework also evaluates the robustness of MARL under faulty or sparse sensing and compares the performance of YOLOv5 and YOLOv8 for vehicle detection. Results show that while better accuracy improves performance, MARL agents can still achieve significant improvements with imperfect detection, demonstrating scalability and adaptability for real-world scenarios.

Sources

Related papers