HSI-Road Relabeled: Surface-Aware Road-Scene Segmentation

arXiv:2609.12151 · cs.CV · Submitted 2026-09-10 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "HSI-Road Relabeled: Surface-Aware Road-Scene Segmentation".

Jane: The HSI-Road dataset provides paired RGB and 25-channel NIR images with binary masks but no surface-level labels,

Tom: First, who's behind it and why it matters.

Title and authors: Jane: Now, let's look at the specific title and the people who wrote this paper, "HSI-Road Relabeled: Surface-Aware Road-Scene Segmentation." It sounds very focused on taking existing data and giving it a new layer of semantic detail related to the surface itself.

Tom: It is! The authors are Imad Ali Shah, Imran Mehmood, Enda Ward, Martin Glavin, Edward Jones, and Brian Deegan from the School of Engineering and Ryan Institute at the University of Galway in Ireland. They bring a solid academic background to this work on HSI-Road Relabeled: Surface-Aware Road-Scene Segmentation.

Lu: Considering the context they're working in, their approach suggests a strong interest in extending deep learning into areas like multispectral image analysis for practical applications, which is very exciting because it moves us closer to real-world sensing capabilities.

Meng: I wonder if this focus on surface annotation means that the resulting data will be immediately usable for industrial tasks without extensive retraining or calibration from our side.

Lalam: The paper's title itself tells us the main goal: moving from just seeing a road to understanding what’s *on* the road, like knowing it's water versus grass, which is a significant step in how we use AI for scene interpretation.

The paper's summary: Tom: So, summarizing what they did in "HSI-Road Relabeled: Surface-Aware Road-Scene Segmentation," the authors introduce this manual six-class taxonomy—Background, Asphalt, Concrete, Dirt, Water, and Grass—and they pair that with an RGB-to-NIR registration pipeline. Essentially, they are taking raw paired RGB and twenty-five-channel NIR images that only had binary masks and adding detailed surface labels <ref:2609.12151#pg0,paired RGB and 25-channel NIR>.

Jane: That means the core contribution is two parts: first, this new set of surface labels created manually using RGBori for finer detail than NIR, and second, a method to register those images so that the information from both modalities can be used together for segmentation.

Lu: The summary points to how they handle the complexity by modifying both the semantic granularity and the local spatial extent of their annotated surface regions, which is a sophisticated way to ensure consistency across the different classes.

Meng: So, if I understand correctly, they are solving a problem where we have limited information by creating a richer labeling scheme that bridges the gap between what we see in RGB and what we measure in NIR.

Lalam: Precisely; they're essentially building a bridge between two different types of visual data to achieve surface awareness, which is a powerful concept for any multimodal research.

The paper's improvements: Tom: Now, let’s talk about the improvements the authors suggest in this work on "HSI-Road Relabeled: Surface-Aware Road-Scene Segmentation." They are evaluating six different semantic-segmentation models, like UNet and SegFormer, across four distinct input configurations: original RGBori, registered low-resolution RGB (RGBreg), NIR, and the channel-stacked RGBNstk.

Jane: The improvement they are highlighting is how the system performs when you change the input resolution or modality; they show that using channel-stacked data, specifically RGBNstk—which is a twenty-eight-channel input—outperforms both NIR and registered low-resolution RGB for all six segmentation models at the matched one hundred ninety-two by three hundred eighty-four spatial resolution <ref:2609.12151#pg1>.

Lu: That comparison shows that even when you reduce the spatial resolution, stacking more channels of registered RGB information provides a substantial performance lift compared to relying on just the NIR data or lower-res RGB.

Meng: From a practical deployment view, that suggests we could potentially use this channel-stacked input if our hardware supports it well, as it seems to provide better results than relying solely on the narrower spectral bands.

Lalam: I also see an important improvement in how they handle the data itself; they show that RGBori gives the highest overall performance but uses twelve times more pixels than their other matched-resolution inputs, which sets a clear baseline for what high-fidelity results look like compared to lower-resolution inputs <ref:2609.12151#pg0,the highest overall performance but>.

Conclusion: Tom: So, to wrap up our discussion on "HSI-Road Relabeled: Surface-Aware Road-Scene Segmentation," the main implication is that spatial resolution really matters when comparing modalities, and that stacking complementary RGB–NIR information can give us measurable benefits depending on the specific surface class we are looking at.

Jane: They conclude that while RGBori has the best scores, once you match resolutions and use the right input combination like RGBNstk, you can get solid results across all six classes. This suggests that our future multimodal research should prioritize matching resolution before comparing different data views.

Lu: The implication for future work is clear: we need to focus on how to adapt the model choice based on deployment constraints rather than just chasing the absolute highest accuracy number.

Meng: I think that points to a practical necessity for us: we need to decide whether maximum segmentation accuracy or low latency is more important for our specific application before we commit resources.

Lalam: I think the paper's contribution is establishing this new framework where surface-aware annotations can guide model selection based on the input configuration, which is a valuable addition to how AI systems are designed for complex scene understanding.

Imad Ali Shah, Imran Mehmood, Enda Ward, Martin Glavin, Edward Jones, Brian Deegan

School of Engineering and Ryan Institute, University of Galway, Ireland

cs.CV

Submitted: 2026-09-10

Updated: 2026-10-05

Code: https://github.com/imadalishah/HSI_Road_relabeled

Importance score: 75/100

The gist: The HSI-Road dataset provides paired RGB and 25-channel NIR images with binary masks but no surface-level labels, and this work introduces a manually labeled six-class taxonomy along with an

Key concepts

Six Semantic Classes
The researchers defined six specific categories for road scenes: Background, Asphalt, Concrete, Dirt, Water, and Grass. These labels were created manually using the high-detail RGBori image to ensure accurate surface representation. This taxonomy allows models to distinguish between different types of road surfaces.
RGB-to-NIR Registration
This process creates a pseudo-NIR grayscale by averaging three specific RGB channels (0, 14, and 23). If the initial correlation is low, the system searches for the best triplet until a high correlation threshold (≥ 0.70) is met. This technique aligns the RGB and NIR data to ensure they correspond to the same physical surface.
Spatial Resolution Effect
The study found that spatial resolution significantly impacts performance. Comparing full-resolution RGBori with lower-resolution inputs like RGBreg or NIR showed substantial drops in accuracy. The impact varied by class, with 'Water' showing the largest performance drop compared to other surfaces.

Terminology

Summary

The HSI-Road dataset provides paired RGB and 25-channel NIR images with binary masks but no surface-level labels, and this work introduces a manually labeled six-class taxonomy along with an RGB-to-NIR registration pipeline to enable surface-aware road scene segmentation.

How it works

  1. Six semantic classes were introduced: Background, Asphalt, Concrete, Dirt, Water, and Grass. These annotations were created manually using RGBori, which provides finer spatial detail than NIR. The relabeling modifies both semantic granularity and the local spatial extent of the annotated surface region, meaning some pixels originally labeled Background are assigned to surface classes like Water and Grass.

  2. A pseudo-NIR grayscale was formed by averaging three channels, initially using 0, 14, and 23. If post-registration Normalized Cross-Correlation (NCC) was less than 0.70, a search for the highest NCC triplet was performed until NCC reached the threshold ≥ 0.70. This process involved contrast-enhanced grayscales matched using ORB with SIFT fallback and RANSAC for initial affine transform estimation, followed by ECC refinement to estimate a residual transform.

Evaluation Configurations

The six semantic-segmentation models (SSMs)—including UNet, UNet-CBAM, DeepLabV3+, SegFormer, and two UPerNet variants—were evaluated under four input configurations:

  1. RGBori: Original dataset resolution RGB (704 × 1280).

  2. RGBreg: Registered low-resolution RGB.

  3. NIR: The 25-channel NIR image (192 × 384).

  4. RGBNstk: Channel-stacked RGBreg–NIR (RGBNstk), which uses a 28-channel input (192 × 384).

Key Performance Findings

The comparison quantifies the effect of spatial-resolution reduction on RGB, as well as evaluation of NIR and RGBNstk.

** RGBori achieves the highest overall performance but contains 12× more pixels than the matched-resolution inputs. At the matched 192×384 resolution, RGBNstk outperforms NIR for all six SSMs and RGBreg for five of six.**

The most consistent gains were observed for the Water class.

For the six-class taxonomy, mIoU ranges from 73.39–82.73% and mF1 from 81.70–89.43%. UPerNetMiT-B0 achieved the highest mIoU for three input configurations (RGBori: 82.73± 0.40%, RGBreg: 77.03±0.27%, and RGBNstk: 79.66±0.46%).

Modality Comparison and Spatial Resolution Effects

The analysis focuses on comparing performance across modalities, noting that spatial resolution has a substantial effect on performance.

Mapping RGBori to RGBreg induces a near 12x difference in the number of pixels, resulting in a performance drop of RGBreg for all classes across the evaluated SSMs.

The reductions in mIoU points are: 4.44 for UNet, 4.01 for UNetCBAM, 4.72 for DeepLabV3+, 4.67 for SegFormer, 5.97 for UPerNetEN-B0, and 5.70 for UPerNetMiT-B0.

The effect is strongly class-dependent: Water drops IoU by 12.68 points and Grass by 8.89 points, compared to 2.95 for Conc., 2.34 for Asphalt, 2.13 for Dirt, and only 0.51 for Bkg.

Conclusion and Implications

The results suggest three practical considerations: (i) modality comparisons should be made at matched resolution, since spatial resolution has a substantial effect on performance; (ii) once resolution is matched, complementary RGB–NIR information can yield a measurable but SSM- and class-dependent benefit; and (iii) SSM–Encoder choice should be driven by deployment constraints as much as by segmentation accuracy. The work establishes that the newly aligned RGB-NIR offers a valuable modality for future multimodal research on HSI-Road.

Index Terms

HSI-Road, Multimodal Segmentation, NIR, Surface-oriented Road Annotations, UPerNet.

Improvements for AI systems

Based on the provided scientific paper, here are specific improvements for AI systems and what those improved systems can achieve:


) 1. Implement a robust, resolution-aware pre-processing pipeline for multi-modal road scene understanding. By adopting the RGBNstk (Channel-stacked RGBreg–NIR) input configuration when matched to the original NIR resolution, the system can leverage complementary spectral information to significantly enhance segmentation performance over single modalities (RGB or NIR).

  1. Integrate a knowledge base that dictates optimal SSM selection based on deployment constraints. The system can dynamically switch between architectures (e.g., favoring SegFormer for low latency/memory or UPerNet variants for maximum accuracy) based on whether the primary goal is high-fidelity mapping (using RGBori) or real-time autonomous driving applications (using RGBreg or NIR).

  2. Develop a class-specific performance calibration module. Since the benefit of stacking varies by road surface type—for instance, Water and Asphalt show significant gains, while Grass shows inconsistent benefits—the system can use input spectral data to predict which SSM/input combination will yield the best result for that specific local scene context, enabling adaptive segmentation accuracy.

  3. Enhance robustness against spatial misalignment artifacts. By incorporating the learned registration pipeline (based on NCC and feature matching) directly into the model architecture or training loop, the system can minimize performance degradation caused by inter-modality misregistration when using lower-resolution inputs (RGBreg or NIR).

  4. Optimize computational efficiency through model pruning and quantization tailored to specific modalities. Since SegFormer offers the lowest parameter count and latency, deploying it for resource-constrained edge devices in autonomous vehicles provides a high accuracy/low-latency trade-off, while larger UPerNet models can be reserved for higher-compute environments where maximum segmentation accuracy is paramount.

  5. Deploy the system as a Surface-Aware Road Scene Segmenter. The improved AI system will move beyond simple drivable area detection (Background/Road) to provide fine-grained material classification, enabling:

  6. Material-Informed Traction Control: Distinguishing between Asphalt, Concrete, and Dirt allows traction control systems to adjust torque and braking force based on the actual road surface condition (e.g., detecting slippery water patches or unstable dirt).

  7. Enhanced Road Condition Monitoring: Automated detection of standing water (Water class) or cracks/potholes (Dirt/Concrete classification) for proactive maintenance alerts, moving beyond simple visual detection.

  8. Improved Autonomous Navigation Safety: Accurate differentiation between drivable Grass and non-drivable surfaces is crucial for path planning, allowing the system to safely navigate grassy areas while avoiding hazardous ones like standing water or deep dirt tracks.

Related papers