HENet++: Hybrid Encoding and Multi-task Learning for 3D Perception and End-to-end Autonomous Driving

arXiv:2511.07106 · cs.CV · Submitted 2025-11-10 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "HENet++: Hybrid Encoding and Multi-task Learning for 3D Perception and End-to-end Autonomous Driving".

Jane: The gist: HENet++ achieves state-of-the-art end-to-end multi-task perception performance on the nuScenes dataset while attaining the lowest collision rate on the nuScenes end-to-end autonomous driving benchmark.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: We're diving into the specifics now on "HENet++: Hybrid Encoding and Multi-task Learning for three dee Perception and End-to-end Autonomous Driving <ref:2511.07106#pg1,HENet++: Hybrid Encoding and Multi-task Learning for 3D Perception and End>." The core thesis is that we need a smarter way to encode the sensor data because large image encoders, high resolution, and long temporal inputs just don't work together well right now.

Jane: They address this with Hybrid Encoding, using a large network for short frames and a smaller one for long sequences so you get both high detail and long context at a lower computational cost.

Lu: And they aren't stopping there; they also propose a multi-task BEV feature encoding specifically tailored for the Bird’s-Eye View models, analyzing task preferences to extract both sparse instance features and dense voxel features.

Meng: So, when you look at the numbers, it says this approach achieves leading multi-task perception performance on the nuScenes dataset. That's a big win if it holds up in real driving conditions.

Lalam: They also introduce a model-merging-based pre-training strategy to further boost that multitask accuracy, which is interesting because it uses weights from different single-task models as a starting point for the main encoder.

Tom: That sounds like they’re trying to leverage existing knowledge instead of training everything from scratch, which makes sense for complex tasks.

Jane: They also extend this into an end-to-end autonomous driving model using an attention-based world prediction module that handles both prediction and trajectory planning at the same time.

Lu: This whole structure, from the hybrid encoding to the multi-task feature extraction, is designed to give the planning module a much more comprehensive set of information than before.

Meng: So, what we see here is a system that’s trying to be smart about its resources while still giving it enough detail for complex driving decisions.

Conclusion: Tom: The paper "HENet++: Hybrid Encoding and Multi-task Learning for three dee Perception and End-to-end Autonomous Driving" shows them using a hybrid approach to handle the trade-off between high resolution and long context in perception tasks <ref:2511.07106#pg1,HENet++: Hybrid Encoding and Multi-task Learning for 3D Perception and End>.

Jane: They are pushing the idea of combining different encoding complexities to efficiently process both short, high-resolution frames and longer sequences, which is key for real driving situations.

Lu: What's important is that they aren't just focusing on one task; they are designing a multi-task BEV feature encoding that separates sparse features for things like object detection and dense voxel features for segmentation or occupancy prediction.

Meng: From an engineering side, the model-merging pre-training strategy suggests a clever way to initialize the encoder so it starts with some useful knowledge already embedded in multiple single-task models.

Lalam: This work implies that by thoughtfully designing how you encode and fuse features across different tasks, you can build a system that performs very well end-to-end on complex driving scenarios like those on the nuScenes dataset.

Tom: So, the big picture is about making perception systems more robust by smartly managing computational constraints while still demanding high quality data from sensors.

Wangxuan Institute of Computer Technology, Peking University · University of California, Merced

cs.CV

Submitted: 2025-11-10

Updated: 2026-10-08

Code: https://github.com/Tsinghua-MARS-Lab/Occ3D

Importance score: 81/100

The gist: The gist: HENet++ achieves state-of-the-art end-to-end multi-task perception performance on the nuScenes dataset while attaining the lowest collision rate on the nuScenes end-to-end autonomous

Key concepts

Hybrid Image Encoding Network
This network uses two different image encoders: a large one for high-resolution short frames and a small one for long frames. This allows the system to process different temporal inputs efficiently, balancing computational needs while maintaining compatibility with existing 3D feature extraction methods.
Temporal Feature Integration
This module fuses features extracted from multiple video frames using both backward and forward processes, including an adjacent frame fusion module (AFFM). This ensures that information from consecutive frames is effectively combined to create a richer representation for planning.
Independent BEV Feature Encoding
To prevent task conflicts, the framework generates separate 3D feature maps for different tasks by assigning different feature sizes. Each task gets its own encoder with independent weights, allowing the system to leverage task-specific preferences and improve overall multi-task performance.

Terminology

Summary

The gist: HENet++ achieves state-of-the-art end-to-end multi-task perception performance on the nuScenes dataset while attaining the lowest collision rate on the nuScenes end-to-end autonomous driving benchmark.

Hybrid Encoding and Multi-task Learning

The framework introduces a hybrid image encoding network that uses a large image encoder for short-term frames and a small one for long-term frames to address computational constraints The proposed architecture maintains compatibility with various existing 3D feature extraction methods and supports multimodal inputs. Furthermore, the framework simultaneously extracts both dense and sparse features, providing more suitable representations for different tasks, reducing cumulative errors, and delivering more comprehensive information to the planning module.

Hybrid Image Encoding Network

The Hybrid Encoding Network utilizes two image encoders of distinct complexities to process different temporal inputs. The first encoder handles high-resolution shortterm frames, passing them through a large image backbone(e.g., VoVNetV2-99 [62]) and a feature pyramid network (FPN) [63], and subsequently applies a complex 2D-to-BEV network to generate high-precision BEV features. The second encoder processes longterm sequences by down-sampling the inputs to low resolution, using a lightweight backbone(e.g., ResNet-50 [64]) with FPN for efficient feature extraction, followed by a simplified 2D-to-BEV module. Both pathways employ BEVPoolv2 [65] to project frustum Algorithm 1 Pseudo-code for Section 3.2.

Temporal Feature Integration

Following the extraction of multi-frame BEV features by the hybrid image encoding network, they are fused using a temporal integration module, which operates through complementary backward and forward processes. This module proposes the adjacent frame fusion module (AFFM) and adopts the temporal fusion strategy with temporal backward and forward processes. The AFFM operation can be formulated as AFFM(fi, fj) =fj + γ × Avg(Atn(⟨fi, fj⟩, fi, fi), Atn(⟨fi, fj⟩, fj, fj)).

Independent BEV Feature Encoding

To mitigate conflicts between tasks and leverage task preferences for different feature sizes, HENet proposes a preliminary solution that provides separate BEV feature maps for different tasks. This involves assigning different-sized BEV features to other tasks after obtaining the fused multi-scale BEV features. Inspired by BEVFusion [67], the proposed encoding process comprises adaptive feature selection and BEV encoding, where the adaptive feature selection fadaptive(F) = σ (Wfavg(F)) · F. The independent encoders for different tasks employ the same architecture but maintain independent weights.

Model-Merging-Based Pre-training

A model-merging-based pre-training strategy is introduced to further enhance multitask accuracy. This strategy consolidates the encoder weights from multiple single-task models into a single encoder and uses it as a pre-trained initialization. Experiments show that when using a 3D object detection model to initialize the encoder, the resulting multi-task model achieves higher accuracy in 3D object detection, while the accuracy of other tasks decreases.

End-to-End Autonomous Driving

The framework is extended into an end-to-end autonomous driving model that employs an attention-based worldprediction module to perform prediction and ego vehicle trajectory planning simultaneously. The prediction and planning decoder employs N Transformer layers for iterative prediction, where the initial query is defined as Q0 = ⟨F0, F1,2,…,K⟩. The total loss is a weighted sum L =αdepthLdepth + αclsLcls+αbboxLbbox + αsegLseg + αoccLocc. The end-to-end model introduces three additional losses: the L1 loss between the ego vehicle’s future trajectory and the ground truth, denoted as Lplan, and the collision constraint Lcol.

Conclusion

HENet++ achieves state-of-the-art end-toend multi-task perception performance on the nuScenes dataset. Based on the HENet++ perception framework, it is the first work that leverages Radar and Camera for end-toend autonomous driving. On the nuScenes dataset, the HENet++ model achieves a lower collision rate compared to existing methods.

How it works

The framework is structured in three stages: I) Hybrid Image Encoding Network, II) Temporal Feature Integration, and III) Independent BEV Feature Encoding. The Hybrid Image Encoding Network uses image encoders of varying complexity to encode long-sequence frames and short-term images, respectively.

Hybrid Encoding for Sparse Instance Feature

The framework innovatively introduces the simultaneous extraction of sparse foreground features and dense background voxel features by extending hybrid encoding to be compatible with both sparse and dense feature extraction. This process involves two backbones and FPNs to extract 2D feature maps from short-term high-resolution images and long-term low-resolution images, further deriving sparse instance features and dense voxel features from them.

Pretrain based on Model Merging

The model merging method is outlined in Algorithm 2, which combines multiple single-task models into a single encoder and uses it as a pre-trained initialization. For the weight matrices in convolutional networks and linear layers, Regression Mean is applied.

Loss for End-toEnd Autonomous Driving of HENet++

The total loss is L =αdepthLdepth + αclsLcls + αbboxLbbox + αsegLseg + αoccLocc. The end-to-end model introduces three additional losses: the L1 loss between the ego vehicle’s future trajectory and the ground truth, denoted as Lplan, and the collision constraint Lcol.

Final Results Comparison

HENet++ outperforms BEVFormer [20] by 11.0 NDS and 15.5 mAP on the 3D object detection task, 8.9 mIoU on the BEV semantic segmentation task, and 15.9 mIoU on the Occupancy task. HENet++ achieves state-of-the-art end-toend multi-task perception performance on the nuScenes dataset. The HENet++ end-toend autonomous driving model retained the top 200 detection results and downsampled the voxels to 50 × 50 × 4. The HENet++ end-toend autonomous driving model achieves a lower collision rate than existing methods.

Ablation Study Summary

The ablation study demonstrates that by combining with BEVDepth4D [13] and BEVStereo [15] through hybrid image encoding, HENet can significantly improve the 3D object detection performance. The proposed backward and forward processes with AFFM achieve the best results in temporal feature integration. The independent adaptive feature selection and BEV encoder for each task can further improve the multi-task performance of 1.7 NDS, 1.1 mAP, and 3.4 mIoU. The model merging strategy improves multi-task performance compared to using single-task detection or occupancy weights. The total loss is L =αdepthLdepth + αclsLcls + αbboxLbbox + αsegLseg + αoccLocc.

Conclusion

In this paper, we first present HENet, an end-toend framework for multi-task 3D perception. We propose a Hybrid Image Encoding Network for BEV and a Temporal Feature Integration Module to handle high-resolution, long-term temporal image inputs efficiently. Based on the HENet++ perception framework, the model leverages Radar and Camera for end-toend autonomous driving. The HENet++ model achieves state-of-the-art end-toend multi-task perception performance on the nuScenes dataset. Based on the HENet++ perception framework, it is the first work that leverages Radar and Camera for end-toend autonomous driving. On the nuScenes dataset, the HENet++ model achieves a lower collision rate compared to existing methods. The framework is compatible with both sparse and dense features, enabling relatively straightforward integration with existing multimodal feature fusion methods.

References

[1] Weng, X., Ivanovic, B., Wang, Y., Wang, Y., Pavone, M.: Para-drive: Parallelized architecture for real-time autonomous driving. In: CVPR, pp. 15449–15458 (2024)

[2] Zheng, W., Song, R., Guo, X., Chen, L.: Genad: Generative end-to-end autonomous driving.

Improvements for AI systems

  1. The HENet++ framework can simultaneously extract sparse foreground features and dense background voxel features, which enables end-to-end prediction for 3D object detection, BEV semantic segmentation, and occupancy semantic segmentation. This allows the system to leverage more comprehensive information for prediction and planning by providing task-specific representations.

  2. The model can perform end-to-end autonomous driving by utilizing an attention-based world-prediction module where "instance features (including the ego vehicle’s features) as queries, and instance features combined with panoramic dense features as key-value pairs, enabling simultaneous iterative prediction of future states and ego planning."

  3. The system can incorporate multimodal inputs by leveraging the framework's compatibility with both sparse and dense features to integrate Radar Encoder from RCBEVDet++ [71] along with the interpolation method described in the paper, allowing for end-to-end autonomous driving that uses both Radar and Camera.

  4. The model can improve multi-task accuracy through a model-merging based pre-training strategy by consolidating encoder weights from single-task models, which is shown to achieve better performance than loading parameters from 2D tasks alone.

  5. The system can achieve lower collision rates on the nuScenes benchmark by employing an end-to-end loss function that includes the L1 loss between the ego vehicle’s future trajectory and the ground truth, denoted as Lplan, alongside prediction and collision constraints (Lcol).

Sources

Related papers