HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting
Xikai Sun, Cangtian Zhou, Kebin Liu, Ke Ma, Xu Wang, Zaishu Chen, Haotian Wang, Li Liu, Yunhao Liu
Tsinghua University · Harbin Institute of Technology · Northwestern Polytechnical University · JD Logistics
cs.AI
Submitted: 2026-08-12
Updated: 2026-08-13
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
Terminology
Summary
Summary
This paper introduces Hugin, a training framework designed to enhance vision-language planning for autonomous logistics sorting systems (ALSS). The authors formulate the core problem as Joint Multi-Scene Understanding (JMSU), which requires joint planning over spatially disjoint camera views. JMSU has two defining characteristics: spatial distribution (observations cover physically disjoint regions with little field-of-view overlap) and decision-level interdependency (different scenes provide complementary evidence, and removing a key image may invalidate the plan).
The paper identifies two major challenges in applying vision-language models (VLMs) to JMSU: (1) substantial domain discrepancy and scarcity of high-quality JMSU data, since VLMs are pre-trained on web-scale data that differs from ALSS observations, and conventional augmentation may break consistency among distributed observations, physical states, and action labels; and (2) JMSU requires holistic reasoning over long visual contexts, where attention is distributed over a larger context as the number of visual tokens increases, making it harder to identify critical evidence.
To address these challenges, Hugin contains two complementary components. First, Endogenous Data Augmentation (EDA) decomposes each annotated sample into verifiable atomic facts and recombines them under domain-specific operational constraints, generating diverse training tasks while preserving consistency among visual observations, physical states, and action labels. EDA consists of three stages: decoupling (semantic decomposition into atomic facts using an LLM-based decoupler), combination (constraint-based task synthesis using a script-based synthesizer that recomputes complete plans under operating constraints), and augmentation (deriving auxiliary tasks for object detection, region identification, and cage-state recognition, plus mixing general VQA samples to reduce catastrophic forgetting).
Second, Global Context Ranking (GCR) is an auxiliary training objective that encourages the instruction representation to align more strongly with the complete visual context than with any partial visual context. GCR exploits the causal attention mechanism's information aggregation property, extracting hidden states at specific anchor positions: local visual context (uniformly sampled intermediate visual boundary), global visual context (final visual boundary), and instruction intent (token preceding trainable answer tokens). GCR optimizes a margin-based Hinge loss with a stop-gradient operator to prevent trivial solutions, using a small margin α=0.1 and a gated weight λ=0.01 that activates only for multi-image samples. GCR is removed after training, so inference cost and the base VLM's interface remain unchanged.
The authors collected 2,000 JMSU training samples over three months from four ALSS workstation layouts and constructed SortingBench with 1,000 held-out evaluation samples. The benchmark requires models to jointly determine the source compartment, localize the target package, select the destination cage, and generate the complete manipulation plan, with test sets including unseen layouts and lighting conditions.
Experiments across five open VLMs (Ovis2.5-2B, Gemma3-4B, MiniCPM-V4 5, Qwen3-VL-4B/8B) show consistent improvements over matched SFT baselines. For example, Qwen3-VL-8B accuracy on SortingBench increases from 63.6% to 78.8% (+15.2%). The improvements are consistent across seen and held-out layouts, with gains of 17.7%, 9.6%, 17.7%, and 16.6% across the four layouts. The paper also compares against RoboBrain2.5-8B-NV (an embodied VLM trained on 12.4M samples), which scores zero without adaptation and reaches 76.1% after the same 2,000-sample SFT, still below Hugin's 78.8% with only 22k samples.
Diagnostic tests show that Hugin is robust to image order shuffling (less than 0.3% variation) and to adding 6–12 distractor images (retaining 76.8% and 69.8% accuracy versus SFT's 45.8% and 18.8%). Ablation studies confirm that removing embodied-related EDA data causes a 5.5% drop on SortingBench, removing general VQA data causes regression in general capabilities (e.g., MMBench/EN drops from 81.4% to 77.25%), and GCR provides a 1.6% improvement on SortingBench and a 3.85% improvement on BLINK/VS cross-image comparison.
The paper also demonstrates positive spillover effects: Hugin improves BLINK/VS from 75.6% to 88.2% and MUIRBench/SU from 64.0% to 64.5% on Qwen3-VL-8B, while preserving general abilities (MME improves from 1699 to 1743, MMBench/CN changes from 81.3% to 81.5%). In real-world deployment testing, the Hugin-optimized Qwen3-VL-8B performed sorting planning for over 15,000 packages and achieved a prediction accuracy of 73.1%.
Improvements for AI systems
Improvements to AI Systems:
- Multi-Scene Joint Reasoning with Decision-Level Interdependency
-
The AI system can now process spatially disjoint camera views simultaneously, generating a single unified plan that requires evidence from all views. It can detect when a critical image is missing and flag the plan as invalid rather than producing a confident but wrong output.
-
Example: In a logistics sorting station, the system uses four non-overlapping camera feeds to determine the source bin, target package location, destination cage, and full manipulation sequence in one pass, even when the package is partially occluded in one view but visible in another.
- Consistent Data Augmentation via Atomic Fact Recombination
-
The system can generate new training tasks from existing annotated samples by decomposing them into verifiable facts (e.g.,
package A is in compartment 3,
cage B is empty
) and recombining them under operational constraints (e.g.,only one package per destination cage
). This preserves physical consistency and action-label validity, unlike naive image transformations. -
Result: The AI can be trained on far fewer real-world samples (e.g., 2,000) while achieving accuracy comparable to models trained on millions of generic samples, and it avoids catastrophic forgetting of general vision-language abilities.
- Global Context Ranking for Long Visual Sequences
-
The AI learns to prioritize the full visual context over any partial subset during training, using a margin-based loss on hidden states at specific anchor positions. This makes the model robust to image order shuffling and distractor images.
-
Result: The system can maintain high planning accuracy (e.g., 69.8%) even when 12 irrelevant images are inserted into the input, whereas a standard fine-tuned model drops to 18.8%. This is critical for real-world deployment where camera feeds may include noise or unrelated scenes.
- Auxiliary Task Integration for Embodied Perception
-
The system is trained with auxiliary objectives for object detection, region identification, and cage-state recognition (full/empty/partial) alongside the main planning task. This improves spatial grounding and reduces errors in downstream manipulation plans.
-
Example: The AI can now explicitly localize a package’s bounding box and identify which cage is available before generating a pick-and-place sequence, reducing planning failures caused by misperception.
- Generalization to Unseen Layouts and Lighting Conditions
-
By using EDA to synthesize diverse task combinations and GCR to enforce global context alignment, the system generalizes to new workstation layouts and lighting conditions without retraining.
-
Result: The AI can be deployed to a new sorting station with different camera angles or illumination and still achieve high planning accuracy (e.g., +16.6% improvement over baseline on a held-out layout).
- Cross-Domain Spillover and Preservation of General Capabilities
-
The training framework improves not only the target sorting task but also general vision-language skills, such as cross-image comparison (BLINK/VS: 75.6% → 88.2%) and multi-image understanding (MUIRBench/SU: 64.0% → 64.5%), while maintaining or slightly improving standard benchmarks (MME: 1699 → 1743).
-
Result: The improved AI system can be used for both specialized industrial planning and general-purpose visual question answering without performance trade-offs.
- Zero-Shot Adaptation for Embodied Agents
-
The system can be fine-tuned on a small, domain-specific dataset (e.g., 2,000 samples) to outperform large embodied models (e.g., RoboBrain2.5-8B trained on 12.4M samples) on the same task, reducing data collection and training costs by orders of magnitude.
-
Example: A robot in a warehouse can be quickly adapted to a new sorting task with only a few hours of annotated data, achieving 78.8% planning accuracy versus 76.1% for a much larger model.
- Robustness to Input Perturbations in Real-Time Operations
-
The system is resilient to image order shuffling (<0.3% accuracy variation) and can handle dynamic camera feeds where the number of images varies (6–12 distractors), making it suitable for real-time deployment in cluttered environments.
-
Result: In a live test, the improved AI successfully planned sorting for over 15,000 packages with 73.1% prediction accuracy, demonstrating operational viability.
Sources
- Pixtral 12B
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
- SimCSE: Simple Contrastive Learning of Sentence Embeddings
- Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
- OpenVLA: An Open-Source Vision-Language-Action Model
- MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models
- Ovis2.5 Technical Report
- RationalVLA: A Rational Vision-Language-Action Model with Dual System
- RoboBrain 2.5: Depth in Sight, Time in Mind
- MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World
- DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
- VLA-R1: Enhancing Reasoning in Vision-Language-Action Models
- MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection