HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting

arXiv:2608.11692 · cs.AI · Submitted 2026-08-12 · Read on arXiv

Xikai Sun, Cangtian Zhou, Kebin Liu, Ke Ma, Xu Wang, Zaishu Chen, Haotian Wang, Li Liu, Yunhao Liu

Tsinghua University · Harbin Institute of Technology · Northwestern Polytechnical University · JD Logistics

cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

Terminology

Summary

Summary

This paper introduces Hugin, a training framework designed to enhance vision-language planning for autonomous logistics sorting systems (ALSS). The authors formulate the core problem as Joint Multi-Scene Understanding (JMSU), which requires joint planning over spatially disjoint camera views. JMSU has two defining characteristics: spatial distribution (observations cover physically disjoint regions with little field-of-view overlap) and decision-level interdependency (different scenes provide complementary evidence, and removing a key image may invalidate the plan).

The paper identifies two major challenges in applying vision-language models (VLMs) to JMSU: (1) substantial domain discrepancy and scarcity of high-quality JMSU data, since VLMs are pre-trained on web-scale data that differs from ALSS observations, and conventional augmentation may break consistency among distributed observations, physical states, and action labels; and (2) JMSU requires holistic reasoning over long visual contexts, where attention is distributed over a larger context as the number of visual tokens increases, making it harder to identify critical evidence.

To address these challenges, Hugin contains two complementary components. First, Endogenous Data Augmentation (EDA) decomposes each annotated sample into verifiable atomic facts and recombines them under domain-specific operational constraints, generating diverse training tasks while preserving consistency among visual observations, physical states, and action labels. EDA consists of three stages: decoupling (semantic decomposition into atomic facts using an LLM-based decoupler), combination (constraint-based task synthesis using a script-based synthesizer that recomputes complete plans under operating constraints), and augmentation (deriving auxiliary tasks for object detection, region identification, and cage-state recognition, plus mixing general VQA samples to reduce catastrophic forgetting).

Second, Global Context Ranking (GCR) is an auxiliary training objective that encourages the instruction representation to align more strongly with the complete visual context than with any partial visual context. GCR exploits the causal attention mechanism's information aggregation property, extracting hidden states at specific anchor positions: local visual context (uniformly sampled intermediate visual boundary), global visual context (final visual boundary), and instruction intent (token preceding trainable answer tokens). GCR optimizes a margin-based Hinge loss with a stop-gradient operator to prevent trivial solutions, using a small margin α=0.1 and a gated weight λ=0.01 that activates only for multi-image samples. GCR is removed after training, so inference cost and the base VLM's interface remain unchanged.

The authors collected 2,000 JMSU training samples over three months from four ALSS workstation layouts and constructed SortingBench with 1,000 held-out evaluation samples. The benchmark requires models to jointly determine the source compartment, localize the target package, select the destination cage, and generate the complete manipulation plan, with test sets including unseen layouts and lighting conditions.

Experiments across five open VLMs (Ovis2.5-2B, Gemma3-4B, MiniCPM-V4 5, Qwen3-VL-4B/8B) show consistent improvements over matched SFT baselines. For example, Qwen3-VL-8B accuracy on SortingBench increases from 63.6% to 78.8% (+15.2%). The improvements are consistent across seen and held-out layouts, with gains of 17.7%, 9.6%, 17.7%, and 16.6% across the four layouts. The paper also compares against RoboBrain2.5-8B-NV (an embodied VLM trained on 12.4M samples), which scores zero without adaptation and reaches 76.1% after the same 2,000-sample SFT, still below Hugin's 78.8% with only 22k samples.

Diagnostic tests show that Hugin is robust to image order shuffling (less than 0.3% variation) and to adding 6–12 distractor images (retaining 76.8% and 69.8% accuracy versus SFT's 45.8% and 18.8%). Ablation studies confirm that removing embodied-related EDA data causes a 5.5% drop on SortingBench, removing general VQA data causes regression in general capabilities (e.g., MMBench/EN drops from 81.4% to 77.25%), and GCR provides a 1.6% improvement on SortingBench and a 3.85% improvement on BLINK/VS cross-image comparison.

The paper also demonstrates positive spillover effects: Hugin improves BLINK/VS from 75.6% to 88.2% and MUIRBench/SU from 64.0% to 64.5% on Qwen3-VL-8B, while preserving general abilities (MME improves from 1699 to 1743, MMBench/CN changes from 81.3% to 81.5%). In real-world deployment testing, the Hugin-optimized Qwen3-VL-8B performed sorting planning for over 15,000 packages and achieved a prediction accuracy of 73.1%.

Improvements for AI systems

Improvements to AI Systems:

  1. Multi-Scene Joint Reasoning with Decision-Level Interdependency
  • The AI system can now process spatially disjoint camera views simultaneously, generating a single unified plan that requires evidence from all views. It can detect when a critical image is missing and flag the plan as invalid rather than producing a confident but wrong output.

  • Example: In a logistics sorting station, the system uses four non-overlapping camera feeds to determine the source bin, target package location, destination cage, and full manipulation sequence in one pass, even when the package is partially occluded in one view but visible in another.

  1. Consistent Data Augmentation via Atomic Fact Recombination
  • The system can generate new training tasks from existing annotated samples by decomposing them into verifiable facts (e.g., package A is in compartment 3, cage B is empty) and recombining them under operational constraints (e.g., only one package per destination cage). This preserves physical consistency and action-label validity, unlike naive image transformations.

  • Result: The AI can be trained on far fewer real-world samples (e.g., 2,000) while achieving accuracy comparable to models trained on millions of generic samples, and it avoids catastrophic forgetting of general vision-language abilities.

  1. Global Context Ranking for Long Visual Sequences
  • The AI learns to prioritize the full visual context over any partial subset during training, using a margin-based loss on hidden states at specific anchor positions. This makes the model robust to image order shuffling and distractor images.

  • Result: The system can maintain high planning accuracy (e.g., 69.8%) even when 12 irrelevant images are inserted into the input, whereas a standard fine-tuned model drops to 18.8%. This is critical for real-world deployment where camera feeds may include noise or unrelated scenes.

  1. Auxiliary Task Integration for Embodied Perception
  • The system is trained with auxiliary objectives for object detection, region identification, and cage-state recognition (full/empty/partial) alongside the main planning task. This improves spatial grounding and reduces errors in downstream manipulation plans.

  • Example: The AI can now explicitly localize a package’s bounding box and identify which cage is available before generating a pick-and-place sequence, reducing planning failures caused by misperception.

  1. Generalization to Unseen Layouts and Lighting Conditions
  • By using EDA to synthesize diverse task combinations and GCR to enforce global context alignment, the system generalizes to new workstation layouts and lighting conditions without retraining.

  • Result: The AI can be deployed to a new sorting station with different camera angles or illumination and still achieve high planning accuracy (e.g., +16.6% improvement over baseline on a held-out layout).

  1. Cross-Domain Spillover and Preservation of General Capabilities
  • The training framework improves not only the target sorting task but also general vision-language skills, such as cross-image comparison (BLINK/VS: 75.6% → 88.2%) and multi-image understanding (MUIRBench/SU: 64.0% → 64.5%), while maintaining or slightly improving standard benchmarks (MME: 1699 → 1743).

  • Result: The improved AI system can be used for both specialized industrial planning and general-purpose visual question answering without performance trade-offs.

  1. Zero-Shot Adaptation for Embodied Agents
  • The system can be fine-tuned on a small, domain-specific dataset (e.g., 2,000 samples) to outperform large embodied models (e.g., RoboBrain2.5-8B trained on 12.4M samples) on the same task, reducing data collection and training costs by orders of magnitude.

  • Example: A robot in a warehouse can be quickly adapted to a new sorting task with only a few hours of annotated data, achieving 78.8% planning accuracy versus 76.1% for a much larger model.

  1. Robustness to Input Perturbations in Real-Time Operations
  • The system is resilient to image order shuffling (<0.3% accuracy variation) and can handle dynamic camera feeds where the number of images varies (6–12 distractors), making it suitable for real-time deployment in cluttered environments.

  • Result: In a live test, the improved AI successfully planned sorting for over 15,000 packages with 73.1% prediction accuracy, demonstrating operational viability.

Sources

Related papers