FigEx2: Visual-Conditioned Panel Detection and Captioning for Scientific Compound Figures
cs.CV, cs.AI, cs.CL
Submitted: 2026-01-12
Updated: 2026-09-05
Code: https://github.com/Huang-AI4Medicine-Lab/FigEx2
License: http://creativecommons.org/licenses/by/4.0/
The gist: Scientific compound figures combine multiple labeled panels into a single image, and downstream pretraining and retrieval require panel-aligned visual-text pairs.
Terminology
Abstract
Scientific compound figures combine multiple labeled panels into a single image, and downstream pretraining and retrieval require panel-aligned visual-text pairs. However, in a PubMed Central (PMC)-scale crawl of 346,567 compound figures, 16.3% have no caption and are discarded by existing caption-decomposition pipelines. We propose FigEx2, a visual-conditioned framework that takes only a compound figure as input and jointly produces labeled panel boxes and panel-wise captions. FigEx2 introduces an Entity-Attention Kullback-Leibler (KL) regularizer that aligns the detector's cross-attention with scientific entities annotated for each panel, providing a stable conditioning signal that also improves localization, and applies Group Relative Policy Optimization (GRPO) with a panel-level Entity-F1 reward to optimize scientific faithfulness. We curate BioSci-Fig-Cap for in-domain supervision and contribute physics and chemistry test suites for cross-disciplinary evaluation. FigEx2 achieves 0.751 mAP@0.5:0.95 on BioSci-Fig-Cap, and outperforms Qwen3-VL-8B by 6.80 Entity-F1 on MedICaT for captioning. It also transfers zero-shot to out-of-distribution domains. The source code is available at https://github.com/Huang-AI4Medicine-Lab/FigEx2.
Sources
- Qwen2.5-VL Technical Report
- Mitigating Open-Vocabulary Caption Hallucinations
- Pix2seq: A Language Modeling Framework for Object Detection
- Fine-grained Image Captioning with CLIP Reward
- VLRM: Vision-Language Models act as Reward Models for Image Captioning
- Open-vocabulary Object Detection via Vision and Language Knowledge Distillation
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- DetGPT: Detect What You Need via Reasoning
- Sequence Level Training with Recurrent Neural Networks
- BERTScore: Evaluating Text Generation with BERT
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models