Visual Jev: Accurate and Efficient Decisions from Shared Visual Context
cs.CV, cs.LG
Submitted: 2026-09-22
Updated: 2026-09-22
Comments: Code: https://github.com/guanxuyu-sv/Visual-Jev
Code: https://github.com/guanxuyu-sv/Visual-Jev
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Many vision applications ask several independent, forced-choice questions about the same image.
Terminology
Abstract
Many vision applications ask several independent, forced-choice questions about the same image. Visual Jev encodes the image and public context once, executes isolated question suffixes as a batch, and reads candidate probabilities from the backbone's language-model head. Across four benchmarks, answer-supervised post-training raises equal-weight macro accuracy from 70.6% to 76.1%, with the gain concentrated on the two task families represented in training. At N=32 questions per image, shared batched execution is 8.9x faster in warm amortized time than independent serial execution and remains 3.4x faster than an already-batched baseline that recomputes the prefix, at the cost of higher peak memory. A matched typed-head control offers no consistent accuracy advantage over the language-model-head readout. The supported design is therefore simple: adapt the backbone for quality, retain the existing readout, and share execution for efficiency.
Sources
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- e-SNLI-VE: Corrected Visual-Textual Entailment with Natural Language Explanations
- Language Models (Mostly) Know What They Know
- DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Visual Entailment: A Novel Task for Fine-Grained Image Understanding
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models