Fusing Visual and Textual Representations via Multi-layer Fusing Transformers for Vietnamese Visual Question Answering
cs.CV
Submitted: 2026-10-01
Updated: 2026-10-01
Terminology
Sources
- SQuAD: 100,000+ Questions for Machine Comprehension of Text
- BEiT: BERT Pre-Training of Image Transformers
- EVJVQA Challenge: Multilingual Visual Question Answering
- BARTpho: Pre-trained Sequence-to-Sequence Models for Vietnamese
- What's so special about BERT's layers? A closer look at the NLP pipeline in monolingual and multilingual models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models