From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding
cs.CV, cs.AI, cs.CL
Submitted: 2026-09-01
Updated: 2026-09-01
Code: https://github.com/expertailab/scigram
Terminology
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Pixtral 12B
- PaLM: Scaling Language Modeling with Pathways
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- The Llama 3 Herd of Models
- SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
- LoRA: Low-Rank Adaptation of Large Language Models
- VisualBERT: A Simple and Performant Baseline for Vision and Language
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning
- KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts Reasoning
- GPT-4 Technical Report
- Bidirectional Attention Flow for Machine Comprehension
- VL-BERT: Pre-training of Generic Visual-Linguistic Representations
- Gemini: A Family of Highly Capable Multimodal Models
- Memory Networks
- Qwen2 Technical Report
- Deep Modular Co-Attention Networks for Visual Question Answering
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models