jina-vlm: Small Multilingual Vision Language Model
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "jina-vlm: Small Multilingual Vision Language Model".
Jane: The paper was written by Andreas Koukounas, Georgios Mastrapas, Florian Hönicke, Sedigheh Eslami, Guillaume Roncari et al. from Jina AI by Elastic.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back, everyone. Today we’re looking at a paper that’s been making waves in the small-model community, and it’s called “JINA-VLM: Small Multilingual Vision Language Model.” Jane, what caught your eye first about this one?
Jane: Oh, Tom, the title alone tells you a lot. It’s small, it’s multilingual, and it’s a vision-language model. That’s three big promises in one package, and usually you have to pick two. The fact that they’re trying to do all three at just two point four billion parameters is a bold move.
Tom: And it’s from Jina AI, which is a name we’ve seen a lot in the embedding and retrieval space. They’re not exactly the first company you’d expect to drop a VLM, but here they are, and they’re claiming state-of-the-art results among open 2B-scale models.
Jane: Right, and that’s the key word—open. They’ve released the weights and the code, so anyone can actually use this thing. That’s huge for researchers who don’t have access to massive GPU clusters.
Lu: I’d add that the “multilingual” part is what really separates this from the pack. Most small VLMs are English-first, and they degrade badly when you switch languages. This paper explicitly tackles that from the start, which is rare at this scale.
Tom: So Lu, you’re saying the title isn’t just marketing? They actually built the multilingual capability into the architecture and training, not as an afterthought?
Lu: Exactly. They paired a SigLIP2 vision encoder with a Qwen3 language decoder, and they trained on data spanning seventeen languages, from Arabic to Vietnamese. That’s not a token effort; that’s a commitment.
Meng: And from an engineering standpoint, the fact that they kept it at 2 point 4B parameters while doing all this is impressive. You can run this on a single consumer GPU, maybe even quantized on a laptop. That’s the kind of accessibility that could really democratize multimodal AI.
Jane: I love that point, Meng. It’s not just about the benchmark scores; it’s about who gets to use the model. If you’re a student in a developing country with a mid-range GPU, you can now experiment with a vision-language model that speaks your language. That changes the game.
Tom: And that’s the hook for the rest of our discussion. We’re going to dig into how they actually trained this thing, what the results look like, and what it means for the field. Stick around.
Summary: Tom: So we’ve set the stage with the title. Now let’s talk about what the paper actually does. Jane, give us the elevator pitch.
Jane: Sure. “JINA-VLM” is a vision-language model that takes an image, tiles it into overlapping pieces, processes each tile through a vision encoder, and then uses a clever pooling mechanism to compress the visual tokens before feeding them to the language model. That compression is the secret sauce—it cuts the token count by four times, which makes the whole thing efficient enough to run at 2 point 4B parameters.
Tom: And that efficiency isn’t just a nice-to-have. The paper reports a three point nine times reduction in prefill FLOPs and a four times reduction in KV-cache memory. Meng, that’s the kind of stuff that makes your eyes light up, right?
Meng: Absolutely. That means you can process higher-resolution images without blowing up your memory budget. They use up to twelve overlapping tiles plus a global thumbnail, and even with all that, the language model only sees about two thousand three hundred sixty-six tokens per image. Without the pooling, it would be over nine thousand. That’s the difference between fitting on a single GPU and not.
Lu: And the results back up the design. On eight standard English VQA benchmarks, they hit an average of seventy-two point three, which beats Qwen2-VL-2B and InternVL3-2B. But the real standout is multilingual performance. On MMMB, they score seventy-eight point eight on average across six languages, which is the best among 2B-scale models.
Jane: That’s the part I find most exciting, Lu. They didn’t just add some multilingual data as an afterthought. They built a two-stage training pipeline where the first stage focuses on cross-language semantic grounding, and the second stage does instruction fine-tuning with a heavy multilingual component. It’s baked in.
Tom: And they also kept text-only data in the mix, which is something a lot of VLM papers skip. That’s why they still perform reasonably on pure language benchmarks like MMLU and GSM-8K, even though there’s some degradation compared to the backbone model.
Meng: The degradation is real, though. MMLU-Pro drops from forty-six point four to thirty point three, which is a big hit. But for a VLM, that’s a trade-off most people are willing to make, especially when the visual performance is this strong.
Lu: And they’re honest about it. The paper acknowledges that the instruction tuning toward concise visual responses hurts multi-step reasoning. That transparency is refreshing.
Tom: So we’ve got a model that’s efficient, multilingual, and competitive on benchmarks. But what’s the bigger picture? That’s what we’ll dig into next.
Improvements: Tom: We’ve covered what the model does. Now let’s talk about what this paper suggests for the field. Jane, what do you think is the biggest improvement they’re proposing?
Jane: I’d say it’s the data mixture analysis. They didn’t just throw a bunch of data at the model and hope for the best. They ran a leave-one-out ablation study, where they removed entire categories of data—VQA, OCR, charts, math, documents, code, text-only, multilingual—and measured the impact. That’s the kind of systematic thinking we don’t see enough of.
Lu: And the findings are genuinely surprising. For example, removing real-world photo data actually improved performance on several unrelated benchmarks. HellaSwag went up by nine point one points, and MMMU went up by seven point three. That suggests those categories were introducing noise or gradient conflicts at this scale.
Meng: That’s a wild result. It means the data mixture isn’t just about adding more; it’s about curating what’s actually useful. For practitioners with limited compute, that’s gold. You don’t have to train for one thousand GPU hours to learn that you should drop certain data categories.
Tom: And they also found that task-specific data, like VQA and OCR, only benefits its own task. It doesn’t transfer broadly. So if you want a model that’s good at charts, you need chart data. There’s no shortcut.
Jane: Right, but the flip side is that text-only and multilingual data impose a cost on multimodal performance. If you don’t need those capabilities, you can drop them and get a boost. That’s a really practical insight for people building specialized models.
Lu: I’d push back slightly on the generalizability, though. These ablations ran at only ten percent of the full training duration. So the effects might not hold at full scale. The paper is careful to call these “diagnostic observations” rather than definitive conclusions.
Meng: Still, it’s a starting point. And it opens up a whole research direction. We need more studies like this, where we systematically understand what data does what, rather than just scaling up and hoping.
Tom: And that’s the real contribution here. It’s not just a model; it’s a methodology for thinking about data in small-model training. That could have a bigger impact than the model itself.
Jane: Absolutely. And it’s a reminder that in the era of massive datasets, sometimes the smartest move is to do less, but better.
Conclusion: Tom: We’ve covered a lot of ground on “JINA-VLM: Small Multilingual Vision Language Model.” Let’s wrap this up. Jane, what’s the one thing you want listeners to remember?
Jane: I want them to remember that small models can be powerful, multilingual, and efficient all at once. This paper shows that with the right architecture and the right data, you don’t need seventy billion parameters to get competitive results. That’s a huge deal for accessibility.
Lu: And I’d add that the data mixture analysis is a blueprint for future work. We need more of this kind of systematic ablation, especially at smaller scales where every compute hour counts.
Meng: From my side, the efficiency gains are the headline. Four times fewer tokens, three point nine times fewer FLOPs, and it all fits on a single GPU. That’s the kind of engineering that makes deployment practical.
Tom: And we can’t forget the multilingual piece. In a world where most VLMs are English-centric, this model speaks seventeen languages and does it well. That’s not just a technical achievement; it’s a step toward more inclusive AI.
Jane: Exactly, Tom. And the authors were transparent about the limitations—multi-image reasoning is weak, and there’s some text-only degradation. But they’ve released the weights, so the community can build on this and push further.
Tom: So we say goodbye to “JINA-VLM” with a sense of optimism. It’s a reminder that progress isn’t always about bigger; sometimes it’s about smarter. Thanks for listening, and we’ll see you on the next paper.
Andreas Koukounas, Georgios Mastrapas, Florian Hönicke, Sedigheh Eslami, Guillaume Roncari, Han Xiao
Jina AI by Elastic
cs.CL, cs.AI, cs.CV
Submitted: 2026-08-17
Updated: 2026-08-18
Comments: 23 pages, 1-10 main content, 11-23 references and appendix
Code: https://github.com/open-compass/VLMEvalKit
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 63/100
The gist: systematically removing task, domain, modality, and language categories—to diagnose which data types are necessary versus redundant and whether task benefits transfer across domains.
Key concepts
- Vision-Language Model (VLM)
- A type of AI model that can process and understand both images and natural language text simultaneously. JINA-VLM is designed to take an image input alongside text prompts to generate relevant responses.
- Multilingual Capability
- The ability of the model to function effectively across multiple languages, rather than being limited primarily to English. This paper explicitly addresses this by training on data spanning seventeen different languages.
- 2.4 Billion Parameters
- Refers to the size and complexity of the AI model. Keeping a VLM small (at 2.4B parameters) allows it to be highly efficient, running on consumer hardware while maintaining strong performance.
- Data Mixture Analysis
- A systematic study where researchers test the impact of removing entire categories of training data (like OCR or charts). This helps determine which types of data are most useful and prevent noise in model training.
Terminology
Summary
Summary
The paper presents jina-vlm, a token-efficient 2.4B parameter vision-language model (VLM) that achieves state-of-the-art multilingual VQA performance among open 2B-scale VLMs. The model couples a SigLIP2 vision encoder with a Qwen3 language decoder and makes use of image tiling and attention-pooling for token-efficient processing of arbitrary-resolution images. To understand the contribution of different training data categories, the authors conduct a leave-one-out data mixture ablation study—systematically removing task, domain, modality, and language categories—to diagnose which data types are necessary versus redundant and whether task benefits transfer across domains. Model weights and code are publicly released at https://huggingface.co/jinaai/jina-vlm.
Introduction states that two challenges limit practical deployment of VLMs: multilingual capabilities often degrade during vision adaptation, and high-quality VLMs remain computationally expensive to train and deploy. jina-vlm addresses both challenges by aligning a SigLIP2-So400M/14-384 vision encoder with Qwen3-1.7B-Base through an attention-pooling connector, trained with a two-stage pipeline that explicitly incorporates multilingual data. Among open 2B-scale VLMs, jina-vlm achieves state-of-the-art multilingual performance on MMMB and Multilingual MMBench. On standard English benchmarks spanning diagrams, charts, documents, and OCR, jina-vlm achieves the highest average score (72.3) across eight VQA benchmarks among 2B-scale VLMs. Token efficiency is achieved via an arbitrary-resolution pipeline that combines overlapping tiling with attention-based token pooling to reduce visual token count by 4×.
Related Work covers VLM architecture and training (PaLI, LLaVA, QwenVL, InternVL, Ovis), efficient resolution-agnostic image processing (tiling, Naive Dynamic Resolution, native-resolution ViTs, token compression methods like HERO), vision-language connectors (Q-Former, Perceiver Resampler, Dense Connector), small VLMs (MiniCPM-V, SmolVLM), multilinguality and text-only performance, and data mixture analysis (MM1, Cambrian-1, Molmo/PixMo, Eagle 2). The authors note their work is complementary to prior data mixture studies: rather than optimizing data proportions or sample weights, they perform leave-one-out ablations that remove entire data categories to measure their marginal contribution and cross-domain transfer effects.
Model Architecture (Figure 1): The model uses overlapping image tiling following Deitke et al. (2025), combined with attention-based token pooling. The vision encoder, SigLIP2-So400M/14-384, is a 27-layer Vision Transformer with 400M parameters that processes 378×378 pixel inputs as 27×27 grids of 14×14 patches. Images are decomposed into overlapping tiles of this size, with a default of 12 tiles during training; a global thumbnail (full image resized to 378×378) provides context. Adjacent tiles overlap by 112 pixels with a stride of 266 pixels between tile origins; a 4×3 grid spans 1176×910 pixels.
The vision-language connector concatenates features from two intermediate layers (third-to-last and ninth-to-last, i.e., layers 24 and 18), following findings that intermediate ViT layers retain spatial detail. The connector applies attention pooling over 2×2 patch neighborhoods, using mean-pooled features as queries, reducing token count by 4× while preserving local structure. A SwiGLU projection layer maps pooled representations to the language model's embedding dimension. Formally, Hconcat = [H(−3); H(−9)] ∈ R N×2dv; queries are computed as the mean of each 2×2 neighborhood; attention pooling is computed as Hpooled = (softmax(QWQ(Hconcat WK)⊤/√dk) Hconcat WV) WO; finally, Hproj = (Swish(Hpooled W1) ⊙ (Hpooled W2)) W3.
The language decoder is initialized from Qwen3-1.7B-Base, which empirically outperformed the instruction-tuned variant. Three special tokens structure visual inputs: and delimit image and thumbnail sequences, while marks row boundaries within the patch grid. Input and output embedding weights are not tied.
Efficiency Analysis (Table 1): With the default 12-tile configuration (plus thumbnail), the unpooled baseline would produce 9,477 visual tokens per image, while 2×2 pooling reduces this to 2,366 tokens. This yields a 3.9× reduction in prefill FLOPs and a 4× reduction in KV-cache memory; overall FLOPs reduction is 2.3× when including the shared ViT cost.
Training proceeds in two stages, with both stages updating all model components (encoder, connector, decoder) without freezing. The combined data comprises samples in multiple languages including English, Chinese, Arabic, German, Spanish, French, Italian, Japanese, Korean, Portuguese, Russian, Turkish, Vietnamese, Thai, Indonesian, Hindi, and Bengali.
Stage 1 (Alignment Training) focuses on cross-language semantic grounding using caption datasets (PixmoCap, PangeaIns) spanning diverse visual domains, plus 15% text-only data from PleiAS/common corpus to mitigate degradation on text-only tasks. The connector uses a higher learning rate and shorter warmup than encoder and decoder.
Stage 2 (Instruction Fine-Tuning) trains instruction-following for VQA and reasoning tasks, combining public dataset collections including LLaVA OneVision, Cauldron, Cambrian, PangeaIns, and FineVision, with text-only instruction data from Singh et al. (2024). The mixture covers academic VQA, document understanding, OCR, mathematics, and reasoning. Given data diversity, single-source batches were found more effective initially; training proceeds for 30K steps with single-source batches, then 30K steps with mixed-source batches.
Hyperparameters (Table 2): Pre-training uses LR ViT 6e-6, LR Connector 2e-4, LR LLM 2e-5, batch size 128, 25K steps, 3.2M samples, 10B tokens, 296 GPU hours. Fine-tuning uses LR ViT 5e-6, LR Connector 5e-6, LR LLM 1e-5, batch size 256, 60K steps, 15.3M samples, 37B tokens, 1,000 GPU hours.
Evaluation compares jina-vlm against lightweight VLMs across six capability areas, using VLMEvalKit with English prompts matching the training format.
General VQA (Table 3): jina-vlm achieves the highest average (72.3) across AI2D (82.0), ChartQA (81.9), TextVQA (83.2), DocVQA (90.6), InfoVQA (71.6), OCRBench (778), SEED-2 Plus (67.2), and CharXiv (32.3/63.5), outperforming Qwen2-VL-2B (66.4), Qwen3-VL-2B (71.6), InternVL3-2B (69.2), and InternVL3.5-2B (71.6).
Multimodal and Real-World Understanding (Table 4): jina-vlm scores 67.4 on multimodal tasks (MME 1965.8, MMB v1.1 75.8, MMStar 56.2) and 61.9 on real-world tasks (RealWorldQA 68.2, MME-RW 50.7, R-Bench 66.7), achieving the best RealWorldQA result.
Multi-Image Reasoning and Hallucination (Table 5): jina-vlm scores 47.3 on multi-image tasks (BLINK 50.1, MuirBench 34.7, MMT 57.2), which is expected given limited multi-image training data, but achieves the best POPE score (90.3), indicating low hallucination rates.
Mathematical Reasoning (Table 6): jina-vlm performs comparably to InternVL3-2B and outperforms Qwen2-VL-2B, with scores of MMMU 45.6, MathVista 59.5, MathVision 19.2, MathVerse 23.9, WeMath 17.1, LogicVista 33.3, overall 33.1.
Text-Only Performance (Table 7): jina-vlm matches or exceeds the backbone Qwen3-1.7B on commonsense reasoning (ARC-C 77.3 vs 73.4, HellaSwag 59.4 vs 59.0) and retains most performance on MMLU (56.1 vs 62.6) and GSM-8K (71.3 vs 75.3). However, MMLU-Pro shows substantial degradation (30.3 vs 46.4), likely because this benchmark emphasizes extended multi-step reasoning that conflicts with instruction-tuning toward concise visual responses.
Multilingual Understanding (Table 8): jina-vlm achieves state-of-the-art multilingual performance among 2B-scale VLMs, with the highest averages on MMMB (78.8) and Multilingual MMBench (74.3), and overall 59.6, outperforming Qwen2-VL-2B (53.8), Qwen3-VL-2B (58.2), InternVL3-2B (57.4), and InternVL3.5-2B (58.0). The evaluation setup uses English system prompts while questions, answer options, and text within images are in the target language.
Data Mixture Analysis (Section 6): The authors conduct a comprehensive ablation study that systematically removes different data types from the training mixture. Starting from the original training mixture, 9 experiments are designed, each removing one category: Xvqa (VQA data), Xocr (OCR data), Xrealworld (real-world photos), Xmath (math/reasoning), Xcharts (charts/diagrams), Xdocs (documents), Xcode (code), Xtext (text-only data), Xmlingual (non-English data). All ablations run for 10% of the full training duration, a known limitation; findings should be treated as diagnostic observations rather than definitive conclusions.
Results (Table 9): Removing VQA data (Xvqa) causes dramatic but highly task-specific degradation on VQA benchmarks (AI2D −9.5, ChartQA −22.1, DocVQA −8.0, InfoVQA −22.8, TextVQA −18.6) with minimal impact on other tasks. Xocr harms OCRBench (−11.0) while improving all other tasks. Domain removals show predictable patterns for Xcharts (ChartQA −39.4%) and Xmath (MathVista −4.6%). Removing real-world (Xrealworld) or document data (Xdocs) improves several unrelated benchmarks (HellaSwag +9.1%, MMMU +7.3%), while code removal (Xcode) degrades HellaSwag (−10.6%). Removing text-only data (Xtext) consistently improves multimodal benchmarks but hurts HellaSwag (−7.3), the only unimodal task. Removing multilingual data (Xmlingual) improves most tasks but drops MTVQA by −16.1%.
The authors draw three conclusions for practitioners under fixed compute budgets: (1) task-specific data (VQA, OCR, charts, math) is necessary for its capabilities but does not transfer broadly; (2) real-world and document data exhibit unexpected interference, as their removal improves several unrelated benchmarks, suggesting noise or gradient conflict at this scale; (3) text-only and multilingual data impose a cost on multimodal performance but are essential for their respective capabilities.
Conclusion states that jina-vlm achieves state-of-the-art multilingual VQA and leading English VQA results among open 2B-scale VLMs. Limitations include: multi-tile processing introduces computational overhead scaling with image resolution; tiling can fragment global spatial context, potentially impairing performance on tasks requiring holistic scene understanding such as object counting or precise spatial reasoning across tile boundaries; safety-critical training or alignment is not emphasized; multi-image reasoning remains weak due to limited training data. The data mixture analysis offers diagnostic observations for resource-constrained training; as ablations ran at 10% of full training, findings should be interpreted as preliminary evidence motivating future systematic investigation.
Improvements for AI systems
Based on the scientific paper, here are specific improvements I can implement in an AI system, along with the resulting capabilities.
-
Improvement: I will replace the standard fixed-resolution image preprocessing with a dynamic tiling pipeline. The system will resize input images to fit a grid of up to 12 overlapping 378×378 tiles (plus a global thumbnail) to preserve fine-grained details (e.g., text, charts) without discarding information. To manage computational cost, I will integrate a 2×2 attention-pooling connector that concatenates features from the 3rd-to-last and 9th-to-last layers of a SigLIP2 vision encoder, reducing visual token count by 4× (from 9,477 to 2,366 tokens for a 12-tile image).
-
Capability: The AI system can now process images of arbitrary resolution and aspect ratio (e.g., long documents, wide screenshots, high-res photos) with significantly lower latency and memory usage (3.9× reduction in prefill FLOPs and 4× reduction in KV-cache memory) while maintaining high accuracy on OCR, document understanding, and chart analysis.
-
Improvement: I will adopt a two-stage training recipe: (1) alignment training on captions (e.g., PixmoCap, PangeaIns) mixed with 15% text-only data to preserve language capabilities, and (2) instruction fine-tuning on a diverse mixture of VQA, OCR, reasoning, and multilingual datasets (covering 17+ languages). I will use a higher learning rate for the connector (2e-4) than the encoder/decoder (6e-6/2e-5) in stage 1, and update all components without freezing.
-
Capability: The AI system will achieve state-of-the-art multilingual VQA performance (e.g., 78.8% on MMMB, 74.3% on Multilingual MMBench) and strong English VQA (72.3% average across 8 benchmarks) among 2B-scale models, while retaining text-only capabilities (e.g., 77.3% on ARC-C, 59.4% on HellaSwag) better than naive fine-tuning.
-
Improvement: Based on the leave-one-out ablation study, I will implement a data curation strategy that prioritizes task-specific data (VQA, OCR, charts, math) for their respective benchmarks, but avoids over-weighting real-world and document data when training for general reasoning, as their removal improved unrelated tasks (e.g., +9.1% HellaSwag, +7.3% MMMU). I will also include text-only and multilingual data only if those capabilities are required, as they impose a measurable cost on multimodal performance (e.g., removing text-only data improved MMBench by +1.1%, RealWorldQA by +6.3%).
-
Capability: The AI system will be trained more efficiently under fixed compute budgets, achieving higher performance on target benchmarks by avoiding data that causes gradient conflicts or noise, and by allocating more training steps to high-value categories (e.g., VQA, OCR) that show strong task-specific gains.
-
Improvement: I will incorporate the model’s architectural choices (attention-pooling, intermediate layer concatenation) and training data (including multi-image reasoning data from LLaVA-OneVision) to improve multi-image understanding. I will also use the observed low hallucination rate (90.3% on POPE) as a baseline and add explicit hallucination-focused instruction data (e.g., HallBench) during fine-tuning to further reduce fabrication.
-
Capability: The AI system will handle tasks requiring comparison across multiple images (e.g., BLINK, MuirBench) with improved accuracy (targeting >50% on BLINK) and demonstrate lower rates of object hallucination, making it more reliable for real-world applications like visual verification and document comparison.
-
Improvement: I will modify the instruction fine-tuning to include a higher proportion of text-only reasoning data (e.g., from Aya dataset) and use a lower learning rate for the language decoder (1e-5) to mitigate degradation on complex text benchmarks. I will also add a small amount of multi-step reasoning data (e.g., from MMLU-Pro) to counteract the observed 16.1% drop on that benchmark.
-
Capability: The AI system will maintain strong text-only performance (targeting >55% on MMLU, >70% on GSM-8K) while excelling at multimodal tasks, making it a versatile single model for both vision-language and pure language applications without needing a separate text-only model.
-
Improvement: I will set up an automated evaluation pipeline using VLMEvalKit that runs the full benchmark suite (8 VQA, 3 multimodal, 3 multi-image, 6 math, 5 text-only, and 3 multilingual benchmarks) after every training checkpoint. This will allow me to detect performance regressions early, especially on multilingual and text-only tasks, and adjust the data mixture or hyperparameters accordingly.
-
Capability: The AI system will have a built-in quality assurance mechanism that ensures stable performance across diverse capabilities, preventing silent degradation during continued training or fine-tuning, and providing actionable insights for data curation decisions.
Abstract
We present jina-vlm, a token-efficient 2.4B parameter vision-language model that achieves state-of-the-art multilingual VQA performance among open 2B-scale VLMs. The model couples a SigLIP2 vision encoder with a Qwen3 language decoder and makes use of image tiling and attention-pooling for token-efficient processing of arbitrary-resolution images. To understand the contribution of different training data categories, we conduct a leave-one-out data mixture ablation study-systematically removing task, domain, modality, and language categories-to diagnose which data types are necessary versus redundant and whether task benefits transfer across domains. Model weights and code are publicly released at https://huggingface.co/jinaai/jina-vlm.
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Qwen2.5-VL Technical Report
- PaliGemma: A versatile 3B VLM for transfer
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- PathVQA: 30000+ Questions for Medical Visual Question Answering
- Measuring Massive Multitask Language Understanding
- Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training
- LLaVA-OneVision: Easy Visual Task Transfer
- SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
- R-Bench: Are your Large Multimodal Model Robust to Real-world Corruptions?
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
- HERO: Rethinking Visual Token Early Dropping in High-Resolution Large Vision-Language Models
- Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
- Ovis: Structural Embedding Alignment for Multimodal Large Language Model
- Ovis2.5 Technical Report
- Multilingual Vision-Language Models, A Survey
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering