Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence
Jakub Pokrywka, Łukasz Grzybowski, Antoni Lasik, Marek Kubis, Jeremi Ignacy Kaczmarek, Wojciech Kusa
Adam Mickiewicz University · ARAAI Poland · NASK National Research Institute · Poznań University of Medical Sciences · T. Marciniak Lower Silesian Specialist Hospital
cs.AI
Submitted: 2026-08-13
Updated: 2026-08-14
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: The paper introduces a Polish-language medical visual question answering (VQA) benchmark built from Polish Board Certification Examination (PES) questions for licensed physicians and dentists
Terminology
Summary
The paper introduces a Polish-language medical visual question answering (VQA) benchmark built from Polish Board Certification Examination (PES) questions for licensed physicians and dentists pursuing specialist certification. The benchmark comprises 286 image-containing questions spanning diverse medical specialties and visual domains, together with a text-only question answering (QA) control set of 480 questions. The authors evaluate Polish-oriented, general-purpose open-weight, and commercial vision-language models. The task remains challenging: "the best model achieves 79.0% accuracy on the full VQA set, and only GPT-5.6 surpasses the approximate human reference on the subset with available candidate responses; all other evaluated models perform worse than humans."
To assess visual grounding, the authors compare complete inputs with configurations omitting the image, the question, or both, and categorize questions by image importance. They find that models derive more useful information from the question text than from the image and perform worse on image-dominant questions.
Across both QA and VQA, models achieve above-chance accuracy from the answer choices alone, showing that non-trivial performance can persist even when key task components are missing.
The dataset is constructed from PES examination materials collected from the Medical Examination Center website, covering examination sessions from 2023 to 2026. A total of 363 examination sheets were processed, with 116 sheets containing visual material. The VQA subset consists of 286 image-containing questions, each containing question text, five answer choices, the correct answer, metadata, and the associated image. The QA control subset consists of 480 text-only questions sampled from specialties with at least 10 image-containing questions. Emergency medicine contributes the largest number of VQA questions (62), followed by maxillofacial surgery (37), orthopedics (28), and pediatric cardiology (28).
The evaluation methodology uses four input configurations for image-containing questions: answer choices only (C), answer choices and question text (C+Q), answer choices and image (C+I), and the full input (C+Q+I). For text-only QA questions, only C and C+Q are applicable. All experiments use Polish prompts, and the prompt instructs models to return only a JSON object with the selected answer. Accuracy is reported as the percentage of questions where the model's predicted answer matches the official answer key, with random guessing corresponding to 20% accuracy.
The models evaluated include three Polish-oriented VLMs (LLaVA-Bielik-11b-v2.6-instruct, LLaVA-PLLuM-12b-nc-instruct-250715, and LLaVA-PLLuM-12b-nc-instruct), three Qwen models (Qwen3.5-397B-A17B, Qwen3.5-9B, and Qwen3.6-27B), Gemma-4-31B-it, and two commercial models (GPT-5.4-nano-2026-03-17 and GPT-5.6-sol). The GPT models were evaluated with both no reasoning and medium reasoning effort.
Results show that on the QA subset, GPT-5.6-sol with reasoning achieves 86.70% accuracy, while human examinees achieve 69.75%. On the VQA subset, GPT-5.6-sol achieves 77.78% accuracy with reasoning and 76.59% without, while human examinees achieve 70.14%. Among open-weight models, Qwen3.5-397B-A17B achieves the highest accuracy on both QA (68.88%) and VQA (63.10%). The LLaVA-based models perform notably worse, with LLaVA-Bielik-11B achieving 43.35% on QA and 39.68% on VQA.
The input-ablation experiments reveal that for VQA, removing the image (C+Q) is less detrimental than removing the question (C+I), which shows that the models make greater use of the question text.
This pattern holds for every model except LLaVA-PLLuM. In the choices-only configuration (C), models perform above the random-guessing baseline of 20%, with GPT-5.6-sol selecting the correct answer in 42.9% of cases without access to the question.
The image-importance categorization assigns each VQA question to one of three categories: image non-essential (5.2% of questions), image and text complementary (45.8%), and image dominant (49.0%). Every evaluated model performs worse on image-dominant questions under both C+Q and C+Q+I configurations. For example, GPT-5.6-sol without reasoning achieves 82.19% accuracy on non-image-dominant questions but only 42.86% on image-dominant questions under C+Q, and 83.56% versus 72.14% under C+Q+I.
The visual domain categorization includes IMAGE (n=102), WAVEFORM (n=95), PLOT (n=23), GRAPHIC (n=32), TABLE (n=20), COMPOSITE (n=14), and OTHER (n=3). WAVEFORM tends to be among the best-performing well-represented domains.
A data contamination analysis using the Data Contamination Quiz framework suggests negligible contamination, with contamination level ranges reported for representative models from each family. For example, LLaVA-Bielik-11B has a range of (16.86, 27.08) on QA and (15.38, 29.02) on VQA, while GPT-5.4-nano has (17.41, 22.92) on QA and (7.69, 18.53) on VQA.
The paper concludes that human examinees achieve nearly identical accuracy on the two subsets, indicating that the tasks are comparable in difficulty,
but most evaluated models perform worse on VQA.
The findings suggest that current models rely more heavily on textual cues than on visual evidence when answering Polish medical examination questions.
The authors also note that high accuracy does not necessarily reflect robust medical or multimodal competence when a task remains partially solvable from incomplete inputs,
which is especially important in medical settings.
Limitations include the selected set of models evaluated, the lack of systematic testing of all reasoning-effort levels and decoding configurations, the exclusion of vision-language models specifically trained for medicine, and the limited dataset size and specialty coverage. The benchmark reflects the structure of Polish board certification examinations rather than the full range of clinical practice.
Improvements for AI systems
Improvements to AI systems:
-
Implement explicit visual-grounding verification modules. The system should be trained to detect when an answer depends critically on image content versus text alone. It can then either (a) abstain from answering image-dominant questions when visual understanding is uncertain, or (b) explicitly cross-check its textual reasoning against visual features before committing to an answer.
-
Add input-ablation training as a regularizer. During fine-tuning, randomly drop the image, the question, or both, and train the model to recognize when it cannot answer reliably without complete inputs. This reduces over-reliance on textual priors and answer-choice heuristics, forcing the model to genuinely integrate multimodal evidence.
-
Introduce a
choice-only leakage
penalty. Since models achieve above-chance accuracy (up to 42.9%) from answer choices alone, the system should be trained with a loss term that penalizes correct answers when the image and question are both ablated. This discourages shortcut learning from answer distribution patterns and improves robustness in high-stakes medical settings. -
Develop a domain-aware confidence calibration mechanism. The system should output calibrated confidence scores separately for text-dominant and image-dominant questions. In practice, it should down-weight its own predictions on image-dominant categories (e.g., radiology images, waveforms) unless its visual encoder shows strong activation patterns, and flag low-confidence cases for human review.
-
Create a bilingual medical VQA adapter. The system can be fine-tuned on Polish medical exam data with a contrastive objective that aligns visual features with Polish medical terminology. This improves performance on under-resourced language-medical pairs and can be extended to other languages with similar board exam structures.
-
Add a
reasoning-effort-aware
inference mode. The system should dynamically allocate more computational steps (e.g., chain-of-thought, multi-pass verification) for image-dominant questions, while using faster inference for text-dominant ones. This optimizes accuracy where it matters most without sacrificing speed on easier items. -
Build a contamination-resistant evaluation harness. The system should include a built-in data-contamination detector (like the Data Contamination Quiz) that flags if training data overlaps with exam-style benchmarks. This ensures that reported accuracy reflects genuine generalization rather than memorization, and the system can self-report reliability scores for deployment.
What the improved AI system can do:
-
Achieve higher accuracy on image-dominant medical questions (e.g., interpreting ECGs, radiology images, pathology slides) by explicitly verifying visual evidence before answering.
-
Refuse or flag answers when it detects that the question cannot be solved from available inputs, reducing silent failures in clinical decision support.
-
Provide calibrated confidence that is lower for image-dominant questions, enabling safer human-AI collaboration in medical diagnostics.
-
Maintain robust performance even when images are missing or corrupted, by recognizing when to fall back on text-only reasoning without overconfident guesses.
-
Generalize better to new medical exams and languages by learning to separate genuine multimodal understanding from answer-choice artifacts.
-
Operate efficiently in real-time clinical settings by allocating reasoning effort based on the estimated visual dependency of each question.
Abstract
We introduce a Polish-language medical visual question answering (VQA) benchmark, built from Polish Board Certification Examination questions for licensed physicians and dentists pursuing specialist certification. The benchmark comprises image-containing questions spanning diverse medical specialties and visual domains, together with a text-only question answering (QA) control set. We evaluate Polish-oriented, general-purpose open-weight, and commercial vision-language models. The task remains challenging: the best model achieves 79.0% accuracy on the full VQA set, and only GPT-5.6 surpasses the approximate human reference on the subset with available candidate responses; all other evaluated models perform worse than humans. To assess visual grounding, we compare complete inputs with configurations omitting the image, the question, or both, and categorize questions by image importance. Models derive more useful information from the question text than from the image and perform worse on image-dominant questions. Across both QA and VQA, they nevertheless achieve above-chance accuracy from the answer choices alone, showing that non-trivial performance can persist even when key task components are missing.
Sources
- LLMzSz{\L}: a comprehensive LLM benchmark for Polish
- Reassessing High-Performing LLMs on Polish Medical Exams: True Competence or Bias-Driven Performance?
- Multilingual Hematology Visual Question Answering Dataset
- GPT-4 passes most of the 297 written Polish Board Certification Examinations
- PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
- Does Language Shift Break Medical Vision-Language Models? Indonesian Radiology Visual Question Answering Case Study
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection