Polish-English medical knowledge transfer: A new benchmark and results
cs.CL, cs.AI
Submitted: 2024-11-30
Updated: 2025-09-12
Journal ref: Findings of the Association for Computational Linguistics: EMNLP 2025, pp. 9042-9063, 2025
DOI: 10.18653/v1/2025.findings-emnlp.480
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large Language Models (LLMs) have demonstrated significant potential in handling specialized tasks, including medical problem-solving.
Terminology
Abstract
Large Language Models (LLMs) have demonstrated significant potential in handling specialized tasks, including medical problem-solving. However, most studies predominantly focus on English-language contexts. This study introduces a novel benchmark dataset based on Polish medical licensing and specialization exams (LEK, LDEK, PES) taken by medical doctor candidates and practicing doctors pursuing specialization. The dataset was web-scraped from publicly available resources provided by the Medical Examination Center and the Chief Medical Chamber. It comprises over 24,000 exam questions, including a subset of parallel Polish-English corpora, where the English portion was professionally translated by the examination center for foreign candidates. By creating a structured benchmark from these existing exam questions, we systematically evaluate state-of-the-art LLMs, including general-purpose, domain-specific, and Polish-specific models, and compare their performance against human medical students. Our analysis reveals that while models like GPT-4o achieve near-human performance, significant challenges persist in cross-lingual translation and domain-specific understanding. These findings underscore disparities in model performance across languages and medical specialties, highlighting the limitations and ethical considerations of deploying LLMs in clinical practice.
Sources
- GPT-4 Technical Report
- PaLM 2 Technical Report
- Fine-Tuned 'Small' LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification
- The Llama 3 Herd of Models
- Gemini: A Family of Highly Capable Multimodal Models
- Measuring Massive Multitask Language Understanding
- Mistral 7B
- PubMedQA: A Dataset for Biomedical Research Question Answering
- Large Language Models: A Survey
- Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering