PragyaDoc: A Universal Document Intelligence Framework for Multilingual Medical Document Understanding in Low-Resource Settings

arXiv:2608.07478 · cs.CV, cs.CL · Submitted 2026-06-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "PragyaDoc: A Universal Document Intelligence Framework for Multilingual Medical Document Understanding in Low-Resource Settings".

Jane: The paper was written by Jagpal Singh Jhala from VIT Bhopal University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the arXiv Review Channel, everyone! I'm Tom, and as always, I'm joined by my brilliant co-host, Jane. Today we're looking at a paper that really grabbed my attention: "PragyaDoc: A Universal Document Intelligence Framework for Multilingual Medical Document Understanding in Low-Resource Settings."

Jane: And Tom, this one is genuinely exciting. The core problem it tackles is something most of us never think about. In India, there are twenty-two official languages, but nearly all medical documents—prescriptions, discharge summaries, diagnostic reports—are written in English. That leaves hundreds of millions of people unable to read their own medical information.

Tom: That's the part that hit me. The paper opens with this story about a farmer in Madhya Pradesh who gets a discharge summary after a cardiac event. Five medicines, dietary restrictions, a follow-up schedule—all in English. His family can't understand any of it. The nearest doctor is forty kilometers away.

Jane: And that's not an edge case. That's daily reality. So the researchers built PragyaDoc, which is a four-layer pipeline. It takes a scanned medical document, runs it through three different OCR engines in parallel, fuses their outputs, structures the text deterministically, and then uses two different LLMs to extract medical entities and generate Hindi explanations.

Tom: The name itself is beautiful—"Pragya" means wisdom in Sanskrit, combined with "Doc" for documents. And the mission is to make medical documents understandable for every Indian. I love that the authors are so clear about the social motivation behind the technical work.

Jane: Right, and it's not just about translation. It's about cultural localization. The system generates explanations at a Grade five reading level, with safety guardrails. It even appends a mandatory disclaimer in Hindi telling patients to always follow their doctor's advice.

Tom: The paper reports ninety-two percent character accuracy and ninety percent medicine extraction accuracy on typed documents. That's a solid result. But I think the most interesting part is how they handle the messy reality of real-world documents—crumpled paper, faded ink, mixed handwriting.

Jane: Exactly. And that's what we'll dig into in the next segment. The technical architecture is genuinely clever, especially the fusion layer that combines three different OCR engines. Let's get into it.

Summary: Jane: So Tom, we're back with "PragyaDoc: A Universal Document Intelligence Framework for Multilingual Medical Document Understanding in Low-Resource Settings." Let's talk about how the system actually works, because the architecture is quite elegant.

Tom: Absolutely. So the first layer runs three OCR engines in parallel—PaddleOCR, EasyOCR, and DocTR. Each one has a different architecture and different strengths. PaddleOCR is tuned for faint ink strokes, EasyOCR handles mixed scripts, and DocTR is great on structured forms.

Jane: And here's the clever part. They run these in parallel using Python threads, not separate processes. That's because the models are already loaded in memory, and OCR inference releases the Global Interpreter Lock during the C++ backend execution. So you get true concurrency without the overhead of re-initializing models.

Tom: Right, and then the fusion layer is where the magic happens. All three engines produce bounding boxes around text regions. But here's the problem—different engines segment the same text differently. One might detect a full line as a single box, while another splits it into two.

Jane: And that's where they introduce this novel metric called Intersection over Minimum, or IoM. The standard metric, Intersection over Union, fails when one box is fully contained within another. Say one engine detects "SMS Hospital" as one box and another detects it as two boxes. The IoU would be low, so the system would treat them as separate detections. But IoM directly measures containment—if one box is inside another, IoM is high.

Tom: That's a genuinely clever fix. And it's backed by an empirical finding—they show that without IoM, a single institution name gets split into two separate text fragments, causing downstream parsing errors. With IoM, it merges correctly.

Jane: Then the fused text goes through a deterministic domain layer. This is important because it applies spatial heuristics before any LLM is involved. It detects document type—structured hospital prescriptions versus free-form clinic notes—and then parses sections using keyword anchoring and Y-axis sorting.

Tom: And the Reading Gravity algorithm is a nice touch. Lines starting with "TAB." or "CAP." are designated as medicine anchors, and all subsequent lines within a vertical tolerance get assigned to the nearest anchor above them. So dosage instructions and frequency notations get grouped with their medicine.

Jane: Only after all that deterministic structuring does the LLM get involved. And that's the key insight—by constraining the LLM's input space, you dramatically reduce hallucination. The first LLM extracts structured JSON with a Rescue Protocol that forces transparent reasoning about ambiguous entities. The second LLM generates the Hindi explanation.

Tom: And the results are solid. On typed documents, they achieve ninety-six point eight percent extraction rate. On handwritten documents, it drops to seven point two percent. That's a huge gap, and we should talk about that. But first, let's bring in Lu and Meng for their take on the architecture.

Lu: I want to build on what Jane said about the deterministic layer. That's the most underappreciated part of this paper. Most systems throw raw OCR output at an LLM and hope for the best. PragyaDoc instead reconstructs the document's logical hierarchy before any generative model runs. That's the difference between a glass box and a black box.

Meng: And from an engineering standpoint, the parallel execution design is smart. Using ThreadPoolExecutor instead of multiprocessing avoids the startup cost of re-initializing models. But I'm curious about the runtime—seventy-seven to one hundred thirty seconds on CPU. That's slow for a real clinic setting.

Tom: That's a fair point, Meng. The paper acknowledges this. On GPU it drops to eight–twelve seconds, but GPU availability in rural clinics is not guaranteed. That's a real deployment constraint.

Jane: And that handwritten gap—seven point two percent extraction rate—is the elephant in the room. We'll tackle that in the next segment, along with the improvements the authors suggest.

Improvements: Tom: So Jane, we're continuing with "PragyaDoc: A Universal Document Intelligence Framework for Multilingual Medical Document Understanding in Low-Resource Settings." And we have to address the big limitation—handwritten prescriptions.

Jane: Right. The paper is very honest about this. On pure handwritten documents, the extraction rate is seven point two percent. That's nearly zero. And the authors identify this as the foundational unsolved problem for their application domain.

Tom: And it's not just about accuracy. The paper documents something fascinating—high confidence scores don't guarantee correctness. They show DocTR returning zero point nine seven eight confidence on "FOUC ACID two point five one MG" which is wrong, while PaddleOCR returns zero point nine zero two on "FOLIC ACID two point five MG" which is correct. The high-confidence output was garbage.

Jane: That's a critical finding. It means confidence scores alone are insufficient for text selection in domain-specific OCR. That's why they add a domain dictionary correction step that snaps common OCR errors to known correct spellings.

Lu: And that finding has implications beyond medical documents. Any OCR system deployed in a specialized domain—legal, financial, technical—needs domain-specific post-processing. The confidence score is a necessary but not sufficient signal.

Meng: I'm also struck by their finding about multilingual OCR interference. When they enabled Hindi language detection in EasyOCR alongside English, the system produced confidence values of zero point zero zero one to zero point zero two on clearly readable English text. The Devanagari script model was misclassifying Latin characters as Hindi script components.

Tom: That's a huge practical insight. The paper's recommendation is that language pre-detection must happen before OCR configuration. You can't just enable all target languages and hope the model figures it out.

Jane: And then there's the finding about Indian prescription structure being bimodal. You have structured hospital prescriptions with labeled section headers, and free-form private clinic prescriptions with no standard format. A single parsing approach fails on a significant portion of real-world documents.

Lu: That's why their domain layer is a plug-in architecture. They designed it so non-medical domains—legal, invoice—can be added without modifying upstream layers. And they actually demonstrate it working on dense, free-form legal documents.

Meng: So what are the proposed improvements? The paper lists several. TrOCR fine-tuning on handwritten medical datasets is the big one. Also a learned fusion scorer using XGBoost instead of heuristic confidence weighting. And language pre-detection before OCR initialization.

Tom: Right. And they want to expand regional language support beyond Hindi to Bengali, Marathi, and Telugu using IndicTrans2 as the translation backbone. Plus an agentic upgrade using LangGraph for conversational follow-up Q andA, so ASHA workers can ask clarifying questions about prescription content.

Jane: And the SLM fine-tuning direction is interesting—fine-tuning a small language model on Indian medical Hindi vocabulary to reduce dependency on API-based LLMs for offline deployment. That would address the GPU availability concern Meng raised earlier.

Meng: That would be a game-changer for rural deployment. If you can run the whole pipeline on a laptop or even a smartphone, the reach becomes completely different.

Lu: And I'd add that the benchmark dataset release they propose—annotating and publicly releasing the fifty-document evaluation set—is essential for reproducible research. This domain desperately needs shared benchmarks.

Tom: So the path forward is clear. The handwritten problem is the biggest technical challenge, but the architectural decisions in this paper—the IoM metric, the deterministic structuring, the dual-LLM pipeline—are solid foundations. Let's wrap up with our final thoughts.

Conclusion: Jane: And that brings us to the end of our discussion on "PragyaDoc: A Universal Document Intelligence Framework for Multilingual Medical Document Understanding in Low-Resource Settings." Tom, what's your final take?

Tom: This paper matters because it tackles a real, urgent problem with genuine technical innovation. The IoM metric is a clever solution to a specific failure mode that would otherwise cause downstream errors. The deterministic structuring before LLM invocation is a design principle more systems should adopt.

Jane: And the seven empirical findings are valuable contributions. The multilingual interference phenomenon, the confidence score unreliability, the bimodal structure of Indian prescriptions—these are insights that will help other researchers avoid the same pitfalls.

Lu: I'd add that the glass-box approach to medical reasoning is important. The Rescue Protocol forces the LLM to surface ambiguous decisions rather than making them silently. In a medical context, that transparency is not optional—it's a safety requirement.

Meng: From an engineering perspective, the parallel execution design and the plug-in domain architecture show careful thinking about real-world deployment. The GPU constraint is a limitation, but the SLM fine-tuning direction addresses it directly.

Tom: The paper is honest about its limitations. Handwritten prescriptions remain unsolved. OpenFDA coverage of Indian brand-name medicines is incomplete. The evaluation dataset of fifty documents is small. But the authors are clear about what works and what doesn't.

Jane: And that's what makes this research credible. They're not overselling. They're building a system for ASHA workers and NGO health volunteers in rural India, and they know exactly where the gaps are.

Tom: So we say goodbye to PragyaDoc. A solid piece of work with real social impact potential. Next up, we have a paper on something completely different—I'm looking forward to that one. Thanks for listening, everyone!

Jane: Take care, and see you next time!

Jagpal Singh Jhala

VIT Bhopal University

cs.CV, cs.CL

Submitted: 2026-06-01

Comments: 7 pages 4 images/figure

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 51/100

Terminology

Summary

Summary

This paper presents PragyaDoc, a four-layer document intelligence framework designed to address the critical accessibility barrier created by India’s 22 official languages, where "the majority of medical documentation exists exclusively in English, yet the patients who most urgently require this information — rural populations, ASHA workers, and patient families — are functionally excluded from understanding it. The system is implemented as a modular pipeline with four layers: a parallel ensemble OCR extraction layer, a geometric-lexical fusion layer, a deterministic domain structuring layer, and a dual-LLM medical reasoning and localization layer."

The extraction layer runs three fundamentally different OCR engines — PaddleOCR (PP-OCRv4), EasyOCR (CRAFT+CRNN), and DocTR (db resnet50+parseq) — in parallel via ThreadPoolExecutor, exploiting GIL release during inference to achieve true concurrency. PaddleOCR is configured with custom binarization thresholds and unclip ratios tuned for Indian prescription paper, running on CPU due to a cuDNN version conflict. EasyOCR uses the CRAFT+CRNN architecture and is initialized with English-only configuration because enabling Hindi caused systematic degradation. DocTR uses a db resnet50 detector with a parseq recognizer, selected for superior performance on structured forms and tables. All engines produce output conforming to a standardized JSON schema: text, confidence, bbox 4point, tokens, line index, source.

The fusion layer introduces a novel application of Intersection over Minimum (IoM) alongside the standard Intersection over Union (IoU) metric to resolve bbox containment errors. The paper states: "Standard IoU fails when one engine reads a complete line as a single bounding box and another engine reads the same line as two shorter detections. In this case, the smaller box is fully contained within the larger one, but IoU scores the overlap poorly because the union is large. IoM detects containment directly: if the intersection equals the smaller box, IoM = 1.0 regardless of the size of the larger box." Two detections are assigned to the same cluster if IoU > 0.5 OR IoM > 0.7. Within each NMS cluster, a secondary lexical alignment step uses character-level Levenshtein similarity (≥ 0.8 AND vertical center distance ≤ 15px) to merge candidates. Majority voting with fuzzy matching selects the best text per cluster, and a domain-specific medical dictionary correction step snaps common OCR errors (e.g., FOUC ACID → FOLIC ACID).

The deterministic domain layer applies spatial heuristics — keyword anchoring, Y-axis hierarchical sorting, and a Reading Gravity algorithm — to reconstruct document structure before any generative model is invoked. Document type detection scans for over 40 medical anchor keywords (e.g., Rx, C/O, Diagnosis, Advice, Follow Up, Chief Complaints, D/D); two or more structured keywords classify the document as 'structured', fewer as 'free form'. For structured documents, a stateful section parser reads lines top-to-bottom, maintaining an active section variable. The Reading Gravity algorithm handles medicine block reconstruction: lines beginning with anchor prefixes (TAB., CAP., SYP., INJ.) are designated as medicine anchors, and all subsequent non-anchor lines within an 8-pixel vertical tolerance are assigned to the nearest anchor above them. For free-form documents, medicine signals are extracted using keyword detection and the entire parsed output is passed to the LLM with explicit instructions to extract structured information from unstructured text.

The dual-LLM pipeline uses Llama 3.1 8B (Groq) for constrained JSON extraction with an explicit Rescue Protocol for misclassified entities, and Llama 3.3 70B (Groq) for culturally localized Hindi explanation generation with enforced safety guardrails. The Rescue Protocol gives the model an 'internal reasoning' field in the JSON schema, instructing it to act as a glass box — to explicitly state its reasoning as it searches for medicine names that may have been misclassified into incorrect sections. The model is also provided with a negative rule set: it must not infer stop instructions, contraindications, or drug interactions that are not explicitly present in the input text. Each extracted medicine name is preprocessed and submitted to the openFDA API for real-world indications, drug class, warnings, and side effects. The 70B model generates Hindi explanations bound by aggressive safety guardrails: it cannot hallucinate drug purposes beyond input or OpenFDA data, must flag OCR-corrected medicine names, and must dynamically assign an urgency level (low/medium/high). Every output appends a mandatory disclaimer: 'yeh jaankari sirf samajhne mein madad ke liye hai. Hamesha apne doctor ki salaah maanen.'

Evaluated on a dataset of 50 real Indian medical documents spanning typed prescriptions, handwritten prescriptions, and mixed-format documents, PragyaDoc achieves 92% Character Accuracy Rate (CAR), 90% medicine extraction accuracy, and 95% dosage schedule accuracy on typed documents, outperforming each individual OCR engine. The key performance metrics on typed documents are: CAR 92%, Medicine Extraction Accuracy 90%, Frequency Accuracy 95%, Advice Extraction Accuracy 95%, Extraction Rate (typed) 96.8%, Extraction Rate (handwritten) 7.2%, F1 Score 85%, CAR improvement over DocTR (best single engine) +4%, Medicine extraction improvement over DocTR +10%, Processing time (CPU) 77–130 seconds, Processing time (GPU estimated) 8–12 seconds. Extraction rates by document type: typed hospital 96.7%, typed legal 96.8%, mixed printed/handwritten 54.4%, pure handwritten 7.2%, mixed typed Ayurvedic 47.9%.

Seven novel empirical findings are documented. Finding 1: EasyOCR Multilingual Interference — enabling Hindi language model simultaneously with English caused near-complete failure on typed English medical documents, with confidence values of 0.001–0.02 on clearly readable English text. Finding 2: Confidence Fusion Naturally Suppresses Noise — confidence-weighted fusion correctly deprioritized EasyOCR garbage outputs without explicit filtering rules. Finding 3: Handwritten Indian Documents Remain Unsolved — all three engines produced low-accuracy outputs on handwritten prescriptions, with DocTR achieving less than 30% useful accuracy despite confidence scores of 0.70–0.95. Finding 4: High Confidence Does Not Guarantee Correctness — DocTR returned 0.978 confidence on 'FOUC ACID 2.51 MG' (incorrect) while PaddleOCR returned 0.902 on 'FOLIC ACID 2.5 MG' (correct). Finding 5: IoM Essential for Split-Line Detection — PaddleOCR detected 'SMS Hospital' at y=8–21, DocTR at y=19–31; IoU = 0.087 (below 0.5 threshold) but IoM = 0.73 (above 0.7 threshold), demonstrating that without IoM these would form separate clusters. Finding 6: Extraction Rate and Accuracy Are Distinct Metrics — at min confidence=0.8, min engines=1, extraction rate is 96.8% on typed documents but 54% on handwritten documents with accuracy below 20%. Finding 7: Indian Prescription Structure Is Bimodal — prescriptions fall into structured (hospital with labeled section headers) and free-form (private clinic and Ayurvedic with no standardized headers) categories, requiring two-mode parsing.

The paper documents several limitations: handwritten prescription accuracy below 30%, incomplete OpenFDA coverage of Indian brand-name medicines, processing time of 77–130 seconds on CPU, an evaluation dataset of only 50 documents, section bleeding issues, and cuDNN version conflicts requiring PaddleOCR to run on CPU. Future work includes TrOCR fine-tuning on handwritten medical data, a learned fusion scorer using XGBoost, legal and invoice domain plugins, regional language expansion to Bengali, Marathi, and Telugu using IndicTrans2, language pre-detection before OCR initialization, benchmark dataset release, an agentic upgrade using LangGraph for conversational follow-up Q&A, and SLM fine-tuning on Indian medical Hindi vocabulary for offline deployment. The system is deployed as a Gradio web interface and targets 50,000+ ASHA workers and NGO health volunteers across rural India.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system, along with the resulting capabilities:


Improvement: Add a fast script/language detection module before OCR initialization. Instead of enabling all target languages simultaneously (which caused systematic degradation), the system will detect the dominant script in the image first, then configure OCR engines with only the relevant language models.

Capability: The system will maintain high accuracy on typed English medical documents (92% CAR) while also correctly handling mixed Hindi-English content without the 99% confidence collapse observed when Hindi was blindly enabled.

Improvement: Replace pure IoU-based clustering with a dual-metric approach (IoU > 0.5 OR IoM > 0.7). This specifically handles the split-line detection failure where one engine reads a full line and another reads a substring, resulting in low IoU despite full containment.

Improvement: Add a post-fusion step that applies a domain-specific medical dictionary to snap common OCR errors (e.g., FOUC ACID → FOLIC ACID). Additionally, implement a rule that does not rely solely on confidence scores for text selection, since high confidence (0.978) was observed on incorrect text.

Improvement: Implement a two-mode parser (structured vs. free-form) that uses keyword anchoring, Y-axis hierarchical sorting, and the Reading Gravity algorithm to reconstruct document structure before any generative model is called. This constrains the LLM's input space and reduces hallucination.

Improvement: Use a smaller LLM (Llama 3.1 8B) for constrained JSON extraction with an explicit internal reasoning field that forces transparent reasoning about misclassified entities. Then use a larger LLM (Llama 3.3 70B) for culturally localized Hindi explanation generation with enforced safety guardrails.

Improvement: Implement a confidence-based routing mechanism that detects when the ensemble cannot reach consensus on handwritten text (extraction rate 54% but accuracy <20%). Route such documents to a human review queue instead of attempting automated extraction.

Improvement: Integrate OpenFDA API for drug information enrichment, with explicit low-confidence annotations when Indian brand-name medicines are not found. Fall back to LLM parametric knowledge only when the API returns no results, and log these instances for future dictionary expansion.

  1. Process typed Indian medical documents (hospital prescriptions, discharge summaries, diagnostic reports) with 92% character accuracy, 90% medicine extraction accuracy, and 95% dosage schedule accuracy.

  2. Handle mixed printed/handwritten documents with 54% extraction rate, routing low-confidence handwritten content to human review rather than producing unreliable automated output.

  3. Generate culturally localized Hindi explanations at a Grade 5 reading level, with enforced safety guardrails and a mandatory disclaimer, making medical information accessible to rural patients, ASHA workers, and NGO health volunteers.

  4. Resolve split-line OCR failures using the IoM metric, correctly merging fragmented text from different engines and preventing downstream parsing errors.

  5. Avoid multilingual interference by detecting document language before OCR configuration, maintaining high accuracy on both English and mixed-script documents.

  6. Provide transparent medical reasoning through the Rescue Protocol, surfacing ambiguous entity classifications and enabling downstream uncertainty flagging for human review.

  7. Enrich extracted medicines with verified drug data from OpenFDA, with explicit fallback logging for Indian brand-name drugs not in the API database.

  8. Operate in low-resource settings with a processing time of 77–130 seconds on CPU (8–12 seconds on GPU), making it deployable in rural clinics without specialized hardware.

Abstract

India's 22 official languages create a critical accessibility barrier: the majority of medical documentation exists exclusively in English, yet the patients who most urgently require this information - rural populations, ASHA workers, and patient families - are functionally excluded from understanding it. This paper presents PragyaDoc, a Universal Document Intelligence Framework that addresses this gap through a four-layer pipeline: a parallel ensemble OCR extraction layer, a geometric-lexical fusion layer, a deterministic domain structuring layer, and a dual-LLM medical reasoning and localization layer

Sources

Related papers