Leveraging Large Language Models for Analyzing Blood Pressure Variations Across Biological Sex from Scientific Literature
Yuting Guo, Seyedeh Somayyeh Mousavi, Reza Sameni, Abeed Sarker
Emory University
cs.CL, cs.AI
Submitted: 2026-08-14
Updated: 2026-08-17
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 55/100
Terminology
Summary
Summary
This paper investigates the use of Large Language Models (LLMs), specifically GPT-3.5-turbo, to automatically extract blood pressure (BP) data—mean and standard deviation of systolic (SBP) and diastolic (DBP)—stratified by biological sex from a large corpus of scientific literature. The authors hypothesize that aggregating BP statistics from published studies can help infer population-wide BP distributions across demographic factors, addressing biases in existing BP measurement standards that do not account for individual characteristics.
Data Collection: The dataset comprised approximately 25 million abstracts from PubMed Central (PMC), downloaded via the PMC FTP service. Abstracts were filtered to those containing the keyword 'blood pressure' (321,226 abstracts), then further refined to those containing 'mmHg' or 'mm Hg' (71,067 abstracts). Only article abstracts, not full texts, were used.
Information Extraction: The authors employed GPT-3.5-turbo in a zero-shot setting (no training data or manual annotation). The prompt instructed the model to extract ten variables: number of males (N male), number of females (N female), mean ± standard deviation of SBP for males and females, and mean ± standard deviation of DBP for males and females. The prompt specified a strict output format for parsing. After generation, a regular expression script extracted the values. The prompt was refined via trial and error.
Model Performance Evaluation: Out of 993 abstracts where the model predicted all ten variables, 44 cases were manually reviewed. In 7 cases where the abstract contained all variables, the model achieved 100% accuracy. However, in cases where abstracts did not contain all variables, the model sometimes generated values anyway, leading to inaccuracies. For example, the model might assign BP values to males and females even when the abstract only reported combined values. The authors also observed that the model sometimes performed its own calculations, such as averaging multiple BP values from an abstract (e.g., averaging SBP for 13-year-old and 17-year-old boys to produce a single mean). They noted that these derived values might originate from full texts available in the model's pretraining data.
BP Variation Analysis: After extraction, post-processing ensured BP values fell within plausible ranges (DBP: 30–120 mmHg, SBP: 60–200 mmHg). Studies with fewer than 100 total BP records were excluded, leaving 582 studies for analysis. The authors used Gaussian mixture models (GMM) to create heatmaps and contour plots comparing BP distributions between males and females. The results showed that males tend to exhibit higher BP values than females, consistent with prior literature. The contour plots revealed distinct but overlapping distributions for each sex.
Key Findings and Limitations: The study demonstrates the viability of using LLMs for large-scale extraction of BP-related information from biomedical literature. However, the authors acknowledge limitations: (1) the evaluation was labor-intensive and not comprehensive due to resource constraints; (2) the LLM is prone to hallucinations
—generating incorrect values not supported by the abstract (e.g., providing standard deviations when none were reported); (3) the study focused solely on biological sex, neglecting other factors like age, race, height, and weight due to cost constraints. The authors suggest future work should expand to other demographic factors and develop better methods for evaluating LLM responses.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems and what the improved system can do:
1. Hallucination Mitigation with Confidence Scoring
-
Add a post-generation validation layer that cross-checks extracted values against the source abstract for explicit mention of each variable (e.g.,
standard deviation
or "±"). -
Implement a confidence score for each extracted value based on textual evidence presence, with a threshold to flag or discard unsupported predictions.
-
Use a two-pass approach: first extract, then re-prompt the LLM to justify each value with a direct quote from the abstract, and only retain values with verifiable quotes.
2. Structured Output Enforcement with Constrained Decoding
-
Replace free-form prompt output with a JSON schema validator that rejects outputs missing required fields or containing non-numeric values.
-
Use grammar-constrained decoding (e.g., via libraries like
outlinesorguidance) to force the LLM to output only valid numeric ranges and units, eliminating parsing errors and reducing random assignment.
3. Multi-LLM Ensemble with Majority Voting
-
Run the extraction task with multiple LLMs (e.g., GPT-3.5-turbo, GPT-4, Claude) and use majority voting across models for each variable.
-
Only accept a value if at least two models agree, reducing single-model hallucinations and improving reliability for downstream analysis.
4. Context-Aware Aggregation for Derived Values
-
Detect when the abstract reports multiple BP values (e.g., age-stratified) and explicitly ask the LLM to either (a) report all values with their subgroups, or (b) compute a weighted average with sample sizes, rather than silently averaging.
-
Add a rule-based preprocessor to flag abstracts with multiple BP mentions and route them to a more detailed extraction prompt.
5. Demographic Factor Expansion with Hierarchical Extraction
-
Extend the prompt to extract age, race/ethnicity, height, weight, and comorbidity status alongside sex, using a hierarchical structure (e.g.,
SBP in male aged 40-49
). -
Use a two-stage pipeline: first extract all demographic factors present, then for each combination, extract BP statistics, enabling multi-dimensional analysis.
-
Provide clinically reliable BP distributions with per-value confidence scores and source quotes, enabling researchers to filter out hallucinations before analysis.
-
Automatically flag and reject abstracts where the LLM cannot verify values, reducing false positives from 44 to near-zero in the
predicted all variables
category. -
Generate multi-demographic BP heatmaps (sex × age × race) from 25 million abstracts, not just sex-based, allowing detection of intersectional biases in BP standards.
-
Perform meta-analysis on-the-fly by aggregating extracted means/SDs with proper weighting (using N), producing population-level estimates with confidence intervals, directly comparable to clinical guidelines.
-
Detect and correct systematic biases in the LLM's averaging behavior by comparing derived values against raw abstract data, and automatically re-extract with explicit instructions when discrepancies exceed a threshold.
-
Scale to full-text analysis (not just abstracts) by using the same constrained decoding pipeline on PMC full-text XML, increasing the yield of usable studies from 993 to potentially tens of thousands, while maintaining verifiability.
Abstract
Hypertension, defined as blood pressure (BP) that is above normal, holds paramount significance in the realm of public health, as it serves as a critical precursor to various cardiovascular diseases (CVDs) and significantly contributes to elevated mortality rates worldwide. However, many existing BP measurement technologies and standards might be biased because they do not consider clinical outcomes, comorbidities, or demographic factors, making them inconclusive for diagnostic purposes. There is limited data-driven research focused on studying the variance in BP measurements across these variables. In this work, we employed GPT-35-turbo, a large language model (LLM), to automatically extract the mean and standard deviation values of BP for both males and females from a dataset comprising 25 million abstracts sourced from PubMed. 993 article abstracts met our predefined inclusion criteria (i.e., presence of references to blood pressure, units of blood pressure such as mmHg, and mention of biological sex). Based on the automatically-extracted information from these articles, we conducted an analysis of the variations of BP values across biological sex. Our results showed the viability of utilizing LLMs to study the BP variations across different demographic factors.
Sources
- A Survey on Blood Pressure Measurement Technologies: Addressing Potential Sources of Bias
- Scaling Vision Transformers to 22 Billion Parameters
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering