Assessing Reliability of BERT-Based Models on Question Answering Tasks

arXiv:2608.10806 · cs.CL · Submitted 2026-08-11 · Read on arXiv

Pooja Yadav, Priyanka Harjule, Basant Agarwal, Marko Robnik Šikonja

Malaviya National Institute of Technology · Central University of Rajasthan · University of Ljubljana

cs.CL

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: Accepted for publication in the Journal of Experimental & Theoretical Artificial Intelligence

DOI: 10.1080/0952813X.2026.2716084

Code: https://github.com/uncertainity-quantification/reliability-estimation-qa1

Project page: https://rajpurkar.github.io/SQuADexplorer

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: This study evaluates the reliability of four BERT-based models—RoBERTa, BERT-Base, DistilBERT, and ALBERT—on question-answering (QA) tasks using two datasets: SQuAD 2.0 and QuAC.

Terminology

Summary

This study evaluates the reliability of four BERT-based models—RoBERTa, BERT-Base, DistilBERT, and ALBERT—on question-answering (QA) tasks using two datasets: SQuAD 2.0 and QuAC. The authors define reliability as the consistency of a model in generating answers to input queries when subjected to controlled perturbations either in the model’s configuration or input. They assess reliability through two complementary methodologies: (1) internal model variations induced via Monte Carlo Dropout (MCD) during prediction, and (2) input perturbations through paraphrasing using a pre-trained BART-based paraphrasing model. For evaluation, they use cosine similarity and F1 score as metrics, comparing predicted answers against ground truth.

For the MCD-based assessment, the authors first conducted a systematic sweep of dropout rates from 0% to 35% and selected 10% as a balanced operating point, noting that dropout rates ≥ 15% are excluded because models increasingly produced blank responses, and that 5% dropout did not introduce sufficient stochastic variation. With 50 stochastic samples per input, they found that enabling MCD during prediction does not disrupt inference dynamics; correlation coefficients between unperturbed and perturbed model outputs were consistently high across all models and datasets (e.g., RoBERTa: 0.8573 cosine similarity and 0.8673 F1 on SQuAD; 0.8607 and 0.8784 on QuAC). The authors state: The empirical results consistently show high correlation coefficients across all models and datasets, indicating that dropout has a minimal impact on prediction stability.

Key findings from the MCD-based evaluation on SQuAD show that DistilBERT achieved the highest average cosine similarity (0.8992) and F1 score (0.8496) for answerable questions, while RoBERTa demonstrated strong overall consistency with average scores of 0.8481 (cosine similarity) and 0.8098 (F1), with standard deviations of 0.1543 and 0.1772. On the QuAC dataset, RoBERTa achieved the best semantic alignment for answerable questions (cosine similarity 0.3641 ± 0.1546, F1 score 0.1942 ± 0.1414), while DistilBERT showed less dispersion overall (cosine similarity 0.4265 ± 0.1470, F1 score 0.2881 ± 0.1360). ALBERT performed poorly on answerable questions, particularly on QuAC, where it recorded an answerable accuracy as low as 0.07%.

For the input perturbation assessment, paraphrased questions were filtered using a semantic similarity interval of 0.75–0.98, empirically validated to maintain balanced semantic preservation and lexical variation. On SQuAD, RoBERTa demonstrated the strongest overall performance, achieving a cosine similarity of 0.829 and F1 score of 0.7503 for answerable questions, and overall scores of 0.4895 (cosine similarity) and 0.3812 (F1). On QuAC, ALBERT led with overall cosine similarity of 0.3217 and F1 score of 0.2333, followed closely by RoBERTa (0.3206 and 0.2135). The authors note: RoBERTa shows the strongest performance on the SQuAD dataset when exposed to inputs, whereas ALBERT performs best on the QuAC dataset.

Statistical analyses included Welch’s t-test comparing RoBERTa and DistilBERT under MCD, which yielded a p-value of 0.0 on SQuAD (significant difference) and 0.4547 on QuAC (no significant difference), indicating that the performance distinction between models is context-dependent. For paraphrased inputs, pairwise Wilcoxon signed-rank tests with Bonferroni correction (α ≈ 0.0021) across 24 comparisons revealed significant performance discrepancies among several model pairs.

Human evaluation was conducted on a stratified random 10% subset of responses, with three annotators achieving substantial inter-annotator agreement (Fleiss’ Kappa κ = 0.7947 for RoBERTa and κ = 0.7671 for DistilBERT). The results showed that a considerable proportion of responses initially categorized as incorrect by automatic metrics were semantically meaningful: for RoBERTa, 45.79% were correct, 6.54% partially correct, and 47.66% incorrect; for DistilBERT, 39.25% correct, 8.41% partially correct, and 52.34% incorrect.

Error analysis revealed instances of hallucination under both perturbation types. For example, under MCD, BERT-Base responded to a question with the question itself, and ALBERT produced blank responses in several cases. Under paraphrasing, models sometimes produced divergent answers, such as DistilBERT answering land-based reinforcements instead of the correct deportation of the French speaking Acadian population from the area.

The authors conclude that accuracy alone is insufficient to represent a model’s reliability and prediction stability, and that RoBERTa, DistilBERT, and AlBERT outperform BERT Base, with performance varying across datasets. They emphasize that "RoBERTa performs superior on both datasets in terms of accuracy and consistency, whereas DistilBERT is more stable in handling internal configuration variations, and ALBERT performs better when subjected to handling input variations, highlighting the scenario dependency. The study also includes an appendix evaluating Tiny Llama, which showed consistently low performance (average cosine similarity and F1 scores below 0.5), leading the authors to conclude that BERT-based models remain relevant candidates for investigating QA tasks, providing stable and accurate outputs that allow meaningful estimation of model stability."

Improvements for AI systems

Improvements to AI Systems:

  1. Reliability-Aware Confidence Scoring: Integrate Monte Carlo Dropout (MCD) at inference time with a dropout rate of 10% to generate multiple stochastic forward passes. Use the variance across these passes to produce a reliability score alongside each answer. The improved system can flag low-confidence responses for human review, reducing silent failures in production QA systems.

  2. Perturbation-Robust Training: Augment training data with paraphrased questions (filtered to semantic similarity 0.75–0.98) to make models invariant to lexical variations. The improved system can maintain answer accuracy when users rephrase queries, as demonstrated by RoBERTa’s 0.829 cosine similarity on SQuAD under paraphrase perturbation.

  3. Model Selection Based on Scenario: Implement a dynamic model router that selects the optimal BERT variant per task context. For internal stability (e.g., noisy model weights), prefer DistilBERT (highest consistency under MCD). For input variation (e.g., user-generated text), prefer ALBERT (best on QuAC paraphrases). The improved system can adaptively choose the most reliable model for a given deployment scenario.

  4. Hallucination Detection via Answer-Question Consistency: Use the finding that models sometimes output the question itself or blank responses under perturbation. Build a post-processing filter that checks if the predicted answer is a verbatim copy of the input question or empty, and re-runs inference with a different dropout seed or paraphrase. The improved system can automatically reject and retry such degenerate outputs.

  5. Human-Aligned Evaluation Metrics: Replace or supplement F1/cosine similarity with a semantic equivalence classifier trained on human annotations (Fleiss’ Kappa ≈ 0.79). The improved system can better judge answer correctness, recovering the 45.79% of RoBERTa’s responses that were semantically correct but scored as incorrect by automatic metrics.

  6. Uncertainty-Guided Answer Abstention: For low-reliability inputs (e.g., high variance under MCD or low similarity across paraphrases), allow the system to abstain from answering rather than providing a likely incorrect response. This is directly supported by the paper’s finding that accuracy alone is insufficient, and that models like ALBERT produce near-zero accuracy on certain answerable questions.

  7. Cross-Dataset Generalization Checks: Before deploying a QA model, run a quick calibration on a secondary dataset (e.g., QuAC) to detect dataset-specific overfitting. The improved system can warn if a model performs well on SQuAD but poorly on QuAC (as ALBERT did), enabling early intervention.

  8. Paraphrase-Augmented Consistency Regularization: Add a training loss term that penalizes divergent answers between original and paraphrased inputs. The improved system will produce more stable answers under rephrasing, reducing cases like DistilBERT’s “land-based reinforcements” error.

Sources

Related papers