What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)

arXiv:2608.06202 · cs.HC, cs.AI · Submitted 2026-08-06 · Read on arXiv

Ro Encarnación, Tina Behzad, Emma Lurie, Danaé Metaxa

University of Pennsylvania · Stony Brook University

cs.HC, cs.AI

Submitted: 2026-08-06

Comments: 18 pages

Code: https://github.com/seerocode/aies-llm-modality-audit

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 77/100

Terminology

Summary

The paper argues that current LLM benchmark evaluation practices contain significant blind spots. As the abstract states: "Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness. Yet most evaluations rely on a single access modality (model APIs), perform a single run per prompt, and report accuracy as the primary outcome metric, without accounting for conditions such as web search that may have effects on model behavior in deployment."

The authors identify three implicit assumptions underlying current benchmarking: "First, they are conducted through a single access modality, treating API-based results as representative of deployed system behavior when users are primarily accessing them through widely deployed consumer-facing chat interfaces. Second, a single response per prompt is treated as adequately reflecting model behavior. Third, similar benchmark accuracy across conditions is often treated as evidence of similar overall model behavior. The paper emphasizes that high benchmark accuracy does not necessarily imply consistent model behavior — a model may answer correctly on average while produc[ing] different answers across repeated runs, ground responses in different information sources, or apply abstention behavior inconsistently across deployment contexts."

The study used a 2×2 factorial audit design comparing two modalities (ChatGPT's consumer chat UI vs. OpenAI's API) and two search conditions (web search enabled vs. disabled), using GPT-5.3 Instant (the default model for logged-out ChatGPT users at the time). The authors note that modality includes the full user-facing system associated with each access method, including system prompts, search, moderation, and other interface behavior that shapes model responses.

Two widely cited safety and bias benchmarks were used as case studies: BBQ (Parrish et al. 2022), covering 11 social bias categories (198 sampled prompts, 18 per category), and SafetyBench (Zhang et al. 2024), covering 7 safety categories (203 sampled prompts, 29 per category). A power analysis targeting Cohen's h = 0.2 indicated a minimum of 197 questions. Each prompt was run three times per condition, producing 4,812 total responses (2,376 for BBQ, 2,436 for SafetyBench), all collected within a one-week window to minimize temporal drift. The chat UI was accessed logged-out to establish a baseline without personalization or account-specific settings, using automated browser sessions, rotating proxies, and staggered delays. The API was queried using default settings including a temperature of 1.0.

The paper evaluated accuracy (proportion matching gold-standard answers, excluding abstentions), consistency (proportion of prompts where all three runs selected the same answer), response similarity (lexical Jaccard similarity and cosine semantic similarity of sentence embeddings), citation rate and source overlap (restricted to search-enabled conditions, at both URL and domain levels), and abstention rate. Statistical analysis used generalized linear mixed-effects models (GLMMs) and linear mixed-effects models (LMEs) with prompt-level random intercepts, with Benjamini–Hochberg false discovery rate correction for primary tests.

Without search, chat UI responses were less accurate than API responses on both benchmarks: for BBQ, chat UI was 79.1% accurate vs. API at 81.9% (a 2.8-point gap); for SafetyBench, chat UI was 85.9% vs. API at 88.5% (a 2.6-point gap). The modality difference was statistically significant on SafetyBench (OR = 0.53, CI [0.30, 0.92], p = 0.025) and borderline on BBQ (OR = 0.66, CI [0.43, 1.02], p = 0.062), meaning chat UI responses are roughly half as likely as API responses to be correct on the same prompt in SafetyBench.

More strikingly, Enabling search further reduced accuracy and reversed the modality gap on SafetyBench. With search enabled, API accuracy on SafetyBench dropped by 7.9 percentage points to 80.6%, while chat UI accuracy fell only 1.5 points to 84.4%, leaving the chat UI more accurate than the API. Statistical analysis confirmed "search reduced accuracy in both modalities (OR = 0.36, CI [0.26, 0.51], p < 0.001), with the decrease significantly larger for API than for the chat UI (modality x search interaction OR = 1.92, CI [1.20, 3.06], p = 0.009). The accuracy decline was steeper on SafetyBench than BBQ, large enough on SafetyBench to reverse the modality effect."

Both modality and search setting contributed to inconsistent outputs. For BBQ, inconsistency ranged from 13.1% (API/no-search) to 21.2% (chat UI/search), with search increasing inconsistency by 6 points for both modalities. SafetyBench showed a reversal pattern: without search, chat UI inconsistency was nearly double that of the API (12.3% vs. 6.4%), but with search, API inconsistency doubled to 12.8% while chat UI inconsistency fell to 8.4% — the only condition of the four benchmark × modality combinations where search decreased inconsistency. A binomial GLMM found search was the stronger driver of inconsistency (OR = 0.51, p = 0.004).

Crucially, the authors distinguished modality-driven variation from model stochasticity by comparing within-modality pairwise disagreement (across repeated runs) with between-modality disagreement (API vs. chat UI on the same prompt). Across all conditions and both benchmarks, between-modality disagreement consistently exceeded within-modality disagreement, ranging from 1.14 to 1.30 times higher than the larger within-modality rate. The logistic mixed-effects model confirmed Between-modality response pairs had significantly higher odds of disagreement than within-modality pairs (OR = 1.51, CI [1.19, 1.92], p = 0.002), indicating modality introduced variation beyond run-level noise.

Even when the API and chat UI selected the same benchmark answer, we found that responses differed more in word choice and in how the answer was explained across modalities than within them. In no-search conditions, BBQ between-modality Jaccard similarity was 0.15, roughly 3–4 times lower than within-modality rates (0.59 API, 0.40 chat); SafetyBench showed a similar pattern (0.16 between vs. 0.77 API and 0.48 chat). The LME confirmed between-modality responses were 0.318 Jaccard units less similar than within-API responses (Cohen's d = −1.67, p < 0.001), and chat UI responses varied more in word choice across runs than API responses.

Semantic (cosine) similarity of explanations showed similarly large effects: Between-modality mean cosine similarity was 0.626 (CI [0.617, 0.635]), compared to 0.802 (CI [0.787, 0.817]) for repeated API responses, and 0.863 (CI [0.857, 0.869]) for repeated chat UI responses, with very large effect sizes (BBQ d = −1.33, SafetyBench d = −1.76). Interestingly, "chat UI responses were also more semantically consistent with each other across runs than API responses... indicating that the chat UI was more self-consistent in how it explained the same answer, even when the words used to do so were different."

The paper also documented stark differences in whether explanatory context was provided at all: "For SafetyBench, the difference was larger with 39.1% of API/no-search responses and 4.9% of API/search responses contained only the answer choice, compared with 1.5% of chat UI/no-search responses and none of the chat UI/search responses. Across all conditions, chat UI responses almost always included additional explanatory context, doing so in 99.6% of responses."

Neither modality cited web sources without search enabled, so citation analyses were restricted to search-enabled conditions. Citation rates reversed direction by benchmark: for BBQ, chat UI cited at least one source in 46.8% of responses vs. 32.2% for the API; for SafetyBench, the API cited more (73.9% vs. 60.6%). The interaction was confirmed as statistically significant (OR = 0.17, [0.12, 0.26], p < 0.001).

When citations were provided, source overlap was minimal: For BBQ, chat UI cited 72% (74% of domains) of all URLs cited across both modalities while the API cited 31% (33%). Only 4% of URLs were cited by both modalities (7% at the domain level). SafetyBench showed an evenly split citation pattern, yet overlap remained nearly the same, at 4% at the URL level and 7% at the domain level. Furthermore, the two modalities had no URL overlap in 37% of BBQ prompts (31% had no domain overlap) and the same was true for 42% of SafetyBench prompts (32% had no domain overlap). The chat UI also returned a broader set of citations through both in-text citations and an additional More Results panel. The example in Figure 3 illustrates a SafetyBench prompt where both modalities answered correctly but the API cited a different source (Netsweeper) than the other two and provided a shorter explanation.

SafetyBench produced no abstentions. On BBQ, only six abstentions occurred across four distinct prompts, all in the no-search condition. Three of the four abstaining prompts triggered an abstention on only one of three runs, and the fourth on two of three runs. All four prompts mentioned protected characteristics across three distinct categories (age, religion (x2), race/ethnicity). Only one prompt triggered abstentions in both modalities. The authors argue this is deeply problematic for safety: If a model guardrail can be bypassed in any interface by simply resubmitting a question and receiving a different output, then the guardrail is functionally useless as a safety mechanism. They note that a model that abstains on one run and answers on the next will score identically to one that answers consistently, leaving safety-relevant inconsistency unaccounted for.

The paper makes three contributions: "First, we show that benchmark outcomes can vary across modality and search conditions in ways that aggregate accuracy fails to capture alone. Second, we demonstrate that consistency, text similarity, citation grounding, and abstention behavior reveal additional dimensions of model behavior that standard benchmark evaluation practices currently overlook. Finally, we propose a modality- and search-aware evaluation methodology for comparing behavioral variation across API and chat-based deployment contexts."

The discussion argues that standard benchmark reporting overstates output reliability, by hiding these sources of variation, instead collapsing repeated interactions into a single aggregate score. The authors draw an analogy: "No one would accept a car safety rating derived from a single test run, under a single condition or simulation, reported as a bare average. No regulator would consider a product safety test conducted once in a single setting and reduced to a single pass rate sufficient. Yet, AI benchmarks are often conducted and interpreted this way."

They call for "modality- and search-aware evaluation practices, including evaluating deployed interfaces alongside APIs; reporting repeated-run measures; separating performance from output measures; and explicitly documenting reasoning, citation, and abstention behavior as necessary for benchmark results to be high fidelity and ecologically valid reflections of deployed system behavior. The paper also highlights an infrastructure gap — chat UI auditing requires substantially more technical complexity (session management, rotating proxies, staggering), meaning If certain access modalities, like chat interfaces, are harder to study, the evaluation literature will overlook it and instead treat other modalities as de facto 'good enough' proxies, and recommends that model developers and government partners" provide stable testbed infrastructure for independent auditing.

The authors acknowledge their results are based on two established AI safety and bias benchmarks, one model family, and a single data collection period. They note that benchmark choice itself may matter, that repeated questions may have been seen in training data, that citation quality was not assessed, and that semantic similarity results are embedding-model dependent. They call for longitudinal evaluation with additional prompt runs and qualitative analyses of source authority, retrieval quality, and user trust implications in future work. The datasets are available at https://github.com/seerocode/aies-llm-modality-audit.

Improvements for AI systems

Here are specific improvements to AI systems motivated by this paper’s findings.

Improvement: Train or fine-tune models to maintain the same answer and safety behavior across deployment contexts (chat UI vs. API) and with/without web search. Add a runtime consistency monitor that compares outputs against a cache of prior responses to the same prompt and flags high-variance cases.

What the improved AI system can do: The system will not silently shift from 88.5% accuracy (API/no-search) to 80.6% accuracy (API/search) on the same prompts. It can detect when its own answer changes merely because the interface or search toggle changed, and either align its behavior or clearly flag uncertainty for the user.

Improvement: Implement safety guardrails as a separate, deterministic layer that is invariant to sampling temperature, repeated runs, and web-search context. Inputs classified as safety-critical should trigger abstention with probability 1, not probabilistically.

What the improved AI system can do: The system will stop exhibiting the paper’s observed failure mode where the same prompt triggered an abstention on one run and a full answer on the next, and where guardrails were entirely bypassed in the other modality. Submitting a question repeatedly or through a different interface will no longer “roll the dice” to get a harmful response.

Improvement: Add a post-search verification step: before returning a final answer when web search is enabled, the model cross-checks its own pre-search answer against retrieved sources. If search materially changes the answer, the system should either reconcile the discrepancy or report a confidence drop.

What the improved AI system can do: This directly addresses the result that search reduced accuracy by 7.9 points on SafetyBench for the API. The system will not blindly over-ride its base knowledge with whatever the first search result says. It will also avoid the observed reversal where the supposedly less-capable chat UI became more accurate than the API simply because it handled search differently.

Improvement: When answering with sources, the system should maintain a consistent citation set across runs and interfaces. Implement a citation-merging mechanism that selects a stable, high-quality subset of sources instead of a different set on each call.

What the improved AI system can do: The system will no longer produce the paper’s finding of only 4% URL overlap between the chat UI and API on the same prompts. Users and auditors will not see contradictory citations for identical questions depending on which interface they use. For disputed or sensitive topics, the system can explicitly state the set of sources it considers authoritative.

Improvement: Enforce a uniform reasoning-and-explanation policy across modalities. If a response contains only the answer choice, it should be flagged as incomplete for safety and bias benchmarks.

What the improved AI system can do: The system will eliminate the observed gap where 39.1% of API/no-search responses contained only an answer choice while chat UI gave full explanations 99.6% of the time. Users in all interfaces will receive the same level of reasoning, and auditors can trust that answer accuracy is not coupled to arbitrary formatting differences.

Improvement: Build an evaluation harness (or a self-evaluation utility) that tests the system in both chat UI and API, with search on/off, using multiple runs per prompt, and reports—beyond accuracy—the paper’s key secondary metrics: consistency, between/within-modality disagreement, citation overlap, and abstention stability.

What the improved AI system can do: When a developer or regulator queries the system for its safety characteristics, the system can report a richer profile: not just “91% accurate on SafetyBench,” but also “91% accurate in API/no-search, 84% in chat UI/search; 12.8% inconsistency; 4% citation overlap.” This makes deployment risks visible and forces any claimed safety/performance score to be conditioned on the actual usage context.

Improvement: Make the system’s confidence calibration aware of the interface it is running in. When accessed through a chat UI with search enabled, the system should apply a stricter confidence threshold before answering because behavior is less controlled.

What the improved AI system can do: This addresses the paper’s finding that chat UI and search conditions change not only accuracy but also inconsistency and abstention patterns. The system will proactively say “I’m less certain here” in high-variability conditions rather than giving a fluent but inconsistent answer, improving safety-relevant behavior in exactly the situations current benchmarks ignore.

Sources

Related papers