From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop

arXiv:2608.11171 · cs.CL, cs.AI, cs.CY · Submitted 2026-08-11 · Read on arXiv

Rahul Gupta, Abhinav Mohanty, Anaelia Ovalle, Anil Ramakrishna, Anubrata Das, Apurv Verma, Jwala Dhamala, Ninareh Mehrabi, Tharindu Kumarage, Yada Pruksachatkun, Yang Trista Cao, Kai-Wei Chang, Aram Galstyan

Amazon AGI · Meta · Autodesk · New Jersey Institute of Technology · Salesforce · University of California, Los Angeles

cs.CL, cs.AI, cs.CY

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: 17 pages, 2 figures, 3 tables. Submitted to ACL ARR August 2026 cycle (EACL 2027)

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

The gist: The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a

Terminology

Summary

The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mechanistic understanding and proactive control of generative systems. We synthesize insights from all 144 proceedings papers, classifying them along six trust dimensions grounded in established frameworks (TrustLLM, DecodingTrust). We observe co-occurrences with capability emergence. The release of the first high-impact chat models activated all trust dimensions simultaneously, while subsequent model generations shifted focus toward truthfulness and safety alignment. Analysis from the classification study reveals that truthfulness is the fastest-growing dimension (absent in 2021–2022, comprising 37% of papers by 2025–2026), fairness remains the most consistent theme, and explainability exhibits a U-shaped trajectory; declining as post-hoc methods lost relevance but resurging in 2026 through mechanistic interpretability. A cross-venue comparison with ACL, NAACL, EACL, and EMNLP (2K papers) in the same period shows that TrustNLP’s topical distribution closely follows the field average. We identify four structural insights and conclude with actionable directions for the research community.

Alongside the 5× increase in archived proceedings papers between 2021 and 2026, the research focus underwent a qualitative transformation in what “trust” means. Before the era of LLM launches (November 2022), the key questions were around fairness and interpretability: can we audit model decisions? After the launch of high-impact chat models, trust also became about reliability and trade-offs: is the LLM consistent across prompts, and how do fairness, performance, and explainability interact? By 2025–2026, with multimodal & agentic systems, trust evolved into controllability: can we constrain model behavior under adversarial conditions and mechanistically verify alignment?

We make three contributions: (1) a quantitative topic analysis of all 144 proceedings papers revealing the fastest-growing and most persistent research themes; (2) a chronological synthesis of technical contributions organized by the dominant paradigm of each phase; and (3) the identification of four structural insights and actionable directions for the research community.

We focus on TrustNLP as a representative case study as three properties make TrustNLP particularly well-suited for this role. First, its publication pipeline is integrated with the ACL ecosystem: accepted papers are peer-reviewed through ARR or direct submission and appear in the ACL Anthology alongside main-conference proceedings at NAACL, ACL, and EACL, ensuring that the work reflects the quality norms and topical pulse of the broader NLP community. Second, the workshop maintains a deliberately narrow scope centered on trustworthiness challenges in generative and interactive language systems, rather than serving as a venue for AI ethics broadly construed; this focus yields a sharper signal about how the field responds to specific capability emergence. Third, TrustNLP offers longitudinal continuity with six consecutive editions (2021–2026) with a stable core of organizers and a consistent programmatic agenda. This enables us to trace thematic evolution over time.

The study of trustworthiness in AI systems spans multiple communities, each contributing distinct methodological perspectives. Several comprehensive surveys and benchmarks have sought to systematize the evaluation of trustworthiness in large language models. DecodingTrust provides a multidimensional assessment of GPT models across toxicity, stereotype bias, adversarial robustness, out-of-distribution generalization, privacy, machine ethics, and fairness, establishing one of the first holistic evaluation protocols for proprietary models. TrustLLM proposes a unified evaluation framework that encompasses truthfulness, safety, fairness, robustness, privacy, machine ethics, transparency, and accountability, and benchmarks over a dozen open-source and proprietary models to reveal systematic gaps between the stated alignment goals and measured behavior. While these frameworks provide valuable snapshots of model trustworthiness at a single point in time, TrustNLP’s longitudinal proceedings capture how trust challenges emerge, intensify, and transform in response to capability shifts, providing a complementary, diachronic perspective that static benchmarks cannot offer.

To study how research topics have evolved across six editions, we classify all 144 archived proceedings papers by trust dimension. We derive our trust taxonomy from TrustLLM, which benchmarks six dimensions of LLM trustworthiness: Truthfulness, Safety, Fairness, Robustness, Privacy, and Machine Ethics. We adopt these as our starting point and refine them using DecodingTrust’s finer-grained distinctions to clarify scope boundaries. DecodingTrust’s Toxicity is folded into Fairness & Bias to merge commonalities where toxic content may disproportionately target particular marginalized groups. DecodingTrust’s three robustness sub-dimensions inform the scope of our Robustness category. We argue that TrustLLM’s Safety and Machine Ethics are methodologically intertwined and merge them into a single Machine Ethics & Safety dimension. In this way, we expand the scope of each class by merging overlapping class labels from the two taxonomies. Finally, neither TrustLLM nor DecodingTrust benchmarks explainability (TrustLLM proposes “Transparency” as a principle but does not operationalize it). However, explainability dominates the TrustNLP corpus in 2021–2022 and appears explicitly in the workshop CfP, so we add it as a corpus-motivated dimension. This yields six trust dimensions: Fairness & Bias, Robustness & Adversarial, Privacy, Machine Ethics & Safety, Truthfulness, and Explainability. Papers may belong to multiple dimensions.

We used three independent annotation sources to classify all 144 papers. First, we defined annotation instructions specifying the scope of each trust dimension. Labels were then obtained from three sources: (i) one author annotated all papers based on their title and abstract following the instructions; (ii) Claude Sonnet 5 (Anthropic) and Amazon Nova Lite 2.0 each independently annotated the same papers using the same instructions as a system prompt. After the first round, 60% of papers (87/144) had perfect agreement across all three annotators, and in 78% of cases (112/144) at least one LLM fully agreed with the human. Agreement is substantial to excellent (κ > 0.7) for four of six dimensions. The two most contested boundaries—Machine Ethics & Safety and Truthfulness—reflect a conceptual overlap (e.g., a paper on LLM safety degradation under attacks touches both robustness and safety). The human annotator then reviewed the remaining 40% of cases with imperfect agreement and decided the final label. A similar classification on trust papers (first filtered via keyword search), shows that TrustNLP’s topical distribution generally matches with the broader field. This is indicative that the workshop mirrors the community’s trust priorities.

Several patterns emerge from the trust dimension distribution: Truthfulness is the fastest-growing dimension, absent entirely in 2021–2022 but rising to 13–14 papers per year in 2025–2026 (38 total, 26% of the corpus). This surge tracks the LLM era: as generative models became the dominant paradigm, hallucination, factuality, and calibration emerged as first-order trust concerns. Fairness & Bias is the most stable dimension, present in every edition at 2–7 papers/year (30 total). However, its nature has shifted: early work focused on standalone bias measurement, while recent work reveals fairness–explainability trade-offs, fairness–performance trade-offs, and internal-vs-output bias mismatches. Fairness is transitioning from a standalone objective to a constraint in multi-objective optimization. Robustness & Adversarial was absent in 2021–2022 but stabilized at 5–8 papers/year from 2023 onward (27 total), driven by jailbreaks, red-teaming, and multimodal attacks. The nature of adversarial work has shifted: pre-LLM work focused on perturbing inputs, while post-LLM work involves co-constructing attacks with the model itself. Machine Ethics & Safety shows steady growth from 0 papers in 2021 to 9 in 2026 (20 total), reflecting the post large LLM (such as ChatGPT) focus around alignment, refusal behavior, and safety guardrails. Explainability peaked early (9 papers across 2021–2022), declined to 2 papers in 2025, but rebounded sharply to 13 in 2026 (29 total). The early peak reflects the pre-LLM focus on understanding model decisions; the 2026 resurgence is driven by mechanistic interpretability—sparse autoencoders, linear probes, and activation steering—representing a shift from post-hoc explanations to causal understanding of model internals. Privacy appears sporadically (concentrated in 2023 with 6 papers), suggesting that privacy remains a specialized concern within the TrustNLP community despite its prominence in the broader trustworthiness literature.

The inaugural workshops focused on fairness, explainability, and privacy. As large pre-trained models became the predominant paradigm around this time, we observed a shift toward research focused on measuring representational harms in generative models. During this period, NLP tasks were predominantly classification and ranking systems, and trust was fundamentally about transparency of decisions. Two complementary research threads defined this phase: establishing interpretability as an auditing requirement beyond simple accuracy reporting, and exposing how demographic and cross-lingual biases are structurally embedded in popular corpora and model representations. Research established interpretability as a test for auditing decisions in consequential domains such as healthcare and legal systems. Early work explored jointly bootstrapping neural relation extractors with explanation decoders to ensure that outputs were accompanied by human-readable rationales. The challenge of accountable error characterization was central: researchers argued that simply reporting accuracy is insufficient for trust; models must quantify their own uncertainty. This led to attention-based attribution frameworks that identify features responsible for both performance and fairness through intervention and weight manipulation. These attribution frameworks provided a mechanism for auditing not just what a model predicts, but why, a distinction that would become increasingly important as models grew more opaque. The emphasis on uncertainty quantification in these early papers anticipated the later focus on truthfulness & confidence calibration that would dominate the 2025 edition. A notable contribution in 2021 was a cross-linguistic study of gender bias in Wikipedia corpora across nine languages, including Arabic, German, Farsi, and Urdu. This work showed that gender bias in NLP corpora is not solely an English-language concern and that bias metrics developed for English require careful adaptation when applied across typologically diverse languages. Related work on word embeddings further demonstrated that measured bias can vary substantially depending on the choice of similarity measure and descriptive statistic, underscoring the need for careful methodological choices when evaluating demographic bias in embedding spaces. In 2022, researchers probed the Allen AI Delphi model to investigate moral reasoning. They found that models mirror the moral principles of the demographic groups involved in the annotation process, raising fundamental questions about whose values are codified in “moral” AI systems. The Encoder Marginalization framework was also introduced for dense passage retrieval, quantifying the contribution of different encoders to in-domain accuracy and providing a mechanism for identifying affecting factors during training. This early work on encoder analysis foreshadowed the mechanistic interpretability research that would emerge in 2026.

The 2023 proceedings mark the workshop’s first major expansion after the release of large LMs like ChatGPT, growing from 8 papers in 2022 to 28 in 2023. The capability emergence activated research across several trust dimensions simultaneously. While fairness remained steady (6 papers), robustness, privacy, and truthfulness all saw first-time emergence—each contributing 6 papers, making 2023 the most evenly diversified edition. With generative models deployed at scale, truthfulness emerged as a first-order concern. Work on abstractive summarization addressed the familiar problem that generated summaries may be unsupported by, or inconsistent with, the source document, placing factuality at the intersection of model training and evaluation. Other papers asked how reliability should be assessed when model behavior depends on prompt wording, when benchmark data may have appeared in training data, and when confidence must be inferred from patterns of agreement across prompts. Together, these papers show the workshop moving from general concerns about unreliable generation toward prompt-robust and contamination-aware evaluation. Robustness went from zero papers in 2021–2022 to six in 2023: work on sample attackability and minority-language robustness treated trust as resistance to perturbation and attack, including for languages and scripts outside the usual English-centered evaluation setting. Privacy similarly surged, with papers examining personal information in training corpora and pseudonymization as a way to preserve data utility while reducing exposure risk. These papers broadened the generative-model discussion by emphasizing the data, associated privacy threat models, and deployment conditions around the system, not only the surface quality of generated text. The workshop’s earlier fairness concerns extended into the generative era. Papers on abuse detection, disability bias, representational-harm metrics, and name-based effects examined how models can treat social groups unevenly and how such behavior can be measured or diagnosed. Fairness maintained its steady presence (6 papers), but the methods shifted from measuring pre-trained embeddings to auditing generative outputs. The overall lesson of 2023 is that capability emergence activated multiple trust concerns in parallel rather than concentrating attention on any single dimension.

With rapid improvements in model quality and progress in open models, the 2024 edition shifted from studying trust dimensions in isolation to understanding how they interact under deployment constraints. The diverse focus from 2023 continued: Fairness (6 papers), Robustness (5), and Truthfulness (5) remained nearly balanced, though Privacy declined from its 2023 peak to zero papers. One study compared shallow learning, LM fine-tuning, and multilingual fine-tuning for detecting machine-generated text. Another showed that detection scores are not determined only by the generated text: prompts and real texts can affect them too. Using causal diagrams, the authors identified backdoor paths and showed that some confounding bias can be partly reduced. This suggests that future detectors should account for prompt and context effects, not only classify outputs in isolation. Safety work became more proactive. Automated Adversarial Discovery framed red-teaming as a search for attacks that both fool a safety classifier and belong to previously unseen harm dimensions. In parallel, Cross-Task Defense showed that instruction tuning with refusal examples can help LLMs handle malicious long documents, such as manuals for illicit activities, while still performing benign NLP tasks. The defining contribution of 2024 was the recognition that trust dimensions conflict. Adapter modules kept text classification accuracy close to full fine-tuning while reducing training time, but their fairness effects varied by sensitive group and could amplify existing bias in some settings. Fairness did not move cleanly with explainability: methods that improved one dimension did not reliably improve the other. A lesson emerged wherein a model should not be called trustworthy because it improves on one axis alone; its deployment costs, group-level effects, and explanations need to be checked together.

The fifth workshop widened the trust problem beyond text-only generation, asking what breaks when models are used with databases, unclear questions, images, audio, and evidence-heavy claims. An IoT smart-building study tested whether models could turn questions into database queries. Models improved at producing queries, but reasoning over returned rows (e.g., whether traffic looks malicious) remained hard, showing that query generation is only part of the job. For question answering, researchers detect ambiguity before answering by measuring disagreement across candidate answers rather than asking the model to label questions as ambiguous, which also tightens confidence calibration on the answerable cases. Audio privacy work on membership inference for CLAP models proposed a text-only detector that uses generated gibberish rather than exposing real audio to the target model, shifting the question from whether a model repeats private data to whether we can tell a speaker was in training without leaking more about them. On the safety side, PBI-Attack jailbreaks vision-language models by combining image and text perturbations in a black-box setting, showing that cross-modal systems can be attacked through interactions between inputs rather than through any single prompt. Several papers moved from finding failures to making outputs easier to check. Self-refine with formatting was proposed as a training-free jailbreak defense that reduced attack success rates even on models without safety tuning, though the authors note it may need several rounds and does not work for every model. For claim verification, Minimal Evidence Groups identify the smallest non-redundant evidence set that fully supports a claim and outperform single-step retrieval. Together, these papers make the 2025 direction concrete: trust is not just about refusing bad requests or giving fluent answers; it is about having a checkable reason for the system’s behavior.

The sixth edition is the largest to date (41 papers) and marks two shifts: (i) explainability resurges via mechanistic interpretability methods that probe model internals rather than generating post-hoc explanations, and (ii) Machine Ethics & Safety reaches its peak as the community confronts alignment fragility at scale. The most striking development in 2026 is the return of explainability research (11 papers, up from 2 in 2025), driven by a fundamentally different methodology than the post-hoc attribution methods of 2021–2022. Linear probes reveal that high classification accuracy in distinguishing reasoning types reflects task format confounds rather than genuine computational differences in model hidden states. Sparse autoencoders trained on multilingual data enable principled layer selection for activation steering across languages. Single-layer activation edits can easily corrupt factual recall but rarely repair it, revealing a fundamental asymmetry in how knowledge is stored. Concept-tracking probes allow monitoring what a model is “thinking” during normal operation. Rather than asking what a model attends to, these works ask what happens when we intervene on specific internal components. Machine Ethics & Safety reaches its highest count (8 papers), with a focus on the fragility of current alignment. Researchers prove that static black-box evaluation cannot guarantee post-update alignment: models can pass all safety tests yet become severely misaligned after a single benign gradient update, with the capacity for latent adversarial behavior growing with model scale. The geometry of refusal is shown to be linearly manipulable—a steerable “safety axis” that serves as both vulnerability and defense primitive. Over-refusal is traced to non-harmful linguistic cues in training data that models learn to associate with refusal. Safety behavior proves domain-dependent, with compliance rates varying from 15% (human trafficking) to 86% (surveillance design) across ethical domains. The emerging lesson is that safety is not a scalar property but a domain-conditional, and geometrically localizable feature of model representations. Truthfulness remains the most represented dimension (13 papers), but the problems grow more subtle. Ghost Context formalizes a new failure mode—misattributed grounding—where models use evidence from the wrong part of a long context, producing errors invisible to standard faithfulness metrics. Quantization’s impact on factual knowledge recall is characterized across methods and model scales. Prospective memory failures show that models drop formatting constraints under cognitive load. In this edition we saw focus where truthfulness research moves from detecting hallucinations to understanding why and where factual knowledge degrades within model architectures.

Our longitudinal synthesis reveals both recurring structural patterns and persistent blind spots in the TrustNLP proceedings. We distill four structural insights that emerge from the six-year arc of the workshop. Insight 1: Trust Topics trail Capability Events. We observe that shifts in TrustNLP’s topical focus follow notable events in the broader AI field. The release of the first high-impact chat models (late 2022) preceded the simultaneous activation of Robustness, Truthfulness, and Privacy in the 2023 edition—dimensions that had been absent or minimal in prior years. The availability of frontier models and open-weight alternatives (2023–2024) preceded the emergence of multi-dimensional trade-off research in 2024, where papers began studying conflicts between fairness, explainability, and performance. The rise of agentic and multimodal systems (2024–2025) preceded the 2025–2026 focus on safety alignment and mechanistic interpretability. In each case, the workshop edition following a capability event showed heightened attention to trust challenges that the new capability made salient. It is natural that the community’s research agenda would respond to newly observable failure modes. Insight 2: Trust Evidence Depends on Interaction Mode. What counts as evidence of trustworthiness has varied with how humans interact with models. Under direct model access (2021–2022), trust evidence meant feature attributions and rationale extraction. Under API-mediated access (2023–2024), it shifted to factuality scores and calibration. Under agentic and multi-step use (2025), it became trajectory-level safety guarantees. However, this progression is not purely linear. The 2026 resurgence of explainability (2 papers in 2025 to 13 in 2026) shows that the field can revisit earlier trust questions with new methods—mechanistic interpretability (sparse autoencoders, activation steering, linear probes) addresses the same “why did the model do that?” question as 2021-era attribution methods, but through causal intervention rather than post-hoc correlation. The interaction paradigm shapes which trust questions are most urgent, but earlier questions do not disappear. They return when new methodology makes them tractable again. Insight 3: Outputs ≠ Internals. Output-level audits, the dominant evaluation paradigm across all six editions, systematically underestimate latent model behaviors. This pattern is visible directly within our corpus: fairness audits based on output distributions miss internal representational biases detectable only through probing; reliability benchmarks missed contamination-driven inflation of benchmark scores; adversarial evaluations missed latent attack surfaces discoverable only through automated red-teaming; and function-calling evaluations missed brittleness to toolkit expansion. The broader literature outside TrustNLP corroborates this gap at a mechanistic level: researchers show a lack of alignment between fairness metrics computed on internal representations and on model outputs; the KEEN probe predicts model factuality directly from internal representations without generating any text; and parametric-knowledge-trace probing reveals that unlearning methods minimally alter concept vectors and primarily suppress them at inference time, leaving the underlying knowledge intact and recoverable. The implication is architectural: trustworthy evaluation must incorporate internal probes, not just behavioral tests, as a first-class component, particularly as agentic systems make autonomous decisions based on latent representations that never surface in user-visible outputs. Insight 4: The Need for a Unified Framework. Despite 144 papers across six editions, TrustNLP lacks a unifying theoretical framework that connects its constituent concerns, such as fairness, robustness, factuality, privacy, and calibration. Existing external frameworks such as DecodingTrust and TrustLLM enumerate dimensions but do not formalize their interactions or trade-offs. However, the proceedings themselves reveal deep interdependencies among these dimensions: adapter modules that improve performance and efficiency can come at the cost of fairness, and improving fairness does not reliably improve explainability, or vice versa. A unified framework would need to explicitly account for these trade-offs rather than treating each dimension in isolation, specify how trust requirements change with the interaction paradigm, and incorporate both output-level and internal-level evidence. Without such a framework, the field risks fragmentation: each new capability emergence produces yet another disconnected line of trust research.

In addition to above insights, we observe that the rapid growth in evaluation papers has produced benchmarks that saturate quickly, with several works already questioning existing metrics. The field needs dynamic, adversarially-maintained benchmarks that evolve with model capabilities, rather than static test sets that become training data for the next generation. Benchmark design should also account for the interaction paradigm shifts identified in Insight 2. For example, a benchmark designed for single-turn generation is inadequate for evaluating multi-turn agentic behavior.

Across six editions and 144 papers, TrustNLP shows a field shaped more by external capability emergence than by internal theoretical development: trust paradigms have shifted from interpretability to reliability to controllability in lockstep with the dominant human–AI interaction mode, with a stable multi-month research lag behind each shock. Several findings go deeper. E.g., output-level audits systematically miss latent model behaviors, so trustworthy evaluation must incorporate internal probes as a first-class component. Second, the field’s trust dimensions, such as fairness, robustness, factuality, privacy, and calibration, interact in ways that no single benchmark currently captures. Third, the persistent absence of a unifying framework means that each new capability emergence produces another disconnected line of trust research rather than cumulative theoretical progress. The growth from 8 to 41 archival papers per edition demonstrates sustained community investment. But documenting failures is not the same as preventing them. The field must now move beyond observation. Research should proactively identify modes when models must not fail, then verify mechanistically and not just behaviorally.

Our analysis is bounded by several scope factors that readers should consider when generalizing the findings. First, our analysis covers a single workshop venue (TrustNLP) over six editions; while we argue this venue is representative of the ACL community’s trust agenda, our findings may not generalize to trustworthy-AI work published at FAccT, AIES, ICLR, NeurIPS, or non-archival venues. Second, our synthesis is restricted to archival proceedings papers; non-archival presentations, posters, and tutorial content are excluded. Third, the paper is anglocentric, the cited literature, the lexical analysis, and the example paradigms are drawn primarily from English-language NLP. The fact that multilingual safety remains underexplored is itself a symptom of this bias in the source corpus. Fourth, we identify structural patterns retrospectively. The reactive delay (Insight 1) is an observation across four capability emergence; whether this regularity holds for future shocks is an empirical question we cannot answer. Finally, this paper does not make an experimental contribution. Our claims are descriptive and analytic, drawing on the proceedings record rather than new measurements. Readers seeking benchmark numbers, model evaluations, or method comparisons should consult the individual papers cited.

Improvements for AI systems

Based on the paper, here are specific improvements for AI systems:

1. Internal-state trust probes as first-class evaluation components

  • Add linear probes, sparse autoencoders, and activation-level monitors to standard evaluation pipelines, not just output-level tests.

  • Improved system can detect latent biases, hidden knowledge remnants after unlearning, and misalignment that behavioral tests miss—e.g., flagging a model that passes safety benchmarks but harbors adversarial capabilities in its hidden states.

2. Mechanistic interpretability for causal verification of alignment

  • Use activation steering, concept-tracking probes, and single-layer intervention analysis to verify why a model refuses or complies, rather than only that it does.

  • Improved system can localize the safety axis in representation space, detect over-refusal caused by spurious linguistic cues, and identify domain-conditional safety (e.g., 15% compliance for human trafficking vs. 86% for surveillance) before deployment.

3. Dynamic, adversarial benchmarks that co-evolve with model capabilities

  • Replace static test sets with adversarially-maintained benchmarks that update as models improve, and include multi-turn, agentic, and multimodal interaction modes.

  • Improved system can resist benchmark saturation and contamination, and can evaluate trust under realistic deployment conditions (e.g., database querying, tool use, long-context reasoning) rather than single-turn prompts.

4. Unified trust framework with explicit trade-off modeling

  • Implement a multi-objective optimization layer that jointly optimizes fairness, robustness, truthfulness, privacy, and explainability, with explicit cost functions for conflicts (e.g., adapters improving efficiency but harming fairness).

  • Improved system can surface trade-offs to users (e.g., this model is 10% more accurate but 5% less fair on group X) and automatically adjust behavior based on deployment context.

5. Truthfulness mechanisms targeting root causes, not symptoms

  • Add ghost-context detection (misattributed grounding), quantization-aware knowledge preservation, and prospective memory safeguards under cognitive load.

  • Improved system can identify where factual degradation occurs in its architecture (e.g., which layers lose knowledge after quantization) and self-correct by routing around damaged components or re-anchoring to correct evidence.

6. Proactive controllability for agentic and multimodal systems

  • Build in mechanistic safety constraints (e.g., refusal geometry manipulation, cross-modal attack detection like PBI-Attack) that are verifiable via internal probes, not just output filters.

  • Improved system can constrain its own behavior under adversarial conditions—e.g., resisting jailbreaks that combine image and text perturbations, or refusing to act on malicious long documents—while maintaining benign task performance.

7. Cross-lingual and cross-modal trust generalization

  • Extend fairness and safety evaluation to typologically diverse languages (e.g., Arabic, Farsi, Urdu) and multimodal inputs, with bias metrics adapted per language rather than translated from English.

  • Improved system can detect and mitigate representational harms in non-English corpora and multimodal settings, avoiding the anglocentric blind spot identified in the paper.

8. Confidence calibration with ambiguity detection

  • Use disagreement across candidate answers to detect ambiguous queries before responding, and calibrate confidence only on answerable cases.

  • Improved system can avoid overconfident wrong answers on unclear questions, and can signal when it lacks sufficient evidence—critical for high-stakes domains like healthcare or legal advice.

Abstract

The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mechanistic understanding and proactive control of generative systems. We synthesize insights from all 144 proceedings papers, classifying them along six trust dimensions grounded in established frameworks (TrustLLM, DecodingTrust). We observe co-occurrences with capability emergence. The release of the first high-impact chat models activated all trust dimensions simultaneously, while subsequent model generations shifted focus toward truthfulness and safety alignment. Analysis from the classification study reveals that truthfulness is the fastest-growing dimension (absent in 2021-2022, comprising 37% of papers by 2025-2026), fairness remains the most consistent theme, and explainability exhibits a U-shaped trajectory; declining as post-hoc methods lost relevance but resurging in 2026 through mechanistic interpretability. A cross-venue comparison with ACL, NAACL, EACL, and EMNLP (2K papers) in the same period shows that TrustNLP's topical distribution closely follows the field average. We identify four structural insights and conclude with actionable directions for the research community.

Sources

Related papers