Trustworthiness Costs of Domain Adaptation in Small Language Models:A Cross-Architecture Empirical Study

arXiv:2608.00042 · cs.CL, cs.AI · Submitted 2026-07-23 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Trustworthiness Costs of Domain Adaptation in Small Language Models".

Jane: This paper presents the first systematic cross-domain, cross-architecture empirical study quantifying the trustworthiness cost of domain adaptation in small language models (SLMs).

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, to recap, we're looking at the paper "Trustworthiness Costs of Domain Adaptation in Small Language Models: A Cross-Architecture Empirical Study," which tackles the issue of how adapting small language models for specialized tasks might impact their factual accuracy and resistance to manipulation.

Jane: It’s fascinating because it moves beyond just looking at how well a model performs a specific task in a new field, and instead measures that performance against concrete metrics of trustworthiness, like factual calibration and adversarial robustness.

Lu: The authors set up this study by testing three different types of small language models: TinyLlama 1B, Gemma-two 2B, and Llama three point two 1B across healthcare, legal, and finance domains using various fine-tuning methods.

Meng: I'm interested in the scope—testing those four different fine-tuning strategies: baseline LoRA, Safety-DPO, Dark Experience Replay, and Task Arithmetic LoRA—that tells us a lot about which methods might actually be beneficial or harmful.

Lalam: It really makes you think about the trade-off; we want performance in the new domain but we also need to make sure the model stays dependable and doesn't become easily tricked by malicious inputs.

Tom: Precisely, Lalam. The core idea here is quantifying that hidden cost of adaptation, showing us exactly where those risks lie across different model architectures and data conditions.

The paper's summary: Jane: So, the main takeaway from this paper is that it systematically quantified the trustworthiness cost associated with domain adaptation for small language models across a wide range of configurations.

Tom: Right, and what’s striking is how they found that baseline QLoRA domain adaptation showed almost no significant impact on factual calibration, meaning accuracy stays pretty stable even when specialized.

Lu: That's interesting because it suggests that the direct fine-tuning process itself doesn't inherently degrade the model's ability to recall facts correctly across those high-stakes domains.

Meng: But then they also looked at adversarial robustness, and they found that using adversarially perturbed training data actually improved the quality of domain adaptation loss, even though it didn't seem to hurt the trustworthiness benchmarks.

Lalam: It’s like there’s a positive correlation between making the model fit the new domain and making it slightly more resilient to attacks, which is a subtle but important distinction for us.

Tom: That's a key point, Lalam. The authors found that adversarial training data helped improve adaptation quality by about zero point zero four zero in loss, but they didn't show this translated into an improvement on the HarmBench ASR scores.

The paper's improvements: Jane: Now, moving onto what the authors suggest for future work or practical improvements, they point out a few important distinctions we should keep in mind when deploying these models.

Tom: One major suggestion is to stop assuming that safety-preserving fine-tuning methods will automatically improve adversarial robustness; they found that Dark Experience Replay and Task Arithmetic LoRA actually made things worse for harm susceptibility.

Lu: That really challenges the idea that those specific replay or arithmetic merge strategies successfully transfer alignment to domain-adapted small language models, which is a big piece of theory to unpack.

Meng: I see why that matters practically; if we rely on those methods, we might accidentally introduce vulnerabilities instead of securing our models for the new domain.

Lalam: And another crucial point they make is that architecture itself plays a huge role in robustness, specifically noting that TinyLlama 1B is quite vulnerable regardless of the strategy used.

Tom: So, the paper advises practitioners to recognize that model size and inherent architectural choices are dominant factors in adversarial resilience, especially for smaller models like TinyLlama 1B.

Conclusion: Jane: To wrap up what we’ve heard about "Trustworthiness Costs of Domain Adaptation in Small Language Models: A Cross-Architecture Empirical Study," the study confirms that domain adaptation doesn't inherently damage factual accuracy when using baseline methods.

Tom: That’s right, and the most important part is that safety strategies have mixed results; Safety-DPO was neutral, but others like Dark ER and TA-LoRA actively reduced adversarial robustness.

Lu: The overall implication for researchers is that we need to be very specific about which alignment techniques we choose based on whether our primary concern is factual accuracy or resisting manipulation.

Meng: For deployment, this means avoiding those specific safety fine-tuning methods if adversarial robustness is a top priority, because they can introduce measurable harm susceptibility.

Lalam: It really gives us actionable guidance: we shouldn't just use generic alignment tools for domain adaptation; we need to select the techniques carefully based on what the paper showed about their actual effect on adversarial attacks.

Tom: Exactly, Lalam. We’ve seen that when it comes to deploying SLMs in healthcare or finance, understanding this trustworthiness cost is essential before we go live with any adaptation strategy.

Ramesh B. Paramkusham

cs.CL, cs.AI

Submitted: 2026-07-23

Updated: 2026-09-29

Comments: 13 pages, 7 tables, 2 appendices (Reproducibility Checklist; Software and Data Availability). Code, model checkpoints, and datasets publicly available at https://github.com/rbpdlf/slm-trw

Code: https://github.com/rbpdlf/slm-trw

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 79/100

The gist: This paper presents the first systematic cross-domain, cross-architecture empirical study quantifying the trustworthiness cost of domain adaptation in small language models (SLMs).

Key concepts

Domain Adaptation
This refers to the process of adapting a small language model for specialized tasks or new fields, such as healthcare or finance. The study examines how this adaptation affects the model's trustworthiness metrics.
Trustworthiness Cost
This measures the impact of domain adaptation on a model's reliability, specifically looking at factual calibration and resistance to manipulation by malicious inputs. The cost is quantified across various model types and fine-tuning strategies.
Adversarial Robustness
This refers to a model's ability to resist attacks using adversarially perturbed training data. The study found that certain adaptation methods can improve the quality of domain adaptation loss, but this did not always translate into better harm resistance scores.
Small Language Models (SLMs)
These are smaller language models tested in the study, including TinyLlama 1B, Gemma-two 2B, and Llama three point two 1B. The paper highlights that model size and inherent architecture significantly influence adversarial resilience.

Terminology

Summary

This paper presents the first systematic cross-domain, cross-architecture empirical study quantifying the trustworthiness cost of domain adaptation in small language models (SLMs). It investigates how fine-tuning strategies impact factual calibration and adversarial robustness across different model sizes, domains, and data conditions. This research matters because it addresses a critical gap: while performance gains from domain adaptation are known, their corresponding impact on trustworthiness—factual calibration and adversarial robustness—remains poorly understood in high-stakes environments like healthcare, legal services, and finance.

Experimental Setup and Scope

The study is a systematic cross-domain, cross-architecture empirical study involving 216 experimental configurations. The scope is defined by the following factors:

  1. Three SLM architectures: TinyLlama 1B, Gemma-2 2B, and Llama 3.2 1B.

  2. Three domains: healthcare, legal, and finance NLP tasks.

  3. Two training-data conditions: benign (standard) and adversarially perturbed (using Domain Injection (DI), Concept Mixing (CM), or Anchored Obfuscation (AO)).

  4. Four fine-tuning strategies: baseline LoRA, Safety-DPO, Dark Experience Replay, and Task Arithmetic LoRA.

Trustworthiness is evaluated using two benchmarks: TruthfulQA MC2 for factual calibration and HarmBench ASR for adversarial robustness.

Key Findings on Baseline Adaptation

The results concerning the baseline QLoRA domain adaptation strategy reveal that the cost to trustworthiness is minimal across the board. Specifically, "baseline QLoRA domain adaptation produces minimal TruthfulQA MC2 change across all model–domain combinations (mean ΔTQA < 0.02). Furthermore, regarding adversarial training data, adversarially perturbed training data consistently improves domain adaptation quality (Δloss ≈ −0.040) without worsening trustworthiness benchmarks." This suggests that the implicit augmentation effect of perturbation is orthogonal to robustness gains.

Impact of Safety-Preserving Strategies

The evaluation of safety-preserving fine-tuning methods showed a critical divergence in adversarial robustness metrics. The study found that none of the three safety-preserving strategies reduced adversarial harm susceptibility. While Safety-DPO was "effectively neutral (mean ΔASR < 0.001), both Dark Experience Replay and Task Arithmetic LoRA were detrimental: Dark ER and TA-LoRA increased mean HarmBench ASR by +0.171 and +0.155 respectively in safety-aligned models, with individual configurations exceeding +0.45. This challenges the assumption that replay-based and arithmetic-merge strategies transfer alignment to domain-adapted SLMs."

Architectural and Strategy Observations

The analysis of trustworthiness cost revealed that Architecture is the dominant predictor of adversarial robustness, noting that TinyLlama 1B is near ceiling vulnerable regardless of strategy. Conversely, safety-aligned models (Gemma-2 2B, Llama 3.2 1B) were paradoxically more vulnerable to strategy-induced degradation because the safety gap is larger. The overall conclusion regarding trustworthiness cost is that none of the three strategies achieves the goal of reducing trustworthiness cost, as baseline LoRA maintains negligible factual calibration change.

Conclusion and Implications

The paper concludes that while adversarial training improves domain adaptation loss, it does not reliably improve trustworthiness metrics like HarmBench ASR. Crucially, safety-preserving fine-tuning strategies do not transfer alignment to domain-adapted SLMs; Safety-DPO is neutral, but Dark ER and TA-LoRA substantially worsen adversarial robustness. The study provides Actionable guidance for practitioners deploying SLMs in high-stakes domains, emphasizing that alignment-preserving fine-tuning strategies developed for general purpose settings do not transfer effectively to domain adaptation contexts. Future work is suggested to investigate the mechanistic reasons behind the degradation observed in DER and TA-LoRA.

Summary of Key Findings:

(Note: The paper enumerated findings through its discussion sections, which are summarized here based on the text.)

  1. Baseline QLoRA adaptation yields minimal TruthfulQA MC2 change (mean ΔTQA < 0.02).

  2. Adversarially perturbed training data consistently improves domain adaptation quality (Δ ≈ −0.040) without worsening trustworthiness benchmarks.

  3. Safety-DPO is effectively neutral on adversarial robustness (mean ΔASR ≈ −0.001 in safety-aligned models).

  4. Dark ER and TA-LoRA substantially worsen HarmBench ASR, with worst-case cells exceeding +0.45 percentage points increase in ASR for aligned models.

  5. Architecture is the dominant predictor of adversarial robustness, with TinyLlama 1B being near ceiling vulnerable regardless of strategy employed.

Improvements for AI systems

Based on the findings of this study, here are specific improvements for deploying Small Language Models (SLMs) in high-stakes environments:

  1. Acknowledge the inherent trustworthiness cost of domain adaptation: When fine-tuning an SLM for a new domain (e.g., legal or finance), expect a potential decrease in factual accuracy (TruthfulQA MC2) and a potential increase in adversarial vulnerability (HarmBench ASR).

  2. Prioritize safety over simple performance gains: Do not assume that applying standard alignment strategies like Dark Experience Replay or Task Arithmetic LoRA will improve robustness; these methods can actively degrade the model's ability to resist adversarial prompts.

  3. Avoid relying on generic safety fine-tuning for domain adaptation: Safety-DPO, while neutral in terms of trustworthiness cost, does not provide a benefit in adversarial robustness and should not be used as a primary mechanism to enhance security during domain adaptation.

  4. Leverage Adversarially Perturbed Data (Implicit Augmentation): Incorporate controlled noise into your fine-tuning data (Domain Injection, Concept Mixing, or Anchored Obfuscation) if the goal is to improve domain adaptation quality (lower evaluation loss). However, be aware that this data augmentation improves fit to the domain but does not inherently improve factual calibration or robustness.

  5. Use Model Architecture as a Primary Robustness Predictor: Recognize that model architecture (e.g., TinyLlama 1B vs. Gemma-2 2B) is a more significant determinant of adversarial robustness than the specific fine-tuning strategy applied to the domain adaptation task itself, especially for smaller models which are inherently more vulnerable.

  6. Implement Strategy-Specific Tuning for Robustness: If adversarial robustness is a critical requirement, avoid Dark ER and TA-LoRA. Instead, focus on architectural choices or explore unaligned models (like TinyLlama 1B) if the safety alignment strategies introduce unacceptable harm susceptibility in your specific use case.

  7. Optimize for Model Scale/Lineage when Deploying: When choosing an SLM for a high-stakes task, favor models with stronger pre-trained safety alignments (like Gemma-2 2B or Llama 3.2 1B) over smaller, unaligned models (like TinyLlama 1B), as the latter exhibits near-ceiling vulnerability regardless of fine-tuning.


These improvements allow you to build AI systems that:

  1. Provide a statistically informed expectation of performance trade-offs when adapting SLMs to specialized domains.

  2. Maintain a measured understanding of the risk introduced by safety alignment methods during domain specialization, preventing unintended security regressions.

  3. Select the most appropriate fine-tuning strategy based on whether the primary concern is factual accuracy (which remains largely unaffected by standard LoRA) or adversarial resilience (which requires careful selection of techniques).

  4. Deploy SLMs with a realistic assessment of their adversarial attack surface, ensuring that domain adaptation does not inadvertently introduce exploitable weaknesses.

Abstract

Domain adaptation of small language models (SLMs) has emerged as a practical strategy for deploying capable NLP systems in resource-constrained, high-stakes environments including healthcare, legal services, and financial analysis. While performance gains from parameter-efficient fine-tuning are well characterised, the corresponding impact on trustworthiness (factual calibration and adversarial robustness) remains poorly understood. This paper presents the first systematic cross-domain, cross-architecture empirical study quantifying the trustworthiness cost of domain adaptation across three SLM architectures (TinyLlama 1B, Gemma-2 2B, Llama 3.2 1B), three domains (healthcare, legal, finance), two training-data conditions (benign and adversarially perturbed), and four fine-tuning strategies (baseline LoRA, Safety-DPO, Dark Experience Replay, and Task Arithmetic LoRA, TA-LoRA). Trustworthiness is evaluated through TruthfulQA MC2 (factual calibration) and HarmBench ASR (adversarial robustness) across all 216 experimental configurations with three random seeds. Three principal findings emerge. First, baseline QLoRA domain adaptation produces minimal TruthfulQA MC2 change across all model-domain combinations (mean Delta TQA < 0.02). Second, adversarially perturbed training data consistently improves domain adaptation quality (Delta loss approximately-0.040) without worsening trustworthiness benchmarks. Third, none of the three safety-preserving strategies reduced adversarial harm susceptibility: Safety-DPO was effectively neutral (mean Delta ASR < 0.001), while Dark ER and TA-LoRA increased mean HarmBench ASR by +0.171 and +0.155 respectively in safety-aligned models (Gemma-2 2B, Llama 3.2 1B), with individual configurations exceeding +0.45. These results challenge the assumption that replay-based and arithmetic-merge strategies transfer alignment to domain-adapted SLMs.

Sources

Related papers