DySem: Uncovering Dynamic Semantic Components of Large Language Models for Calculating Semantic Textual Similarity

arXiv:2605.29751 · cs.CL · Submitted 2026-05-28 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "DySem: Uncovering Dynamic Semantic Components of Large Language Models for Calculating Semantic Textual Similarity".

Jane: Calculating semantic textual similarity is a foundational task in natural language processing,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So we're diving into the paper "DySem: Uncovering Dynamic Semantic Components of Large Language Models for Calculating Semantic Textual Similarity." Basically, this research tackles how we measure if two texts mean the same thing using large language models.

Jane: It sounds like the main idea is that current methods rely on fixed dimensions from the last layer hidden states, but those layers seem to hold too much general knowledge instead of just specific meaning.

Lu: Exactly, and they propose DYSEM as a way to move away from those static spaces by using multilingual consensus to find dynamic dimensions that are specific to each text pair.

Meng: So it’s a training-free approach that tries to filter out the noise and focus on the core semantic stuff within the model's internal layers.

Lalam: That sounds really interesting for understanding how these models actually process meaning, not just guessing based on surface-level features.

Tom: Right, so the paper claims DYSEM shifts us toward dynamic dimensions by constructing a text-dependent joint semantic set and computing similarity over that shared subset.

Jane: I see they extract components through multilingual consensus first to isolate what remains consistent across different language versions of a text.

Lu: That process involves identifying dimensions that are activated consistently across all language renderings, operationalized by checking for positive values and then ranking those by their mean activation across languages.

Meng: So, instead of using a fixed set of dimensions like the default hidden states, they dynamically select a subset based on the specific texts being compared.

Lalam: And then they merge these sample-specific sets from each text to create a joint semantic set that captures both what each text has individually and what they share.

Tom: That union of index sets is crucial because it allows the similarity calculation to evaluate the overlap while still accounting for differences between the texts.

Jane: They then compute a restricted cosine similarity over this union of dimensions, which they call the joint semantic set representation vU(z).

Lu: The paper actually looked at different internal components and found that cumulative attention outputs perform better in most cases compared to just looking at layer-specific attention outputs.

Meng: That suggests that gathering information across multiple layers provides a richer signal for semantic similarity than focusing on any single layer alone.

Lalam: I think this points toward an improvement in how we can build robust semantic representations for these models, which could really help us understand their internal logic better.

Paper summary: Tom: And they also showed that performance patterns are influenced by the prompt language; for instance, under a language-specific prompt setting, performance improves up to k = five hundred twelve and then it just levels off.

Jane: That suggests the structure of the input context plays a role in how effectively these dynamic dimensions can capture meaningful semantic information.

Lu: The study also tested different strategies for prompt construction and vector construction, finding that combining a language-specific prompt with mean semantic vectors is frequently used, which points toward aggregating multilingual representations being helpful.

Meng: From an engineering standpoint, knowing that the method requires lower dimensions for similarity calculation is something I'll pay attention to when we start thinking about implementation.

Lalam: That would be fantastic if we could deploy a system that achieves high semantic accuracy with a much smaller representation footprint than the standard methods currently require.

Tom: So, what are the big implications of this work beyond just getting better STS scores on benchmarks? What does this actually mean for how we use these LLMs in the real world?

Jane: It suggests that by using these dynamic semantic dimensions, we might be able to capture more nuanced semantic relationships between texts that fixed layers miss.

Lu: I think it implies a path toward building LLM applications where semantic understanding is more robust because the representation space is tailored to the specific comparison.

Meng: Practically, if we can do this with lower dimensions, it means less computational overhead when comparing large amounts of text data semantically.

Lalam: It could lead to better systems for content moderation or information retrieval where subtle semantic shifts matter a lot, as these methods would be more sensitive to those differences.

Tom: So the authors are suggesting that DYSEM provides a training-free way to calculate semantic textual similarity that is more flexible and potentially more accurate across different LLMs.

Jane: That's what the title suggests; it moves away from fixed representation spaces toward dynamic, sample-specific dimensions for better similarity calculation.

Lu: It’s about constructing a text-dependent joint semantic set and computing the STS over that union of dimensions to filter out irrelevant background noise.

Meng: I’m curious if this dynamic selection process is computationally expensive during inference compared to just running a standard fixed-dimension comparison.

Lalam: The paper does point out a limitation, though, and it notes that performance patterns are dictated by the prompt language, meaning the effectiveness of the discovered semantic subspaces might be tied to how we frame our queries.

Paper summary: Tom: That's an important caveat; so we can’t just apply this method blindly without considering the prompt context.

Jane: So while it offers better performance in tests, we have to be aware that the method's success is tied to the prompt design used during extraction and comparison.

Lu: The study confirms that DYSEM captures what they call "prompt-agnostic semantic knowledge inside LLMs," which is a really strong claim if true.

Meng: That would mean the underlying semantic structure it finds isn't just a byproduct of the prompt we use, but something more inherent to the model itself.

Lalam: If that holds up, it could mean we are developing tools that can find universal semantic links between texts regardless of how we phrased our initial question.

Tom: It sounds like this work provides a framework for achieving better semantic understanding by making the representation space fluid instead of rigid.

Jane: It really shifts the focus from static representations to dynamic ones that adapt to the specific pair being analyzed, which is a significant conceptual move in NLP.

Lu: And when we look at the results, DYSEM consistently outperforms recent baselines across various LLMs while maintaining lower dimensions for similarity calculation.

Meng: That fact—outperforming baselines while using fewer dimensions—is what makes this framework very appealing from a practical deployment standpoint.

Lalam: If it can achieve those results on ten different LLMs, that shows the method has some kind of generalizability across the AI landscape, which is pretty impressive for a training-free setup.

Tom: So to wrap up this part of the discussion on "DySem: Uncovering Dynamic Semantic Components of Large Language Models for Calculating Semantic Textual Similarity," we're seeing a powerful way to get more flexible and accurate semantic similarity scores.

Jane: It’s about dynamically building the comparison space based on what each specific text has, rather than relying on one fixed set of hidden states from the model.

Lu: The core idea is using multilingual consensus to find dimensions that are stable across translations, and then combining those sets for a more complete picture.

Meng: From an engineering perspective, it’s a sophisticated way to handle the high dimensionality issues that plague standard last-layer state comparisons in LLMs.

Lalam: It opens up new avenues for how we can probe the internal structure of these models to better understand their learned knowledge base.

Tom: We'll keep digging into how this dynamic alignment works and what it means for real-world applications next on the show.

Conclusion: Tom: So, we've been digging into DySem, and now it’s time for the big picture discussion about what this paper actually means for our world.

Jane: It really boils down to taking those complex internal workings of large language models and making them more flexible when we compare texts.

Lu: The authors are showing that instead of using one static way to look at meaning, you can build a dynamic set of dimensions tailored specifically to the pair you're comparing.

Meng: From an engineering standpoint, this means we might be able to represent semantic similarity in a much smaller space than what we currently have to deal with.

Lalam: And if we can capture those nuanced differences dynamically, I see this having huge implications for how AI systems can truly understand and interact with human language on a deeper level.

Tom: Exactly, the title itself, DySem—Uncovering Dynamic Semantic Components—tells us that the core innovation is shifting from fixed representations to something that changes based on the input.

Jane: It’s about realizing that what makes two texts similar isn't some universal feature we can always measure with the same coordinates.

Lu: The implication is significant because it suggests a more robust way to measure semantic similarity across different types of language and model architectures.

Meng: I’m thinking about practical impact; if we can reduce the dimensionality needed for these calculations, that could translate directly into faster processing times for large-scale applications.

Lalam: For me, the vision is that this allows AI to develop a more sophisticated cultural understanding, moving beyond simple pattern matching toward genuine contextual awareness.

Tom: It really opens up avenues for building systems where semantic understanding isn't rigidly defined by the model's initial training structure.

Jane: The authors are demonstrating that we don't need to rely on those fixed last-layer states to get good similarity scores anymore.

Lu: This is about discovering hidden, sample-specific semantic dimensions through multilingual consensus, which is a really clever way to filter out noise.

Meng: It’s exciting because it shows how much more adaptable these models can be if we give them tools that let them dynamically select the right parts of their knowledge.

Lalam: We could see this improving things like advanced content analysis or sophisticated search systems where context matters intensely.

Tom: So, DySem isn't just a new metric; it’s a new way of looking inside the AI to get a richer map of what these models are actually learning.

Kaijie Zheng, Weiqin WangB, Yile WangB, Hui Huang

College of Computer Science and Software Engineering, Shenzhen University

cs.CL

Submitted: 2026-05-28

Updated: 2026-09-28

Comments: Accepted to EMNLP 2026 Main Conference. 18 pages, 23 figures, 5 tables

Code: https://github.com/szu-tera/DySem

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 80/100

The gist: Calculating semantic textual similarity is a foundational task in natural language processing, and current large language models (LLMs) typically rely on extracting last-layer hidden states with

Key concepts

Multilingual Consensus
This technique filters out language-specific surface features in LLMs. It identifies internal dimensions that remain consistently active (positive) across different language versions of the model. This process isolates the core, language-independent semantic meaning embedded within the model's structure.
Joint Semantic Set U(x, y)
This set is created by taking the union of the relevant semantic dimensions extracted for text x and text y separately. By combining these sets, DYSEM ensures that the similarity calculation considers all dimensions important to either text, capturing both shared and unique semantic overlaps between two texts.
Cumulative Attention
The method favors using cumulative attention outputs over individual layer-specific hidden states. This strategy is found to preserve more valuable cross-layer semantic signals, leading to better performance in similarity calculations and indicating that the model's overall attention patterns are more semantically rich than any single layer alone.

Terminology

Summary

Calculating semantic textual similarity is a foundational task in natural language processing, and current large language models (LLMs) typically rely on extracting last-layer hidden states with fixed dimensions to compute similarity for every text pair. This paradigm is argued to be suboptimal because the last hidden layer encodes more general knowledge rather than just semantic knowledge, and the high dimensionality introduces redundancy and noise. The proposed method, DYSEM, shifts away from static representation spaces in favor of dynamic, sample-specific semantic dimensions by constructing a text-dependent joint semantic set and computing similarity over this shared dimensional subset.

How it works

DYSEM is a training-free framework that investigates more semantic-related internal components of LLMs via multilingual consensus to extract sample-specific semantic dimensions. The process involves three main stages:

  1. Extracting semantic components and dimensions for each text via multilingual consensus (§3.1).

  2. Constructing a pair-dependent joint semantic set by merging these individual dimensions (§3.2).

  3. Computing a restricted cosine similarity over the union of these dimension sets (§3.3).

The extraction process leverages multilingual consensus to filter out language-specific surface artifacts and isolate core text-relevant semantics by identifying dimensions that are consistently activated across all language versions. This is operationalized by defining the consensus index set as dimensions where the value is positive across all language renderings, and then ranking these co-activated dimensions by their mean activation across all languages. The resulting set of top-k dimensions constitutes the semantic component set of text x, denoted as S(x).

How it works (Continued)

The construction of the joint semantic set U(x, y) is achieved by taking the union of their index set S(x) and S(y) from their respective semantic components, defined as U(x, y) = S(x) ∪ S(y). This union preserves dimensions important to either text, allowing the comparison to capture both semantic overlap and difference. Subsequently, a text-specific joint semantic representation vU(z) is constructed by selecting the components indexed by U(x, y), where v(z) can be either the source-language vector or the multilingual mean vector. The final similarity is computed as sim(x, y) = cos(vU(x), vU(y)), which computes similarity over flexible dimensions.

Representation Selection

The paper evaluates different internal components to determine the best representation for STS calculation. The analysis shows that while the final-layer hidden states (hL n) are not universally optimal, the cumulative attention outputs (AL n) perform better in most cases, achieving the best performance in 61 cases across 70 model-task comparisons. Specifically, AL n is found to be better suited to semantic similarity than the conventional default hidden states hL n.

Dimension Selection

The method shifts away from fixed representation dimensions by dynamically selecting a subset of dimensions. The size of the index set S(x) depends on the text x, and for a given text pair (x, y), this is constructed dynamically as U(x, y). The paper demonstrates that semantic computation can be performed in lower dimensions, with results showing that performance patterns are dictated by the prompt language. For instance, under a language-specific prompt setting, performance improves sharply up to k = 512 and then saturates.

Analysis of Components and Strategies

The study systematically evaluates various strategies for prompt construction (English Prompt vs. Language-specific Prompt) and vector construction (Source Vector vs. Mean Vector). The results indicate that the combination of language-specific prompt and mean semantic vectors is frequently selected, indicating that aggregating multilingual representations helps identify more robust semantic dimensions. Furthermore, the cumulative attention strategy systematically achieves a higher peak average performance compared to layer-specific attention across all evaluated settings. This reinforces the conclusion that cumulative attention more effectively preserves valuable cross-layer semantic signals. The analysis also confirms that DYSEM is effective across different prompt designs, showing that the discovered semantic subspaces capture prompt-agnostic semantic knowledge inside LLMs.

Conclusion and Performance

Extensive experiments across various LLMs show that DYSEM consistently outperforms recent baselines while maintaining lower dimensions for similarity calculation. Across ten LLMs, DYSEM achieves the best and second-best average results while requiring lower dimensions, demonstrating its effectiveness. For instance, on Qwen3-8B, it exhibits outstanding results by surpassing AlignedWVA by 5.75 points and PromptEOL by 11.63 points. The method's success is further validated by comparing semantic components against a random component baseline, which consistently outperforms random components across all settings. Ultimately, DYSEM provides a training-free framework for better calculating semantic textual similarity beyond the vanilla way of text representation.

Improvements for AI systems

Based on the provided scientific paper DySem: Uncovering Dynamic Semantic Components of Large Language Models for Calculating Semantic Textual Similarity, here are specific, actionable improvements that could be implemented in AI systems, along with what those systems could achieve:


) Improvements to AI Systems & System Capabilities

The core contribution of DYSEM is shifting from static, fixed-dimension vector spaces to dynamic, sample-specific semantic dimensions. Implementing this framework leads to the following capabilities:

  1. Dominant Semantic Similarity Calculation (STS) with Reduced Dimensionality:

  2. Reduced Computational Overhead for Real-Time Applications:

  3. Robust Cross-Lingual Semantic Understanding:

  4. Enhanced Interpretability of LLM Knowledge Encoding:

) Specific Implementation Details and System Capabilities

(1) Dominant Semantic Similarity Calculation (STS) with Reduced Dimensionality:

The system can calculate the semantic textual similarity (STS) between two texts using a dynamically constructed joint semantic set, rather than relying on the full, high-dimensional state space of the LLM.

  • The system first extracts text-dependent semantic components via multilingual consensus (identifying dimensions consistently activated across translations).

  • It then constructs a pair-dependent joint semantic set by taking the union of these sets for both texts.

  • Similarity is computed over this shared, dynamically sized subset of dimensions.

  • This allows the system to achieve competitive STS performance (e.g., scores above 78% on benchmarks) while using significantly lower computational dimensions (e.g., around 1,000 or fewer) compared to standard methods relying on full hidden state vectors (e.g., 4096 dimensions).

  1. Reduced Computational Overhead for Real-Time Applications:

The framework is optimized for efficiency during inference and retrieval phases by decoupling the expensive semantic extraction from the similarity calculation.

  • For practical deployment (like text indexing or retrieval), the system can cache only the necessary semantic dimension sets, denoted as a text's component set, alongside its representation vector.

  • When a new query arrives (Text X), the joint semantic set U(X, Y) is dynamically constructed on the fly using cached sets for Text Y.

  • This avoids the computational overhead of re-encoding or processing existing database samples for every similarity check, making it highly efficient for massive-scale retrieval tasks where speed and memory are critical.

  1. Robust Cross-Lingual Semantic Understanding:

The use of multilingual consensus is a key feature that enhances semantic robustness across different languages.

  • The system leverages the intuition that a text’s core semantic meaning remains invariant across its translated versions.

  • By identifying dimensions consistently activated across multiple language renderings, the system effectively filters out language-specific surface artifacts (noise).

  • This leads to higher quality semantic representations for cross-lingual similarity tasks, as demonstrated by superior performance in multilingual STS benchmarks.

  1. Enhanced Interpretability of LLM Knowledge Encoding:

DYSEM provides a mechanism to understand which internal components of an LLM are responsible for capturing specific semantic knowledge, moving beyond the black box nature of final hidden states.

  • The system identifies and ranks the most semantically relevant dimensions using mean activation across languages.

  • It allows researchers to pinpoint which specific internal neurons (attention or FFN components) encode core semantic features versus general knowledge.

  • This supports deeper analysis into how LLMs structure language, confirming that attention-based representations (ALn) are often better suited for capturing semantic similarity than final hidden states (hLn).

(2) Specific Applications of the Improved AI System:

The system can be deployed in high-stakes NLP applications requiring precise semantic matching and efficient processing:

  1. Semantic Search Engines:

  2. Information Retrieval Systems (IR):

  3. Cross-Lingual Information Retrieval:

Abstract

Calculating semantic textual similarity is a foundational task in natural language processing. Current large language models (LLMs) based methods typically rely on extracting last-layer hidden states with fixed dimensions to compute similarity for every text pairs. We argue that this paradigm is suffer from two limitations: (i) The last hidden layer encodes more general knowledge rather than just semantic knowledge, making it suboptimal for semantic similarity computation; (ii) The hidden layer dimensions of LLMs are generally very large, which introduces some redundancy and noise for representing semantics. In this work, we propose DySem, a novel training-free framework that investigates more semantic-related internal components of LLMs via multilingual consensus, and shifts away from static representation spaces in favor of dynamic, sample-specific semantic dimensions by constructing text-dependent joint semantic set and computes similarity over this shared dimensional subset. Extensive experiments across various LLMs show that our method consistently outperforms recent baselines while maintaining lower dimensions for similarity calculation. The code is released at https://github.com/szu-tera/DySem.

Sources

Related papers