NMIXX: Domain-Adapted Neural Embeddings for Cross-Lingual eXploration of Finance

arXiv:2507.09601 · cs.CL, cs.AI, q-fin.CP · Submitted 2025-07-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "NMIXX: Domain-Adapted Neural Embeddings for Cross-Lingual eXploration of Finance".

Tom: General-purpose sentence embedding models often struggle to capture specialized financial semantics—especially in low-resource languages like Korean—due to domain-specific jargon, temporal meaning shifts, and misaligned bilingual vocabularies.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So, we’re talking about the paper titled "NMIXX: Domain-Adapted Neural Embeddings for Cross-Lingual eXploration of Finance," and looking at the authors, we see a team from various financial institutions and research labs working together on this specific problem. It clearly shows they’re tackling how to make AI better at understanding finance when dealing with different languages like Korean and English simultaneously.

Jane: That makes sense, Tom; when I look at the title, it suggests they are not just making a general language model more knowledgeable about money, but specifically engineering embeddings that handle the unique way financial concepts change across those two languages. It’s about building a bridge between Korean finance and English finance representations.

Lu: The authors are clearly pointing out that existing models fall short because they don't account for the specific jargon or the way time affects language in financial reporting, which is a big issue when dealing with low-resource languages like Korean twelve twenty-six. They are responding to the fact that general embeddings just aren't sensitive enough to those kinds of linguistic shifts.

Meng: That need for specialization is significant. I wonder how they managed to pull together that massive multilingual corpus and then filter it down into the exact data needed for this fine-tuning process before they could even begin training the core NMIXX model itself. It sounds like a lot of careful preparation was required upfront.

Lalam: The focus on cross-lingual exploration really opens up big potential for making financial services accessible globally; it means they are building a system that understands the relationship between financial ideas when you switch between Korean and English contexts.

The paper's summary: Tom: Moving into the actual summary of "NMIXX: Domain-Adapted Neural Embeddings for Cross-Lingual eXploration of Finance," the main point is that they introduced NMIXX as a suite of cross-lingual embedding models that they tuned using eighteen point eight K high-confidence triplets. These triplets are specifically crafted to pair sentences that are contextually similar, alongside deliberately tricky ones, including exact Korean to English translations and hard negatives based on semantic shifts.

Jane: That sounds incredibly systematic, Tom; they aren't just training on random text randomly. They use a very specific method of pairing similar financial sentences with intentionally challenging ones that have slight meaning shifts or different formal tones, which helps the model learn the subtle differences in financial language structure.

Lu: What I find really interesting is how they structured their training data around a fine-grained typology of semantic shifts based on four core document types: temporal variation from news, perspectival framing from research reports, structural formality from regulatory disclosures, and logical semantics from legal texts twelve twenty-six.

Meng: That level of detail in defining the shift patterns is impressive; it shows they didn't just look at surface words but tried to map out the actual cognitive and contextual changes that happen when a financial statement moves from one format to another. That’s where I see the practical application for real-world analysis.

Lalam: It’s really about capturing the intent behind the words, not just matching keywords; by understanding those different ways meaning can shift, like distinguishing between "strong growth" and "solid growth," the embedding model gets a much richer picture of what's actually happening in the financial story.

The paper's improvements: Tom: When we look at what NMIXX actually improves, it’s clear they designed this two-part system—a specialized model and a benchmark called KorFinSTS—to tackle the issue of domain specificity in low-resource settings. They used triplet fine-tuning combined with multilingual positive examples and domain-balanced training data to achieve these specific results.

Jane: The improvement they highlight is that this framework allows for high performance gains specifically on the KorFinSTS benchmark, showing a Spearman’s rho gain of plus zero point two two compared to earlier baselines, which is a fairly significant number when you are working with specialized data like Korean finance texts.

Lu: They also pointed out that the process for creating those high-quality training triplets involved mining them at scale using an LLM pipeline that used GPT-4o tagging for semantic shift patterns, generating hard negatives based on those patterns, and then validating pairs where an LLM judge scored them on a zero to ten scale twelve twenty-six.

Meng: The methodology for creating those high-quality triplets is what really catches my eye; using an LLM as a judge to validate semantic divergence before selecting the best pairs suggests a very rigorous way to filter out low-quality data and focus the fine-tuning effort on the most informative examples.

Lalam: That systematic construction of those eighteen point eight K high-confidence triplets is what makes this work so solid; it shows they prioritized quality over just having a lot of data, which is essential when you’re trying to teach a model deep domain knowledge.

Conclusion: Tom: So, wrapping up our discussion on "NMIXX: Domain-Adapted Neural Embeddings for Cross-Lingual eXploration of Finance," the main point is that by carefully designing and fine-tuning embeddings with this specific triplet method against diverse financial texts, they've managed to capture specialized meanings in Korean finance effectively where general models struggle. This gives us a tool that actually works where others fail.

Jane: It really boils down to having a model that can accurately compare two different pieces of Korean financial text and tell you if their core meaning aligns, even when the language or the specific terminology is slightly different across those texts. It’s about understanding the underlying financial message, not just matching words.

Lu: The implications for cross-lingual finance are substantial because this work demonstrates a workable path for taking knowledge gained in one market context and applying it effectively to another, overcoming the usual hurdles of language misalignment in specialized fields.

Meng: From a practical standpoint, this means that financial analysis tools built on these embeddings can become much more robust when dealing with various global data sources because they won't get confused by the linguistic noise coming from different regional reports.

Lalam: For us, the impact is cultural; it means we can build AI assistants that truly understand the complex way people discuss money and markets in Korean, leading to more sophisticated and helpful interactions for users in that language.

Tom: Fantastic discussion on NMIXX today! It’s clear this work provides a concrete method for improving cross-lingual financial understanding. We’ll be keeping an eye on how this specialized embedding approach plays out as we look at the next set of papers coming from arXiv.

Hanwool Lee, Sara Yu, Yewon Hwang, Jonghyun Choi, Heejae Ahn, Sungbum Jung, Youngjae Yu

FinancialNLPLab, MODULABS (Shinhan Securities) · FinancialNLPLab, MODULABS (KT) · EMRO (Seoul University) · FinancialNLPLab, MODULABS (Samsung Fire & Marine Insurance KB Securities)

cs.CL, cs.AI, q-fin.CP

Submitted: 2025-07-13

Updated: 2026-09-30

Comments: 8 pages, 6 figures

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 76/100

The gist: General-purpose sentence embedding models often struggle to capture specialized financial semantics—especially in low-resource languages like Korean—due to domain-specific jargon, temporal

Key concepts

KorFinSTS
This is an enhanced Semantic Textual Similarity benchmark specifically designed to test sentence understanding within the financial domain in Korean. It covers various document types like news and regulations, helping researchers expose subtle semantic differences that standard benchmarks often miss.
Semantic Shift Typology
This framework categorizes how meaning changes across different financial texts. It includes shifts based on time (news), perspective (research reports), structure (regulations), and logic (legal texts). This allows the model to understand context-specific nuances in finance.
High-Confidence Triplets
These are carefully curated training examples consisting of a source sentence, a semantically equivalent positive paraphrase, and a hard negative. They were generated using advanced LLMs to ensure high quality, teaching the model precise financial relationships.
Tokenizer Importance
The study found that the success of adapting models for low-resource languages heavily depends on the tokenizer. Models with extensive native vocabulary coverage (like Korean tokens) performed significantly better, indicating that vocabulary richness is a key architectural factor.

Terminology

Summary

General-purpose sentence embedding models often struggle to capture specialized financial semantics—especially in low-resource languages like Korean—due to domain-specific jargon, temporal meaning shifts, and misaligned bilingual vocabularies.

The gist: NMIXX is a suite of cross-lingual embedding models fine-tuned with 18.8 K high-confidence triplets that pair in-domain paraphrases, hard negatives derived from a semantic-shift typology, and exact Korean↔English translations to achieve Spearman’s ρ gains of +0.22 on KorFinSTS while revealing a modest trade-off in general STS performance.

The Problem Addressed

Sentence representation learning faces challenges when applied to specialized fields like finance, where domain-specific linguistic phenomena and vocabulary shifts can substantially degrade embedding quality. This degradation is particularly pronounced in low-resource language contexts. Prior work shows that off-the-shelf embedding models frequently underperform in specialized settings due to issues such as market jargon, temporal semantic drift, and regulatory formality. The gap exists because general embeddings often lack sensitivity to these domain-specific nuances.

The NMIXX Framework

NMIXX is a suite of cross-lingual embedding models designed to capture the unique semantics of financial texts in a Korean-English context. This framework is built upon two main components:

  1. A specialized embedding model, Neural eMbeddings for Cross-Lingual Exploration of Finance (NMIXX), which was informed by semantic shift patterns extracted from a large-scale multilingual financial corpus.

  2. An enhanced Semantic Textual Similarity (STS) benchmark called KorFinSTS, designed to expose nuances that general benchmarks miss by spanning news, disclosures, research reports, and regulations.

Data Construction and Training Strategy

The construction of the training data for NMIXX involved several rigorous steps:

  1. Collecting a comprehensive corpus of financial-domain texts encompassing diverse register variations from six openly licensed corpora (totaling 2.46M document–level records prior to filtering).

  2. Performing Domain-balanced augmentation by scraping Korean and U.S. regulatory filings and commissioning GPT-4o rewrites to create synthetic documents mirroring formal disclosure style, resulting in a pool of 46.1k sentences after quality control.

  3. Mining high-quality training triplets at scale using an LLM pipeline involving:

(1) Axis identification via GPT-4o tagging of semantic shift patterns.

(2) Hard-negative generation by synthesizing semantically divergent yet lexically similar variant[s] guided by the chosen pattern.

(3) Hard-negative validation where GPT-4.5 (LLM-as-judge) scores pairs on a 0–10 scale, retaining only those with scores ≥ 8.

(4) Positive generation and validation where GPT-4o produces a semantically equivalent paraphrase, kept only if the LLM-as-judge scores the pair ≥ 9.

This process yielded 18.8k high-confidence triplets (source, positive, hard negative) for supervised fine-tuning using a temperature-scaled triplet negative-log-likelihood loss.

Semantic Shift Typology

The model's training data was structured around a fine-grained typology of semantic shifts derived from four core document types:

  1. Temporal variation from Financial News (modeling real-time market sentiment and evolving narratives).

  2. Perspectival framing from Investment research reports (targeting distinctions such as micro vs. macro analysis and facts vs. opinions).

  3. Structural formality and consistency from Regulatory disclosures (targeting patterns like intensified sentiment or plan realization).

  4. Logical and rule-based semantics from Legal & regulatory texts (focusing on shifts like legal interpretation shifts or shifts in sanction application).

Evaluation and Key Findings

The models were evaluated against seven open-license baselines, including the multilingual bge-m3 variant. The results demonstrated significant performance gains:

(1) FinSTS Gains:

(2) KorFinSTS Gains:

The multilingual bge-m3 achieved a Spearman’s ρ gain of +0.0998 on FinSTS and +0.2220 on KorFinSTS compared to pre-adaptation baselines, which is the highest average improvement among all tested models.

Tokenizer Importance

A critical finding highlighted the importance of tokenizer design for low-resource, bilingual adaptation: models with more extensive native-language vocabulary (in this case, Korean) were substantially more responsive to our fine-tuning method. Specifically, bge-m3 contained over 5,400 full Korean tokens (2.17% of its vocabulary), whereas models like gte-Qwen2-1.

Improvements for AI systems

Here are specific, actionable improvements for AI systems based on the NMIXX framework and KorFinSTS benchmark:

  1. Improved Cross-Lingual Financial Semantic Understanding:

A domain-adapted embedding model (like the multilingual bge-m3 variant fine-tuned with NMIXX) can be deployed to perform high-fidelity sentence similarity tasks in Korean financial documents, achieving significant performance gains (+0.22 on KorFinSTS). This allows AI systems to accurately determine if two different Korean financial statements, news articles, or regulatory filings convey the same core meaning despite variations in jargon or framing.

  1. Enhanced Low-Resource Language Financial QA and Retrieval:

The system can be specifically fine-tuned on the KorFinSTS benchmark (comprising news, disclosures, reports, and regulations) to answer complex financial questions posed in Korean. Unlike general models that fail due to domain jargon or low-resource limitations, this specialized model will retrieve relevant financial documents with high precision from a corpus of Korean text.

  1. Nuance-Aware Sentiment Analysis in Finance:

By leveraging the semantic shift typology (temporal variation, perspectival framing, structural formality), the system can move beyond simple keyword matching to understand nuanced sentiment. For example, it can differentiate between strong growth and solid/steady growth in Korean news by recognizing the specific contextual patterns learned during training.

  1. Robust Cross-Lingual Transfer for Financial Data:

AI systems trained with NMIXX can effectively transfer knowledge from English financial datasets (like FinSTS) to Korean financial datasets (like KorFinSTS). This capability is crucial for building global financial models where insights discovered in one language market are directly applicable to another, mitigating the performance degradation typically seen when directly translating or transferring non-domain-adapted models.

  1. Optimized Tokenization Strategy for Non-English Models:

For developing new multilingual embedding architectures (e.g., future LLM2Vec variants), the research indicates that incorporating a higher proportion of full Korean tokens into the vocabulary (as seen with bge-m3's 2.17% coverage) is critical for achieving stable cross-lingual alignment and robust performance in low-resource settings. This suggests an architectural improvement where tokenizers are explicitly designed or augmented to favor comprehensive coverage of target language vocabulary.

Abstract

Financial text embeddings must distinguish changes in event status, perspective, and obligations even when passages share similar wording. NMIXX adapts existing encoders through 18.8k source-linked triplets: paraphrases and Korean-English translations preserve meaning, while targeted financial rewrites introduce semantic contrasts. We examine this recipe across seven backbones on English and Korean financial and general-domain semantic textual similarity (STS), and analyze the composition and passage lengths of KorFinSTS. BGE-M3 attains the highest adapted financial correlations in this comparison, improving from 0.1969 to 0.2967 on FinSTS and from 0.0512 to 0.2732 on KorFinSTS. Its general English and Korean correlations decrease by 0.0391 and 0.0463. Across the seven models, five improve their mean financial correlation, but all reduce their mean general-domain correlation. Per-language comparisons and benchmark-weight sensitivity analysis reveal differences obscured by a single aggregate score. The study contributes a finance-specific supervision design and evidence for evaluating adaptation jointly with retained general semantic capability; direct cross-language retrieval remains outside its evaluation scope.

Sources

Related papers