RT-SFT: Text Style Transfer from Non-Parallel Corpora by Roundtrip Translation

arXiv:2602.15013 · cs.CL · Submitted 2026-02-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "RT-SFT: Text Style Transfer from Non-Parallel Corpora by Roundtrip Translation".

Jane: This study proposes a novel method for Text Style Transfer (TST) that adapts Large Language Models (LLMs) to transfer text from an arbitrary domain to a target style using only…

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, looking at the conclusion of "RT-SFT: Text Style Transfer from Non-Parallel Corpora by Roundtrip Translation," it really boils down to using roundtrip translation to create a style-neutral pseudo-parallel corpus from monolingual text, which then allows for supervised finetuning of LLMs for Text Style Transfer <ref:2602.15013#pg1>.

Jane: The authors emphasize that this method shows consistent superiority over existing approaches like zeroshot prompting and fewshot ICL when measured by both BLEU scores and style accuracy scores across four investigated domains <ref:2602.15013#pg0>.

Lu: They also highlight the role of the RAG integration in enhancing robustness, particularly when it comes to handling unseen or complex style domains during inference <ref:2602.15013#pg2>.

Meng: It seems like they’ve successfully built a pipeline that addresses the scarcity of parallel corpora by creating its own synthetic training material, and using parameter efficient LoRA finetuning makes it accessible for smaller models <ref:2602.15013#pg0>.

Lalam: I think the real implication here is that we can start building models that are inherently more adaptable because they’ve learned a universal style space first, which could really improve how AI interacts with different human communication patterns <ref:2602.15013#pg0>.

Tom: Exactly; the title itself, "RT-SFT: Text Style Transfer from Non-Parallel Corpora by Roundtrip Translation," perfectly captures the essence of how they solved the data problem using translation as a bridge <ref:2602.15013#pg0>.

Jane: In simple terms, they've shown that you don't need perfectly aligned parallel text to teach an LLM to change style; you can generate it from existing text and use it for supervised learning instead <ref:2602.15013#pg1>.

Lu: The finding that RT-first inference significantly enhances performance when dealing with stylistically diverse queries shows how crucial the pre-processing step is for achieving high quality output <ref:2602.15013#pg2>.

Meng: From a practical standpoint, the stability they found in similarity-based finetuning compared to random n-shot finetuning suggests this method is more reliable for deployment than some other training strategies <ref:2602.15013#pg0>.

Lalam: The paper also points out that their reliance on Marian NMT models for the roundtrip translation pipeline is a limitation, and they admit that post-processing steps are needed to mitigate semantic drift and error propagation from the synthetic data <ref:2602.15013#pg0>.

Tom: That's a fair point, Lalam; acknowledging those limitations about semantic drift means we have to be careful when applying this in production because the translation quality feeds directly into the style transfer result <ref:2602.15013#pg0>.

Jane: So, while the method is promising for overcoming data scarcity, we do need to keep an eye on those potential errors introduced during that roundtrip translation process when using RT-SFT <ref:2602.15013#pg0>.

Conclusion: Tom: So, we've been looking at how these authors tackled the challenge of teaching Large Language Models to change text styles without having massive amounts of parallel data, and now we're getting to the conclusion of this RT-SFT paper.

Jane: I think focusing on the title, "RT-SFT," really gets to the heart of what they did—using roundtrip translation as a shortcut instead of needing huge datasets.

Lu: Exactly; by synthesizing those 'style-neutral' texts through translation, they managed to create a synthetic parallel corpus that we could then use for supervised finetuning. That's incredibly creative thinking from the authors.

Meng: From an engineering standpoint, it’s interesting how they managed to freeze the main model weights and only train those low-rank matrices with LoRA; that makes it much more feasible for smaller models to adapt.

Lalam: I see this as a major cultural shift because if we can effectively transfer text between drastically different registers—say, formal academic language to casual social media tone—it opens up ways for AI to interact much more naturally across all human communication settings.

Tom: That's the big picture, Lalam; it’s not just about technical accuracy, it's about making AI communication more versatile and less rigid.

Jane: And when you look at the authors who put this together, they clearly understood the data bottlenecks in current style transfer research and found a clever way around them.

Lu: Their methodology of using NMT models for that initial roundtrip translation step is quite elegant; it's a clever way to bridge those stylistic gaps using existing bilingual resources.

Meng: But we have to keep an eye on those limitations they mentioned, especially concerning semantic drift and error propagation from the synthetic data they created, which will dictate how robust this method is in real-world applications.

Lalam: Those limitations are important; understanding where the synthetic data might introduce inaccuracies helps us design better mitigation strategies moving forward.

Tom: So, we've seen how they built this system from scratch using translation and then fine-tuned an LLM on it, and now we have a clearer sense of what this means for practical style transfer applications.

Jane: It really shows that even when the ideal data isn't available, creative use of existing tools like machine translation can lead to powerful new training techniques.

Lu: The implications are huge because it lowers the barrier for applying advanced LLM capabilities to specialized stylistic domains that we currently struggle with.

Meng: I wonder how much computational overhead these retrieval augmentation steps add compared to just fine-tuning on raw data, but the results suggest it pays off in terms of quality when dealing with unseen styles.

Lalam: Ultimately, this work suggests we can move toward a future where AI can communicate not just *what* to say, but precisely *how* to say it for any given context.

Ruoxi Liu, Philipp Koehn

Johns Hopkins University

cs.CL

Submitted: 2026-02-16

Updated: 2026-10-01

Importance score: 83/100

The gist: This study proposes a novel method for Text Style Transfer (TST) that adapts Large Language Models (LLMs) to transfer text from an arbitrary domain to a target style using only monolingual corpora

Key concepts

Roundtrip Translation
This involves using two neural machine translation models—one for each direction—to translate a monolingual text into a pivot language and then back again. This process is used to create paired examples: the original in-style text is paired with its style-neutral equivalent derived from the roundtrip process, effectively synthesizing parallel data where none existed.
Pseudo-parallel Corpus
This is a synthetic dataset created using roundtrip translation. By translating monolingual texts into and out of a neutral pivot language, the method generates pairs of text that are stylistically consistent but not directly parallel in the original languages. This allows an LLM to learn style transfer patterns from this newly constructed, artificial parallel data.
Retrieval Augmented Generation (RAG) for TST-LLM
RAG is integrated into the finetuning and inference stages to improve the LLM's performance. During training, it retrieves similar target-side examples to create better instruction pairs. At inference, a 'sketch-first' method uses initial random examples to guide a second retrieval step, ensuring the model finds highly relevant context for generating the final style-transferred output.
Parameter Efficient LoRA Finetuning
LoRA is a technique used to finetune large language models efficiently. Instead of retraining all model weights, it freezes the main pre-trained weights and introduces small, trainable low-rank decomposition matrices. This allows for effective finetuning of large models (like 7B or 8B) using significantly less computational power and memory.

Terminology

Summary

This study proposes a novel method for Text Style Transfer (TST) that adapts Large Language Models (LLMs) to transfer text from an arbitrary domain to a target style using only monolingual corpora and roundtrip translation. This approach addresses the scarcity of parallel corpora by synthesizing 'neutralized' texts, creating a shared input style at training-time and inference-time, which is shown to consistently outperform state-of-the-art methods like fewshot In Context Learning and Automatic Post-Editing.

The gist

Roundtrip translation allows for the synthesis of a style-neutral to target domain pseudo-parallel corpus from monolingual in-style corpora, enabling supervised finetuning of LLMs for TST tasks where bitext is lacking.

How it works

  1. A pair of neural machine translation (NMT) models is trained using a large-scale generic bilingual parallel corpus between English and a selected pivot language to constitute the roundtrip translation pipeline.

  2. This pipeline is then used to roundtrip translate a monolingual, stylistically consistent corpus using the pretrained NMT models to construct a style-neutral to target-domain pseudo-parallel corpus. This synthetic parallel dataset pairs each in-style text with its style-neutral equivalent.

  3. A Large Language Model (LLM) is then finetuned on this synthetic parallel dataset specifically for the task of MT-destylized text to in-style text.

Retrieval Augmentation for TST-LLM

The framework integrates Retrieval Augmented Generation (RAG) into both the finetuning and inference stages to enhance the LLM’s adaptability.

  1. During finetuning, similarity-based retrieval is used where we take its target-side text, search for top-k most similar examples excluding itself, and look up the source-side counterparts of these retrieved sentences to form example transfer sentence pairs. This focuses on retrieving examples with relevant target-sides for better instruction quality.

  2. At inference time, a sketch-first example retrieval method is employed. The model first performs few-shot inference with randomly selected examples to generate a sketch output, which is then used as the query to retrieve examples with high similarity from the vector bank for a second-round inference yielding the refined output.

  3. Terminology and name retrieval involves constructing a domain-specific termpair bank by prompting an LLM twice for each instance in the synthetic parallel corpus: first, to identify domain-specific terms or names from the source side, and second, to find their counterparts in the target side sentence. This results in adding relevant domain-specific term instructions to prompts when some trigger words are present.

Parameter Efficient LoRA Finetuning

The method utilizes Low-Rank Adaptation (LoRA) for parameter-efficient finetuning of LLMs. This approach involves freezing the pre-trained model’s weight matrices and introducing trainable low-rank decomposition matrices into the model’s layers, allowing for finetuning of models like 7B and 8B LLMs with limited computational resources. The training hyperparameters include a learning rate of 2e-4, a rank for the low-rank approximation set to 512, and a scaling factor of 256.

Evaluation Metrics and Results

Performance is evaluated using two primary metrics: BLEU scores to assess content preservation and BERT-based style classifiers trained on held-out in-domain data to yield a style classification accuracy score. Experiments across four distinct styles (IRS, Literary, Treasury, NCBI) compared the proposed method against baselines such as Few-shot In Context Learning (ICL) and Automatic Post-Editing (APE). Results show that the RT-first inference method yields noticeably better generation quality when dealing with unseen text styles, bringing considerable improvements facing stylistically diverse and complex queries. Furthermore, similarity-based finetuning was found to be much more stable than random n-shot finetuning, yielding up to a 12.22 increase in BLEU score and 0.191 increase in style classification accuracy across the four tested domains. The integration of terminology retrieval demonstrated a 7.29% average improvement on the Acc. score for 5-shot finetuning.

Limitations

The main limitations identified are semantic drift and error propagation, as inaccuracies from roundtrip translation can be embedded in the synthetic parallel corpus used for finetuning. Additionally, the study's reliance on Marian NMT models for roundtrip translation is noted, and alternative workflows involving LLMs for machine translation were not tested due to time constraints. Finally, experiments are conducted on six style domains, which may not fully capture the range of stylistic variations in real-world scenarios. The authors also note that post-processing steps to mitigate such effects are beyond the scope of this study.

Improvements for AI systems

Here are the specific improvements to AI systems based on this scientific paper, detailing what the improved system can achieve:


  1. The proposed method allows for Text Style Transfer (TST) in domains where parallel corpora are scarce, using only monolingual data and roundtrip translation to synthesize pseudo-parallel datasets.

  2. The LLM fine-tuning process is made parameter-efficient using Low-Rank Adaptation (LoRA), enabling finetuning of large models (e.g., 7B and 8B LLMs) on consumer or accessible hardware with reduced computational overhead compared to full fine-tuning.

  3. The system can perform style transfer tasks—rephrasing text while preserving core semantics and intent—across diverse stylistic attributes such as formality, attitude, verbosity, and preferred terminology (e.g., Shakespearean vs. Informal).

  4. The roundtrip translation pipeline acts as a neutralizer, transforming input text from various in-domain styles into a shared, stylistically neutral representation before it reaches the finetuned LLM. This makes the LLM robust to unseen or complex style domains during inference.

  5. The system can integrate Retrieval-Augmented Generation (RAG) for enhanced robustness:

  6. During finetuning, RAG is used via similarity-based retrieval to select relevant example pairs based on semantic similarity (searching with target-side answers rather than source-side questions).

  7. During inference, a multi-stage RAG approach is utilized:

  8. A sketch-first method retrieves similar examples from a vector bank to generate an initial sketch, which is then used to refine the final output (similar 5-shot finetuning).

  9. Terminology and name retrieval can be integrated into the prompting pipeline by dynamically appending domain-specific term instructions when trigger words are present in a query, significantly improving terminology correctness and long-term consistency.

  10. The improved system can exhibit superior performance over state-of-the-art baselines like Few-shot In Context Learning (ICL) and Automatic Post-Editing (APE), achieving higher BLEU scores for content preservation and more stable style classification accuracy, especially in complex domains like literary or governmental texts.

The improved AI system can perform the following specific tasks:

  1. Generate text in a specific target style from a source text, even when no direct parallel data exists for that style.

  2. Adapt an LLM to transfer text between highly disparate domains (e.g., administrative documents to literary prose) with high fidelity regarding semantic content and intent.

  3. Maintain consistent terminology and proper naming conventions across generated texts when dealing with domain-specific jargon or character names, even when the input query uses different vocabulary or style markers.

  4. Handle out-of-domain queries—inputs that the model has never seen during training—by first neutralizing the input style via roundtrip translation, allowing the finetuned LLM to perform accurate transfer based on its learned neutral representation.

  5. Provide highly reliable, factually consistent stylistic outputs in complex scenarios (e.g., literary analysis or legal drafting) by leveraging retrieved domain-specific term guidance during inference.

Sources

Related papers