Semantic Flow Regularization: Teaching LLMs to Generate Diverse Yet Coherent Responses
cs.CL, cs.AI
Submitted: 2026-05-27
Updated: 2026-08-31
License: http://creativecommons.org/licenses/by/4.0/
The gist: When large language models are fine-tuned to generate persona- or tone-conditioned responses, their output diversity is severely limited--a failure we term Cross-Style Collapse.
Terminology
Abstract
When large language models are fine-tuned to generate persona- or tone-conditioned responses, their output diversity is severely limited--a failure we term Cross-Style Collapse. We trace this collapse to the cross-entropy objective, which under shared representations tends to suppress diverse continuations. We propose Semantic Flow Regularization (SFR), a lightweight auxiliary objective that supervises the backbone with continuous sentence-encoder embeddings of future segments via conditional flow matching. The stochastic flow source preserves multi-modality by construction; the flow-matching head is discarded at inference, adding zero deployment cost. On a large-scale industrial dialogue dataset (Qwen3-32B, 9 personas), SFR improves output diversity, style fidelity, and response quality over SFT. We further validate on the public LiveCodeBench-v5 (Qwen2.5-Coder-7B-Instruct), where SFR consistently improves pass@k, confirming generality beyond stylized dialogue. A controlled comparison on MBPP reveals Multi-Token Prediction to be a degenerate special case of SFR.
Sources
- OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs
- Program Synthesis with Large Language Models
- DeepSeek-V3 Technical Report
- One Agent to Serve All: a Lite-Adaptive Stylized AI Assistant for Millions of Multi-Style Official Accounts
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- CTRL: A Conditional Transformer Language Model for Controllable Generation
- Large Concept Models: Language Modeling in a Sentence Representation Space
- Qwen2.5-Coder Technical Report
- Rectified Flow: A Marginal Preserving Approach to Optimal Transport
- rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset
- The Chameleon's Limit: Investigating Persona Collapse and Homogenization in Large Language Models
- Qwen3 Technical Report
- Embarrassingly Simple Self-Distillation Improves Code Generation
- CharacterGLM: Customizing Chinese Conversational AI Characters with Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering