Adapting Technical-Service LLM Agents with Latent Logic Augmentation, Robust Noise Reduction, and Hybrid Reward Modeling
cs.LG, cs.AI, cs.IR, stat.AP
Submitted: 2026-03-18
Updated: 2026-09-07
Comments: 36 pages, 6 figures, 14 tables. Camera-ready version accepted to the EMNLP 2026 Industry Track. Title, author order and metadata, experiments, analysis, references, and appendices have been updated; the set of authors is unchanged
License: http://creativecommons.org/licenses/by/4.0/
The gist: Technical-service LLM agents are entering production workflows, where value depends on whether engineers adopt generated replies.
Terminology
Abstract
Technical-service LLM agents are entering production workflows, where value depends on whether engineers adopt generated replies. Service tickets hide decision logic, contain noisy single-reference responses, and make reward evaluation costly, making standard post-training brittle. Existing post-training and LLM-as-a-Judge approaches improve grounding or feedback, but do not jointly model latent decision logic, response diversity, and reward cost. We address this gap by coupling latent logic augmentation, robust noise reduction, and hybrid reward modeling. The framework augments supervised fine-tuning data with Planning-Aware Trajectory Modeling and Reasoning Augmentation, builds dual-filtered Multiple Ground Truths, and trains the policy with a hybrid reward that combines a Reranker with an LLM-as-a-Judge. On real Cloud technical-service tasks, the adapted Qwen3-4B achieves the highest Multi-ECS (0.441), lower reward cost, and the highest production adoption rate (46.63%).
Sources
- Kimi K2: Open Agentic Intelligence
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
- LongCat-Flash Technical Report
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
- Scaling Agents via Continual Pre-training
- What and Whose Knowledge? Measuring Epistemic Diversity in Large Language Models
- Fine-Tuning Language Models from Human Preferences
- Proximal Policy Optimization Algorithms
- Automated Optimization Modeling via a Localizable Error-Driven Perspective
- Intelligently Weighting Multiple Reference Models for Direct Preference Optimization of LLMs
- Self-Rewarding Language Models
- A new strategy to optimize complex absorbing potentials for the computation of resonance energies and widths
- The Perfect Blend: Redefining RLHF with Mixture of Judges
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- Qwen3 Technical Report
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks