QoNext: Towards Next-generation QoE for Foundation Models

arXiv:2509.21889 · cs.CL · Submitted 2025-09-26 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "QoNext: Towards Next-generation QoE for Foundation Models".

Jane: The paper was written by Yijin Guo, Zicheng Zhang, Ye Shen, Farong Wen, Junying Wang et al. from Shanghai Jiao Tong University, Shanghai AI Laboratory, Fudan University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: We're kicking off today by looking at a fascinating paper called "QoNext: Towards Next-generation QoE for Foundation Models," which is a huge step forward in how we judge AI quality. The authors, Yijin Guo and the team, are suggesting that we need to stop just relying on benchmarks and look at the actual user experience instead of just correctness.

Jane: What struck me when reading the introduction is how they are moving away from that single "perfect score" idea, focusing instead on how different parts contribute to a holistic feeling when talking to a real person.

Lu: They’re essentially proposing that QoE, or Quality of Experience, needs to be treated as a composite score that takes into account multiple inputs influencing the user' perception throughout the entire dialogue. This is a massive theoretical shift for the AI community.

Meng: When they mention evaluating parameters like output speed and latency, it forces us engineers to think about infrastructure constraints alongside model performance, which is a huge practical shift in focus for deployment.

Tom: Right! It’s not just about how smart the model *can* be; it's about how smoothly it *delivers* that intelligence to you in real time when the user interacting with the service.

Jane: Think of it like talking to a person versus talking to a text-to-speech bot—even if the information is identical, the feeling is completely different based on timing and flow. The human element matters.

Lu: The paper highlights that traditional evaluation methods often fail because they can't model that complex interplay between these operational dimensions, which makes their proposed framework so powerful for understanding the user experience.

Meng: And this brings up a really interesting point: if we can quantify the impact of latency—say, comparing one second versus nine seconds—we're giving engineers concrete targets for optimization and QOS management.

Lalam: It reminds me that human communication isn't instantaneous; there are natural pauses, and those pauses carry meaning. By quantifying this artificial timing, they are helping us build AI that understands the rhythm of conversation itself.

Tom: So it’s a complete overhaul of the metrics we use to judge conversational AI quality, moving from a single score to a dynamic look at how deep does this understanding go? We need to talk about what they actually improve upon next.

Summary: Jane: Following up on that summary, the paper gets into specifics about the core mechanism of QoNext, and it really shows how robust their proposed model is supposed to be. They aren't just proposing vague ideas; they are testing a defined framework.

Tom: Exactly! They established this dual-layered framework where QoE integrates content quality—how good the information is—and service quality, which focuses on the operational performance of the delivery mechanism itself.

Lu: This dichotomy is key because it acknowledges that users judge both aspects simultaneously; they don't separate content from how quickly it arrives, which is what makes their approach so grounded in user reality.

Meng: The authors define five specific dimensions to cover this dual nature: information density, content accuracy, output speed, latency position, and duration. This moves the discussion away from vague concepts toward measurable units.

Jane: And it wasn't just about listing those; they used a systematic approach that brings up how the model handles both factual correctness and how "dense" or concise the information is.

Lu: The paper suggests that content accuracy emerges as the most critical factor, which is a powerful finding—it means users fundamentally prioritize substance over smooth delivery.

Meng: From an engineering standpoint, seeing specific metrics like v for output speed makes sense. If we know exactly where users are sensitive, we can optimize resource allocation immediately.

Lalam: What I found fascinating was how they used the concept of information density—it's about being concise but also being rich, capturing that human preference for quality over mere volume of data.

Tom: It sounds like a complete overhaul of the metrics we use to judge conversational AI quality, giving us a blueprint for how deep this measurement goes into the structure itself.

Improvements and Methodology: Jane: Moving beyond the summary, let's talk about how they actually conducted this research because that's where the real proof lies. They went beyond simply stating that user experience matters; they actually ran a massive experimental setup.

Tom: Absolutely! They didn't just run quick tests; they designed a controlled experiment using fifty-four questions, pairing them with answers of varying quality and then subjecting them to specific QoS conditions defined in Table one.

Lu: That ability to systematically vary parameters is monumental, as it means the model isn't just memorizing patterns; it's learning underlying principles about user perception across the entire dataset.

Meng: The fact that they tested wildly varying parameters—like setting output speed to.03 seconds or even.06 seconds per word—tells us this framework is highly adaptable for different deployment environments and infrastructure constraints.

Jane: And it wasn't just speed; they looked at latency duration, testing both really short times of one second and much longer gaps of nine seconds, to see how the human brain holds up under pressure.

Lu: The results show that even when you mess with these parameters dramatically, the predictive performance remains stable across all dimensions, which is a major technical achievement for demonstrating reliability.

Meng: Stability is key for industry adoption. If we know that changing our infrastructure slightly won't cause our QoE prediction model to completely break down, that's a massive win for deployment reliability in any real-world scenario.

Lalam: And what I found fascinating was how they used an independent group of raters who had never been involved in the original data construction, which lends incredible credibility to the findings about human perception.

Tom: So, the paper basically proves that this model can predict user ratings even when we throw entirely new knobs and dials at it, demonstrating strong generalization ability. But how do we apply this to real-world design goals?

Conclusion and Implications: Jane: To wrap up our discussion on "QoNext: Towards Next-generation QoE for Foundation Models," we're looking at a huge shift in how we judge AI quality, moving from just pure correctness to a whole package that includes interaction feel.

Tom: I think the most important thing is that the paper shows us we can actually measure user satisfaction by accounting for those time-based factors like latency and speed, which is something researchers had been struggling to quantify before this really sophisticated.

Lu: It’s truly exciting because it suggests that the future of AI isn't just about models getting smarter; it’s about making them feel more human in their delivery and interaction rhythm. We are designing better experiences for us as users.

Meng: From an engineering standpoint, this means we can start building feedback loops where real-world performance metrics directly influence how we tune the model, which is a huge step for operationalizing AI at scale.

Lalam: I believe this research helps us create AI that doesn't feel like a cold machine but something genuinely helpful and trustworthy, making it a more empathetic part of our daily lives.

Tom: Lalam’s point about trust resonates with me; we want to use these tools without feeling frustrated by poor timing or clunky responses, and this gives us the data to achieve that.

Jane: Exactly, Tom, and that's interesting because the personalized analysis showed us how users are sensitive in different ways across different personality types like MBTI. It’s not a one-size-fits-all approach.

Lu: The PCA results show we aren're not just building better algorithms; we’re designing better interaction strategies based on data, which is a huge leap forward for the theoretical side of AI.

Meng: And by focusing on those specific, measurable dimensions like accuracy and speed, we can move away from vague user complaints and towards hard, data-driven optimization goals.

Lalam: The paper gives us a roadmap to build an AI that is both powerful and genuinely pleasant to use for everyone.

Tom: It's great to see this level of detail in the work—we are definitely seeing major shifts coming in how we view AI quality metrics. We'll be right back after the break, but stick around for our next deep dive into another groundbreaking paper!

Shanghai Jiao Tong University, Shanghai AI Laboratory, Fudan University

cs.CL

Submitted: 2025-09-26

Updated: 2026-09-04

Code: https://github.com/aiben-ch/LMM-Evaluation-Survey

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 82/100

The gist: Existing evaluations of foundation models fail to capture "user’s experience during interaction," often treating evaluation as a matter of output correctness alone.

Key concepts

Quality of Experience (QoE)
QoE is a composite score used to judge AI quality. It moves away from the idea that relying on benchmarks is sufficient, instead focusing on how different operational inputs contribute to the user's holistic feeling throughout an entire dialogue.
Dual-Layered Framework
This framework structures QoE by acknowledging two aspects: content quality (how good the information is) and service quality (the operational performance of the delivery mechanism). It recognizes that users judge both aspects simultaneously.
Operational Dimensions
The authors define specific, measurable units to quantify user experience. These dimensions include information density, content accuracy, output speed, latency position, and duration. This allows for data-driven optimization.

Terminology

Summary

Existing evaluations of foundation models fail to capture user’s experience during interaction, often treating evaluation as a matter of output correctness alone. This paper introduces QoNext, the first framework that adapts Quality of Experience (QoE) principles from networking and multimedia to the assessment of foundation models. QoNext addresses this critical gap by providing a unified, fine-grained analytical framework that connects measurable system parameters with subjective human perception, enabling a comprehensive, human-centric evaluation paradigm for productized AI services.

The QoNext Framework

QoNext is designed to systematically integrate the dual nature of user experience—service quality and content quality—into model assessment. The framework identifies five key factors that shape this experience:

  • Content accuracy (alpha): Measures correctness and logical coherence.

  • Information density (rho): Measures conciseness and focus of information.

  • Output speed (v): The rate of generating responses.

  • Latency position (l pos): Where a delay occurs in the output stream (e.g., start, middle, end).

  • Latency duration (l time): How long the delay lasts (e.g., 3s, 5s, 7s).

Experimental Design and Data Collection

To build a comprehensive database for this framework, researchers conducted a large-scale human evaluation experiment involving over 70 participants. The design required controlling these parameters to systematically examine the effect of output speed on user perception. Participants were tasked with rating model outputs on a 1–5 integer scale across three dimensions: Overall Impression, Content Quality, and Perceived Responsiveness. This meticulous process allows for the construction of a human-labeled dataset that links specific model parameters to subjective ratings.

Key Findings and Analysis

The analysis of the collected data reveals several critical insights into user perception. First, content quality is found to outweigh service quality in shaping overall user experience, with content accuracy emerging as the most critical factor. Second, while different personality types exhibit consistent structures, they differ in their sensitivities to specific factors. For instance, extraverted users showed higher sensitivity to Latency Duration compared to introverted users. These findings are further supported by Principal Component Analysis (PCA), which suggests that user ratings are shaped by the interplay of multiple factors rather than a single dominant variable.

Predictive Modeling

The QoNext framework utilizes the constructed database to train regression models that predict user-perceived experience from measurable system parameters. The primary metric used is SRCC (Spearman's Rank Correlation Coefficient), which quantifies the consistency between predicted scores and human-annotated scores. The ablation study, where parameters were systematically removed, confirmed that content accuracy plays a central role in shaping user judgments, with its absence causing a drastic performance degradation. This validates the feasibility of using QoNext as a scalable method for automated user-experience assessment.

Improvements for AI systems

The following recommendations are derived directly from the QoNext framework and its experimental findings, aimed at transforming how foundation models are evaluated, optimized, and deployed for human interaction. These changes move beyond simple task-based benchmarking toward holistic user experience management.


Current State: AI systems are primarily judged by objective scores (e.g., MMLU, GSM8K), which measure task accomplishment but fail to capture subjective user satisfaction.

Improvement: Implement the QoNext Multi-Dimensional Regression Model as the primary evaluation metric alongside traditional benchmarks.

  • Actionable Change: Integrate a controlled testing environment that systematically manipulates five critical parameters: Content Accuracy (alpha), Information Density (rho), Output Speed (v), Latency Position (l pos), and Latency Duration (l time).

  • System Capability: The system can automatically run these configurations through the trained regression model (e,g., LightGBM) to generate a predicted Mean Opinion Score (MOS), providing a quantitative, scalable proxy for human-centric evaluation without relying solely on costly manual labor.

Current State: LLMs prioritize maximizing token accuracy, often leading to verbose or poorly paced outputs that frustrate users.

Improvement: Optimize the generation pipeline using the insights from the ablation study, prioritizing content quality while dynamically adjusting interaction rhythm.

  • Actionable Change (Prioritization): Establish Content Accuracy (alpha) as the primary optimization constraint. Any generation where alpha falls below a defined threshold must be flagged for revision, regardless of speed or density.

  • Actionable Change (Rhythm): Implement dynamic pacing controls based on user context:

  • For complex knowledge retrieval (high rho), allow for slightly longer pauses (l time in [3s, 5s]) to signal deep thought and maintain perceived quality.

  • For fast-paced creative tasks, prioritize high Output Speed (v > 0.5 tokens/s) to maintain fluency and avoid cognitive load.

  • System Capability: The the AI system can now dynamically modulate its response cadence (speed and pause placement) to maximize perceived smoothness, ensuring that interaction rhythm complements the semantic quality of the generated text.

Current State: Latency is treated as a monolithic technical failure point; users do not perceive latency in isolation.

Improvement: Implement Contextual Latency Management to mitigate disruptive delays based on the temporal location of the pause.

  • Actionable Change (Early/Mid-Turn): If a pause occurs early (l pos = 0 or 0.25, e.g., initial token generation), the system must provide immediate visual feedback (e.g., Thinking... indicator) to prevent users from perceiving the delay as a system failure, even if l time is moderate.

  • Actionable Change (Mid/Late-Turn): If a pause occurs mid-flow (l pos = 0.5), the system must ensure that the preceding and succeeding text segments are highly coherent to minimize the user's train of thought disruption, even if l time increases slightly.

  • System Capability: The AI system can proactively manage its own streaming output, inserting calculated pauses or providing visual cues to actively reduce user impatience associated with specific temporal windows.

Current State: AI interaction is treated monolithically; all users are assumed to have the same tolerance and preferences.

Improvement: Deploy Personality-Aware Interaction Profiles to tailor the experience based on user characteristics.

  • Actionable Change (E-I/J-P): For users identified as Introverted/Judging (I/J), prioritize high Information Density (rho) and stable pacing, as these users show a higher tolerance for content depth.

  • Actionable Change (S-N): For users identified as Intuitive (N), subtly increase the frequency of concepts or abstract connections within the response to better align with their preference for conceptual breadth.

  • System Capability: The AI system can dynamically adjust its response style—for example, increasing conciseness for users who value efficiency (J) or expanding context for users who value depth (N)—to achieve personalized alignment with user cognitive processing styles.

Abstract

Existing evaluations of foundation models predominantly focus on output correctness, treating interaction as a static exchange of information. However, such perspectives overlook the essence of the LLM-driven conversational experience, which is determined not only by content quality but, crucially, by dynamic service attributes such as generation velocity and latency patterns. To address this gap, we introduce QoNext, the first framework that adapts Quality of Experience (QoE) principles from networking and multimedia to the holistic assessment of human-AI interaction. QoNext identifies experiential factors that shape user experience and incorporates them into controlled experiments in simulated interaction scenarios, where human ratings are collected under diverse configurations. From these studies we construct the QoNext Database and train the QoNext Model, a neural predictor that estimates user experience directly from measurable system parameters. Our results demonstrate that QoNext effectively decodes the underlying mechanisms of user satisfaction and enables precise prediction of human sentiment across varied service conditions.

Sources

Related papers