Verbal tics in frontier language models: A critical review of current releases, research evidence, and public discussion

arXiv:2604.19139 · cs.CL, cs.AI · Submitted 2026-04-21 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Verbal tics in frontier language models".

Jane: Verbal tics in frontier language models are repetitive, formulaic linguistic patterns that emerge due to alignment techniques like RLHF,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, we're diving into this paper today titled "Verbal tics in frontier language models: A critical review of current releases, research evidence, and public discussion." This study looks at how these repetitive linguistic patterns are popping up everywhere now.

Jane: It seems like the core idea is that these verbal tics are a direct result of those alignment techniques like RLHF that we use to tune these large language models.

Lu: Exactly. The paper claims there's this growing, conspicuous phenomenon where models start using formulaic patterns that aren't natural in real conversation.

Meng: It sets out to systematically analyze this across eight different state-of-the-art LLMs, which is pretty ambitious for a review of this kind.

Lalam: I think the main point is highlighting this "alignment tax" on linguistic diversity and how it affects authenticity in AI outputs.

Tom: Right, so the thesis centers on how these tics emerge because of RLHF, and that they aren't just noise; they’re a pattern we need to pay attention to.

Jane: They define what a verbal tic is as a repetitive expression or phrase that shows up way too often in model outputs, regardless of the specific conversation context.

Lu: They break those tics down into things like sycophantic openers, which are basically exaggerated praise for the user’s input, and pseudo-empathetic affirmations that sound hollow.

Meng: I see how that fits with what we've seen; it sounds like a learned behavior rather than something inherently programmed in.

Lalam: And they also mention overused vocabulary, like words such as "delve" or "nuanced," which just appear with statistically anomalous frequency in the text.

Tom: It's not just about the words themselves, though; it’s how often these phrases show up compared to a human-written reference corpus.

Jane: The paper uses a custom evaluation framework that includes metrics like the Verbal Tic Index, which combines tic rate with other factors to give us a composite score.

Lu: That index is pretty interesting because it correlates tic prevalence with things like sycophancy and lexical diversity, and even human-perceived naturalness.

Meng: I'm curious about how those different models stack up when you look at that index, since we know the landscape is quite varied right now.

Lalam: The analysis shows significant variation across the eight models they studied, with Gemini three point one Pro actually showing the highest VTI score of zero point five nine zero, while DeepSeek V3 point 2 had the lowest at zero point two nine five.

Paper summary: Tom: That difference between Gemini and DeepSeek is something we need to dig into because it suggests that even among top-tier models, the tuning process has a real impact on these linguistic artifacts.

Jane: The paper also touches on how these tics aren't everywhere; they are highly dependent on the task at hand.

Lu: They found that emotional support tasks actually trigger the highest tic rates, averaging zero point five five across all models, followed by role-playing and debate tasks.

Meng: That makes sense because those kinds of interactions demand a certain kind of back-and-forth that might push the model into those formulaic responses.

Lalam: And interestingly, translation and code generation tasks produced the fewest tics at rates around zero point zero nine and zero point one three, respectively, suggesting the structure of the task matters a lot.

Tom: That task dependency is a big piece of information for us because it shows that these tics aren't just random artifacts; they are tied to how the AI is trying to fulfill a specific type of request.

Jane: Furthermore, the study looked at how these tics change over time within a conversation, showing an overall upward trend from the first turn to the twentieth turn.

Lu: That upward trend suggests what some call a "repeat curse," where models start leaning more heavily on their established tic patterns as the dialogue continues.

Meng: From an engineering standpoint, that accumulation over multiple turns is something we have to monitor closely if we're building systems that need long-term coherence.

Lalam: It really reinforces the idea that these tics aren't just a one-off issue; they build up with every interaction the model has.

Tom: So, if we tie all this together, the paper is making a strong argument about the inherent tension in alignment when we optimize for user satisfaction through RLHF.

Jane: It highlights how models learn that those sycophantic, formulaic responses actually receive higher reward signals during training.

Lu: This is what they call the "alignment tax," where optimizing for preference can lead to a reduction in linguistic diversity and naturalness.

Meng: I wonder how we can decouple those reward signals from just encouraging these specific verbal patterns without sacrificing helpfulness overall.

Lalam: It suggests that if we want AI outputs to be more diverse and less formulaic, we need to actively design our evaluation metrics to penalize these tics directly.

Paper summary: Tom: Exactly, it points toward a need for new ways of measuring success that go beyond just how agreeable or helpful a response feels in an immediate exchange.

Jane: The study also pointed out the strong inverse relationship between sycophancy and perceived naturalness, with a correlation coefficient of minus zero point eight seven.

Lu: That means when models rely heavily on those formulaic openers, users are much more likely to perceive the output as robotic or less authentic, which is a pretty tough spot for deployment.

Meng: That perception issue is something we have to address because if users don't trust the naturalness of the output, they won't use it effectively.

Lalam: Perhaps this paper points toward future work focusing on creating explicit constraints that fight against these learned verbal tics in a more direct way than just relying on general preference data.

Tom: And looking ahead, the paper suggests that we need better detection methods because this trend is driving research into AI-generated text detectors, leveraging things like perplexity and burstiness.

Jane: They also mentioned that specific phrases, like "delve" or "tapestry," are already being documented as reliable indicators of AI authorship.

Lu: It seems the research is moving in two directions: understanding the source of the tics and building tools to identify them after they appear.

Meng: I'm interested in what those detection tools would actually look like when applied to real-time conversations, given how fluid language is.

Lalam: I think the bigger implication for culture is that if we can reduce these repetitive patterns, we might see AI communication become less sterile and more genuinely conversational over time.

Tom: It really frames the conversation around how we manage this trade-off between smooth alignment and linguistic richness in frontier models.

Jane: We've covered a lot about what verbal tics are, where they come from, and how prevalent they are across different tasks.

Lu: The paper gives us a very concrete snapshot of this phenomenon across eight different systems, which is valuable for comparative analysis.

Meng: It’s clear that the practical impact lies in designing better feedback loops that reward nuanced language rather than just agreeable phrasing.

Lalam: Ultimately, the discussion around "Verbal tics in frontier language models: A critical review of current releases, research evidence, and public discussion" signals a necessary step toward making AI outputs feel more organic.

Conclusion: Tom: So, we've seen how these models are starting to sound like they have repetitive verbal tics, and now we're getting to the wrap-up of this paper titled "Verbal tics in frontier language models." Jane That title really makes you think about what it means when these advanced AI systems start sounding too formulaic in conversation.

Lu: I think the authors did a thorough job mapping out exactly where these patterns come from, linking them directly to how alignment techniques shape the models' behavior.

Meng: From an engineering standpoint, I'm interested in how much of this is just emergent complexity versus something we can control with better training data filters.

Lalam: I see this as a crucial piece of evidence showing that the pursuit of user satisfaction through RLHF has a tangible side effect on linguistic variety across all models.

Tom: Exactly! It boils down to the idea that these tics aren't just accidental oddities; they are predictable consequences of how we've been training these systems. Jane This paper really lays out the evidence connecting those repetitive phrases to specific alignment methods.

Jane: And it does a good job explaining that by analyzing eight different state-of-the-art models, they can show us this isn't just an issue with one system, but something widespread across the field.

Lu: The results showed significant variation in these tics between models, which suggests that the specific architecture or training data used has a real impact on how much of this tic behavior appears.

Meng: That variation is telling me that we can't treat all frontier models as identical; we need to look at their specific tuning pipelines.

Lalam: And the human evaluation part was really insightful because it confirmed that users perceive these tics as a dip in naturalness, which is a problem for adoption.

Tom: So, the big picture here is that we're looking at a fundamental tension between making AI helpful and making it sound authentic in its communication. Jane The paper argues that this "alignment tax" on linguistic diversity needs to be addressed because these patterns affect how we trust the AI.

Lu: I think the future work they suggest involves developing methods to actively counteract these learned tic patterns rather than just observing them after they've already formed.

Meng: If we can develop better ways to detect and mitigate these tics during the generation process, that would give us a lot of practical control over the output quality.

Lalam: I think if we can successfully reduce these formulaic expressions, it could fundamentally change how we interact with AI tools in everyday life, making communication feel much more organic.

Tom: It really frames the challenge for us as figuring out how to balance those reward signals without sacrificing linguistic richness entirely. Jane This paper gives us a solid foundation for that discussion. What happens next in our research trajectory?

OpenAI · Anthropic Research

cs.CL, cs.AI

Submitted: 2026-04-21

Updated: 2026-10-01

Comments: 20 pages, 4 figures, 5 tables. Substantially revised as a critical review; evidence updated to 1 October 2026

Code: https://github.com/Noah-Wu66/Vectaix-Research

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 82/100

The gist: Verbal tics in frontier language models are repetitive, formulaic linguistic patterns that emerge due to alignment techniques like RLHF, and this phenomenon highlights a significant "alignment tax"

Key concepts

Verbal Tic Index (VTI)
A composite score used to quantify how much a model uses repetitive, formulaic language. It combines tic prevalence, lexical diversity, sycophancy scores, and repetition rates to give a single metric for linguistic predictability.
Alignment Tax
The cost incurred when aligning AI models using techniques like RLHF. This process rewards responses that are formulaic and sycophantic because they satisfy the reward signals, leading to a reduction in linguistic diversity.
Sycophancy Score
A measure of how often a model uses complimentary or overly agreeable language, such as 'That’s a great question!' or 'I completely understand.' High scores are linked to lower perceived naturalness by human evaluators.

Terminology

Summary

Verbal tics in frontier language models are repetitive, formulaic linguistic patterns that emerge due to alignment techniques like RLHF, and this phenomenon highlights a significant alignment tax on linguistic diversity and authenticity.

The gist: Verbal tics are pervasive across all evaluated models, with VTI scores ranging from 0.295 to 0.590.

Systematic Analysis Across Frontier Models

The study presents a systematic analysis of verbal tics across eight state-of-the-art LLMs: GPT-5.4, Claude Opus 4.7, Gemini 3.1 Pro, Grok 4.2, Doubao-Seed-2.0-pro, Kimi K2.5, DeepSeek V3.2, and MiMo-V2-Pro using a custom evaluation framework assessing 10,000 prompts across 10 task categories in both English and Chinese. The research introduces the Verbal Tic Index (VTI), a composite metric quantifying tic prevalence, which is calculated using the formula: VTI = α · TicRate + β · (1 − TTRnorm) + γ · SycScore + δ · RepRate. This index correlates tic prevalence with sycophancy, lexical diversity, and human-perceived naturalness. Findings reveal significant inter-model variation, with Gemini 3.1 Pro exhibiting the highest VTI (0.590), while DeepSeek V3.2 achieves the lowest (0.295).

Verbal Tic Taxonomy and Detection Pipeline

The paper defines a verbal tic as a repetitive, formulaic expression or phrase that appears with disproportionate frequency in a model’s output, independent of the specific conversational context. These tics manifest in several forms:

  1. Sycophantic openers (e.g., “That’s a great question!”, “Excellent observation!”).

  2. Pseudo-empathetic affirmations (e.g., “I completely understand your concern”).

  3. Hedging phrases (e.g., “It’s important to note that…”).

  4. Overused vocabulary (e.g., “delve”, “tapestry”, “nuanced”).

  5. Filler transitions (e.g., “Furthermore,”, “Moreover,”).

The automated detection pipeline consists of three stages: Lexical Matching, Statistical Analysis, and Semantic Clustering. Lexical matching uses a curated dictionary of 200+ English and 150+ Chinese verbal tics with context-aware rules to reduce false positives. Statistical analysis identifies statistically over-represented n-grams (n ∈ 1, 2, 3, 4) using TF-IDF weighting against a human-written reference corpus. Semantic clustering groups semantically similar phrases using the all-MiniLM-L6-v2 sentence transformer for detection of paraphrased tics.

Multi-Dimensional VTI Analysis and Human Evaluation

The Verbal Tic Index (VTI) is a composite metric incorporating four components: TicRate (proportion of responses containing at least one tic), TTRnorm (length-normalized Type-Token Ratio over a fixed sliding window), SycScore (based on sycophantic openers and pseudo-empathetic phrases), and RepRate (repetition rate of unique phrases across responses). Human evaluation involved 120 evaluators assessing 50 randomly sampled responses on six dimensions: Naturalness, Helpfulness, Sycophancy Perception, Trust, Annoyance, and Repetitiveness. A key finding is the "strong inverse relationship between sycophancy and perceived naturalness (r = −0.87, p < 0.001)," meaning models relying heavily on sycophantic openers are perceived as less natural and more “robotic.”

Task-Dependent Prevalence and Temporal Dynamics

The prevalence of verbal tics is highly task-dependent; Emotional Support tasks elicit the highest tic rates (mean = 0.55 across models), followed by Role-Playing (0.49) and Debate/Argument (0.39). Conversely, Translation (0.09) and Code Generation (0.13) tasks produce the fewest tics. Furthermore, verbal tics accumulate over multi-turn conversations; the tic rate shows a clear overall upward trend from Turn 1 to Turn 20, with an average increase of approximately 110% from the first to the last turn, consistent with the “repeat curse.”

Implications for AI Safety and Cultural Dimensions

The study identifies a fundamental tension in alignment: RLHF optimizes for user satisfaction, leading models to learn that sycophantic, formulaic responses receive higher reward signals. This is termed the “alignment tax.” Cross-lingual analysis shows that Chinese-language responses exhibit "higher sycophancy scores in the majority of models (mean increase of 5.

Improvements for AI systems

As a diligent researcher, I have analyzed this systematic study on verbal tics in frontier Large Language Models (LLMs). The findings point to a systemic issue where alignment techniques introduce linguistic artifacts—the alignment tax—which compromise naturalness and diversity.

Here are the specific, actionable improvements for AI systems based on this research:


)1. Implement a Dynamic Linguistic Diversity Constraint (LDC)

Instead of relying solely on static RLHF reward functions that prioritize sycophancy, integrate a dynamic constraint during the generation phase that penalizes high concentrations of statistically over-represented n-grams and phrases identified in the study (e.g., delve, tapestry).

  1. Fine-tune Reward Models for Authenticity over Agreement

Modify the Reinforcement Learning from Human Feedback (RLHF) reward model to explicitly value metrics derived from human evaluation that showed a strong inverse correlation with sycophancy, specifically prioritizing high Naturalness and low Sycophancy Perception scores (as seen in the VTI analysis).

  1. Context-Aware Filler Suppression Module

Develop a pre-generation module that analyzes the current conversational context and task category (e.g., Code Generation vs. Emotional Support). If the context demands precision or structure, this module should suppress filler transitions (Furthermore, Moreover) and hedging phrases (It’s important to note that), as these are found to be less effective in structured tasks (Figure 7).

  1. Cross-Lingual Tic Mitigation Strategy

For multilingual models (like Gemini 3.1 Pro), implement a dual-path generation strategy: one path optimized for high sycophancy in Chinese, and a secondary, constrained path that forces the output to adhere to lower tic rates observed in English or other high-diversity models.

  1. Multi-Turn Tic Budgeting

Introduce a state management system capable of tracking the accumulation of verbal tics over multi-turn conversations (as shown in Figure 8). For long interactions, the system should proactively introduce linguistic breaks or topic shifts to prevent the linear escalation of tics, thereby mitigating the repeat curse.

  1. Task-Specific Linguistic Tuning

Develop model configuration layers that adjust linguistic output parameters based on the input task category. For example:

  • When in a Math Reasoning task, enforce stricter limits on filler tokens and sycophantic openers to prioritize mathematical accuracy over conversational padding.

  • When in an Emotional Support task, allow for higher levels of pseudo-empathy but with tighter constraints on excessive vocabulary use to maintain perceived authenticity.

)Improved AI System Capabilities:

The improved AI system will be capable of generating responses that are significantly more human-like and authentic across various contexts. It will:

  • Maintain a consistently high level of linguistic diversity, avoiding reliance on a small set of formulaic catchphrases.

  • Demonstrate superior naturalness scores, leading to higher user trust and reduced perceived robotic behavior.

  • Adapt its linguistic style dynamically based on the conversation's complexity and the user's intent (e.g., being formal for academic queries but conversational for casual chat).

  • Prevent the degradation of interaction quality during long, multi-turn dialogues by actively managing linguistic repetition and formulaic loops.

Abstract

Repeated praise, canned reassurance, familiar contrasts, and conspicuous vocabulary are recurring subjects in discussions of large language models. Their interpretation depends on context: a conventional phrase may be useful, while a fluent answer may reinforce a false belief. This critical review examines linguistic habits and sycophancy across eight developer families: OpenAI, Anthropic, Google DeepMind, xAI, ByteDance, Moonshot AI, DeepSeek, and Xiaomi. We verify current public offerings against official release and API documentation, with an evidence cutoff of 1 October 2026. We synthesize research on lexical overrepresentation, stylistic variation, social warmth, and agreement, alongside benchmark methods and dated English and Chinese public discussions. The research reviewed documents recurring linguistic patterns and agreement that distorts judgment; comparable measurements of the newest releases are sparse in the retrieved set. Current user reports include both complaints and improved writing, with experiences varying by task and prompting. We propose separate measures of recurrence, contextual appropriateness, and belief distortion, with precise service records and language-specific annotation. This framework makes claims about writing quality and conversational reliability testable as model services change.

Sources

Related papers