Rhetorical Questions in LLM Representations: A Linear Probing Study

arXiv:2604.14128 · cs.CL, cs.AI, cs.LG · Submitted 2026-04-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Rhetorical Questions in LLM Representations".

Jane: Rhetorical questions are asked to persuade or signal stance rather than seek information, and understanding how large language models internally represent these questions remains unclear.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So we’ve touched on what this paper is about: "Rhetorical Questions in LLM Representations: A Linear Probing Study" investigates how large language models internally represent rhetorical questions, aiming to understand why they aren't organized along a single linear axis. The central claim is that rhetorical content emerges early and is best captured by last-token representations, and it proves they are linearly separable from information-seeking questions within datasets.

Jane: And what makes this important for us right now is the discovery of the divergence between discriminative performance and representational alignment; even though the signals are transferable, probes trained on different distributions produce different rankings with little overlap in the top or bottom ranks. This shows that transferability doesn't mean a single shared representation exists.

Lu: I think it’s crucial to remember that they are focusing on real-world contexts, using datasets like RQ and SRAQ where the discourse level can vary significantly, which grounds the analysis in actual usage rather than just abstract linguistic theory.

Meng: That context variation is key for practical deployment because if a model only learns one type of rhetorical question structure, it will fail completely when presented with a different conversational setting. We need robustness across contexts to make this useful for any real application.

Lalam: The paper’s qualitative analysis is what really makes me lean into the impact angle: it demonstrates that rhetorical meaning is inherently heterogeneous, spanning discourse-level rhetorical stance and localized, syntax-driven interrogative acts rather than one unified dimension. This suggests a richer internal model for language understanding.

Tom: So, to summarize what we’ve covered about this study, the authors systematically analyzed these signals across two social media datasets and concluded that while rhetorical questions are separable and transferable, they are encoded by multiple non-collinear directions rather than one simple vector. This points toward a more complex internal structure for handling persuasive language.

Jane: That's right, Tom. The paper’s core contribution is showing that the way models process rhetorical questions isn't neatly packaged into one representation space, which has big implications for how we design and train next generation AI systems to handle nuanced human communication.

Conclusion: Tom: Thinking about the overall conclusion of "Rhetorical Questions in LLM Representations: A Linear Probing Study," we’re looking at the work by Louie Hong Yao, Vishesh Anand, Yuan Zhuang, and Tianyu Jiang. The main takeaway is that we shouldn't treat strong probing performance or successful cross-dataset transfer as proof that a single shared representational dimension exists for rhetorical questions in LLMs.

Jane: That’s a very important distinction to make for listeners. What this study really tells us is that rhetorical questions are encoded by multiple linear directions, each emphasizing different cues, reflecting a structure that is inherently heterogeneous and context-sensitive.

Lu: I see it as suggesting we need to stop looking for one perfect feature and start looking at the whole landscape of features that contribute to rhetorical intent in the model’s internal workings. That opens up avenues for much more creative architectures where different parts of the network handle different aspects of persuasion.

Meng: From an engineering standpoint, this means our goal shifts from finding a single magical feature to identifying and controlling these various distinct cues identified by the probes. It gives us a map of what we need to target if we want specific rhetorical behaviors in our models.

Lalam: And for culture, this suggests that AI can be designed to be more sophisticated in how it signals its position, moving beyond simple factual recall into nuanced stances based on context and structure. That’s a deep level of capability I find really compelling for the future of AI interaction.

Tom: So we’ve explored how these different datasets and probes revealed this complexity, and the final conclusion is that rhetorical questions are encoded by multiple linear directions emphasizing distinct cues, which means we need to treat them as a collection rather than a single entity.

Jane: It’s about recognizing that the structure of human persuasion isn't reducible to one simple mathematical vector within the AI's brain, which is a pretty fundamental realization for anyone building these systems.

Lu: That complexity is where the fun starts, because it means we can start designing models that are more expressive in capturing different layers of meaning simultaneously. That’s where true creative AI possibilities lie.

Meng: I just hope that when we start targeting these specific directions, the implementation remains feasible and scalable for real-world use, because theoretical complexity doesn't always translate smoothly into robust engineering solutions.

Lalam: But if we can manage that control, the potential to shape how information is presented and accepted in society is immense. That’s what I’m really looking forward to seeing realized through this kind of research.

Louie Hong Yao, Vishesh Anand, Yuan Zhuang, Tianyu Jiang

University of Cincinnati

cs.CL, cs.AI, cs.LG

Submitted: 2026-04-15

Updated: 2026-10-02

Code: https://github.com/ruyi101/rq-representation-probing

Importance score: 91/100

The gist: Rhetorical questions are asked to persuade or signal stance rather than seek information, and understanding how large language models internally represent these questions remains unclear.

Key concepts

Linear Probing
A technique used to test if specific features or concepts (like rhetorical questions) can be separated linearly within a model's internal mathematical structure. Researchers project high-dimensional data into lower dimensions to see if a simple straight line can distinguish between different categories.
Last-Token Representation
A specific way of looking at the final word generated by an LLM. This representation is often used in decoder models and is analyzed here because it provides a stable signal for testing whether rhetorical signals are encoded in the model's output structure.
Non-collinear Directions
In geometry, this means that different vectors or features are not pointing in the exact same straight line. The paper found that rhetorical questions aren't encoded by one single feature vector but by several directions spread out in a multi-dimensional space, indicating complexity.
Cross-Dataset Transfer
Testing whether a model trained on one type of data (like Twitter) can successfully identify the same concept (rhetorical questions) when applied to a different type of data (like Reddit). The results showed that while some separability exists, the alignment between representations differs across datasets.

Terminology

Summary

Rhetorical questions are asked to persuade or signal stance rather than seek information, and understanding how large language models internally represent these questions remains unclear. This study analyzes rhetorical signals in LLM representations using linear probes on two social-media datasets with different discourse contexts, finding that rhetorical content is linearly separable but encoded by multiple, non-collinear directions rather than a single shared representation.

The gist

Rhetorical questions are not organized along a single linear axis, but reflect multiple linguistic features that are emphasized differently depending on context and data.

Datasets and Representations

The analysis utilizes two real-world rhetorical question datasets: RQ, which consists of Twitter questions annotated as rhetorical or informational with conversational context; and SRAQ, drawn from Reddit conversation threads where examples are annotated as rhetorical or informational. For representation choice, the study focuses on two sequence-level representations: the last-token representation (hT), commonly used in decoder-only models, and mean-pooled representations (h̄ = 1/T Σ t=1 ht). To ensure fair comparison across probes, all representations are projected into a PCA space with k = 64 dimensions, defined separately for each dataset and model.

Linear Probing Framework

The researchers employ three types of linear probes to analyze rhetorical separability:

  1. A training-free population-level diffMean probe, which estimates a direction by subtracting class-conditional means (wDM = µ+ − µ−).

  2. Two trained discriminative probes based on logistic regression and linear support vector machine, which learn weight vectors (w) to minimize cross-entropy loss or hinge loss.

Evaluation metrics used include:

((3) Rank Agreement):

The Spearman’s rank correlation (ρs) measures the agreement between the rankings induced by two different probes or a probe trained on two datasets.

((4) Overlap at the tails):

The Jaccard index (J(Ap, Bp)) assesses whether different probes retrieve similar top- and bottom-ranked examples.

Findings on Separability and Alignment

The study finds that rhetorical content is linearly separable from information-seeking questions within a given context and remains detectable under cross-dataset transfer, with AUROC around 0.7–0.8 for transferred probes. However, the results reveal a divergence between discriminative performance and representational alignment: probes learned from different data distributions induce substantially different rankings with little overlap in the upper and lower ends of the ranking. Qualitative analysis demonstrates that rhetorical meaning is inherently heterogeneous, spanning discourse-level rhetorical stance and localized, syntax-driven interrogative acts rather than a single, unified representational dimension.

Geometric Characterization of Differences

To characterize dataset differences geometrically, the researchers compare RQ and SRAQ subspaces using two measures:

  1. Geodesic distance on the Grassmann manifold, which provides a global measure of subspace alignment.

  2. Mean cosine similarity between corresponding principal components, which is sensitive to the alignment of individual PCA directions and implicitly depends on their ordering.

These geometric comparisons show that while subspaces may not be maximally separated, their dominant directions are often largely misaligned, suggesting that rhetorical questions are encoded in a context-sensitive and heterogeneous manner, rather than along a single linear feature. Furthermore, causal steering experiments confirm this by showing that the identified linear directions capture rhetorical question intent in the model’s internal representations.

Representation Choice and Transferability

The analysis of representation choices shows that while last-token representations provide more stable signals than mean pooling, mean pooling retains useful lexical information at early layers. Cross-dataset transfer results indicate that probe directions retain meaningful separability while showing limited alignment, with Jaccard overlap between top-ranked examples often falling below 0.2, suggesting that similar accuracy can reflect different underlying representations. This implies that rhetorical signals are encoded in a set of distinct, noncollinear directions.

Conclusion

The findings caution against treating strong probing performance or successful cross-dataset transfer as evidence of a single shared representational dimension. Instead, the study concludes that rhetorical questions are encoded by multiple linear directions emphasizing different cues, reflecting a structure that is inherently heterogeneous and context-sensitive. Future work should focus on defining these distinct features and exploring whether the identified rhetorical signals are controllable.

Limitations

The empirical analysis is restricted to two social media datasets, which limits the generalizability of the findings to other domains lacking comparable annotation reliability or granularity. Additionally, the methodology focuses exclusively on linear probing, excluding signals encoded through nonlinear interactions or mechanisms not linearly separable in the representation space considered.


(Self-Correction: The prompt requires a single integer rating based on a provided question and context, which is outside the scope of summarizing the paper. I will proceed with generating the summary as requested.)

Rating:

(No rating required for this task; only summary generation is requested.

Improvements for AI systems

Here are the specific improvements that could be made to AI systems based on this research, along with what those improved systems could achieve:

  1. Improve rhetorical intent detection in LLM outputs by leveraging a multi-directional probing approach.

  2. Develop more robust rhetorical classification/detection models by incorporating representations from different pooling strategies (last-token vs. mean-pooled) and comparing them across datasets to identify contextually stable cues versus noise.

  3. Enhance cross-dataset transferability of rhetorical understanding by training probes on one domain (e.g., RQ dataset) and applying them to another (SRAQ dataset), while explicitly monitoring the resulting divergence in ranking agreement to understand the heterogeneity of rhetorical signals across contexts.

  4. Improve interpretability by moving beyond single-direction analysis to characterize rhetorical meaning as a set of distinct, non-collinear linear features, allowing researchers to distinguish between discourse-level stance taking and localized, syntax-driven interrogative acts within the model's internal space.

  5. Create more stable rhetorical classification systems by prioritizing the use of last-token representations over mean-pooled representations for sequence summarization tasks where rhetorical intent is crucial.

  6. Develop a method for detecting when a successful linear probe (high AUROC) does not imply alignment with a single, shared representational dimension, thereby improving the reliability assessment of model interpretability tools.

  7. Implement controllable rhetorical generation by using the identified probing directions as steering vectors to modify the internal representation during inference and predict subsequent rhetorical behavior (e.g., generating a follow-up question that is more strongly rhetorical or informational).

The improved AI systems could achieve:

  1. A system capable of accurately identifying whether a given text segment in an LLM output functions rhetorically (persuading/signaling) versus informatively (fact-seeking), even when the context shifts between social media domains.

  2. A more reliable mechanism for summarizing or extracting rhetorical intent from long texts, ensuring that the summary relies on stable, late-stage representations rather than noisy aggregate information.

  3. A model that can generalize its understanding of rhetorical questions from one type of conversational context (e.g., Twitter) to a different one (e.g., Reddit), while simultaneously quantifying the degree to which this generalization is accurate versus merely superficial transferability, leading to a deeper understanding of cross-domain rhetoric.

  4. A diagnostic tool that reveals the rhetorical fingerprint of an LLM representation—showing that rhetorical intent isn't encoded by one feature, but by a manifold of distinct features (e.g., one feature for discourse stance and another for syntactic interrogation).

  5. More stable and interpretable classification layers that are less susceptible to the dilution of signal caused by token averaging, leading to higher accuracy in rhetorical detection tasks on complex inputs.

  6. A superior method for auditing model interpretability, allowing researchers to distinguish between probes that are discriminative (high AUROC) and those that capture a truly shared structural property versus those that capture idiosyncratic features specific to the training data distribution.

  7. A generative system capable of being steered during text generation to specifically enforce or modulate rhetorical behavior in the output, enabling the creation of targeted persuasive arguments or critiques.

Sources

Related papers