Structuring the Space of Perspectives

arXiv:2608.12113 · cs.CL · Submitted 2026-08-12 · Read on arXiv

Agnese Daffara, Sebastian Padó, Tanise Ceron

University of Stuttgart · Bocconi University

cs.CL

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: Under review for TACL (editor decision: b)

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 50/100

The gist: The same event can be reported from different perspectives depending on the experiences, background, and beliefs of the writer or speaker.

Terminology

Summary

The same event can be reported from different perspectives depending on the experiences, background, and beliefs of the writer or speaker. A variety of NLP areas engage with perspectives, spanning from text analysis to algorithm optimization. A wide range of operative concepts (such as stances, sentiment, frames, and arguments) has been used to capture perspectives in texts, however the precise relationships among those concepts remain unclear. Arguably, a deeper theoretical understanding of these concepts would empower more effective research on perspectives. In this paper, we address this gap by reviewing the space of perspectives in NLP and defining a set of properties that help distinguishing perspective-related concepts. Our analysis leads us to posit a hierarchy which organizes these concepts linearly along a single axis. Finally, we show how this principled conceptual hierarchy can help researchers navigate the field and select operationalizations of perspective that align with their specific research objectives.

The paper asks two research questions: RQ1: What concepts are used to identify perspectives in text? RQ2: How are these concepts related? To answer these questions, the authors first introduce and characterize the space of perspective-related concepts based on the literature (§3). They then organize and structure the conceptual space with a property annotation, clustering analysis and Principal Component Analyis (PCA) (§4). The outcome of the analysis is the model of perspective shown in Figure 1: they find that clusters of perspective-related concepts form a hierarchy along a linear scale of conceptual and linguistic specificity. To make this insights actionable, they propose a decision tree (Figure 5) to help researchers make informed choices among the discussed concepts that match their specific research goals (§5).

The paper is not a full survey, but rather a conceptual organization, supported by a synthesized literature review with the goal of capturing the space of perspective-related concepts together with their definitions and characteristic properties. To understand the concepts related to perspectives in NLP, the authors first conduct a regular expression based search on the ACL Anthology bibliography and select 60 papers. Since this search did not capture some foundational work that contributed to the conceptualization of perspective in NLP, neither the literature from neighbor areas (e.g, communication science), they continue collecting papers manually on Google Scholar, by following citation trails in references and related works. Papers are included in the analysis if they: i) provide a definition of perspective; ii) use perspective in a conceptual way that distinguishes it from similar work; iii) establish a conceptual grounding for a certain term. After some iterations, they organize all papers in a bottom-up fashion, which resulted in a final set of relevant concepts. The final number of collected papers is 227.

Based on the analysis of the extant literature, the authors collect a total of 15 concepts relevant for perspectives: Morals and Values, Ideology, Stances, Sentiment and Emotions, Opinions, Claims and Arguments, Frames, and Topics. For each concept, they characterize it and briefly discuss relevant textual features and methods used in computational modeling (Perspective Signals). Discriminative properties of each concept are marked in bold; linguistic levels are indicated in italics.

The authors include morals and values because they are increasingly studied in NLP as the foundation of perspective given their proximity with ideological positions. Values motivate arguments by influencing the positions adopted and the justifications offered for them. They form the basis for all processes of evaluation, being the intrinsic goods or ideals that individuals pursue or cherish, such as FREEDOM and EQUALITY. Morals are frequently studied together in the NLP literature, but they origin from distinct frameworks; they are societal and encode a shared judgment of what is right or wrong, e.g., the moral principle that causing harm is wrong is accepted across cultures.

While morals and values provide the overarching beliefs guiding decisions, ideology is the coherent system that organizes them into a stable set of ideas shared by a group. Ideology is stable because it is anchored in a shared system of core values that function as evaluative criteria across topics and contexts; just like the grammar of a language, which is the reference point for its rules. Because it is anchored in these group-level commitments, ideology operates at a higher level of abstraction than stance or sentiment (which are granular and variable), and can be interpreted as a cluster of aligned opinions. In NLP, ideology is used in political domains. It manifests as (i) a position in a debate (e.g. PRO-ISRAEL vs. PRO-PALESTINE), (ii) a political leaning on a spectrum (e.g. LEFT vs. RIGHT), or (iii) a party affiliation (e.g. DEMOCRATICS vs. REPUBLICANS).

Stance indicates an ideological position towards a target. It is usually treated together with sentiment as an evaluative and affective concept. However, it may also be epistemic if there is no affective component. Typical labels are AGAINST/NEGATIVE, NEUTRAL/NEITHER, and PRO/FAVOR/POSITIVE. While binary ideology detection is sometimes referred to as stance detection, there are important distinctions between the two: (i) ideology is a stable overarching system, whereas stance is individually variable, and (ii) ideology is target-generic, whereas stance is target-specific, requiring one or more explicit or inferable targets: entities (e.g. Donald Trump), policy issues (e.g. border control), events (e.g. the approval of a new policy), or claims (e.g. we need to increase border control to ensure national security). Stance also overlaps with entity-based or aspect-based sentiment analysis, which seeks to identify emotional attitude toward a target; however, stance reflects evaluative alignment rather than emotional tone. As Hasan and Ng (2012) note, the same document may have negative sentiment expressions but a positive stance.

In contrast to ideology and stance, sentiment and opinions indicate a subjective response with an affective component. In fact, they are traditionally linked to the area of subjectivity detection, where they were originally theorized as private states, i.e., mental states that are not accessible to objective observation or verification. Sentiment can be interpreted in two ways: (i) in a general sense, it refers to instances of perspective expressed in a text, which is why sentiment analysis and opinion mining are traditionally treated as equivalent tasks; (ii) in a more fine-grained view, it reflects the polar orientation (POSITIVE, NEUTRAL, NEGATIVE) of an opinion, whereas an opinion represents the full perspective expression. Emotions are affective states that go beyond polar orientation to specify the type of emotional response (such as FEAR, ANGER, or SADNESS) typically grounded in psychological models such as Ekman's basic emotions or Plutchik's wheel.

Opinion is a structured construct that can be decomposed into various sub-components. There is a shared intuition that opinion involves some degree of uncertainty, as it is not factual but instead tied to personal beliefs, just like sentiment. The key distinction between sentiment and opinion seems to lay in the facts that (i) an opinion can lack a sentiment, like in the sentences Bin Laden is hiding in Pakistan or I believe the word is flat and (ii) an opinion can be identified with the linguistic expression itself (e.g., Mary said the dress is beautiful), while sentiment tends to denote an abstract attitude mapped onto polar labels (e.g. POSITIVE). According to Kim and Hovy (2004), an opinion consists of four elements: (i) the topic, (ii) the holder, (iii) the claim, and (iv) optionally, the sentiment.

Argumentation lies at the core of perspective and can be understood as a basis for representing it. The authors include the notions of claim (a statement functioning as the minimal unit of argumentation) and argument (a set of statements composed of premises and conclusions). While opinions describe what people think about a topic or product, arguments explain why they hold these opinions. Although opinion mining and argument mining are distinct tasks, their boundaries can be blurry, because opinion mining may also involve identifying argumentative motivations behind a sentiment. In unsupervised approaches, a perspective can be seen as an aggregation of arguments. Clusters of similar arguments (and therefore similar opinions) can reveal overarching belief systems and be interpreted as ideological groups or can group texts by frame, where a frame is defined as a set of arguments that shares an aspect.

According to Boydstun et al. (2013), framing means portraying an issue from one perspective to the necessary exclusion of alternative perspectives. In some work, frame and perspective are used as synonyms. When we frame something, we do three things: (i) selecting: choosing what to present and what not to; (ii) focussing: highlighting or emphasizing some parts; (iii) embedding: presenting some information as the part and some other as the whole. These processes take place at the cognitive level (mental representations of the world), the semantic level (choosing what linguistic structures to use), and the communicative level (impact on the audience). However, defining framing is notoriously slippery because there are various ways of analyzing it: one can use, for example, topic-like dimensions such as media frames, semantic patterns such as semantic frames and connotation frames, narrative dimensions such as narrative frames (e.g. HERO, VICTIM), or morality dimensions such as morality frames (e.g. CARE/HARM). Media frames or communication frames are framing dimensions found in media. They can be issue-generic or issue-specific and the labels can be defined in an inductive or deductive fashion. A famous annotation framework is the Media Frame Corpus, including 15 labels (e.g., MORALITY, ECONOMIC, HEALTH AND SAFETY), widely adopted in media studies. Semantic frames, on the other hand, study framing through linguistic structures and semantic roles (e.g. KILLING: the KILLER or CAUSE causes the death of a VICTIM). The theory of semantic frames was introduced by Fillmore (1976) and operationalized as FrameNet.

The choice of which topics to present contributes to perspective, because it is inherently linked to selection and framing bias. Besides this, topic modeling has been used for perspective detection, leveraging unsupervised methods such as LDA to uncover latent semantic structures. Here, a perspective is seen as an aggregation of documents or text segments that share topical distributions. This approach is conceptually related to argument clustering, but relies primarily on lexical patterns rather than argumentative structures and fine-grained discourse signals. The ideological dimension is obtained from the topic representation as a latent variable: texts are assigned one weight for their topic and another one for their ideology, so that the latter can be isolated. Despite their decent performance, topic models risk oversimplifying perspective by focusing too heavily on lexical distributions rather than higher-level linguistic features.

The authors summarize how the concepts can be combined for perspective detection and analysis. Political ideology is the prominent conceptualization of perspective in NLP. It is an overarching belief system rooted in a shared system of values that aggregates multiple fine-grained viewpoints on entities, topics, and issues. For example, a CONSERVATIVE ideology will be the sum of specific positions on various topics (e.g., AGAINST migration, AGAINST public health, etc.). These positions, intended as stances, sentiments or opinions, can serve as proxies for ideology detection. In contrast to ideology, these concepts are tied to specific situations and are variable across time and topics; e.g., a structured opinion will have a specific time, location, holder, and be linked to a specific event. Such fine-grained beliefs are encoded in argumentation, making this dimension the core nucleus of perspective and a concrete level where perspective can be identified. For example, to fully understand why someone has certain thoughts towards migration (opinion), a certain position towards the topic (stance), and a certain affective attitude (sentiment), we must consider the reasons deeply encoded in argumentation. In parallel, the choice of what information to present and how to present it is a good proxy for perspective. Media frames present a text under a particular light, and semantic frames induce perspectivization. Both can be combined with the abstract beliefs discussed above to represent perspectives holistically, aggregating the ideological belief and their concrete realizations in language and cognition.

To investigate the organization of the perspective-related concepts (RQ2), the authors carry out an analysis to induce a property-driven structure over these concepts from expert judgments. They first manually annotate each concept along four properties that they characterize during the literature review. The definitional properties discussed in §3, such as polar and affective, are binary and concept-specific: they apply only to subsets of concepts and are not gradable, making them unsuitable for a comparison across all concepts. Therefore, they inductively derive four additional dimensions from the literature that are both universal (applicable to every concept in the survey) and gradient (varying continuously across concepts). These properties are: (i) strength of linguistic cues: how strongly the concept is associated with specific linguistic elements; (ii) granularity (scope): the typical scope or localization of the concept within a text; (iii) entity-specificity: how strongly the concept is tied to specific entities; (iv) number of discrete classes: in classification, how many classes are used. This property annotation is performed independently by the three authors in their capacity as experts for the literature discussed above. Based on a codebook, they locate concepts on a Likert scale from 1 to 5 with respect to each property.

The property with the highest IRR is number of discrete classes (ρ = 0.80), followed by granularity (ρ = 0.77). Strength of linguistic cues and entity-specificity have significantly lower agreement (ρ = 0.31 and 0.26). Arguments and claims are difficult to agree on in strength of linguistic cues because they are not systematically associated with certain recurrent linguistic cues, but at the same time they are identified with the linguistic expression itself. Stances in turn can also be ambiguous because they can be interpreted either as the ideological position toward an issue or as its concrete instantiation (similar to opinions). Entity-specificity raises problems for concepts that can optionally refer to specific entities or not; for example, semantic frames involve entities and their roles, but they may or may not be necessary for annotation and detection.

To group the concepts, the authors apply agglomerative hierarchical clustering on annotator-aggregated scores, specifically average linkage with Euclidean distance. This results in four compact groups: values and ideology, the most abstract and global level, comprising overarching belief systems and political ideologies; sentiment and stances, including specific beliefs that express personal positions; topics and media frames, including topical dimensions; argumentation, including arguments and claims, semantic frames, and opinions, all identifiable with linguistic structures. The spider chart in Figure 4 shows the average scores for the clusters on each property. The chart demonstrates the properties the authors consider are strongly correlated at the cluster level: the ranking of the clusters regarding the different properties agrees almost perfectly. The only exception is that the sentiment and stances cluster shows a lower class count than expected. This is presumably the case because the cluster includes concepts whose interpretation varies depending on their theoretical definition and operationalization, notably emotion.

Given the strong correlation between the four properties, the authors investigate whether the perspective-related concepts can be reduced to a single axis by running a Principal Component Analysis (PCA) of the annotations. They find that this is largely the case: the first principal component (PC1) explains 62% of the variance and shows a positive loading with each of the properties. When they represent all concepts purely in terms of their value on the dimension formed by PC1, they recover the four clusters almost perfectly, with the only outlier a swap between topics and stances. This result supports their interpretation that there is a latent linear ordering underlying the concepts. The axis identified by PC1 captures both linguistic and conceptual aspects, and can be interpreted as a dimension of (generic) specificity. Figure 1 is informed by this analysis and shows their model: the space is represented as a set of concentric circles, ranging from ideological beliefs (outer) to linguistic instantiations (inner). At one end, we find generic concepts, such as ideology, which are not bound to specific situations, but underlie other fine-grained perspectives; they emerge throughout the document mainly with lexical cues, and map onto a few labels. At the other end, we find argumentation-related concepts, which are instead more specific both in what they express, i.e., precise situations and entities, and in how they are expressed: well localized in language spans, signalled by semantic and syntactic patterns, and mapping to a wide or open-ended class range.

The authors also consider additional factors not considered in the main model: information about the writer, annotator, or media source, even though they are indicators of perspective. They include them in a separate box, since they describe the (extralinguistic) context of the text rather than its content. Indeed, metadata may not match the perspectives expressed in the text, and it is important to distinguish between grouping emerging from texts and groupings based on external data. They identify three types of extra-textual factors, based on the perspective holder: (i) the author's characteristics, (ii) the annotator's characteristics and (iii) the media source. These include socio-demographic, political and cultural background (e.g., political orientation, gender, country, social affiliations, editorial stance), as well as annotation-related information (e.g., IAA).

The authors compare their model with previous hierarchies. Doan and Gulla (2022) divide the methods for perspective detection into: (i) political ideologies/leaning/party detection, (ii) political stance/framing detection, and (iii) political viewpoint extraction. Klebanov et al. (2010) distinguish four levels of perspective, from less to more abstract: (i) opinions, (ii) stances on specific issues, (iii) ideological positions, and (iv) demographic factors and life of the author (e.g., place of birth, religion, culture, political tradition). Most similar to their proposal is the hierarchy by Van Der Meer (2024), comprising three levels of abstraction: (i) stances, (ii) arguments, and (iii) values. Their model makes a number of contributions: (i) they include subjectivity-related concepts that have traditionally been treated separately; (ii) they include the linguistic level, following the idea that perspectives can be identified through arguments and semantic patterns; (iii) they annotate conceptual properties, providing an empirical support for the hierarchy; (iv) they distinguish content from context-related factors.

The goal of the study was to clarify and structure the space of concepts relevant for research on perspectives in NLP. They ask two research questions: RQ1, what concepts are used in perspective identification. They identify 15 concepts and characterize both their definition and operationalization based on a literature analysis. In RQ2, they ask how these concepts are related. The analysis of their expert annotations of four properties found that perspective-related concepts can be organized along a single dimension of linguistic and conceptual specificity that captures most of the variance between the concepts.

Figure 1 orders the concepts in terms of specificity but does not provide guidance for choosing which one(s) to use in a hypothetical application. In Figure 5, the authors present a decision tree that leverages the features from §3 to guide concept selection. The following scenarios illustrate how the tree can be used.

Scenario 1: I am studying a corpus of news that is fully topic- and issue-agnostic. I want to detect political perspectives for building a diverse news recommender. Following the decision tree, I decide to consider perspectives that emerge from the text. I aim to capture generic beliefs (emerges from the text > denotes a mental state or belief), without identifying explicit targets or relying on affective features, as the dataset is unstructured and mostly comprises factual news. The proposed operationalization is political ideologies. If I want to derive perspectives bottom-up from concrete language patterns (... > denotes a mental state or belief = NO > is about information selection = NO), I could start from opinion mining: as discussed, perspectives can be inferred by aggregating minimal positions on smaller topics. For diversification purposes, these opinions should then be reduced to a small number of meaningful clusters or categories.

Scenario 2: I am analyzing a corpus of news articles from different outlets covering the same event, with the goal of comparing how it is presented across sources. In this case, I have more flexibility, as I am not constrained by a fixed topic or application, and I do not need to cluster articles. My focus is on how information is organized and empathized rather than on what content is conveyed. Following the tree (... > is about information selection > involves emphasis), the proposed device is frames. In case I care about linguistic patterns and event structures, I could work with semantic frames.

Scenario 3: I am building a politically-aligned LLM-based persona. I could leverage metadata about authors' demographics from a corpus to guide the alignment (emerges from the text = NO). Otherwise, the persona can be aligned with a broader political leaning which emerges from a consistent pattern of opinions across multiple issues. If my analysis is more granular, I could control for stances toward specific issues or targets (... > requires a target) (e.g., PRO migration, PRO same-sex marriage, AGAINST gun control). If I care about the affective tone (... > is affective), I may consider sentiment or emotions as complementary dimensions.

The authors discuss future research directions. Newspapers make editorial decisions at multiple levels of perspective, including how to frame events, which arguments to use, what topics to cover, and which stances to adopt. Despite lacking such deliberate mechanisms, LLMs convey perspectives emerging from training data in a comparable way. These viewpoints, embedded in textual choices, often go unnoticed by readers. The authors claim that detecting, controlling, and communicating these layers with transparency is worth-while to support people's access to information and promote critical engagement. By surveying perspective concepts, they have shown how analyzing argument structures jointly with information selection and presentation can map specific opinions to beliefs at different levels of granularity, up to ideology and values, while framing can reveal hidden over-emphasizing signals. Exploring how these levels can be integrated into a coherent representation is a promising research direction in support of critical social analysis. Possible outcomes of this paper include a comprehensive annotation scheme, a modeling recipe, or an evaluation protocol for perspectives in text.

On a more operational level, the conceptual hierarchy also carries direct implications for how we evaluate and audit language models. Rather than treating perspective bias as a monolithic property, the specificity axis offers a diagnostic lens: bias in LLMs may manifest differently at different levels, from systematic skews in ideological framing that pervade entire outputs, to more localized choices in argumentation structure or semantic framing that subtly shift responsibility or salience. Benchmarks for perspective diversity in generated text could be designed to probe each level independently, yielding a richer picture of where training data or alignment procedures introduce distortions.

At the same time, the normative framing underlying much of this work – that greater perspective diversity is inherently desirable – deserves scrutiny. Diversity of perspectives is a meaningful democratic value when it reflects the genuine range of informed viewpoints on a contested issue; it becomes problematic when operationalized in ways that treat fringe or harmful positions as simply another point on a spectrum to be represented. The hierarchy may help draw this distinction more precisely: diversity at the level of values and ideology calls for different normative criteria than diversity at the level of claims or arguments, where factual accuracy and logical coherence impose additional constraints beyond mere representational balance. Navigating this tension – between pluralism and epistemic responsibility – is as important as the development of methods to evaluate generated text.

Improvements for AI systems

Improvements to AI systems based on this paper:

  1. Multi-level perspective detection and generation: Instead of treating perspective as a single monolithic attribute, AI systems can be designed to detect and generate text at multiple hierarchical levels simultaneously—from abstract ideology and values (outer layer) down to specific arguments, claims, and semantic frames (inner layer). This enables systems to produce or analyze text with explicit control over each level independently (e.g., generating a conservative-leaning article that uses particular semantic frames and argument structures).

  2. Perspective-aware evaluation and auditing of LLMs: The specificity axis provides a diagnostic framework for auditing bias in language models. Rather than a single bias score, AI systems can be evaluated separately for: (a) ideological skew at the document level, (b) stance distribution toward specific targets, (c) argumentation patterns, (d) framing choices (e.g., semantic roles, salience), and (e) sentiment/emotion profiles. This yields a granular perspective fingerprint for a model, enabling targeted mitigation of distortions at specific levels.

  3. Controllable text generation with perspective parameters: The decision tree (Figure 5) can be operationalized as a control interface for text generation. Users specify constraints like emerges from text, requires a target, is affective, or involves emphasis, and the system selects the appropriate perspective concept(s) to condition generation on—e.g., generating news with a specific media frame, or a persona with a consistent ideological stance but variable opinions on sub-issues.

  4. Cross-level consistency checking: AI systems can be trained to verify that perspectives expressed at different levels are coherent—e.g., that an opinion (inner level) aligns with the stated ideology (outer level), or that arguments support the claimed stance. This improves factuality and logical consistency in generated persuasive or opinionated text.

  5. Perspective-diverse summarization and recommendation: For news aggregation or recommendation systems, the hierarchy enables generating multiple summaries of the same event, each conditioned on a different level of the perspective axis (e.g., one ideological summary, one argument-focused summary, one frame-focused summary). This supports diverse viewpoint presentation without requiring explicit stance labels.

  6. Metadata-aware perspective modeling: The paper's separation of content-based perspective from extra-textual factors (author, annotator, media source) allows AI systems to explicitly model and control for these confounds—e.g., separating what the text says from who wrote it, enabling more robust cross-source comparison and debiasing.

  7. Hierarchical prompt engineering for LLMs: The linear axis (from abstract values to concrete linguistic instantiations) can guide prompt design: prompts can be structured to elicit perspectives at a chosen specificity level (e.g., argue for this position using specific claims and semantic frames vs. express a general ideological leaning), improving output controllability.

  8. Perspective-aware data annotation and training: The property annotation scheme (strength of linguistic cues, granularity, entity-specificity, number of classes) can be used to create richer training datasets where each text is labeled at multiple levels of the hierarchy, enabling multi-task learning that jointly predicts ideology, stance, argument structure, and framing—improving overall perspective understanding.

  9. Normative diversity calibration: The paper's distinction between diversity at the values/ideology level vs. claims/arguments level allows AI systems to be calibrated differently: ensuring representational balance for ideological diversity while enforcing stricter factual/logical constraints at the argument level, preventing the amplification of harmful fringe positions under the guise of diversity.

  10. Perspective-aware dialogue and persona systems: For conversational AI, the hierarchy enables dynamic adjustment of perspective granularity—e.g., a system can maintain a stable ideological persona while adapting stances on specific topics, and can shift between argument-based and sentiment-based responses depending on user needs, improving realism and user alignment.

Abstract

The same event can be reported from different perspectives depending on the experiences, background, and beliefs of the writer or speaker. A variety of NLP areas engage with perspectives, spanning from text analysis to algorithm optimization. A wide range of operative concepts (such as stances, sentiment, frames, and arguments) has been used to capture perspectives in texts, however the precise relationships among those concepts remain unclear. Arguably, a deeper theoretical understanding of these concepts would empower more effective research on perspectives. In this paper, we address this gap by reviewing the space of perspectives in NLP and defining a set of properties that help distinguishing perspective-related concepts. Our analysis leads us to posit a hierarchy which organizes these concepts linearly along a single axis. Finally, we show how this principled conceptual hierarchy can help researchers navigate the field and select operationalizations of perspective that align with their specific research objectives.

Sources

Related papers