Conflict or Strategy? Asymmetric Role Framing of La France insoumise and Rassemblement National in French News Headlines, 2022-2025

arXiv:2608.09936 · cs.CL · Submitted 2026-05-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Conflict or Strategy? Asymmetric Role Framing of La France insoumise and Rassemblement National in French News Headlines, 2022–2025".

Jane: The paper was written by Amr Sobhy from Le French News Lab.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everybody. Today we're digging into a paper that's got a mouthful of a title — "Conflict or Strategy? Asymmetric Role Framing of La France insoumise and Rassemblement National in French News Headlines, two thousand twenty-two–two thousand twenty-five." Jane, I'm going to need you to help me unpack that one.

Jane: Happy to, Tom. So the title is really asking a simple question: when French news headlines talk about the far-left party, La France insoumise, and the far-right party, Rassemblement National, do they treat them the same way? And the answer the paper gives is a pretty clear no.

Tom: And that's not just about whether one gets nicer headlines than the other, right? It's about the kind of role each one is given.

Jane: Exactly. The paper found that LFI tends to show up in headlines as an aggressor — attacking, clashing, denouncing. RN, on the other hand, gets framed more through strategy — electoral competition, tactical positioning, alliance math. Same media system, two very different roles.

Tom: So it's not a question of who's painted as the villain more. It's that they're painted as different kinds of political animals altogether.

Jane: Right. And that's a much more interesting finding than just "the press is biased against one side." It suggests the press has a kind of template for how to cover each party.

Tom: And the authors — Amr Sobhy from Le French News Lab — they looked at over twenty-eight thousand headlines from twenty-five different outlets over four years. That's a serious corpus.

Jane: It is. And the fact that the pattern holds across so many different outlets, from left-leaning Libération to right-leaning Le Figaro, makes it harder to dismiss as just one editorial line.

Tom: So the title is really asking us to think about framing not as good or bad, but as a set of roles. Who's the fighter, who's the chess player.

Jane: And that's the lens we're going to keep coming back to as we dig into the actual findings. Stick around.

Summary: Tom: So we've got the title unpacked. Now let's talk about what the paper actually found. Jane, walk us through the core result.

Jane: So the headline finding — pun intended — is this asymmetry. LFI headlines had a conflict frame thirteen point eight percent of the time, versus eight point seven percent for RN. Meanwhile, strategic-game framing showed up in forty-five point nine percent of RN headlines, versus thirty-seven point two percent for LFI. Both gaps are statistically solid.

Tom: And the paper's careful to say this isn't just about negativity. It's about the kind of political actor each party is made to look like.

Jane: Precisely. And they also looked at a secondary layer — moral accounting. Who's blamed, who's legitimized, who's cast as a victim. And here's where it gets really interesting: on those measures, the aggregate numbers basically cancel out. No net difference between the parties.

Tom: But that's not because there's no pattern. It's because the pattern is split by outlet.

Jane: Exactly. Left-leaning outlets like Libération delegitimize RN far more — forty-eight point nine percent of their RN headlines versus twenty point four percent for LFI. Right-leaning outlets like JDD do the opposite. They delegitimize LFI more. So the pooled average hides a really polarized structure underneath.

Tom: And the paper makes a big deal about the difference between these two layers — the behavioral layer and the moral-accounting layer.

Jane: Right. The behavioral layer — who attacks, who strategizes — is consistent across outlets. The moral layer — who's blamed, who's legitimate — is where editorial politics kick in.

Tom: So you've got shared agreement on the "what" but sharp disagreement on the "who's at fault."

Jane: That's the deepest result in the paper, honestly. It means the press isn't just picking sides wholesale. It's sharing a common narrative about what each party does, but then fighting over what that means.

Tom: And that's a much more nuanced picture than "the media is biased against the left" or "against the right."

Jane: Much more. And it's why the paper argues we need to move beyond simple sentiment analysis. You can't capture this with a positive-negative score.

Tom: So what does that mean for how we study political communication going forward?

Jane: It means we need to think about roles — aggressor, strategist, victim, legitimate player — as separate dimensions. And that's exactly what this paper tries to do.

Improvements: Tom: So we've got the findings. Now let's talk about what this paper suggests we should do differently. Jane, what's the big methodological contribution here?

Jane: The big one is that they built a three-model annotation pipeline. Three different large language models — GPT-OSS, Llama, and Mistral — each independently labeled every headline, and then they took a majority vote.

Tom: And why is that better than just using one model?

Jane: Because any single model might have its own biases. By using three from different providers and architectural families, you reduce the chance that one model's quirk is driving the results.

Tom: But they didn't just trust the models, right?

Jane: No. They did a human validation study — four hundred headlines coded by two human annotators, blind to what the models said. And they found something really important: not all the fields were equally reliable.

Tom: So some of the labels are more trustworthy than others?

Jane: Exactly. The behavioral fields — conflict and strategic-game framing — had strong human agreement, with kappa scores of zero point seven six and zero point eight one. But the moral-accounting fields, like delegitimization, were much weaker — fifty-three point four percent unanimous agreement among the models, and lower human agreement too.

Tom: So they're basically saying, "Trust our conflict findings, but be more careful with the delegitimization ones."

Jane: Right. And that's a really honest way to do research. Instead of presenting everything as equally solid, they grade their own claims. They call the behavioral findings "foreground claims" and the moral-accounting ones "structured tendencies."

Tom: That's refreshing. A lot of papers would just present all their results with the same confidence.

Jane: And they also ran a bunch of robustness checks. They did a permutation test with ten thousand shuffles — zero of those shuffles reproduced their core findings. They did cluster bootstrapping to account for the fact that headlines from the same outlet aren't independent.

Tom: So the methods are really rigorous.

Jane: They are. And the paper argues this should be a template for other researchers. If you're going to use LLMs to annotate political text, you need to stratify your claims by how reliable the construct actually is.

Tom: So it's not just about the French case. It's about how we do computational social science.

Jane: Exactly. The specific finding is about France, but the methodological framework — the construct-stratified reliability — is portable to any political system.

Conclusion: Tom: Alright, we're wrapping up our look at "Conflict or Strategy? Asymmetric Role Framing of La France insoumise and Rassemblement National in French News Headlines, two thousand twenty-two–two thousand twenty-five." Jane, give us the final take.

Jane: The core finding is that French headlines don't treat LFI and RN as symmetric extremes. They cast LFI as a fighter and RN as a strategist. And that role asymmetry is stable across outlets and across most of the two thousand twenty-two–two thousand twenty-five period.

Tom: And the moral-accounting layer — blame, legitimacy, victimhood — that's where the outlets diverge.

Jane: Right. The behavioral layer is shared; the moral layer is polarized. That's the paper's deepest contribution.

Tom: And methodologically, they've given us a roadmap for using LLM annotation responsibly.

Jane: Three models, majority vote, human validation, and a clear hierarchy of which claims to trust. That's a model for the field.

Tom: Any caveats before we say goodbye?

Jane: One big one. The paper is observational. It can't tell us whether the press is constructing these roles or just accurately reporting what the parties actually do. LFI really did lead the pension-reform opposition. RN really did pursue a normalization strategy. So the headline pattern could reflect reality.

Tom: So we can't say the press is distorting anything.

Jane: We can't. We can only say the roles are asymmetric. Whether that's fair or accurate is a separate question the data can't answer.

Tom: Fair enough. That's a good note to end on. Thanks for joining us, and we'll be back with the next paper soon.

Jane: Take care, everyone.

Amr Sobhy

Le French News Lab

cs.CL

Submitted: 2026-05-29

Comments: 19 pages, 3 figures, includes appendices

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 66/100

The gist: This study examines whether French news headlines frame left- and right-populist challengers as symmetric "extremes" or as fundamentally different political adversaries, analyzing 28,592 headlines

Key concepts

Asymmetric Role Framing
The study's central finding that news coverage treats LFI and RN differently. LFI is consistently depicted as an aggressor or conflict-oriented political force, whereas RN is framed through strategic concepts like tactical positioning and electoral competition.
Behavioral Layer
This layer analyzes the observable actions of the parties in headlines—specifically, who attacks and who strategizes. The research found that this behavioral framing (aggressor vs. strategist) was consistent across different news outlets, indicating a shared narrative about what each party does.
Moral Accounting Layer
This layer examines how headlines assign blame, legitimize parties, or cast them as victims. Unlike the behavioral layer, these moral assessments varied significantly depending on the editorial slant of the outlet (e.g., left-leaning vs. right-leaning).
LLM Annotation Pipeline
The methodology used a three-model annotation pipeline involving GPT-OSS, Llama, and Mistral. By having multiple large language models independently label headlines and then taking a majority vote, the researchers reduced the chance that any single model bias drove the results.

Terminology

Summary

This study examines whether French news headlines frame left- and right-populist challengers as symmetric extremes or as fundamentally different political adversaries, analyzing 28,592 headlines about La France insoumise (LFI) and Rassemblement National (RN) published by 25 French-language outlets between 2022 and 2025, annotated through a three-model LLM pipeline validated against a stratified human audit.

The clearest finding is role asymmetry rather than valence asymmetry: conflict framing and strategic-game framing are more robust across models and time than delegitimization, with AGGRESSOR serving as corroborating role syntax. LFI appears in headlines more often through a conflict register and RN through a strategic-electoral register. This role gap is direction-stable across all three annotation models, survives bootstrapping and permutation tests, and persists across outlet families and most of 2022–2025. Because the design is observational, the paper reports how headlines distribute these roles, not evidence that the press constructs them independently of what the parties did.

A secondary moral-accounting layer (who is blamed, legitimized, or cast as a victim) is structured by outlet rather than party, producing aggregate nulls that conceal some of the corpus’s most polarized patterns. Methodologically, the annotation pipeline reveals a two-tier reliability profile: conflict and strategic-game framing achieve the strongest human validation and cross-model stability; actor role is direction-stable but treated as corroborating because its audit reliability is lower; normative-judgment constructs (legitimacy, blame) are weaker. The paper contributes political-role assignment as a target for computational framing research that decomposes what valence-based measures conflate, and establishes a construct-stratified reliability framework for calibrating majority-vote LLM annotation pipelines in political text tasks.

The corpus consists of 28,592 French news headlines drawn from a source database of 902,111 headlines published by 25 French-language outlets between January 2022 and December 2025. Headlines were included if they contained at least one surface reference to LFI, RN, LFI-led electoral coalitions, or selected high-salience party leaders: Jean-Luc Mélenchon, Mathilde Panot, and Manon Aubry for LFI; Marine Le Pen and Jordan Bardella for RN. Of the 28,592 headlines, 12,741 primarily frame LFI, 14,770 primarily frame RN, and 1,081 frame both parties jointly or set them in explicit opposition (coded BOTH). The corpus spans four calendar years, with the heaviest coverage in the 2024 election year (n = 9,697) and the lightest in the non-election year 2023 (n = 4,652). BOTH headlines are excluded from the main LFI-versus-RN regressions so each observation has a single primary party target; a three-way stability check confirms the exclusion does not inflate or suppress the core role asymmetry, with all eight primary outcomes shifting by at most ∆OR ≤ 0.004 with no direction reversals.

The annotation scheme has a strict hierarchy: conflict and strategic-game framing carry the strongest behavioral claims; actor role supplies corroborating syntax; legitimacy supplies a supporting moral-accounting layer; blame, target, reference form, personalization, and party target are diagnostic fields. The foreground constructs are the binary conflict frame and binary strategic-game frame. The four-way actor-role field (AGGRESSOR / TARGET / NEUTRAL / MIXED) is retained as corroborating syntax: its direction is stable and party-stratified audit errors are symmetric, but its overall validation is weaker than the two binary behavioral fields. Legitimacy frame (NORMALIZING / DELEGITIMIZING / NONE) distinguishes behavioral role assignment from moral accounting; lower model agreement (53.4% unanimous) and contested human-audit performance mean it is treated as a structured tendency rather than a stable property.

Each headline was independently annotated by three open-weight large language models served through Regolo, a cloud inference API: GPT-OSS-120B, Llama-3.3-70B-Instruct, and Mistral Large 2 (123B), all queried with an identical annotation prompt specifying the codebook definitions and tie-break rules. For binary fields, the final label is the majority vote of the three models. For categorical fields with more than two values, the final label is the plurality label; in the small proportion of cases where all three models assign different categorical values (three-way splits), the label is flagged and treated as ambiguous in all downstream analyses. The binary behavioral constructs show the highest model unanimity (79.6–85.5%). legitimacy frame is the most contested field at 53.4% unanimous (42.8% of its corpus-level labels are 2-of-3 majority labels), which is why delegitimization claims rest on a narrower evidentiary base than the behavioral-role findings.

A stratified 400-headline audit stress-tests the pipeline at its weakest points: disagreement-heavy cases (n = 150), high-agreement controls (n = 150), and theory-targeted oversamples (n = 100). Two annotators annotated all 400 headlines blind to model output after a calibration session on 30 held-out practice headlines. Four of five fields reach κ ≥ 0.60; blame frame (κ = 0.523, moderate) is treated as a supporting indicator throughout. Human–MV agreement is above threshold for all audited fields; final three-way legitimacy frame splits remain unreliable (human–MV = 31.2% on 16 split rows), so delegitimization is retained as a structured tendency.

Framing rates are compared as proportions (chi-squared tests) and through logistic regression with outlet and year fixed effects, estimating party-level gaps net of compositional differences. For each binary outcome Y, the pooled model is logit P(Y = 1) = α + β RN + γoutlet + δyear; reported odds ratios are exp(β) from standard statsmodels maximum-likelihood logit fits. The interaction models replace the pooled party term with party × outlet or party × year terms and use likelihood-ratio tests against the pooled fixed-effect model.

The strongest stable finding is not generalized negativity but a split between two registers. LFI is covered more through a behavioral conflict register, appearing more often than RN in a conflict frame (13.8% vs. 8.7%; +5.1 pp; adjusted RN OR = 0.614, p < 0.001). RN is covered more through a strategic-electoral register (45.9% vs. 37.2%; adjusted RN OR = 1.414, p < 0.001). These two primary behavioral estimates survive the stricter outlet-cluster bootstrap that relaxes the within-outlet independence assumption: conflict 95% CI [0.572, 0.652], strategic-game [1.226, 1.635]. AGGRESSOR provides corroborating role syntax: LFI is more likely than RN to be cast as the aggressor in a confrontation (8.6% vs. 6.3%; adjusted RN OR = 0.732, bootstrap CI [0.636, 0.814]), but this label is interpreted cautiously because actor-role validation is weaker than the two binary behavioral fields and audit errors are symmetric by party. Crucially, AGGRESSOR measures syntactic and behavioral posture (narrating the party as actively initiating a challenge or critique), rather than serving as an editorial verdict of hostility; 91.4% of LFI headlines and 93.7% of RN headlines carry no aggressor label.

The pooled estimate shows RN headlines more likely to carry a delegitimizing frame (31.2% vs. 25.7%; adjusted RN OR = 1.291, p < 0.001), but four considerations qualify it. First, legitimacy judgments are more sensitive to model disagreement than the paper’s behavioral-role findings, with only 53.4% unanimous agreement in the final ensemble. Second, the human validation audit shows that legitimacy is among the more contested audited fields between human coders (κ = 0.644) and the lowest-agreement field between humans and the final majority vote (74.5% overall, 31.2% on final three-way model splits). Third, the cluster-bootstrapped confidence interval for delegitimization (Boot 95% CI: [0.999, 2.003]) is 7.17× wider than the Wald interval and nearly includes 1.0. Fourth, the temporal decomposition shows the gap reversing year-on-year. The paper therefore frames RN delegitimization as a structured tendency organized by outlet and electoral period, not as a stable property of French press coverage.

The main asymmetries do not behave uniformly over time. RN’s strategic advantage is strongest in 2023 (RN-LFI gap = +0.150; OR = 2.205) and remains visible in 2024, but it essentially disappears in 2025 (gap = +0.022; OR = 1.087, ns). By contrast, LFI’s conflict excess is present in every year and is strongest in 2025 (gap = −0.087; OR = 0.464). Delegitimization is the least temporally uniform result: it reverses in 2023 before returning to an RN-positive gap in 2024 and 2025. The 2023 reversal is best understood as a crisis-specific interruption: the pension-reform conflict moved LFI into a high-volume legitimacy register as it led parliamentary and street opposition to the government’s use of Article 49.3. The 2025 strategic-game compression has a different source: Marine Le Pen’s March 2025 conviction shifted RN coverage toward judicial and eligibility registers, displacing ordinary electoral-positioning language.

The pooled blame and target coefficients are near zero or non-significant (blame frame adjusted RN OR = 0.966, p = 0.426; actor role = TARGET adjusted RN OR = 1.049, ns) and should not be reported as simple aggregate party differences. But they are not empty nulls; they are structural aggregation nulls: the pooled coefficient cancels left- and right-leaning outlet logics of comparable magnitude pointing in opposite directions. Normalization is RN-positive but slight (43.7% LFI vs. 45.1% RN; RN OR = 1.080, p < 0.01) and is a secondary finding. The personalization row runs opposite: leader-centric or mixed headlines are more common for LFI (51.4% vs. 47.7%; RN OR = 0.907, p < 0.001), consistent with LFI coverage routed more through Mélenchon and other named leaders.

The sharpest structural contrast is between aggressor and the target/blame pair. The aggressor finding has a much weaker outlet interaction (LR(8) = 17.4, p = 0.027) than target (LR(8) = 123.7, p < 0.001) or blame (LR(8) = 129.8, p < 0.001), which remain among the most outlet-structured findings in the dataset. This yields the paper’s deepest result: broad agreement on who attacks, but sharp editorial disagreement on who suffers and who is responsible. The per-model robustness analysis independently validates this two-tier structure: the behavioral tier survives in every model; the moral-accounting tier does not. The data show consistent LFI conflict excess across outlet families but diverge sharply on the moral consequence of that behavior: left outlets frame RN as the party under siege and bearing responsibility for political conflict, while right outlets frame LFI as unfairly targeted by a hostile establishment.

An accurate-coverage explanation remains plausible: LFI, under Mélenchon and then Panot, led Article 49.3 opposition, censure motions, and street mobilizations against pension reform, while the FN/RN under Marine Le Pen pursued a dédiabolisation strategy aimed at normalization and broader electoral appeal. The observational design cannot distinguish behavioral supply from framing construction. For delegitimization, the JDD editorial transition is the most suggestive within-outlet discontinuity: after Bolloré’s summer-2023 consolidation of editorial control, JDD’s delegitimization direction reversed. Because there is no matched control outlet, the shift coincides with the high-conflict pension-reform period, and the pre-Bolloré window is thin, this is treated as suggestive rather than quasi-experimental.

Limitations include: LLM annotator dependence and correlated biases (the three annotation models are not statistically independent annotators; majority vote reduces random disagreement but not shared systematic bias); legitimacy field reliability (the weakest field in the human-versus-majority validation study); keyword retrieval and syntactic sampling (the corpus was constructed by retrieving headlines containing party-name keywords, which necessarily oversamples headlines in which the party or its leaders are mentioned by name); observational design and causal scope (the design cannot adjudicate between behavioral reality, journalistic convention, and editorial ideology as explanations); generalizability (the corpus covers 25 French-language outlets in metropolitan France between 2022 and 2025, and generalization to other national press systems requires separate empirical work); and headline-only scope (headlines are high-salience framing units but are not representative of full article content).

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems and what the improved systems can do:


Improvement: Implement a multi-model majority-vote LLM annotation system that separates constructs into evidentiary tiers based on cross-model stability and human validation, rather than treating all annotation fields as equally reliable.

What the improved system can do:

  • Automatically classify annotation fields into foreground (high agreement, direction-stable), structured tendency (moderate agreement, context-sensitive), and diagnostic (low reliability) tiers.

  • Flag three-way model splits as ambiguous rather than forcing a plurality label, preventing false confidence in contested cases.

  • Apply per-model robustness checks to detect when one model reverses direction on a construct (as Mistral did on delegitimization), triggering automatic downgrade of that construct's evidentiary status.

Net capability gain: The improved system can detect asymmetric political role assignment in news headlines with calibrated confidence, distinguishing stable behavioral patterns from outlet-structured moral accounting, and resisting both statistical and annotation-method artifacts that would mislead simpler pipelines.

Abstract

Do French news headlines frame left- and right-populist challengers as symmetric ``extremes,'' or as fundamentally different political adversaries? We examine 28,592 headlines about La France insoumise (LFI) and Rassemblement National (RN) published by 25 French-language outlets between 2022 and 2025, annotated through a three-model LLM pipeline validated against a stratified human audit. The clearest finding is role asymmetry rather than valence asymmetry: conflict framing and strategic-game framing are more robust across models and time than delegitimization, with AGGRESSOR serving as corroborating role syntax. LFI appears in headlines more often through a conflict register and RN through a strategic-electoral register. This role gap is direction-stable across all three annotation models, survives bootstrapping and permutation tests, and persists across outlet families and most of 2022-2025. A secondary moral-accounting layer (who is blamed, legitimized, or cast as a victim) is structured by outlet rather than party, producing aggregate nulls that conceal some of the corpus's most polarized patterns. Methodologically, the annotation pipeline reveals a two-tier reliability profile: conflict and strategic-game framing achieve the strongest human validation and cross-model stability; actor role is direction-stable but treated as corroborating because its audit reliability is lower; normative-judgment constructs (legitimacy, blame) are weaker. The paper contributes political-role assignment as a target for computational framing research that decomposes what valence-based measures conflate, and establishes a construct-stratified reliability framework for calibrating majority-vote LLM annotation pipelines in political text tasks.

Sources

Related papers