Who Speaks Matters: Authority-Aware Multi-View RAG over Italian Parliamentary Proceedings

arXiv:2608.13410 · cs.AI · Submitted 2026-08-13 · Read on arXiv

Mirko Tritella, Riccardo Pozzi, Matteo Palmonari

University of Milano-Bicocca

cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: Accepted at ISWC 2026 In-Use Track. Please cite the ISWC version

Code: https://github.com/Emeierkeio/thesis-ParliamentRAG

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: ParliamentRAG is a Retrieval-Augmented Generation (RAG) system designed for the Italian Chamber of Deputies that addresses three specific risks in applying RAG to parliamentary transcripts:

Terminology

Summary

ParliamentRAG is a Retrieval-Augmented Generation (RAG) system designed for the Italian Chamber of Deputies that addresses three specific risks in applying RAG to parliamentary transcripts: dominance of the most frequent speakers, inability to weight speakers according to topical expertise, and citation misattribution in politically sensitive text. The system's core contribution is a topic-dependent authority model that estimates each speaker's authority as a function of the current query, combining interpretable components such as profession, education, and previous interventions. Given a user query, the system retrieves relevant speech chunks, identifies topic-relevant experts across parliamentary groups, and generates a summary synthesizing their perspectives, accompanied by supporting quotations.

The system operates over a knowledge graph stored in a single Neo4j instance, integrating parliamentary proceedings, metadata, and legislative activity. The graph includes 387 deputies, 64 government members, 10 parliamentary groups, 80 committees, 608 sessions, 6,010 debates, 6,515 phases, 40,416 speeches, and 27,576 legislative acts, with speeches segmented into 151,073 chunks that serve as atomic retrieval units. The architecture uses a dual-channel retrieval strategy: a dense channel performing vector similarity search over speech chunks, and a graph channel based on parliamentary acts and their signatories. Retrieved evidence is reranked using a weighted score combining relevance, diversity, coverage, authority, and salience. The authority score aggregates semantic similarity between query embeddings and speaker attributes (profession, education, committees, roles) with time-decayed activity-based signals from legislative acts and speech interventions, with weights set empirically (profession 0.15, education 0.10, committee 0.25, legislative acts 0.20, speech interventions 0.25, institutional role 0.05).

The generation pipeline follows four stages—Analyze, Generate, Integrate, Cite—that enforce coverage, grounding, and coherence by design. In the final Cite stage, "quotations are resolved deterministically into verbatim quotations from the parliamentary record. The language model never generates quotation text: it only inserts placeholders with character offsets that are later replaced with the text from the original speeches," preventing hallucinations by construction.

The system was evaluated against Google NotebookLM on 15 policy topics using a two-level protocol combining automated metrics and blind A/B human evaluation by six domain experts. Automated results show ParliamentRAG achieves near-perfect quotation-level group coverage (GQ = 0.97 vs. 0.95), Quotation Faithfulness (QF = 1.00 by design, whereas NotebookLM reaches 0.95), and slightly higher Mean Authority of cited speakers (0.53 vs. 0.52). Human evaluation results show overall satisfaction is nearly identical (4.24 vs. 4.27), with complementary strengths: NotebookLM achieves higher ratings on prose-oriented dimensions like Answer Quality (4.30 vs. 4.04) and Answer Clarity (4.51 vs. 4.27), while ParliamentRAG consistently outperforms on source-related dimensions including Source Relevance (4.07 vs. 3.84), Source Authority (4.21 vs. 4.00), and Source Coverage (4.64 vs. 4.39). Pairwise preferences reinforce this pattern, with ParliamentRAG preferred more frequently on Source Relevance (33% vs. 19%), Source Authority (30% vs. 13%), and especially Source Coverage (25% vs. 5%). The paper concludes that "fluency limitations can often be mitigated through stronger language models or lightweight rewriting stages, whereas guarantees such as verbatim quotation faithfulness and systematic per-group coverage require architectural support and cannot be reliably enforced through prompting alone."

Improvements for AI systems

Improvements to AI Systems:

  1. Implement a Topic-Dependent Authority Model – Replace static or uniform source weighting with a dynamic authority score that adapts per query. The system computes authority by combining interpretable components (profession, education, committee membership, legislative activity, speech frequency) with time-decay, weighted empirically (e.g., committee 0.25, speech interventions 0.25). This allows AI to prioritize speakers with genuine topical expertise over frequent but less relevant voices.

  2. Add a Dual-Channel Retrieval Strategy – Combine dense vector similarity (semantic search over text chunks) with a graph-based channel that traverses legislative acts and signatories. This ensures retrieval captures both semantic relevance and structural/relational context (e.g., who co-sponsored a bill), improving recall of diverse, authoritative evidence.

  3. Enforce Verbatim Quotation by Construction – Prevent hallucinated citations by having the language model output only placeholders with character offsets, which are then resolved deterministically against the original source text. This guarantees 100% quotation faithfulness, eliminating fabricated or misattributed quotes in sensitive domains.

  4. Introduce a Multi-Stage Generation Pipeline (Analyze → Generate → Integrate → Cite) – Structure generation to enforce coverage, grounding, and coherence by design. The Analyze stage identifies topic-relevant experts and groups; Generate synthesizes perspectives; Integrate merges evidence; Cite resolves quotes deterministically. This reduces unsupported claims and improves source coverage systematically.

  5. Add a Reranking Score with Weighted Multi-Factor Criteria – Rerank retrieved chunks using a composite score of relevance, diversity, coverage, authority, and salience. This balances topical fit with representativeness across political groups, preventing dominance by any single speaker or faction.

  6. Incorporate Time-Decayed Activity Signals – Weight recent legislative acts and speech interventions more heavily than older ones, making authority estimates contextually current and reducing the influence of outdated expertise.

What the Improved AI System Can Do:

  • Answer policy questions with balanced, multi-perspective summaries that explicitly cover all major political groups (per-group coverage score of 0.97), avoiding echo chambers or majority-speaker bias.

  • Guarantee that every cited quote is verbatim and correctly attributed to the original speaker and session, with zero hallucinated quotations—critical for legal, parliamentary, or journalistic applications.

  • Identify and prioritize the most topic-relevant experts for any given query, even if they are not frequent speakers, by leveraging their committee roles, education, and legislative track record.

  • Retrieve evidence from both semantic similarity and relational graph structures, capturing connections (e.g., bill signatories, committee memberships) that pure text search would miss.

  • Produce source-grounded answers with higher authority and coverage than general-purpose RAG systems (e.g., 4.64 vs. 4.39 on source coverage in human evaluation), while maintaining comparable overall satisfaction.

  • Support transparent auditing of why a speaker was considered authoritative, since the authority score is composed of interpretable, weighted components rather than a black-box neural model.

  • Mitigate fluency limitations by separating factual grounding (architectural guarantees) from prose generation, allowing future upgrades to stronger language models without risking citation integrity.

Abstract

Parliamentary proceedings are a primary record of democratic deliberation, yet their volume and fragmentation make multi-perspective access difficult for citizens, journalists, and researchers. Applying Retrieval-Augmented Generation (RAG) to parliamentary transcripts introduces three specific risks: dominance of the most frequent speakers, inability to weight speakers according to topical expertise, and citation misattribution in politically sensitive text. We present ParliamentRAG, a RAG system for the Italian Chamber of Deputies that addresses these risks jointly. Its core contribution is a topic-dependent authority model that estimates each speaker's authority as a function of the current query, combining interpretable components such as profession, education, and previous interventions. Given a user query, the system retrieves relevant speech chunks, identifies topic-relevant experts across parliamentary groups, and generates a summary synthesizing their perspectives, accompanied by supporting quotations. ParliamentRAG is evaluated against Google NotebookLM on 15 policy topics via a two-level protocol combining automated metrics and blind A/B human evaluation by six domain experts. The system achieves higher coverage across political groups (0.97 vs. 0.95), perfect quotation faithfulness (1.00 vs. 0.95), and stronger expert preferences on source-related dimensions, while NotebookLM remains stronger on prose-oriented dimensions.

Sources

Related papers