THIVLVC: Retrieval Augmented Dependency Parsing for Latin
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "THIVLVC: Retrieval Augmented Dependency Parsing for Latin".
Jane: The paper was written by Luc Pommeret, Thibault Wagret and Jules Deret from Université Paris-Saclay and CNRS and LISN and École Normale Supérieure de Lyon and HISOMA.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Summary: Tom: Now, let’s break down how this two-stage pipeline works, because the system itself is quite elegant in its simplicity.
Jane: The core idea is that the AI doesn't work in a vacuum; first, there is an information retrieval phase designed to find relevant linguistic parallels before we even get to the LLM.
Lu: The structural retriever looks at things like sentence length and POS bigrams to match our query sentence against seven hundred sixty-two sentences in the CIRCSE database. It’s not just a keyword match at this initial stage.
Meng: That's a very specific filtering mechanism; it ensures that the physical and grammatical structure of the retrieved example is similar enough to guide our search, which is crucial for accuracy.
Tom: The second stage takes all this information—the baseline parse from UDPipe, the official UD guidelines, and these structural examples—and passes them to Gemini-three-flash. It's a massive amount of data input.
Jane: It’s a sophisticated prompting process where the LLM acts as a "Latin Chief Annotator," comparing all these inputs against the established rules we provided in the prompt.
Lu: This is powerful because we are teaching the AI not just to follow rules blindly, but to act like it is making a highly informed decision based on established precedents from other authors.
Meng: The system gets five similar examples, which keeps the input manageable while ensuring high relevance for every single sentence in the test set we want to analyze.
Tom: So, instead of relying on one big neural network that's a mix of everything it read during training, we are using a targeted approach to find specific solutions.
Jane: We are selectively pulling in highly relevant historical context to guide the way we interpret the syntax of Latin text, making sure the LLM has all its tools.
Lu: This ensures that even if the initial automatic parser makes a common mistake, there is a high probability that similar examples from correcting patterns will teach us the correct pattern.
Meng: The practical impact here is that we are using AI to synthesize information across multiple valid approaches, not just making a single prediction based on static weights.
Lalam: This process allows us to see how the collective knowledge of historical linguistic solutions can be integrated into a machine's decision-making logic for cultural preservation.
Improvements: Tom: Moving into the results, the performance gains are quite impressive and depend heavily on which genre of text you're looking at.
Jane: The paper shows that for Classical poetry with Seneca, THIVLVC achieves a significant improvement of +seventeen points in CLAS when comparing it to the UDPipe baseline. That’s a huge jump in quality!
Lu: That large gain suggests the structural retriever is exceptionally good at identifying subtle syntactic nuances specific to poetic language that standard models miss them.
Meng: It’s interesting because while the gain on poetry was huge, they noted that for philosophical prose, the improvement was much more modest at just +one point five CLAS.
Tom: That makes sense if Thomas Aquinas's writing is already very structured and easy for a baseline parser to handle without much help from external context.
Jane: But even with that smaller gain in the prose category, the consistent improvement from making that RAG step shows how helpful context is across different types of texts.
Lu: The fact they used Gemini-three-flash also suggests they found a balance between powerful reasoning and practical efficiency for this complex parsing task, which is rare to find.
Meng: We need to look at the error analysis because it shows that out of three hundred divergences, annotators unanimously agreed on one hundred sixty-seven cases, and fifty-three point three percent of those unanimous decisions favored THIVLVC's output.
Tom: That is a huge finding for the field and indicates that many differences between the gold standard are not actually errors at all.
Jane: It really highlights that some of these disagreements are just variations in annotation convention or subtle linguistic ambiguity, rather than one definitive correct way to mark a relationship.
Lu: The taxonomy of disagreement they provided really breaks down why we need to rethink the way we measure success in this field, showing specific types of label confusion.
Meng: This is a crucial insight for us because it shows that if we use AI tools on other historical languages, we must be prepared to deal with inconsistent human standards as well.
Lalam: Seeing fifty-three percent of unanimous decisions favoring the system allows us to imagine a future where machine-assisted interpretation becomes a recognized part of our cultural understanding.
Conclusion: Tom: So, wrapping up our discussion on "THIVLVC: Retrieval Augmented Dependency Parsing for Latin," it’s a powerful system that uses the combination of structural similarity and LLM reasoning to achieve high accuracy.
Jane: It shows that the future work in this area needs to be more flexible with how we define correctness in linguistic analysis, moving beyond the idea of a single ground truth.
Lu: I hope we start seeing more systems that can handle these multiple valid interpretations rather than forcing one definitive answer onto them, allowing us to explore the full range of linguistic possibility.
Meng: My takeaway is that for large-scale AI applications dealing with historical data, this approach offers a practical pathway to incorporate rich contextual knowledge into complex reasoning tasks.
Lalam: For the cultural impact, it allows us to explore the nuances of ancient texts with more computational rigor than we have ever achieved before.
Tom: It’s a fantastic example of combining traditional linguistic knowledge with cutting-edge AI techniques in "THIVLVC: Retrieval Augmented Dependency Parsing for Latin."
Jane: We can't wait to see how other people approach this, especially when the next challenge comes along for historical texts.
Lu: I believe that is the direction we should all be looking toward, moving beyond a single gold standard and into a richer understanding of what makes language work.
Meng: It provides a practical blueprint for applying advanced AI methods to historical data sets that are inherently messy or inconsistent, making this valuable for any large dataset.
Lalam: It allows us to see the history of language not as a fixed object, but as something dynamic, constantly evolving through interpretation and context.
Conclusion: Tom: So, we've spent a lot of time digging into this research, but it’s clear that "THIVLVC: Retrieval Augmented Dependency Parsing for Latin" is far more than just another parsing tool; it represents a huge shift in how we treat historical texts.
Jane: It really shows that the biggest breakthrough isn't just achieving high accuracy, though—it’s acknowledging the reality of annotation inconsistency, which is something almost unavoidable when dealing with ancient sources.
Lu: I think this points toward a future where we don't just search for one single correct answer in AI models, but where they can gracefully handle and present multiple valid interpretations based on evidence.
Meng: And that means the technical challenge for implementation is huge, because we have to build systems that are designed not only to parse but also to manage and retrieve a vast library of potential historical solutions.
Lalam: From my perspective, this work allows us to preserve the full complexity of Latin literature, not just its most easily defined structure, which is a massive win for cultural understanding.
Tom: It’s truly impressive how this entire system bridges the gap between rigorous historical linguistics and cutting-edge AI architecture.
Jane: We’ve seen that the system works beautifully on poetry and competitive on prose, but what's more than that is the potential for human error to guide our machine learning process.
Lu: It suggests we are moving toward a more nuanced understanding of linguistic ambiguity rather than trying to force every piece of data into a rigid box.
Meng: I think the biggest practical implication will be how this allows us to apply these complex RAG strategies not just to Latin, but across various other challenging historical languages.
Lalam: It empowers us to see the evolution of language as a dynamic process, respecting the context and history embedded in every sentence.
Tom: It’s a powerful combination of structure and reasoning that we're excited to witness.
Jane: We hope this paper helps clear the path for future work in "THIVLVC: Retrieval Augmented Dependency Parsing for Latin," setting a new standard for what we can achieve with LLMs in historical analysis.
Tom: That’s all the time we have left for today, but I think this is just the beginning of a lot of conversations about how AI and ancient knowledge will interact.
Luc Pommeret, Thibault Wagret, Jules Deret
Université Paris-Saclay · CNRS · LISN · École Normale Supérieure de Lyon · HISOMA
cs.CL
Submitted: 2026-04-07
Updated: 2026-08-20
Code: https://github.com/l-pommeret/THIVLVC
Importance score: 83/100
The gist: THIVLVC is a system designed for the EvaLatin 2026 Dependency Parsing task, which involves parsing Latin texts across two genres: "Classical poetry with Seneca, and philosophical prose of Thomas
Key concepts
- Retrieval Augmented Dependency Parsing
- A system that enhances parsing by first finding relevant linguistic parallels (retrieval) before using an LLM. It guides the AI's decision-making process using established historical context to interpret syntax.
- Structural Retriever
- The initial phase of the system that analyzes a query sentence by looking at physical and grammatical structures, such as POS bigrams and sentence length. It matches these features against a large database (CIRCSE) to find similar examples.
- LLM (Gemini-three-flash)
- A Large Language Model used in the second stage. It acts as an 'Annotator,' receiving structural examples, baseline parses, and guidelines to synthesize information and make highly informed decisions about Latin syntax.
- Annotation Ambiguity
- The reality that human annotators may disagree on the correct way to mark linguistic relationships in ancient texts. The system's analysis highlights that these disagreements are often variations in convention rather than errors.
Terminology
Summary
THIVLVC is a system designed for the EvaLatin 2026 Dependency Parsing task, which involves parsing Latin texts across two genres: Classical poetry with Seneca, and philosophical prose of Thomas Aquinas.
The motivation for this approach stems from the limitations of traditional supervised neural models, which learn whatever patterns the training data contains, including, inevitably, annotation choices that predate current guidelines.
The system is described as a two-stage pipeline
consisting of: (1) retrieval of structurally similar sentences from CIRCSE, and (2) generation,
where a Large Language Model (LLM) refines the baseline parse.
Methodology and Architecture:
- Information Retrieval: Given an input sentence, the system retrieves k=5 most similar sentences from the training set of the CIRCSE treebank (762 sentences). This is achieved using a
structural retriever,
which calculates similarity based on a weighted combination of normalized length difference and Jaccard similarities over POS bigrams and trigrams. The similarity function is defined as:
simstruct (q, s) = 0.33 flen + 0.33 fbi + 0.34 ftri
- Generation: The retrieved examples, along with the
official UD annotation guidelines and the baseline parse from UDPipe,
are passed to gemini-3-flash (Google DeepMind). The LLM is prompted to act as a “Latin Chief Annotator,” tasked with comparing the baseline parse against the retrieved examples and guidelines, ultimately outputting arefined CoNLL-U block
for the input sentence.
The system was tested in two configurations: THIVLVC 1 (which uses LLM + UDPipe + UD guidelines, without retrieval
) and THIVLVC 2 (the same configuration, but with RAG on CIRCSE).
Evaluation and Results:
The system demonstrated significant performance gains over the UDPipe baseline. On poetry (Seneca), THIVLVC improves CLAS by +17 points over the UDPipe baseline.
On prose (Thomas Aquinas), the gain was more modest (the gain is +1.5 CLAS
). Table 2 further shows that THIVLVC 2 brings a consistent improvement, especially on prose, achieving a gain of +6.9 CLAS with subtypes
compared to UDPipe.
Error Analysis and Findings:
A detailed qualitative analysis was conducted on 300 divergences between the system and the gold standard. This analysis revealed that among unanimous annotator decisions, 53.3% favour THIVLVC,
indicating that a significant proportion of discrepancies are not parsing errors but inconsistencies in annotation practices.
The taxonomy of disagreements identified several categories:
-
Contradictions between CIRCSE and EvaLatin: An example was given where the error type
advmod:lmod(used in CIRCSE) versusadvmodaccounts for a significant portion of errors, illustratinga legitimate annotation divergence
between the two corpora. -
Internal gold inconsistency: Discrepancies arose within EvaLatin 2026 itself, such as using
oblinstead ofobl:arg, suggesting thatEvaLatin 2026 contains a degree of internal annotation noise.
-
Clear-cut human errors: Instances where the
relation mark is reserved in the UD guidelines for subordinating conjunctions... whereas prepositions governing noun phrases are annotated as case,
allowing THIVLVC to correctly identifycaseinstead ofmark. -
Undecidable ambiguity: Cases that
reflect genuine linguistic ambiguity that the UD guidelines do not fully resolve,
highlighting a structural limitation of single-analysis treebank annotation.
Improvements for AI systems
The following improvements address critical limitations in THIVLVC, focusing on enhancing robustness, generalizability, efficiency, and linguistic fidelity—factors essential when dealing with expensive, high-stakes AI systems.
Improvement: Replace the ad-hoc comparison of models with a systematic evaluation framework that includes Model Agnostic Prompting (MAP) and Chain-of-Thought (CoT) refinement. We will implement a structured prompt that forces the LLM to justify its decision against both the UD guidelines and specific retrieved examples, rather than just comparing them.
What the Improved System Can Do:
-
Robust Model Selection: The system can automatically select the optimal LLM (e.g., balancing performance vs. cost/latency) based on a pre-trained benchmark set, ensuring reproducible and justifiable deployment across different cloud providers or hardware configurations.
-
Justified Refinement: The LLM will not merely output a corrected parse; it will generate a brief
Rationale
field alongside the CoNLL-U block, detailing why the retrieved examples or UD guidelines supersede the baseline parse (e.g., "Subtype:lmodis required to denote spatial origin, as per Guideline 3.2").
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering