On Measuring Semantic Preservation in Legal Ontology Learning

arXiv:2608.12326 · cs.CL · Submitted 2026-05-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "On Measuring Semantic Preservation in Legal Ontology Learning".

Jane: The paper was written by Albert Sadowski and Jarosław A. Chudziak from Warsaw University of Technology.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's got a title that just rolls off the tongue — "On Measuring Semantic Preservation in Legal Ontology Learning." Jane, I gotta say, even the title makes me lean in. It's about whether we lose meaning when we turn legal documents into structured data.

Jane: Oh, absolutely, Tom. And that's such a huge question, because we're at this moment where everyone's using AI to organize information. But the paper is asking something really basic: when you compress a contract into a structured format, does the important stuff survive? It's like photocopying a photo — you get a copy, but you might lose the color.

Tom: Right, and the authors, Albert Sadowski and Jarosław Chudziak from Warsaw University of Technology, they're not just theorizing. They actually tested it. They took real merger agreements, turned them into ontologies using three different methods, and then asked six different language models to answer questions from both the original text and the transformed versions.

Jane: And the results were pretty striking. Every single method lost some meaning. Some lost a lot. The worst case, they saw over twenty-seven percentage points drop in accuracy. That's not a small hiccup, that's a serious degradation.

Lu: I have to jump in here, because what excites me is the methodology. They're using the LLM itself as the measuring stick. Instead of some abstract metric, they're saying, "Look, if the model can answer questions from the original text, that's our baseline. If it can't answer from the ontology, we've lost something." That's clever because it's practical.

Meng: Yeah, but as an engineer, my first thought is, "Well, of course you lose information when you abstract." The real question is whether you gain something else — like the ability to reason formally, or to query the data in a structured way. The paper acknowledges that, but they're measuring the loss side, which is what's been missing.

Tom: And that's the gap they're filling. Before this, people were checking if ontologies were structurally correct — did they have the right hierarchy, the right axioms. But nobody was checking if the meaning survived. It's like checking that a bridge is painted but not checking if it can hold cars.

Jane: Exactly. And the paper's contribution is giving us a way to measure that. It's not perfect, but it's a starting point. And the fact that they found such consistent loss across all models and methods tells us this is a real problem, not just a quirk of one system.

Lu: I'd push back a little on the consistency, though. They found that the model-method pairing matters a lot. Gemini with one method preserved eighty-eight percent of the meaning on a hard task, while another pairing preserved only thirteen percent. So it's not just "ontologies are bad" — it's "you need to pick the right tool for the right job."

Meng: And that's the kind of practical insight I can actually use. If I'm building a legal knowledge system, I need to know which combination works. This paper gives me a framework to test that, rather than just hoping it works.

Jane: So we've got a problem — semantic loss — and we've got a measurement tool. But what does that mean for the future? That's what we're going to dig into next.

Summary: Tom: So we've established that "On Measuring Semantic Preservation in Legal Ontology Learning" is tackling a real problem. Now let's talk about what they actually did, because the setup is pretty elegant.

Jane: Right. They used a dataset called MAUD — that's the Merger Agreement Understanding Dataset from LegalBench. It's got thirty-four different tasks about merger agreements, things like "what counts as a material adverse effect" or "when can the board change its recommendation." They took sixty-nine examples from each task, so about two thousand three hundred total.

Tom: And then they ran three different ontology learning methods. The first is LLMs4OL, which is like a parallel approach — it does three tasks at once: classifying terms, finding hierarchies, and extracting relationships. The second is NeOn-GPT, which is a structured five-stage pipeline. And the third is NeOn-CoT, which is a simplified version that uses chain-of-thought reasoning in one go.

Lu: What I find interesting is the difference in philosophy. LLMs4OL is completely domain-agnostic — it doesn't know it's looking at legal documents. NeOn-GPT and NeOn-CoT are explicitly told, "Hey, this is a merger agreement, here's the legal context." And that domain awareness seems to matter.

Jane: It really does. NeOn-GPT had the best preservation overall, with losses between about nine and eighteen percentage points. LLMs4OL was the worst, with losses up to twenty-seven points. So knowing what you're looking at helps you not lose it.

Meng: But here's the thing that caught my eye — the numbers aren't uniform across models. Claude Sonnet four lost more than Gemini two point five Flash, even with the same method. So the model's own architecture affects how much meaning survives the transformation.

Tom: And that's a big deal, because it means you can't just pick a method and assume it works everywhere. You have to test the combination. The paper even found that Gemini paired with NeOn-CoT was a standout — it kept eighty-six to eighty-eight percent of the meaning on some of the hardest tasks, while other pairings dropped to thirteen percent.

Jane: Yeah, that interaction effect is the hidden story here. It's not just "ontologies lose information" — it's "some combinations lose way more than others." And the paper gives us a way to find the good combinations.

Lu: I also appreciate that they separated success cases from failure cases. When the baseline model already got the answer right, the ontology kept most of that — sixty to eighty-three percent for NeOn-GPT. But when the baseline model got it wrong, the ontology rarely helped — only nine to thirty-six percent accuracy. So ontologies aren't adding new knowledge, they're just reorganizing what's already there.

Meng: That's a sobering finding for anyone hoping ontologies would magically fix model limitations. They don't. They just rearrange the furniture.

Tom: And that's the summary — a clear problem, a clear measurement, and a clear warning that you need to test before you trust. But what does the paper suggest we do about it? That's next.

Improvements: Tom: Alright, so we know the problem — semantic loss — and we know how they measured it. But what does "On Measuring Semantic Preservation in Legal Ontology Learning" actually suggest we do differently? What's the path forward?

Jane: Well, the paper is pretty honest that this isn't a problem you can just algorithmically fix. The loss comes from the abstraction itself. When you take a phrase like "more likely than not" and turn it into a generic concept like "likely violation of fiduciary duties," you've lost the precision that makes legal language work.

Lu: That's the deep point, Jane. It's not a bug in the code — it's a fundamental property of abstraction. You're compressing rich, nuanced language into categories, and categories are lossy by nature. The paper calls this "intrinsic to the abstraction process." So the improvement isn't "make better ontologies," it's "know when ontologies are the right tool."

Meng: And that's where I see the practical improvement. They're suggesting selective transformation. Don't turn everything into an ontology. Some content — like the simple binary questions, "is this a stock deal or an asset deal?" — survives fine. Other content — like fiduciary duty standards with all their modal language — loses too much. So transform the simple stuff, keep the complex stuff in natural language.

Tom: That's a really practical takeaway. It's not all-or-nothing. You can have a hybrid system where the ontology handles the structured parts and the LLM handles the nuanced parts.

Jane: And they also point to model-method optimization. Since Gemini paired with NeOn-CoT worked so well, the suggestion is to systematically explore which model works best with which method for which type of task. Instead of assuming one approach works everywhere, you test and find the best pairing.

Lu: I'd add that the paper's evaluation framework itself is the improvement. By giving us a way to measure semantic preservation, they're enabling evidence-based decisions. Before this, you'd build an ontology, check that it's structurally valid, and assume it's fine. Now you can actually ask, "Did the meaning survive?" and get a number.

Meng: And that number is actionable. If I'm building a legal AI system, I can run this test on my specific documents and my specific model, and see whether the ontology is helping or hurting. That's a huge improvement over guessing.

Tom: So the improvements are: measure semantic preservation, use that measurement to decide when to transform, and optimize the model-method pairing. But what does this look like in practice? What's the actual first page of the paper telling us?

Jane: Good question. Let's look at the opening arguments and see what they're really claiming.

First Page: Tom: So we're digging into the first page of "On Measuring Semantic Preservation in Legal Ontology Learning." Jane, what stood out to you when you read the opening?

Jane: The framing, honestly. They start by saying that natural language works great for humans but poorly as a communication protocol between computers. It's ambiguous, it's inconsistent. That's why we build ontologies — to give machines a controlled vocabulary they can actually reason with.

Lu: And that's the tension they're highlighting. Ontologies exist to make information machine-readable, but in doing so, they can strip away the meaning that made the information valuable in the first place. The first page sets up this trade-off very clearly.

Meng: I also noticed they're positioning this as a complement to existing evaluation methods. Structural metrics — checking axioms, hierarchy depth, logical consistency — those are still useful. But they don't tell you if the ontology actually supports the reasoning tasks it was built for. That's the gap this paper fills.

Tom: Right, and they make a specific claim: an ontology can pass all structural tests and still fail to support the reasoning it was designed for. That's a bold statement, and it's backed up by their results.

Jane: They also introduce the idea of using LLM performance as a proxy for semantic accessibility. It's a clever move because it's practical — you're measuring what a model can actually extract from the representation, not some abstract notion of meaning.

Lu: I think the key insight on that first page is that they're measuring relative preservation, not absolute semantic content. They're not claiming to know the "true" meaning of a document. They're saying, "Here's what the model can extract from the original, and here's what it can extract from the ontology. The difference is the loss." That's a defensible, testable claim.

Meng: And it's a claim I can work with. I don't need a philosophical definition of meaning. I need to know if my system performs better or worse after transformation. This gives me exactly that.

Tom: So the first page sets up the problem, the methodology, and the promise. But I want to step back — what does this mean for the broader world? Who should care about this?

Jane: Anyone building legal tech, for starters. Contract analysis, due diligence, regulatory compliance — these are all areas where ontologies are being used. And this paper says, "Check whether you're actually preserving the meaning before you trust the system."

Lu: I'd extend that beyond law. Any domain where precise language matters — medicine, finance, engineering standards — faces the same issue. If you're abstracting clinical guidelines or safety regulations into structured formats, you need to know what you're losing.

Meng: And the framework is domain-agnostic. You can apply it to any text and any ontology learning method. That's the real contribution — a general tool for a general problem.

Conclusion: Tom: Alright, we've spent a good chunk of time with "On Measuring Semantic Preservation in Legal Ontology Learning," and I think we've got a clear picture. Jane, want to wrap it up?

Jane: Absolutely. The paper gives us two big things. First, a framework for measuring semantic preservation — using LLM performance on source text versus transformed text to quantify what's lost. Second, empirical evidence that semantic loss is real and varies dramatically based on which model and which ontology method you pair together.

Lu: And the numbers back it up. Losses ranged from about eight to twenty-seven percentage points across all combinations. NeOn-GPT was the best overall, LLMs4OL was the worst, and the model-method interaction mattered more than anyone expected.

Meng: For me, the practical takeaway is clear: don't assume ontologies are helping just because they're structurally valid. Test them. Use this framework to measure whether the meaning survives for your specific use case.

Tom: And that's the legacy of this paper — it moves ontology evaluation from structural correctness to semantic utility. It gives practitioners a way to make evidence-based decisions instead of hoping for the best.

Jane: We should also note the limitations. It's one legal subdomain, multiple-choice tasks, and the framework relies on LLM performance as a proxy. But as a starting point, it's a solid one.

Lu: And it opens up a lot of future work — selective transformation, hybrid systems, model-method optimization. There's a whole research agenda here.

Tom: Well said. So we're saying goodbye to "On Measuring Semantic Preservation in Legal Ontology Learning." Thanks to Sadowski and Chudziak for the thoughtful work. We'll be back next time with another paper, but for now, this is Tom and Jane signing off.

Jane: Take care, everyone. Keep asking whether the meaning survives.

Albert Sadowski, Jarosław A. Chudziak

Warsaw University of Technology · Warsaw University of Technology

cs.CL

Submitted: 2026-05-30

Comments: Accepted for publication at the 30th International Conference on Knowledge-Based and Intelligent Information & Engineering Systems (KES 2026)

Code: https://github.com/albsadowski/ontology-learning-eval

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 59/100

The gist: "Ontology learning transforms unstructured text into structured representations for automated reasoning.

Key concepts

Semantic Preservation
This refers to whether the original meaning of a document survives after it has been transformed into a structured format. The paper measures this by comparing an LLM's ability to answer questions using the original text versus its performance on the resulting ontology.
Ontology Learning
This is the process of converting unstructured text, such as legal contracts, into a structured knowledge base. The paper tests three methods—LLMs4OL, NeOn-GPT, and NeOn-CoT—to see how well they preserve meaning during this conversion.
Semantic Loss
The degree to which information is lost when transforming natural language into categories or structured data. The study found that semantic loss could be as high as twenty-seven percentage points, indicating a serious degradation of original meaning.

Terminology

Summary

Summary

The paper On Measuring Semantic Preservation in Legal Ontology Learning addresses a critical gap in evaluating ontology learning systems. The authors state: "Ontology learning transforms unstructured text into structured representations for automated reasoning. Yet structuring information risks losing it, and current evaluation methodologies cannot detect such loss, focusing on structural correctness while failing to measure whether meaning survives transformation. They propose an evaluation methodology that addresses this: comparing LLM task performance on source documents against performance on transformed representations, with the difference quantifying semantic loss."

The methodology is described as follows: Our approach treats baseline LLM performance on source documents as a model-specific upper bound on accessible semantic content, then measures how much of that content remains accessible after ontological transformation. The authors formalize this as: SemanticLoss = BaselineP erf − T ransf ormedP erf. They emphasize that our methodology measures semantic preservation relative to each LLM’s capabilities, accounting for model-specific baseline variations.

The paper evaluates three ontology learning approaches: LLMs4OL [1], NeOn-GPT [10], and a simplified NeOn-CoT method we introduce. LLMs4OL decomposes ontology construction into three parallel tasks: term typing, taxonomy discovery, and non-taxonomic relation extraction and is designed to be domain-agnostic. NeOn-GPT implements a structured five-stage pipeline: requirements specification, competency question generation, conceptual model creation, ontology draft implementation, and enrichment with instances and annotations and "explicitly incorporates domain awareness by identifying the text as mergers and acquisitions (M&A) contract fragments. NeOn-CoT consolidates the NeOn methodology principles into a single chain-of-thought prompting strategy while maintaining domain awareness by explicitly referencing the legal context and M&A domain."

The experimental setup uses the Merger Agreement Understanding Dataset (MAUD) from LegalBench [12, 29] containing 34 multiple-choice tasks focusing on merger agreement analysis. They selected 69 examples from each task (the minimum available across all tasks), yielding 2,346 total examples. Evaluation spans six language models: o4 mini (o4-mini-2025-04-16), GPT 4.1 mini (gpt-4.1-mini-2025-04-14), Claude Sonnet 4 (claude-sonnet-4-20250514), Gemini 2.5 Flash (gemini-2.5-flash-preview-05-20), DeepSeek v3, and Llama 4 Maverick. Each model generates its own ontologies, and Temperature is set to 0 where supported.

The results reveal substantial semantic loss across all ontology learning approaches, with performance degradation ranging from 8.4 percentage points (pp) to 27.4pp compared to baseline accessibility levels. Specifically: "LLMs4OL exhibits the most severe semantic loss (13.9-27.4pp), likely due to information fragmentation during parallel task decomposition. NeOn-GPT demonstrates the best preservation characteristics (8.8-18.1pp.), suggesting that explicit domain awareness and sequential processing better preserve legal document semantics. NeOn-CoT shows intermediate performance with notable variability (8.4-25.3pp.), trading some preservation benefits for computational efficiency."

The authors report: "When baseline models successfully extract information, ontology learning methods retain 45-83% of accessible content, with NeOn-GPT achieving the highest success preservation rates (60-83%). Analysis of failure cases reveals limited compensatory benefits from ontological transformation, with most approaches showing only 9-36% accuracy on baseline failures, indicating that ontology learning primarily reorganizes existing accessible content rather than making previously inaccessible information available."

Cross-model consistency is noted: "The performance hierarchy NeOn-GPT > NeOn-CoT > LLMs4OL holds across all models tested, despite architectural differences. However, absolute sensitivity varies considerably across models. Claude Sonnet 4 shows higher sensitivity (45-61% preservation rates), while Gemini 2.5 Flash demonstrates more robust preservation (57-83% rates). All differences were statistically significant (Wilcoxon signed-rank tests, p < 0.001), with effect sizes ranging from medium (d = 0.53 for NeOn-GPT) to large (d = 0.85 for LLMs4OL)."

Task-level analysis reveals: "The most severe semantic loss occurs with nuanced legal standards and modal language defining procedural requirements. Task t13 (fiduciary duty standards for superior offer recommendations) shows the highest average loss at 65%, followed by t12 (fiduciary duty standards for intervening event recommendations) at 53%. The authors explain: ontological representations abstract this to generic concepts like LikelyV iolationOf F iduciaryDuties, losing precise probabilistic language (more likely than not vs. reasonably likely vs. could reasonably be expected) that distinguishes legal standards. Conversely, tasks involving straightforward binary determinations demonstrate near-perfect preservation across all combinations. Task t19 (whether financial point of view is sole consideration for superior offers) achieved best preservation with only 1% average loss."

Critically, the authors identify significant model-method interactions: Gemini paired with NeOn-CoT achieves 88% semantic retention on t13, and on t12 Gemini with NeOn-CoT maintains 86% baseline performance compared to just 13% with LLMs4OL. They note: Gemini consistently outperforms other models with NeOn methods, winning 9 out of 10 challenging tasks analyzed.

The discussion states: "Semantic loss appears intrinsic to the abstraction process - when ontological representations transform 'more likely than not' into generic concepts like LikelyV iolationOf F iduciaryDuties, critical legal distinctions vanish irreversibly. This is not a technical problem to be solved through better algorithms, but a fundamental characteristic of structural abstraction that practitioners must acknowledge. The authors also note: Despite systematic semantic loss during transformation, ontologies remain important in the LLM era for bridging neural and symbolic systems."

The contributions are stated as: "(1) an evaluation framework for measuring semantic preservation in ontology learning, and (2) empirical evidence that semantic loss varies dramatically with model-method pairing, providing guidance for selecting optimal configurations in legal knowledge systems."

Limitations acknowledged include: The evaluation covers a single legal subdomain, merger and acquisition contract analysis through MAUD, The task format is multiple-choice, Coverage of ontology learning approaches is limited to three methods, and the framework does not credit benefits that would materialize only through formal reasoning such as SPARQL queries or automated inference. Future work directions include Domain generalization studies, Computational context evaluation, Selective transformation approaches, and systematic model-method optimization.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems, along with what the improved system can do:


Improvement 1: Add a semantic-preservation evaluation layer to ontology-learning pipelines

  • What to change: After any ontology learning method (LLMs4OL, NeOn-GPT, NeOn-CoT, or similar) transforms source text into a structured representation, run a task-based evaluation: have the same LLM answer a set of downstream questions from (a) the original text and (b) the ontology. Compute the accuracy difference as a semantic loss score and report it alongside structural metrics (axiom counts, hierarchy depth, etc.).

  • What the improved system can do: Automatically flag transformations that lose critical meaning (e.g., legal standards like more likely than not being abstracted to generic concepts). Practitioners can reject or revise an ontology before deployment if the semantic loss exceeds a threshold, rather than discovering the failure after integration into a reasoning system.

Improvement 2: Implement model–method pairing selection for ontology learning

Improvement 3: Add a nuance-preservation module for modal and probabilistic language

Improvement 4: Add a selective-transformation decision mechanism

Improvement 5: Integrate a semantic loss dashboard into knowledge-system monitoring

Summary of what the improved AI system can do:

  • Measure whether meaning survives ontological transformation, not just whether the structure is valid.

  • Select the best model–method combination for a given domain automatically.

  • Preserve critical legal nuance by attaching original phrasing to ontological concepts.

  • Decide when to transform and when to keep text as-is, based on predicted task difficulty.

  • Monitor semantic preservation over time and alert on degradation.

These changes directly address the paper's core finding: semantic loss is systematic, varies dramatically with model–method pairing, and is invisible to current structural evaluation. Implementing them turns ontology learning from a blind transformation into an evidence-based, quality-controlled process.

Abstract

Ontology learning transforms unstructured text into structured representations for automated reasoning. Yet structuring information risks losing it, and current evaluation methodologies cannot detect such loss, focusing on structural correctness while failing to measure whether meaning survives transformation. We propose an evaluation methodology that addresses this: comparing LLM task performance on source documents against performance on transformed representations, with the difference quantifying semantic loss. We demonstrate this approach on legal merger agreement analysis, a domain chosen for its complex language and precise semantic requirements, comparing direct LLM application against three ontology learning methods across six language models. The results reveal systematic semantic loss with significant variation based on reasoning complexity and model-method interactions. Our contributions are: (1) an evaluation framework for measuring semantic preservation in ontology learning, and (2) empirical evidence that semantic loss varies dramatically with model-method pairing, providing guidance for selecting optimal configurations in legal knowledge systems.

Sources

Related papers