On Measuring Semantic Preservation in Legal Ontology Learning

summary

Video file (mp4)

The gist

"Ontology learning transforms unstructured text into structured representations for automated reasoning.

In short

The episode discusses a paper titled "On Measuring Semantic Preservation in Legal Ontology Learning." The authors developed a framework to measure whether meaning is lost when converting legal documents, like merger agreements, into structured data (ontologies). They found that semantic loss is real and varies significantly depending on the chosen AI model and ontology method.

Key concepts

Semantic Preservation
This refers to whether the original meaning of a document survives after it has been transformed into a structured format. The paper measures this by comparing an LLM's ability to answer questions using the original text versus its performance on the resulting ontology.
Ontology Learning
This is the process of converting unstructured text, such as legal contracts, into a structured knowledge base. The paper tests three methods—LLMs4OL, NeOn-GPT, and NeOn-CoT—to see how well they preserve meaning during this conversion.
Semantic Loss
The degree to which information is lost when transforming natural language into categories or structured data. The study found that semantic loss could be as high as twenty-seven percentage points, indicating a serious degradation of original meaning.

Terminology used across episodes

This episode discusses

The paper

On Measuring Semantic Preservation in Legal Ontology Learning · Read on arXiv

Albert Sadowski, Jarosław A. Chudziak

Warsaw University of Technology · Warsaw University of Technology

Ontology learning transforms unstructured text into structured representations for automated reasoning. Yet structuring information risks losing it, and current evaluation methodologies cannot detect such loss, focusing on structural correctness while failing to measure whether meaning survives transformation. We propose an evaluation methodology that addresses this: comparing LLM task performance on source documents against performance on transformed representations, with the difference quantifying semantic loss. We demonstrate this approach on legal merger agreement analysis, a domain chosen for its complex language and precise semantic requirements, comparing direct LLM application against three ontology learning methods across six language models. The results reveal systematic semantic loss with significant variation based on reasoning complexity and model-method interactions. Our contributions are: (1) an evaluation framework for measuring semantic preservation in ontology learning, and (2) empirical evidence that semantic loss varies dramatically with model-method pairing, providing guidance for selecting optimal configurations in legal knowledge systems.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "On Measuring Semantic Preservation in Legal Ontology Learning".

Jane: The paper was written by Albert Sadowski and Jarosław A. Chudziak from Warsaw University of Technology.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's got a title that just rolls off the tongue — "On Measuring Semantic Preservation in Legal Ontology Learning." Jane, I gotta say, even the title makes me lean in. It's about whether we lose meaning when we turn legal documents into structured data.

Jane: Oh, absolutely, Tom. And that's such a huge question, because we're at this moment where everyone's using AI to organize information. But the paper is asking something really basic: when you compress a contract into a structured format, does the important stuff survive? It's like photocopying a photo — you get a copy, but you might lose the color.

Tom: Right, and the authors, Albert Sadowski and Jarosław Chudziak from Warsaw University of Technology, they're not just theorizing. They actually tested it. They took real merger agreements, turned them into ontologies using three different methods, and then asked six different language models to answer questions from both the original text and the transformed versions.

Jane: And the results were pretty striking. Every single method lost some meaning. Some lost a lot. The worst case, they saw over twenty-seven percentage points drop in accuracy. That's not a small hiccup, that's a serious degradation.

Lu: I have to jump in here, because what excites me is the methodology. They're using the LLM itself as the measuring stick. Instead of some abstract metric, they're saying, "Look, if the model can answer questions from the original text, that's our baseline. If it can't answer from the ontology, we've lost something." That's clever because it's practical.

Meng: Yeah, but as an engineer, my first thought is, "Well, of course you lose information when you abstract." The real question is whether you gain something else — like the ability to reason formally, or to query the data in a structured way. The paper acknowledges that, but they're measuring the loss side, which is what's been missing.

Tom: And that's the gap they're filling. Before this, people were checking if ontologies were structurally correct — did they have the right hierarchy, the right axioms. But nobody was checking if the meaning survived. It's like checking that a bridge is painted but not checking if it can hold cars.

Jane: Exactly. And the paper's contribution is giving us a way to measure that. It's not perfect, but it's a starting point. And the fact that they found such consistent loss across all models and methods tells us this is a real problem, not just a quirk of one system.

Lu: I'd push back a little on the consistency, though. They found that the model-method pairing matters a lot. Gemini with one method preserved eighty-eight percent of the meaning on a hard task, while another pairing preserved only thirteen percent. So it's not just "ontologies are bad" — it's "you need to pick the right tool for the right job."

Meng: And that's the kind of practical insight I can actually use. If I'm building a legal knowledge system, I need to know which combination works. This paper gives me a framework to test that, rather than just hoping it works.

Jane: So we've got a problem — semantic loss — and we've got a measurement tool. But what does that mean for the future? That's what we're going to dig into next.

Summary: Tom: So we've established that "On Measuring Semantic Preservation in Legal Ontology Learning" is tackling a real problem. Now let's talk about what they actually did, because the setup is pretty elegant.

Jane: Right. They used a dataset called MAUD — that's the Merger Agreement Understanding Dataset from LegalBench. It's got thirty-four different tasks about merger agreements, things like "what counts as a material adverse effect" or "when can the board change its recommendation." They took sixty-nine examples from each task, so about two thousand three hundred total.

Tom: And then they ran three different ontology learning methods. The first is LLMs4OL, which is like a parallel approach — it does three tasks at once: classifying terms, finding hierarchies, and extracting relationships. The second is NeOn-GPT, which is a structured five-stage pipeline. And the third is NeOn-CoT, which is a simplified version that uses chain-of-thought reasoning in one go.

Lu: What I find interesting is the difference in philosophy. LLMs4OL is completely domain-agnostic — it doesn't know it's looking at legal documents. NeOn-GPT and NeOn-CoT are explicitly told, "Hey, this is a merger agreement, here's the legal context." And that domain awareness seems to matter.

Jane: It really does. NeOn-GPT had the best preservation overall, with losses between about nine and eighteen percentage points. LLMs4OL was the worst, with losses up to twenty-seven points. So knowing what you're looking at helps you not lose it.

Meng: But here's the thing that caught my eye — the numbers aren't uniform across models. Claude Sonnet four lost more than Gemini two point five Flash, even with the same method. So the model's own architecture affects how much meaning survives the transformation.

Tom: And that's a big deal, because it means you can't just pick a method and assume it works everywhere. You have to test the combination. The paper even found that Gemini paired with NeOn-CoT was a standout — it kept eighty-six to eighty-eight percent of the meaning on some of the hardest tasks, while other pairings dropped to thirteen percent.

Jane: Yeah, that interaction effect is the hidden story here. It's not just "ontologies lose information" — it's "some combinations lose way more than others." And the paper gives us a way to find the good combinations.

Lu: I also appreciate that they separated success cases from failure cases. When the baseline model already got the answer right, the ontology kept most of that — sixty to eighty-three percent for NeOn-GPT. But when the baseline model got it wrong, the ontology rarely helped — only nine to thirty-six percent accuracy. So ontologies aren't adding new knowledge, they're just reorganizing what's already there.

Meng: That's a sobering finding for anyone hoping ontologies would magically fix model limitations. They don't. They just rearrange the furniture.

Tom: And that's the summary — a clear problem, a clear measurement, and a clear warning that you need to test before you trust. But what does the paper suggest we do about it? That's next.

Improvements: Tom: Alright, so we know the problem — semantic loss — and we know how they measured it. But what does "On Measuring Semantic Preservation in Legal Ontology Learning" actually suggest we do differently? What's the path forward?

Jane: Well, the paper is pretty honest that this isn't a problem you can just algorithmically fix. The loss comes from the abstraction itself. When you take a phrase like "more likely than not" and turn it into a generic concept like "likely violation of fiduciary duties," you've lost the precision that makes legal language work.

Lu: That's the deep point, Jane. It's not a bug in the code — it's a fundamental property of abstraction. You're compressing rich, nuanced language into categories, and categories are lossy by nature. The paper calls this "intrinsic to the abstraction process." So the improvement isn't "make better ontologies," it's "know when ontologies are the right tool."

Meng: And that's where I see the practical improvement. They're suggesting selective transformation. Don't turn everything into an ontology. Some content — like the simple binary questions, "is this a stock deal or an asset deal?" — survives fine. Other content — like fiduciary duty standards with all their modal language — loses too much. So transform the simple stuff, keep the complex stuff in natural language.

Tom: That's a really practical takeaway. It's not all-or-nothing. You can have a hybrid system where the ontology handles the structured parts and the LLM handles the nuanced parts.

Jane: And they also point to model-method optimization. Since Gemini paired with NeOn-CoT worked so well, the suggestion is to systematically explore which model works best with which method for which type of task. Instead of assuming one approach works everywhere, you test and find the best pairing.

Lu: I'd add that the paper's evaluation framework itself is the improvement. By giving us a way to measure semantic preservation, they're enabling evidence-based decisions. Before this, you'd build an ontology, check that it's structurally valid, and assume it's fine. Now you can actually ask, "Did the meaning survive?" and get a number.

Meng: And that number is actionable. If I'm building a legal AI system, I can run this test on my specific documents and my specific model, and see whether the ontology is helping or hurting. That's a huge improvement over guessing.

Tom: So the improvements are: measure semantic preservation, use that measurement to decide when to transform, and optimize the model-method pairing. But what does this look like in practice? What's the actual first page of the paper telling us?

Jane: Good question. Let's look at the opening arguments and see what they're really claiming.

First Page: Tom: So we're digging into the first page of "On Measuring Semantic Preservation in Legal Ontology Learning." Jane, what stood out to you when you read the opening?

Jane: The framing, honestly. They start by saying that natural language works great for humans but poorly as a communication protocol between computers. It's ambiguous, it's inconsistent. That's why we build ontologies — to give machines a controlled vocabulary they can actually reason with.

Lu: And that's the tension they're highlighting. Ontologies exist to make information machine-readable, but in doing so, they can strip away the meaning that made the information valuable in the first place. The first page sets up this trade-off very clearly.

Meng: I also noticed they're positioning this as a complement to existing evaluation methods. Structural metrics — checking axioms, hierarchy depth, logical consistency — those are still useful. But they don't tell you if the ontology actually supports the reasoning tasks it was built for. That's the gap this paper fills.

Tom: Right, and they make a specific claim: an ontology can pass all structural tests and still fail to support the reasoning it was designed for. That's a bold statement, and it's backed up by their results.

Jane: They also introduce the idea of using LLM performance as a proxy for semantic accessibility. It's a clever move because it's practical — you're measuring what a model can actually extract from the representation, not some abstract notion of meaning.

Lu: I think the key insight on that first page is that they're measuring relative preservation, not absolute semantic content. They're not claiming to know the "true" meaning of a document. They're saying, "Here's what the model can extract from the original, and here's what it can extract from the ontology. The difference is the loss." That's a defensible, testable claim.

Meng: And it's a claim I can work with. I don't need a philosophical definition of meaning. I need to know if my system performs better or worse after transformation. This gives me exactly that.

Tom: So the first page sets up the problem, the methodology, and the promise. But I want to step back — what does this mean for the broader world? Who should care about this?

Jane: Anyone building legal tech, for starters. Contract analysis, due diligence, regulatory compliance — these are all areas where ontologies are being used. And this paper says, "Check whether you're actually preserving the meaning before you trust the system."

Lu: I'd extend that beyond law. Any domain where precise language matters — medicine, finance, engineering standards — faces the same issue. If you're abstracting clinical guidelines or safety regulations into structured formats, you need to know what you're losing.

Meng: And the framework is domain-agnostic. You can apply it to any text and any ontology learning method. That's the real contribution — a general tool for a general problem.

Conclusion: Tom: Alright, we've spent a good chunk of time with "On Measuring Semantic Preservation in Legal Ontology Learning," and I think we've got a clear picture. Jane, want to wrap it up?

Jane: Absolutely. The paper gives us two big things. First, a framework for measuring semantic preservation — using LLM performance on source text versus transformed text to quantify what's lost. Second, empirical evidence that semantic loss is real and varies dramatically based on which model and which ontology method you pair together.

Lu: And the numbers back it up. Losses ranged from about eight to twenty-seven percentage points across all combinations. NeOn-GPT was the best overall, LLMs4OL was the worst, and the model-method interaction mattered more than anyone expected.

Meng: For me, the practical takeaway is clear: don't assume ontologies are helping just because they're structurally valid. Test them. Use this framework to measure whether the meaning survives for your specific use case.

Tom: And that's the legacy of this paper — it moves ontology evaluation from structural correctness to semantic utility. It gives practitioners a way to make evidence-based decisions instead of hoping for the best.

Jane: We should also note the limitations. It's one legal subdomain, multiple-choice tasks, and the framework relies on LLM performance as a proxy. But as a starting point, it's a solid one.

Lu: And it opens up a lot of future work — selective transformation, hybrid systems, model-method optimization. There's a whole research agenda here.

Tom: Well said. So we're saying goodbye to "On Measuring Semantic Preservation in Legal Ontology Learning." Thanks to Sadowski and Chudziak for the thoughtful work. We'll be back next time with another paper, but for now, this is Tom and Jane signing off.

Jane: Take care, everyone. Keep asking whether the meaning survives.

More episodes

← Home