NL2SHACL-Bench: A Benchmark Suite for Natural Language to SHACL Translation

arXiv:2608.07530 · cs.AI, cs.CL, cs.DB · Submitted 2026-08-22 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "NL2SHACL-Bench: A Benchmark Suite for Natural Language to SHACL Translation".

Jane: The paper was written by Yuchen Zhou, Niels Bobet and Maribel Acosta from Technical University of Munich.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title and Authors: Tom: Welcome back to the show, everyone! Today we’re digging into a fresh arXiv paper that’s got me genuinely pumped. It’s called “NL2SHACL-Bench: A Benchmark Suite for Natural Language to SHACL Translation,” coming out of the Technical University of Munich.

Jane: And Tom, I have to say, the title alone tells a story. SHACL is this W3C standard for validating knowledge graphs — basically making sure the data in a graph database follows the rules it’s supposed to. But writing those rules is notoriously technical. This paper wants to let people just say what they need in plain English and have the computer turn that into SHACL.

Tom: Right, and the authors — Yuchen Zhou, Niels Bobet, and Maribel Acosta — they’re tackling a real gap. There are tools that generate SHACL from ontologies or from existing data, but nobody had built a proper benchmark to test whether an AI can go from natural language to SHACL. That’s what this suite is for.

Jane: And it’s not just about generating any SHACL. The paper makes a big point that two different SHACL shapes can look completely different but mean the exact same thing. So you can’t just compare text strings and say, “close enough.” You need to check whether the generated rules actually validate the same data the same way.

Tom: Exactly. And that’s why they built three things: a framework to create datasets, a dataset with two hundred forty human-verified natural-language-to-SHACL pairs across six domains, and a set of metrics that measure validity, structure, and semantics. It’s a complete toolkit.

Jane: I love that they pulled shapes from real-world sources — chemistry, public procurement, invoices, healthcare, even DBpedia. That’s not a toy setup. These are actual constraints people use in production.

Tom: And the implications? If this works well, domain experts who know their data but don’t know RDF or SHACL can finally write validation rules themselves. That’s a huge deal for data quality in industry.

Jane: It really is. And the authors are careful to say this is the first dedicated benchmark for this task. So they’re setting the standard for everyone else to build on.

Tom: We’re just getting started. Next up, we’re going to look at the core ideas in the paper’s abstract and how they frame the whole problem.

Paper Summary: Jane: So Tom, we’ve got the title and the authors down. Now let’s talk about what the paper actually claims. The abstract says that current large language models are really good at producing syntactically valid SHACL, but they still struggle with semantically equivalent constraints for complex logical and structural patterns.

Tom: And that’s a really honest assessment. They tested four state-of-the-art LLMs — Gemini three point one Pro, Claude Opus four point seven, Qwen3 point 5, and GLM five point one — and the validity rates were sky-high. Gemini got perfect scores on all three validity metrics. But when it came to whether the generated shapes actually enforced the same rules as the reference shapes, things got messy.

Jane: Right. The paper defines the task formally: given a natural language description and an ontology that provides the vocabulary, produce a SHACL shapes graph that encodes those constraints. And correctness means the generated graph is equivalent to a reference graph — meaning any data graph that conforms to one also conforms to the other.

Tom: That’s a clean definition. And they built their benchmark around it. The dataset has two hundred forty records, each with a natural language description, a reference SHACL shape, and an ontology snippet. They cover domains like chemical entities, DCAT metadata, e-procurement, electronic invoices, healthcare information management, and DBpedia.

Jane: What I found really clever is how they got the natural language descriptions. Some shapes already had descriptions written by humans, but most didn’t. So they used an LLM to draft descriptions, then had two human annotators review and fix them in two rounds. That’s a solid semi-automated pipeline.

Tom: And they didn’t just trust the LLM. The reviewers checked for hallucinations, missing constraints, overly technical phrasing — all the ways a description could mislead a downstream model.

Jane: The results are what really matter though. On the semantic equivalence metric, which is the gold standard, the models did great on simple datasets like DBpedia and ePO. But on Invoice and SNIK, only about half the records were semantically equivalent. That’s a big gap.

Tom: And the paper digs into why. The models struggle with inverse paths, sequential paths, and logical constraints using sh:or and sh:not. They sometimes replace sh:hasValue with sh:class, which changes the meaning entirely. And they hallucinate non-standard vocabulary like sh:condition.

Jane: So the takeaway is clear: LLMs can write valid SHACL, but they don’t always understand the semantics of complex constraints. That’s exactly the kind of insight a benchmark is supposed to surface.

Tom: Now, let’s get into the improvements the paper suggests. That’s where the future gets interesting.

Improvements Suggested: Tom: Jane, the paper doesn’t just stop at “here’s a benchmark and here’s how models fail.” It actually proposes concrete ways to make things better. And one of the most exciting is converting SHACL shapes into a common logical representation, like first-order logic.

Jane: Right, they mention SHACL2FOL specifically. That’s a tool that translates SHACL into first-order logic, which would let you check equivalence with mathematical certainty instead of relying on sampling data graphs. The current semantic metric generates synthetic data and compares which nodes violate the constraints — but that’s an approximation.

Tom: Exactly. And they’re honest about the blind spot. If a generated shape has a wrong target declaration and just happens to produce no violations on the sampled data, the metric might wrongly say it’s equivalent. Converting to logic would close that gap.

Jane: They also want to extend the benchmark to cover SHACL-SPARQL, which is the more expressive part of SHACL that uses SPARQL queries for constraints. Right now they exclude it because translating natural language to SPARQL is a whole separate research problem. But it’s a natural next step.

Tom: And there’s a practical angle too. They want to evaluate NL2SHACL under more realistic settings. Right now, the models are given the exact ontology terms and an in-context example. In the real world, you might not have that. You might have to figure out the right vocabulary yourself.

Jane: That’s a great point. The paper acknowledges that their current setup is somewhat idealized. Real users might describe constraints using different words than the ontology uses. So the benchmark could evolve to test that harder scenario.

Tom: And they want to extend the framework to ShEx, which is another shape language for RDF validation. The framework is designed to be language-independent, so that’s a natural fit.

Jane: I also appreciate that they plan to publish the framework as a PyPI package. That lowers the barrier for other researchers to adopt it and contribute their own datasets.

Tom: So the improvements are about rigor — making the evaluation provably correct — and about breadth — covering more expressive features and more realistic settings.

Jane: And that sets us up nicely to look at the first page of the paper itself, where they lay out the motivation and the problem statement in detail.

First Page Discussion: Tom: So Jane, let’s actually walk through the opening page of “NL2SHACL-Bench.” The authors start by reminding us that SHACL is a W3C recommendation for validating RDF knowledge graphs. It’s used for compliance checking, metadata validation, all that serious stuff.

Jane: And then they hit the core problem: creating SHACL shapes requires expertise in both the application domain and semantic web technologies. Most domain experts don’t have that. So constraints get expressed in natural language, and then someone has to manually translate them into SHACL. That’s slow and error-prone.

Tom: They also point out that existing SHACL resources are mined from structured data or designed for specific domains. None of them come with aligned natural language descriptions. So there was no way to even test NL2SHACL systematically.

Jane: And they make a really important point about evaluation. You can’t just compare strings because semantically equivalent shapes can differ in serialization and structure. Two shapes might use different blank node structures or different ordering, but validate exactly the same data.

Tom: So they propose three contributions: a benchmark suite, an extensible framework, and a multi-domain dataset with human-validated descriptions. Plus a set of metrics that capture syntactic, structural, and semantic equivalence.

Jane: And they frame the whole thing around a formal definition of the task. Given a natural language specification and an ontology, produce a shapes graph that encodes the constraints. Correctness means equivalence with a reference shape.

Tom: I love that they ground it in the validation semantics. It’s not about producing something that looks like SHACL. It’s about producing something that behaves like the reference shape when you actually validate data.

Jane: And the first page sets up the structure of the paper — related work, foundations, the benchmark details, experiments, and conclusions. It’s a well-organized roadmap.

Tom: The motivation is really clear: domain experts shouldn’t need to learn RDF and SHACL to enforce data quality. They should be able to say what they need and have the system handle the rest.

Jane: And that’s a vision worth getting excited about. Let’s bring in Lu and Meng to hear their takes.

Lu: I think the formal definition is the strongest part of the paper. By tying correctness to validation equivalence, they’ve made the benchmark future-proof. Any new model can be tested against the same standard.

Meng: And from an engineering standpoint, the fact that they release the framework as open source with a modular design means we can plug in our own models and datasets without rewriting everything. That’s practical.

Tom: Great points from both of you. Now let’s wrap this up.

Conclusion: Tom: Alright, we’ve spent a good chunk of time with “NL2SHACL-Bench: A Benchmark Suite for Natural Language to SHACL Translation,” and I think we can all agree this is a significant contribution.

Jane: Absolutely. It’s the first dedicated benchmark for natural language to SHACL translation, and it brings together a framework, a dataset, and a set of metrics that actually measure what matters — whether the generated shapes enforce the same constraints as the reference.

Tom: The experiments with four LLMs showed that they’re great at producing valid SHACL but still struggle with complex logical and structural patterns. That’s a clear signal for where the research needs to go.

Jane: And the paper is honest about its limitations. The semantic metric is sampling-based, so it has blind spots. But they’ve proposed concrete improvements, like using SHACL2FOL for provable equivalence checking.

Lu: I’d add that the benchmark’s design makes it extensible. Adding more expressive features like SHACL-SPARQL or extending to ShEx are natural next steps.

Meng: And the open-source framework means the community can build on this immediately. That’s how benchmarks become standards.

Tom: So we’re saying goodbye to this paper, but not to the problem. The path forward is clear: better benchmarks, better models, and eventually, a world where domain experts can write validation rules just by describing what they need.

Jane: And that’s a future worth working toward. Thanks for joining us, everyone. Next up, we’ll be looking at another paper that pushes the boundaries of what’s possible with knowledge graphs. See you then!

Tom: Take care, folks!

Yuchen Zhou, Niels Bobet, Maribel Acosta

Technical University of Munich

cs.AI, cs.CL, cs.DB

Submitted: 2026-08-22

Updated: 2026-08-25

Comments: 18 pages, 8 figures, 2 tables

Code: https://github.com/DE-TUM/NL2SHACL-Framework

Project page: https://de-tum.github.io/NL2SHACL-Bench

License: http://creativecommons.org/publicdomain/zero/1.0/

Importance score: 72/100

The gist: "Given a natural language specification T describing constraints over a domain, and an ontology O providing the vocabulary (classes and properties) referenced by T.

Key concepts

SHACL
SHACL is a W3C standard used for validating knowledge graphs. It defines the rules or constraints that data within a graph database must follow to ensure data quality and compliance.
NL2SHACL Translation
This is the core task of converting a natural language description—a plain English explanation of a rule—into generating a formal SHACL shape graph that encodes those specific constraints.
Semantic Equivalence
This metric measures correctness by checking if two different SHACL shapes validate the exact same data. It ensures equivalence even if the generated shapes look structurally different, meaning they behave identically.
First-Order Logic (FOL)
A common logical representation proposed as a tool to check equivalence between SHACL shapes with mathematical certainty. This moves beyond current sampling methods to ensure rigorous validation.

Terminology

Summary

Summary

The paper introduces NL2SHACL-Bench, the first benchmark suite specifically designed for the task of translating natural language requirements into SHACL (Shapes Constraint Language) shapes, a W3C recommendation for validating RDF knowledge graphs. The authors motivate the work by noting that authoring SHACL shapes requires technical expertise that most domain experts lack, and that while constraints are often expressed in natural language by domain experts, they are then manually translated into SHACL, a process that is time-consuming and error-prone. They further state that there is currently no dedicated benchmark, standardized dataset, or evaluation methodology for systematic NL2SHACL evaluation, and that evaluating generated shapes requires methods beyond string comparison, as semantically equivalent shapes can differ in serialisation and structure.

The paper formally defines the NL2SHACL translation task (Definition 1): "Given a natural language specification T describing constraints over a domain, and an ontology O providing the vocabulary (classes and properties) referenced by T. The task is to produce a shapes graph S over O s.t. S encodes the constraints expressed in T." Correctness is defined as the generated shapes graph being equivalent to a reference shapes graph, meaning that for all data graphs over O, the data graph conforms to the generated shapes if and only if it conforms to the reference shapes.

The benchmark suite comprises three main components:

  1. NL2SHACL-Framework: An extensible, modular Python pipeline for dataset construction and evaluation, consisting of three modules. The Dataset Construction Module is a semi-automated pipeline that extracts individual shape fragments from reference shapes graphs, performs structural completeness checks, extracts ontology term URIs, and generates natural language descriptions. Descriptions are either curated from existing sh:description annotations or generated by an LLM prompted to act as a domain expert unfamiliar with RDF/SHACL. All descriptions undergo a two-round human review process using a custom GUI, with annotators checking for six categories of issues including incorrectly described constraints, missing or redundant information, overly technical phrasing, unnatural language, exposure of SHACL-specific terminology, and content hallucinated by the model. The NL2SHACL Translator Module is model-agnostic, allowing users to plug in arbitrary translation systems. The Shapes Evaluation Module computes the predefined metrics.

  2. NL2SHACL-Dataset: A multi-domain dataset containing 240 manually verified NL–SHACL pairs (296 data points total across 6 sub-datasets), each with a unified structure of a natural language description, a reference shapes graph, and an ontology snippet. The sub-datasets cover CHEMROF (chemical entities), DCAT (metadata for data portals), ePO (public procurement), Invoice (EDIFACT-based invoicing), SNIK (healthcare information management), and DBpedia (cross-domain). The dataset covers 21 of the 33 SHACL Core constraint components (excluding SHACL-SPARQL). The authors note that the data sources were selected from the Shapes of You index based on criteria including each represents a well-defined application domain; an associated ontology is available...; and each contains a sufficient number of shapes (more than 10 5).

  3. NL2SHACL-Metrics: A set of eight metrics across three dimensions. Validity metrics are a layered pipeline: RDF Parsing Validity, SHACL Specification Validity (validated against W3C SHACL-SHACL meta-shapes), and SHACL Vocabulary Validity (checking for hallucinated non-standard vocabulary like sh:if, sh:then, sh:condition). Structural metrics measure graph-level similarity after canonicalization: Exact Matching (binary) and Partial Matching (F1-score based on triple overlap). Semantic metrics assess equivalence by generating a synthetic RDF data graph using RDFGraphGen, validating it against both the generated and reference shapes, and comparing the sets of violating nodes. A binary Semantic Equivalence score is assigned if the violating node sets are identical.

The authors evaluate four state-of-the-art LLMs (Gemini 3.1 Pro, Claude Opus 4.7, Qwen3.5-397B, GLM 5.1) on 234 records, using a prompt template with a system prompt, a one-shot example, and the test input with ontology snippet. The results show that all four models achieved very strong validity performance, with Gemini 3.1 Pro obtaining perfect scores on all three metrics. However, structural scores vary substantially across datasets, with DBpedia and Invoice achieving higher Exact Match Rate (EMR) and Partial Match Score (PMS), while SNIK, DCAT, and EPO show near-zero EMR. The authors attribute this to variations in the structural organization of the reference shapes, noting that LLMs typically generate a single node shape with nested constraints instead of multiple interconnected shapes.

Semantic results show high Semantic Equivalence Rate (SER) for CHEMROF, DBpedia, DCAT, and EPO, but substantially lower scores for Invoice and SNIK. The paper identifies representative semantic error patterns, including Inverse Path Omission, Sequential Path Truncation, Constraint Semantics Substitution (e.g., replacing sh:hasValue with sh:class), Conditional Constraint Omission, Non-standard Vocabulary Hallucination (e.g., sh:condition), and Datatype Omission. The authors conclude that LLMs perform well on simple SHACL shapes, but complex logical and structural patterns remain challenging.

The paper discusses limitations, noting that the SER metric has a narrow blind spot: if a generated shape has an incorrect target declaration and produces no violations alongside the reference shape on the sampled data, SER may incorrectly judge them as equivalent. Future work directions include converting shapes to a common logical representation via SHACL2FOL, extending the benchmark to SHACL-SPARQL, evaluating under more practical settings, and extending the framework to ShEx. The resource is available on GitHub (MIT license) and Zenodo (CC-BY-SA 4.0).

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system, along with what the improved system can do:

  • Improvement: Implement the paper's semantic metrics (SER) using validation-based comparison instead of string matching. The system will generate synthetic RDF data graphs via RDFGraphGen, run SHACL validation on both reference and generated shapes, and compare non-conforming node sets.

  • What it can do: Detect when two structurally different SHACL shapes enforce identical constraints. For example, it can recognize that sh:path [sh:inversePath rdf:type] with sh:targetClass is semantically equivalent to a different structural arrangement, avoiding false negatives from exact-match comparisons.

  • Improvement: Add a dedicated post-processing check that verifies inverse paths (sh:inversePath) and sequential paths (e.g., sh:path (rdfs:subClassOf rdf:type)) are not truncated or simplified. The system will parse generated shapes and flag any reduction of complex paths to simpler target declarations.

  • What it can do: Prevent the common error where LLMs reduce sh:path [sh:inversePath rdf:type] to just sh:targetClass dcat:Dataset, which changes validation semantics. It will automatically correct or flag such outputs.

  • Improvement: Implement the paper's Vocab-Validity check as a hard filter. The system will maintain a whitelist of all 33 SHACL Core constraint components and reject any output containing invented predicates like sh:if, sh:then, sh:condition, or sh:xone.

  • What it can do: Automatically reject and regenerate shapes that use hallucinated rule-like constructs instead of valid SHACL patterns (e.g., replacing sh:or/sh:not combinations with non-existent conditional components).

  • Improvement: Implement the paper's fragment extraction logic as a generation constraint. The system will decompose complex natural language descriptions into individual node shapes with explicit sh:property blocks, rather than generating a single monolithic node shape.

  • What it can do: Produce shapes that match reference structures for datasets like SNIK, DCAT, and ePO, where multiple interconnected node shapes are expected. This improves EMR and PMS scores while maintaining semantic correctness.

  • Improvement: Add specialized handling for sh:or with sh:not combinations, which the paper identifies as a major failure point. The system will use a template-based approach for these logical patterns, ensuring proper nesting and avoiding simplification.

  • What it can do: Correctly generate shapes like sh:or ([sh:not (...)] [...]) without omitting the conditional logic, addressing the Conditional Constraint Omission error pattern in Table 2.

  • Improvement: Implement a mandatory datatype check that ensures sh:datatype declarations are never dropped during translation. The system will cross-reference the ontology snippet for known datatypes (e.g., xsd:float) and verify their presence in the output.

  • What it can do: Prevent the error where Qwen 3.5 omitted sh:datatype xsd:float in Invoice cases, ensuring all type constraints from the natural language description are preserved.

  • Improvement: Implement the paper's layered validity checks (RDF parsing → SHACL specification → vocabulary whitelist) as a sequential gating mechanism. Outputs failing any stage are automatically flagged for regeneration with error-specific feedback.

  • What it can do: Catch and fix low-level formatting issues (e.g., missing rdfs: prefix declarations, malformed string literals, incorrect escaping) before semantic evaluation, reducing wasted computation on invalid outputs.

  • Improvement: Implement the paper's ontology snippet extraction to ground all generated shapes in the provided ontology. The system will resolve every URI in sh:path, sh:targetClass, sh:node, sh:hasValue, and sh:class against the ontology index, rejecting any unresolved terms.

  • What it can do: Ensure generated shapes reference only valid domain vocabulary, preventing hallucinated class or property URIs that would cause validation failures on real knowledge graphs.


  1. Generate SHACL shapes from natural language with 100% syntactic validity across RDF parsing, SHACL specification, and vocabulary checks (matching Gemini 3.1 Pro's perfect scores).

  2. Achieve near-perfect semantic equivalence (SER > 95%) on complex datasets like Invoice and SNIK, where current LLMs only reach 50% equivalence, by preserving inverse paths, sequential paths, and logical sh:or/sh:not structures.

  3. Automatically detect and correct its own errors using the layered validity pipeline, reducing manual post-editing effort from hours to minutes.

  4. Produce structurally consistent shapes that match reference organization (multiple node shapes with explicit property blocks) for enterprise datasets like DCAT and ePO, improving maintainability and composability.

  5. Provide confidence scores based on the paper's metrics (EMR, PMS, SER), allowing users to know when a generated shape is semantically verified versus when it requires manual inspection.

  6. Handle edge cases like negation constraints and conditional logic without hallucinating non-standard vocabulary, addressing the most common failure modes identified in the paper's evaluation.

Abstract

SHACL is a core technology for validating the conformance of RDF knowledge graphs (KGs). Yet, authoring SHACL shapes requires technical expertise that most domain experts lack. Translating natural language requirements into SHACL (NL2SHACL) would lower this barrier. However, there is no dedicated benchmark for NL2SHACL, and evaluating generated shapes requires methods beyond string comparison, as semantically equivalent shapes can differ in serialisation and structure. To tackle these challenges, we present NL2SHACL-Bench, a benchmark suite for natural language to SHACL translation. Using NL2SHACL-Bench, we evaluate four state-of-the-art large language models (LLMs) for this task. Our results show that current LLMs are highly capable of generating syntactically valid SHACL, but still struggle to produce semantically equivalent constraints for complex logical and structural patterns. This indicates that NL2SHACL-Bench provides a meaningful basis for measuring advances in the NL2SHACL state of the art.

Sources

Related papers