Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements

arXiv:2605.22079 · cs.CL · Submitted 2026-05-21 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements".

Jane: The paper was written by Ryo Kanazawa, Tomoki Ando, Shuhei Saitoh, Koyo Hidaka, Chenguang Wang et al. from ONESTRUCTION Company and AWS GenAI Innovation Center.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary: Tom: Now, let's dive into the summary of *Ishigaki-IDS-Bench*. The researchers developed this benchmark by taking eighty-three practical scenarios and expanding them into one hundred sixty-six detailed examples, which are then available in both Japanese and English. It’s a huge dataset.

Jane: The core idea is that for each input—whether it's a checklist or a paragraph of instructions—the AI must produce a "gold standard" IDS file, which is the perfect version according to the industry rules. The benchmark tests how close the AI-generated output gets to this gold standard.

Lu: It’s not just about matching words; it’s about matching semantic intent and structure. If an LL can't capture that specific intent, even if the XML looks fine, it fails the test of understanding domain requirements.

Meng: The use of two different metrics is crucial here: a formal audit using tools like IDSAuditTool and then a content fidelity check using "Facet F1." This ensures we are looking at both the technical correctness and the practical usefulness of what’s generated.

Lalam: This method allows us to measure how well an AI understands the *intent* behind a requirement, not just its ability to string together XML tags. That is incredibly powerful for machine learning research.

Tom: It’s a highly practical approach because, in construction, if the data doesn't match the requirements exactly, you can’t build it correctly. You need that granular agreement between what was asked and what was delivered.

Jane: So we have this massive library of one hundred sixty-six examples that forces us to look at both sides: the technical validity and the content accuracy simultaneously. It’s a comprehensive test of adherence to industry standard XML generation.

Lu: And Meng is right, we are testing the rigor. If an AI only passes the structure check but misses a key requirement, it’ fails in practice, and this benchmark catches that difference perfectly.

Meng: That combination of auditing and content checking is how I would justify using this in a pipeline—it provides the necessary safety net for complex generation tasks.

Lalam: We are building tools that respect established professional standards, which is a positive step for our industry's cultural evolution toward automation.

Improvements & Methodology: Tom: The researchers designed a very specific evaluation protocol to make sure the testing is rigorous, going beyond just one simple pass/fail test. They use two stages to analyze the failure modes of LLMs on *Ishigaki-IDS-Bench*.

Jane: First stage involves running the generated IDS through an audit tool—we check if it's processable, if its structure is sound, and whether it adheres to the constraints of the IDS standard and IFC vocabulary. This is a technical hurdle for Stage one.

Lu: The second stage is where we measure content agreement using Facet F1. This measures how well as many individual data points or requirements in the generated output match those found in our perfect "gold" standard file, regardless of the XML structure.

Meng: This two-stage approach is brilliant because it separates syntactic success from semantic success. An LL could be great at making valid XML but still miss a critical requirement, which Stage two would flag as a failure.

Lalam: I think this method helps us train our models to achieve high fidelity. It's teaching the AI that even if the output is technically perfect, it must match the expected content to be successful.

Tom: The findings show us some interesting trends in how these LLMs perform on *Ishigaki-IDS-Bench*. We see that multi-turn conversation examples are actually handled much better than single requests for certain models.

Jane: That suggests that when the AI is allowed to interact or iterate, it becomes more successful at capturing complex requirements. It's like the AI understands the context of a sustained discussion better than a one-off prompt.

Lu: And another observation is that when tabular inputs, like CSV data, are used, the LLMs tend to align with facets much better than when natural language is used.

Meng: That aligns with what I've seen in real-world scenarios; tables force structure on messy data, which helps the AI follow a consistent logic path for generation.

Lalam: We are seeing evidence that complexity requires more nuanced models, and we are creating a framework to understand *how* those nuances translate into reliable performance.

Findings & Implications: Tom: The results from zero-shot evaluation on *Ishigaki-IDS-Bench* show that while some LLMs perform very well in certain areas, the overall success rate for satisfying both the standard and the content is quite low.

Jane: The best performers, like GPT-five point five, achieved a high Facet F1 score of sixty-five point six percent, which is impressive for technical data extraction. However, their Content pass rate was only about thirty-three percent.

Lu: That disparity is the most important thing to me. It shows that even if an AI can extract and format the individual components correctly, it still struggles with the holistic agreement of the overall requirement—the big picture integrity.

Meng: The failure modes are quite specific too. We see LLMs struggling with property facets, particularly when property set names or constraints become complicated, which is a major roadblock for implementation.

Lalam: This highlights that we need AI that doesn't just parrot information but truly understands the *constraints* of engineering disciplines to produce trustworthy outputs.

Tom: And we also see issues where natural language inputs are highly ambiguous, causing the AI to either miss requirements or generate too many unnecessary facets in those more conversational examples.

Jane: It’s a clear indication that this industry-specific task requires more than just general knowledge; it needs specialized training on how to structure and formalize requirements.

Lu: We must also consider how the scale of this benchmark is valuable, moving beyond small samples to provide a robust test against the real-world complexity of one hundred sixty-six scenarios.

Meng: Scaling this means we can start thinking about deploying AI that can handle diverse project inputs, not just one kind of data source.

Lalam: This research opens the door for building a culture where technical documentation is generated with high confidence and low ambiguity.

Conclusion & Wrap-up: Tom: So, as we wrap up our discussion on *Ishigaki-IDS-Bench*, the main takeaway is that generating industry-standard specifications like IDS requires more than just good at making XML; it demands a deep understanding of domain vocabulary and content fidelity.

Jane: It’s a powerful tool for teaching us that the AI needs to pass two distinct tests: one of technical compliance, and another of matching the actual information requested.

Lu: The fact that we have such detailed data on how different models handle complex inputs shows us exactly where our next breakthroughs in large model reasoning need to happen.

Meng: And I think this provides a much clearer picture for the engineering side, showing where we can improve training and what challenges remain before we can fully automate these tasks.

Lalam: We are moving toward a world where AI assists in creating documentation that is not just functional, but truly professional and culturally aligned with industry standards.

Tom: I think it’s clear that the future of building complex systems relies on models like *Ishigaki-IDS-Bench* are designed to train.

Jane: It's an exciting area to watch, showing a real progress in how we handle highly structured, professional information exchange.

Lu: I'm confident this framework allows us to push the boundaries of what AI can achieve in specialized domains like construction.

Meng: I hope this is the start of a more robust set of tools that will integrate into our industry workflows soon.

Lalam: We are hopeful for a future where machines and collaboration work together to deliver specifications with absolute clarity and adherence to standard excellence, building a better culture in the process.

Ryo Kanazawa, Tomoki Ando, Shuhei Saitoh, Koyo Hidaka, Chenguang Wang, Atomu Kondo, Teppei Miyamoto, Dayuan Jiang, Koki Arakawa, Takayuki Kato, Naofumi Fujita, Daiho Nishioka

ONESTRUCTION Company · AWS GenAI Innovation Center

cs.CL

Submitted: 2026-05-21

Updated: 2026-08-25

Code: https://github.com/buildingSMART/IDS-Audit-tool

Importance score: 100/100

The gist: This paper introduces Ishigaki-IDS-Bench, the "first publicly released benchmark for IDS generation from BIM information requirements." It addresses the difficulty of evaluating large language models

Key concepts

Ishigaki-IDS-Bench
This is a large dataset of 166 detailed examples, derived from eighty-three practical scenarios. The benchmark tests how closely an AI generates a perfect Information Delivery Specification (IDS) file based on the input requirements provided.
Gold Standard IDS File
This refers to the perfect version of the Information Delivery Specification according to industry rules. The benchmark measures how close an AI-generated output gets to this standard, ensuring it meets both technical and semantic requirements.
Facet F1
This is a content fidelity check used in the second stage of evaluation. It measures how well individual data points or requirements found in the AI's output match those present in the perfect 'gold' standard file, regardless of XML structure.
Two-Stage Evaluation
This rigorous protocol separates syntactic success from semantic success. Stage one runs an audit tool to check technical compliance (structure and constraints), while stage two uses Facet F1 to check content agreement, flagging failures where the AI misses critical requirements.

Terminology

Summary

This paper introduces Ishigaki-IDS-Bench, the first publicly released benchmark for IDS generation from BIM information requirements. It addresses the difficulty of evaluating large language models (LLMs) on tasks where outputs must not only be syntactically valid but also conform to industry-standard data formats, domain-specific vocabularies, version constraints, and audit results from external validation tools.

Benchmark composition

The benchmark consists of 166 examples derived from 83 practical scenarios, authored in both Japanese and English by six BIM/IDS experts. The researchers prioritized expert-judged practical representativeness over mechanical IFC-schema coverage, ensuring each example is paired with a gold IDS file and metadata covering input format, turn setting, target IFC versions, and construction domain.

The metadata taxonomy includes:

  • Input format (CSV or Natural language)

  • Language (JA or EN)

  • Turn setting (Single-turn or Multi-turn)

  • Target IFC versions (IFC2X3, IFC4, or IFC4X3)

  • Construction domain (Architecture, Structural, MEP, or General)

Evaluation protocol

The study employs a two-stage protocol for evaluating generated IDS from both formal and content perspectives. In the first stage, the IDSAuditTool verifies IDS extractability, schema compliance, and conformance to the IDS standard and IFC vocabulary constraints. This stage evaluates three specific metrics: Processability, Structure, and Content.

In the second stage, the protocol measures content fidelity scored by facet-level macro-F1 against the gold IDS. This stage is designed to capture generations that are formally valid but differ from the input document by comparing IFC-version labels, entity, attribute, property, and the dataType/cardinality of properties against the expert-authored gold standard.

Baseline results

In a zero-shot evaluation of 10 LLMs, the results indicated that current models struggle to stably generating outputs that satisfy the IDS standard and IFC vocabulary constraints. While GPT-5.5 achieved the highest Facet F1 of 65.6%, the highest Content pass rate was only 33.1%, achieved by Claude Opus 4.5. Several performance trends were observed:

  • Multi-turn examples reached significantly higher Content and Facet F1 scores than single-turn examples.

  • CSV inputs achieved higher Facet F1 than natural-language inputs, suggesting tabular requirements facilitate alignment.

  • Scale alone does not guarantee robust processability or standards-compliant content generation.

Failure tendencies

The researchers identified three primary failure modes during the generation process. First, many outputs can be extracted as IDS candidates but still violate the IDS XSD, IDS standard, or IFC vocabulary constraints. Second, while entity and attribute facets are relatively easy to recover, errors are common in property facets involving property set names, cardinality, and value constraints. Finally, natural-language inputs often create ambiguity between requirement targets and value constraints, leading to more facet omissions or overgeneration.

Improvements for AI systems

The current research trajectory points toward the necessity of moving Large Language Models (LLMs) from general knowledge retrieval tools to highly constrained, standard-compliant reasoning engines within the Architecture, Engineering, and Construction (AEC) domain. The improvements must bridge the gap between linguistic fluency and rigorous digital modeling standards (like IFC).


  • Improvement: Develop a novel decoding mechanism that enforces adherence to formal industry schemas (e.g., IFC, bSDD, or specific national building codes) during the generation process, rather than simply validating the output afterward. This requires integrating formal grammar masking techniques ([20]) with domain-specific ontological knowledge graphs ([28]).

  • Improved AI Capability: The system can generate model view definitions (MVDs) or compliance reports that are syntactically and semantically guaranteed to conform to established standards. For example, if asked to specify a structural element, the AI will not only provide the correct textual description but will output structured code fragments or data payloads that pass immediate validation against the target IFC schema, eliminating hallucinated metadata.

  • Improvement: Construct a unified LLM framework capable of simultaneously ingesting and reasoning across three distinct modalities: Natural Language Requirements to Geometric/Spatial Data (BIM) to Formal Data Standards (IFC/bSDD). This moves beyond simple text-to-text or text-to-database queries.

  • Improved AI Capability: The system can perform complex, multi-step design verification tasks. For instance, given a textual requirement (The fire exit path must maintain a minimum clearance of 2.5 meters and be visible from the main lobby), the AI can:

  1. Parse the text into measurable constraints (Natural Language to Logic).

  2. Query the underlying BIM model to identify all relevant geometric elements (Spatial Data Retrieval).

  3. Verify that every identified element satisfies all constraints, and then generate a detailed, standards-compliant report citing the specific conflicting elements and the relevant code section ([19], [14]).

  • Improvement: Create an advanced benchmarking framework that simulates real-world regulatory audits. This framework must evolve past simple Question Answering (QA) to include Adversarial Compliance Testing, where the model is intentionally provided with conflicting or ambiguous inputs designed to force failure points in standard adherence.

  • Improved AI Capability: The system can be rigorously evaluated not just on accuracy, but on its Robustness of Compliance. A specialized benchmark (akin to an Ishigaki-IDS-Bench for compliance) would test:

  1. Conflict Detection: Identifying instances where two parts of the design violate a single code article.

  2. Traceability Mapping: Tracing a specific physical feature in the model back through its documentation, standards definitions, and originating requirement text, providing an auditable chain of custody for every data point ([27]).

  • Improvement: Integrate a specialized, dynamic Knowledge Graph (KG) layer that maps relationships not only between concepts (e.g., Wall to Fire Rating) but also between standards versions and industry best practices ([26]). This KG must be queryable by the LLM's attention mechanism.

  • Improved AI Capability: The AI can provide context-aware recommendations that account for temporal changes in standards. If a user references an outdated design manual, the system will flag the discrepancy, identify the current required standard version (e.g., This element requires compliance with ISO 19650:2018, not 2015), and dynamically adjust its reasoning path accordingly.

Sources

Related papers