ConstructCIE: A Dataset for Extracting Causal Information from Construction Accident Narratives
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "ConstructCIE: A Dataset for Extracting Causal Information from Construction Accident Narratives".
Jane: The paper was written by Hung Nguyen, Jaehoon Lee, Namgyun Kim and Kuan-Hao Huang from Department of Computer Science, Texas A&M University and Department of Construction Science, Texas A&M University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Discussion of the Core Problem: Tom: The paper, "ConstructCIE: A Dataset for Extracting Causal Information from Construction Accident Narratives," shows they are fundamentally changing how we view these accident narratives.
Jane: Before this work, you often just focus on who was involved or what equipment was present, but the researchers realized that simply focusing on "what happened" wasn' isn't enough.
Lu: The problem is that causality in a construction site isn't just one thing; it’s often a long chain of subtle factors that are spread across many paragraphs.
Meng: They found existing event extraction tools aren't designed to handle this kind of deep, distributed reasoning, which is really where the practical difficulty lies.
Lalam: It seems like the current AI models struggle with this kind of multi-step, implicit logic because they are trained on simpler patterns rather than complex causal webs.
Tom: So, if we're looking at a fractured spine accident described in the report, it might not be the tool itself but the fact that workers were riding unrestrained in a van that was't equipped for the kind of hazards involved.
Jane: That’s exactly what they found; it requires understanding how various factors contribute to build up to what is often just called "struck-by."
Lu: They are moving beyond localized triggers and looking at the entire context, which is a huge theoretical shift for how we model complex failure modes.
Meng: We need systems that can grasp these long-span relationships, not just local connections, so I think this sets up a real engineering challenge for future AI development.
Lalam: This really shows how AI needs to evolve past simple pattern matching if it’s going to actually help safety managers understand the full scope of risk in industry.
Discussion of the Dataset and Methodology: Tom: Now, let's talk about how they solved this problem with "ConstructCIE." They aren't just throwing raw text at us; they’ve built a hierarchical structure around it.
Jane: They created a taxonomy that breaks down every possible cause, from the high-level accident type down to specific sub-causes.
Lu: This schema is brilliant because it forces a categorization that maps directly onto real-world safety science theories we already know, like HFACS.
Meng: From an implementation standpoint, this structure helps us define exactly what we are looking for when training a machine learning model to find the right pieces of evidence.
Lalam: It’s about structuring the knowledge, so that AI isn't guessing what it should be looking for; it knows it must find a sub-cause under its main factor.
Tom: They are using two main approaches: supervised tagging, like TagPrime-C and TagPrime-CR, and these LLMs using in-context learning.
Jane: It’s interesting to see the differences in how they approach the problem—the structured tagging versus the flexibility of seeing how an LLM handles it.
Lu: The way they frame this hierarchically, it allows us to build a much richer causal knowledge graph than previous event extraction datasets allowed.
Meng: We need to know if we can scale this specific hierarchical approach to other industries or if we' are limited by the scope of the construction industry itself.
Lalam: The goal is to teach AI how human experts think about causation, and using that structure is a powerful way to do it, making sure the machine understands the context.
Discussion of Results and Improvements: Tom: The results are quite revealing when looking at performance. They found that most models are really good at predicting the overall accident type.
Jane: That’s a big win for general understanding, but the real story is in the sub-causal factors, where things get much more nuanced and challenging.
Lu: The models can recover broad causal meaning—they know *why* an accident happened generally—but they often struggle with precise evidence extraction at the fine-grained level.
Meng: This gap between broad understanding and accurate span selection is a major pain point for me; if we don't pinpoint the exact text, the data is less useful for root cause analysis.
Lalam: It seems like AI understands the *intent* of finding a cause, but it isn't always good at extracting the specific boundaries that prove that intent in real-world language.
Tom: They found that JHE, or Joint Hierarchical Extraction, generally performs better on exact and soft matching compared to IHE.
Jane: That suggests grouping the information together helps the models maintain context better than processing each factor separately, which is a key methodological finding.
Lu: This points to a direction for future work: we need to find ways to leverage joint reasoning without increasing computational load too much.
Meng: From an engineering perspective, optimizing for accuracy in span boundary detection seems like the most critical area of focus moving forward.
Lalam: We need AI that doesn't just tell us *a* cause, but one that can pinpoint the exact text supporting the cause to truly improve safety protocols and build better industry knowledge.
Conclusion and Wrap-up: Tom: As we wrap up this deep dive into "ConstructCIE: A Dataset for Extracting Causal Information from Construction Accident Narratives," it’s clear that this work is a major step forward.
Jane: It gives us a structured way to teach AI how to handle the complex, often implicit causal stories found in construction sites.
Lu: The hierarchical approach really provides the framework we need to model how human error and environmental factors interact, which is crucial for safety science.
Meng: We've seen that while LLMs are powerful, they still struggle with precise span extraction, meaning we have a very clear direction for improvement in AI implementation.
Lalam: This has the potential to fundamentally change how industry knowledge is stored and applied, moving us toward truly data-driven safety culture.
Tom: It’s a challenge that requires more than just general language understanding, as the authors showed us.
Jane: We appreciate all of you joining us on this fascinating topic.
Lu: I'm looking forward to seeing how the next generation of models tackles this specific problem-solving structure.
Meng: I’m already thinking about how to build a pilot system using these data, so it sounds like a great deal of work ahead for practical impact.
Lalam: I believe that with the right AI tools, we can make safety information accessible and actionable for everyone in the future, building on what "ConstructCIE" provides.
Hung Nguyen, Jaehoon Lee, Namgyun Kim, Kuan-Hao Huang
Department of Computer Science, Texas A&M University · Department of Construction Science, Texas A&M University
cs.CL
Submitted: 2026-08-06
Updated: 2026-08-25
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 81/100
The gist: The paper introduces C ONSTRUCT CIE, a construction-specific dataset designed for Causal Information Extraction (CIE) from OSHA accident narratives.
Key concepts
- ConstructCIE
- This is the name of the dataset discussed in the episode. It was created to provide a structured way for AI to extract causal information from construction accident narratives by building a hierarchical structure around potential causes.
- Causality in Construction
- The hosts explain that causality on a construction site is not simple; it involves long chains of subtle, distributed factors across multiple paragraphs. The goal is to move beyond localized triggers to understand how various factors contribute to an accident.
- Joint Hierarchical Extraction (JHE)
- This is a methodological finding where the hosts discuss that grouping information together helps AI models maintain context better than processing each factor separately. JHE generally performed better in performance testing.
- Span Extraction
- This refers to the specific challenge of AI pinpointing the exact boundaries within text that prove a cause. While models can understand the *intent* of a cause, they often struggle with accurately extracting these precise textual boundaries.
Terminology
Summary
The paper introduces C ONSTRUCT CIE, a construction-specific dataset designed for Causal Information Extraction (CIE) from OSHA accident narratives. The research addresses the challenge that while construction accident narratives contain rich causal information,
this evidence is often implicit, long-span, and distributed.
Existing Event Extraction (EE) datasets are deemed insufficient because they focus on explicit event triggers rather than the complex, multi-sentence causal chains required for this domain.
The CIE task requires a hierarchical extraction process where the model must:
-
Predict the high-level Accident Type (e).
-
Identify broad Causal Factors (seven primary contributors).
-
Determine or extract specific Sub-causal Factors, which are either classified or extracted based on the factor's nature.
The C ONSTRUCT CIE schema is derived from construction safety literature and defines seven primary causal factors:
-
Working Circumstances: Describe the physical and operational context of the accident.
-
Object Involved: Captures
the equipment, structure, material, or object that physically interacted with or contributed to the accident mechanism.
-
Managerial Factors: Capture organizational or supervisory deficiencies.
-
Working Condition Factors: Capture physical work-environment contributors (e.g., Weather Condition and Workspace Condition).
-
Equipment Factors: Capture equipment-related contributors (e.g., Protective Equipment Condition and Work Equipment Condition).
-
Behavioral Factors: Capture task-level human actions or lapses (e.g., Inattentive Behavior and Noncompliant Behavior).
-
Consequences: Capture accident outcomes (e.g., Severity and Affected Body Part).
The dataset was constructed from 530 Occupational Safety and Health Administration (OSHA) accident investigation summaries published between 2011 and 2023. The curation process involved a two-stage screening: record-level screening to remove duplicates, and content-level screening to retain only reports whose narratives explicitly described both accident circumstances and causes.
The distribution of the data reveals several patterns:
-
Accident Types:
Fall accidents are the most common, followed by Caught-in/between and Struck-by; Electrocution is least common.
-
Causal Factors:
Working Circumstances, Consequences, and Object Involved... appear in 100.0%, 99.2%, and 94.5% of reports,
respectively. Conversely,Managerial Factors are much less frequent, appearing in only 6.4%.
-
Sub-causal Factors:
Workspace Condition is the most frequent at 34.3%, followed by Noncompliant Behavior at 28.7% and Work Equipment Condition at 22.1%.
The data was independently annotated by two domain experts, and all date information was removed to prevent models from learning false temporal patterns.
The study evaluates supervised sequence taggers (TagPrime variants) and Large Language Models (LLMs), including Llama3.2B/11B/70B, Qwen3.5-9B, and Qwen3.5-27B.
Two primary extraction strategies are compared:
-
End-to-End Joint Hierarchical Extraction (JHE):
extract[s] all causal factors in a single grouped output and then extracts extraction-based sub-causal factors jointly while predicting classification-based sub-causal factors individually.
-
End-to-End Individual Hierarchical Extraction (IHE):
processes each causal factor separately by first determining whether it is present and then extracting or classifying its associated sub-causal factors.
The models are evaluated using three metrics:
-
Exact String Match (E): Measures correctness only if the predicted span exactly matches the annotated gold span.
-
Soft String Match (S): Relaxes exact match by measuring similarity using a threshold of 0.8 via the Gestalt string matching algorithm.
-
Keyword Match (K): Evaluates whether the predicted span
contains all required keywords associated with the corresponding gold span.
The study also examines Few-Shot Scaling, testing in-context learning with k in 0, 5, 10, 20, 30 examples.
Overall Performance:
Results show that most evaluated models achieve strong accident-type prediction and recover broad causal meaning but remain limited in precise span-level extraction.
Strategy Comparison (JHE vs. IHE):
-
JHE generally achieves stronger exact and soft matching, while IHE sometimes achieves higher keyword F1.
-
For instance, at k=30, Qwen3.5-27B obtained
48.0% exact and 60.1% soft F1 with JHE, compared with 41.7% exact and 52.1% soft F1 with IHE.
-
The difference in performance is particularly pronounced for Llama3.2-3B, where JHE achieved
19.2% exact F1... compared with only 1.2% exact F1 under IHE.
Factor-Level Performance:
While models perform well on explicit information (e.g., Construction Trade, Severity), they struggle with fine-grained evidence: "The evaluated models perform better on frequent and explicit information... but struggle with more fine-grained causal evidence, such as Workspace Condition, Protective Equipment Condition, and several managerial and behavioral sub-causal factors."
Error Analysis:
Table 8 details the error distributions across strategies. A consistent pattern is observed between keyword and exact matching: "Higher keyword F1 indicates that the models often identify the central causal content of the annotation, even when the predicted span does not fully match the gold evidence. In contrast, lower exact F1 shows that the evaluated models frequently fail to reproduce the complete annotated span with the required boundaries."
Conclusion:
The study concludes that reliable construction accident CIE requires accurate high-level accident understanding together with more precise and better-grounded evidence extraction.
Improvements for AI systems
As a diligent AI researcher operating under these high-stakes conditions, I require the scientific paper itself to proceed with any meaningful analysis. My recommendations are entirely dependent on identifying novel methodologies, architectural weaknesses, or untapped theoretical potential within the source material.
Please provide the arXiv paper.
Once provided, my analysis will be structured to deliver actionable engineering and algorithmic improvements. My response will strictly adhere to the following format and depth of detail:
Core Improvement Focus: [A high-level summary of the primary theoretical or computational gap I identified in the paper.]
-
Specific Change: [e.g., Replace the current attention mechanism with a Sparse Transformer block to reduce quadratic complexity.]
-
Technical Justification: [Explain why this change is necessary, referencing limitations in the paper's methodology (e.g.,
The original model fails to capture long-range dependencies efficiently because...
).] -
Enhanced Capability: The resulting system will gain the ability to process context windows exceeding [X units], allowing for comprehensive analysis of [specific complex task].
-
Specific Change: [e.g., Implement a multi-modal contrastive learning pipeline that fuses text embeddings with geometric feature maps before training.]
-
Technical Justification: [Address the input limitations (e.g.,
The paper treats modalities independently; fusion is necessary to resolve ambiguity between visual and textual cues.
).] -
Enhanced Capability: The system will achieve superior robustness in noisy, real-world environments, specifically improving accuracy when inputs are partially corrupted or ambiguous across different data types.
-
Specific Change: [e.g., Integrate a knowledge graph embedding layer during fine-tuning to ground abstract concepts in verifiable relationships.]
-
Technical Justification: [Focus on reliability and interpretability (e.g.,
The current architecture is black-box; this refinement forces the model to generate traceable reasoning paths, mitigating hallucination.
).] -
Enhanced Capability: The improved system will not only predict but also explain its predictions with high fidelity, providing source attribution for every major decision point—a critical feature for deployment in regulated industries.
I am prepared to analyze the paper with the utmost rigor. Awaiting the input document.
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering