Automated Event Log Generation from Unstructured Text Using Finetuned LLMs

arXiv:2609.01320 · cs.AI · Submitted 2026-09-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Automated Event Log Generation from Unstructured Text Using Finetuned LLMs".

Jane: The paper was written by J. Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li et al. from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary: Tom: So, we were talking about how crucial it is to get accurate event logs from messy text, and the paper provides a detailed summary of their approach.

Jane: If you look at the summary, they aren't just relying on a basic prompt; they’re detailing a structured process that makes sure every piece of extracted data adheres to specific constraints.

Lu: What I found really interesting in the summary was how they framed this not as just text processing, but as an information extraction challenge with high temporal dependency, which adds another layer of complexity.

Meng: From an engineering point of view, the fact that they quantify performance metrics so rigorously—it suggests their pipeline is modular and testable, which is exactly what a commercial product needs.

Lalam: It’s about moving beyond mere pattern recognition; the summary implies a deep contextual understanding, which mirrors how humans build narratives in our own conversations.

Tom: Right, because the summary really hammers home that general LLMs can hallucinate or wander when pressed for specific facts, and they had to guide the model with more than just natural language instructions.

Jane: They must be providing a kind of guardrail system for the AI so it doesn't just make things up or get distracted by irrelevant details in the source document.

Meng: So, when they talk about their data preparation steps, are they essentially creating highly curated training datasets that specifically target common failure modes of existing models?

Lu: Absolutely. They aren't just feeding it random text; they're giving it examples of *bad* extraction and showing the model how to correct those errors itself.

Tom: That focus on the input quality and the specific training structure is key, isn’t it? It elevates this from a simple API call to a sophisticated machine learning task.

Lalam: This capability has massive implications for fields like digital humanities, where researchers sift through thousands of historical letters or diary entries that are inherently messy.

Jane: It means that the sheer volume of historical text no longer represents a barrier to understanding; the AI can help us process it

Paper discussion segment 2: Tom: So what this paper really shows us is that we can automate turning messy human language into precise digital timelines of events using specialized AI models.

Jane: Exactly! Think about any industry that generates tons of text—healthcare, legal services, insurance claims—it’s all just unstructured narrative right now, and the manual work of extracting key actions is exhausting.

Lu: But the fact that they used *fine-tuned* LLMs is what blows my mind; it means they didn't just give the AI a general prompt, but they trained it specifically on the nuances of event extraction within certain domains.

Meng: That specificity is critical because if you use a general-purpose model, it might get confused by jargon or ambiguous phrasing that a domain expert would instantly understand. I wonder about the computational overhead of maintaining those fine-tuned versions across multiple clients.

Lalam: It’s incredible how this shifts the focus from human labor to machine understanding; imagine the sheer volume of historical data we could finally process and analyze for cultural insights, like tracking public sentiment changes over time based on news archives.

Tom: Right, Lalam brought up a huge point about history—it’s not just about logging events *now*; it's about creating a searchable, structured record of everything that has ever been written down.

Jane: That structure allows researchers to move beyond simply knowing *what* happened and start understanding the complex sequence of *why* it happened in relation to previous steps.

Lu: We could build entirely new forms of digital archaeology, reconstructing entire operational histories just from text snippets that were never meant to be analyzed that way.

Meng: From an implementation standpoint, if we could guarantee high accuracy across different input formats—like switching from dictated notes to handwritten transcripts—the real-world efficiency gain would be staggering.

Lalam: And this isn't just about efficiency; it’s about accessibility, Tom. It gives power back to people who were previously drowned out by the sheer volume of data, allowing narratives and processes to finally speak in a computable language.

Tom: So we're talking about fundamentally changing how knowledge is managed and retrieved across massive organizations, which is a huge deal.

Jane: It suggests that the bottleneck isn't really the information itself, but our ability to translate it into a format that machines can follow step by step.

Lu: I bet this opens up whole new fields of cross-modal AI—combining text extraction with visual data or audio recordings automatically.

Meng: We’d need robust pipelines for data governance, though; if the AI is building these critical process logs, we have to make sure the data lineage and privacy compliance are baked into the architecture from day one.

Lalam: It truly elevates how we perceive information itself; it's not static text, but a dynamic flow of cause and effect waiting to be revealed.

Tom: Given all this potential for deep historical analysis, I wonder what the next major challenge in event logging will be?

Paper discussion segment 3: Jane: So, just to recap, this paper shows that using finetuned Large Language Models makes turning messy documents into clean process logs a much more reliable process.

Tom: And what’s really exciting about this isn't just that it works, Jane; it's how fundamentally it changes what we can analyze in the real world.

Lu: Exactly, Tom! We aren't talking about just generating a log anymore; we're enabling semantic understanding across vastly different operational domains that previously required specialized human knowledge to structure.

Jane: So, what Lu means is that instead of needing an expert who knows exactly how a hospital or a factory works to map out the steps, the AI can learn those complex rules from just reading thousands of reports.

Meng: But if it learns from varied reports—say, combining medical notes with billing records—how do we guarantee that the LLM won't hallucinate a sequence or conflate two different procedures into one single, incorrect step?

Tom: That’s a huge question, Meng; it brings up the issue of trust and fidelity in these generated logs.

Lu: We need rigorous validation frameworks that test not just grammatical accuracy, but deep causal relationships between events to ensure the resulting process model is genuinely sound.

Jane: It's about adding a layer of critical thinking to the AI output, making sure the structure it proposes actually makes sense given real-world constraints and rules.

Meng: And from an engineering standpoint, if we're talking about massive enterprise deployments—think thousands of documents arriving every hour—what’s the computational overhead for running these finetuned models at that scale?

Tom: That scalability concern is critical, Meng; it implies that the improvements aren't just academic parlor tricks but genuinely deployable tools for industry.

Lalam: Considering how much complex human experience is currently trapped in unsearchable text, this capability fundamentally democratizes institutional knowledge, giving power back to data analysis by making every piece of writing actionable.

Jane: It truly means that the insights previously reserved for highly paid process architects are now accessible through a sophisticated AI pipeline.

Tom: I wonder what happens when we combine this log generation with predictive models—can we not only see *what happened*, but also predict *what should happen* next, and then automatically flag deviations?

Conclusion: Tom: Wow, what a deep dive into how LLMs are changing process mining! It really feels like we’ve seen a glimpse of how much more automated and natural this field is becoming.

Jane: Exactly. If you look at the sheer volume of unstructured text out there—medical records, customer service chats, police reports—it’s overwhelming. This paper, "Automated Event Log Generation from Unstructured Text Using Finetuned LLMs," gives us a clear path to making sense of that mess.

Lu: I keep thinking about this potential for historical analysis; imagine applying this methodology not just to modern processes, but to decades of archived documents in fields like medicine or law. The data suddenly becomes actionable, which is a huge paradigm shift for knowledge retrieval.

Meng: But Lu, while the scope is massive, I'm wondering about the real-world implementation complexity. How do you handle domain drift when you finetune these models? We need robust pipelines that can adapt to completely new types of jargon without requiring a full retraining cycle every time.

Lalam: You raise a critical point about adaptability, Meng. The true impact here isn't just the technology itself, but how it democratizes access to deep data insights. Improving our ability to structure chaotic information improves our cultural understanding of complex systems—whether it's city infrastructure or human interaction.

Jane: It really does boil down to that: turning narrative into structure. Tom, so if we had to sum up the main takeaway for our listeners who might be intimidated by process mining, what should they walk away thinking?

Tom: They should think that the days of manual data entry and tedious rule-writing are fading away. The future is about feeding AI massive amounts of raw text and letting it do the heavy lifting of figuring out *what happened* and *in what order*.

Lu: And to build on Tom's point, this isn't just a better tool; it’s an entirely new research frontier that bridges NLP with operational science in ways we haven't seen before.

Meng: For the industry side, I think the immediate implication is a massive reduction in time-to-insight. Companies won’t wait months for data scientists to manually annotate logs; they can get near real-time process maps using this approach.

Lalam: And from a societal view, if we can automate event log generation this effectively, it means we can analyze systemic failures and biases across entire populations much faster, leading to better policy making globally.

Tom: So that’s the wrap-up on "Automated Event Log Generation from Unstructured Text Using Finetuned LLMs." It’s clear this is going to be a massive area of growth.

Jane: We definitely feel energized by this one, and we can't wait to talk about what comes next in AI research.

J. Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Weizhu Chen

cs.AI

Submitted: 2026-09-01

Updated: 2026-09-01

Code: https://github.com/omseeth/Automated-Event-

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 79/100

The gist: The paper addresses the critical challenge of transforming vast quantities of unstructured textual data—such as customer service transcripts, medical notes, or incident reports—into structured,

Key concepts

Event Log Generation
The process of converting large volumes of messy, unstructured human language (like reports or letters) into a structured, chronological record of actions and events. This creates a searchable digital timeline.
Finetuned LLMs
Large Language Models that have been trained specifically on niche datasets or tasks, rather than general text. This specialization allows the AI to accurately extract domain-specific information and adhere to complex constraints.
Unstructured Text
Raw human language data that does not follow a predefined format, such as medical notes, historical letters, or customer service chats. This type of text is difficult for machines to process manually.
Process Mining
A field that uses data analysis to discover how business processes actually operate. The paper's methodology automates the creation of the structured logs needed for this analysis.

Terminology

Summary

The paper addresses the critical challenge of transforming vast quantities of unstructured textual data—such as customer service transcripts, medical notes, or incident reports—into structured, machine-readable event logs suitable for process mining analysis. This capability is vital because traditional process mining techniques require sequential records of activities (event logs) to model workflows; however, real-world data often exists only in narrative form. By leveraging fine-tuned Large Language Models (LLMs), the authors propose a robust methodology that automates this extraction, thereby enabling deeper quantitative insights into complex organizational processes that were previously inaccessible due to data format limitations.

The Core Problem and Motivation

The primary limitation in applying process mining remains the data gap between narrative text and structured event logs. Traditional methods struggle with the inherent variability and ambiguity found in natural language. The authors argue that simply using off-the-shelf LLMs is insufficient, necessitating a specialized fine-tuning approach to ensure high fidelity. The goal is not merely extraction, but semantic understanding—identifying the who, what, when, and where of an action sequence from disjointed sentences. Key phrases highlighted include the necessity of moving from macro to micro analysis and achieving semantic-aware process mining tasks.

Methodology: Fine-Tuning for Event Extraction

The proposed system builds upon a foundational LLM architecture, which is then adapted using domain-specific datasets. The fine-tuning process is crucial because it grounds the general capabilities of the LLM in the specific vocabulary and syntactic structures of the target domain (e.g., healthcare or IT incident response). This customized adaptation allows the model to perform highly accurate Information Extraction (IE). The methodology involves training the model on pairs of (unstructured text, structured event log), thereby teaching it to map narrative flow directly onto process elements.

Structured Output Generation and Schema Alignment

A critical component of the framework is ensuring that the extracted events adhere to a predefined process schema. The LLM must not only identify an action but also correctly classify it according to established business rules. The system employs a multi-stage decoding process: first, it identifies potential entities and activities; second, it establishes temporal ordering based on linguistic cues (e.g., subsequently, afterward); and finally, it structures these findings into standard event log tuples (case id, activity, timestamp). This rigorous alignment process ensures the output is immediately usable by established process mining toolkits.

Evaluation and Benchmarking

To validate the system's efficacy, the authors detail a comprehensive evaluation protocol that moves beyond simple accuracy metrics. The performance is benchmarked against established baselines, focusing specifically on metrics related to temporal coherence and schema compliance. The paper emphasizes testing robustness across diverse data sources, including those characterized by low signal-to-noise ratios. They present several key evaluation dimensions:

  1. Temporal Accuracy: Assessing the correct ordering of events extracted from non-linear text passages.

  2. Entity Resolution: Verifying that ambiguous references (e.g., pronouns) are correctly mapped back to specific actors or resources mentioned earlier in the text.

  3. Process Completeness: Measuring the ability to reconstruct an entire process flow, even if certain steps are only implied rather than explicitly stated in the text.

This rigorous testing confirms that fine-tuned LLMs significantly outperform general-purpose NLP models when tasked with complex, multi-step event log generation from unstructured text.

Improvements for AI systems

Given that I am operating under the assumption of high stakes—where errors carry significant financial risk—my improvements must prioritize rigor, verifiability, and domain constraint enforcement over sheer generative fluency. The current state-of-the-art LLM applications in Process Mining (PM) often suffer from hallucination or a failure to adhere to the strict formalisms required by BPMN or Petri Net models.

Based on the synthesis of these references—which heavily point toward leveraging LLMs for unstructured data extraction, event log construction, and process discovery—I propose moving beyond simple extraction and building a multi-stage, constrained reasoning architecture.


The CPDE is not simply an LLM wrapper; it is an orchestrated pipeline that uses the LLM as a powerful hypothesis generator and an external, formal process model validator as the final arbiter of truth.

1. Improvement: Integrating Formal Constraint Solvers (The Validator Layer)

  • Problem Addressed: LLMs are probabilistic and can generate syntactically correct but logically impossible or inconsistent process steps (e.g., skipping a mandatory approval step).

  • Improvement: Implement a mandatory, two-stage verification loop. After the LLM generates a candidate event log sequence or process model structure, this output is passed to an external, deterministic constraint solver (like those used in Petri Net analysis or formal grammar checking).

  • Mechanism: The system must be trained not just on what the text says, but on what is possible according to the target domain's ontology (e.g., A patient cannot move from 'Triage' to 'Discharge' without passing through 'Initial Consultation'). If the generated sequence violates a hard constraint, the LLM must be prompted with a specific failure message and forced to regenerate until compliance is achieved.

2. Improvement: Multi-Dimensional Semantic Graph Generation (Beyond Triples)

  • Problem Addressed: Simple NLP extraction often yields linear triples (Subject-Predicate-Object), losing crucial contextual relationships or temporal dependencies inherent in the narrative text (e.g., "The manager approved the request after the system flagged an error").

  • Improvement: The LLM must be guided to generate a Temporal Semantic Graph. Instead of just extracting entities and relations, it extracts nodes, edges, and associated temporal/causal weightings.

  • Mechanism: This requires fine-tuning the model (or using advanced prompt engineering) to identify causality markers (because, following, due to) and temporal ordering markers (subsequently, prior to). The output is a graph structure ready for visualization and formal analysis, providing richer input for the PM engine than raw event logs.

3. Improvement: Retrieval-Augmented Process Modeling (RAG-PM)

  • Problem Addressed: General LLMs lack deep, niche domain knowledge (e.g., specific hospital protocols or internal regulatory compliance steps). Fine-tuning on massive, diverse datasets is prohibitively expensive and requires data that may not exist.

  • Improvement: Implement a specialized RAG layer built on curated, domain-specific knowledge bases (KBs), industry standards documents, and successful historical process models (the golden standard processes).

  • Mechanism: When processing a new document, the system first queries this RAG layer to retrieve relevant contextual snippets. These snippets are then prepended to the LLM prompt as high-priority context. This forces the LLM's reasoning pathway to be grounded only in verified industry best practices before generating any output, dramatically reducing hallucination risk in critical applications.

The Constrained Process Discovery Engine (CPDE) can perform the following high-value, mission-critical tasks:

  1. Automated Compliance Auditing from Unstructured Text:
  • Function: Ingest thousands of pages of qualitative data (e.g., customer complaints, incident reports, medical discharge summaries).

  • Output: It doesn't just summarize; it generates a validated, traceable process deviation report. For every step found in the text that violates the pre-loaded compliance ontology (e.g., HIPAA rules), it flags the exact sentence fragment and details why it is a violation, providing immediate remediation instructions.

  1. Zero-Shot Process Model Generation:
  • Function: Analyze a completely novel type of business workflow described only in narrative text (e.g., How to onboard a new vendor in Country X).

  • Output: It generates a complete, formal BPMN 2.0 compliant model diagram definition file (.bpmn) and an associated event log schema, all guaranteed to be logically consistent with the known domain constraints (via the Validator Layer).

  1. Root Cause Analysis with Process Mapping:
  • Function: Given a failure event (e.g., Project delayed by 3 weeks), the CPDE analyzes associated textual logs from multiple

Abstract

Process mining (PM) provides a powerful framework for discovering and optimizing operational processes from event data. However, the efficacy of PM techniques is strictly predicated on the availability of structured event logs. Thus far, event logs have often been laboriously created by domain and process mining experts. This costly effort causes large portions of organizational knowledge, including incident tickets, manuals, and textual reports, to remain underutilized. We address this bottleneck by investigating the efficacy of Large Language Models (LLMs) as automated data translators. We propose a scalable framework that leverages LLMs as data translators to bridge the gap between unstructured textual resources and structured event data. We finetune LLMs on a newly created text-to-log dataset, demonstrating that the resulting models can extract high-fidelity event logs from unstructured resources. Our results show that this finetuning approach outperforms few-shot or zero-shot prompting by a large amount, highlighting finetuning as a necessary pre-condition for generating reliable event data. We conclude that our method provides a promising pipeline for making previously unused data available to process mining ecosystems, effectively expanding the possibilities of using PM to further investigate organizational workflows.

Sources

Related papers