Event Abstraction for Enterprise Collaboration Systems to Support Social Process Mining

arXiv:2308.04396 · cs.LG, cs.AI · Submitted 2023-08-09 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Event Abstraction for Enterprise Collaboration Systems to Support Social Process Mining".

Jane: The paper was written by Jonas Blatt, Patrick Delfmann and Petra Schubert from University of Koblenz.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we’re digging into a paper that’s got a real mouthful of a title: “Event Abstraction for Enterprise Collaboration Systems to Support Social Process Mining.” Jane, I’m going to need your help unpacking that one, because it’s dense.

Jane: Happy to, Tom. So the title is really three parts. First, “Event Abstraction” — that’s about taking super detailed, messy records of what people do on a computer and turning them into bigger, meaningful actions. Think of it like turning every single keystroke into “wrote an email” instead of “pressed the letter E.”

Tom: Right, and the second part, “Enterprise Collaboration Systems” — those are tools like HCL Connections, which is the one they actually studied. These are platforms where teams share files, write wiki pages, post blogs, comment on stuff. Basically the digital office.

Jane: Exactly. And the third part, “Social Process Mining” — that’s the exciting bit. Process mining usually looks at structured business software like ERP systems to find how work flows. But here, they want to do that for collaborative, social software. The goal is to see patterns in how people actually work together, not just how they process orders.

Tom: And that’s where the problem starts. These collaboration systems log everything at a super fine level. One user action, like uploading a file, can trigger three or four log entries. And those entries might not even happen in the same order every time. So if you run standard process mining on that, you get what they call a “spaghetti model” — just a tangled mess of lines.

Jane: Yeah, I’ve seen those. They look like a plate of spaghetti thrown at a wall. Completely unreadable.

Tom: So this paper’s whole point is to fix that. They built a method called ECSEA — that’s short for Enterprise Collaboration System Event Abstraction — which learns how to group those tiny log entries into meaningful high-level activities. And once you have that cleaner log, you can actually do social process mining and see real collaboration patterns.

Jane: And the cool part is they don’t need a domain expert to sit there and manually define what each log entry means. They train the model by observing real user clicks and comparing them to the system logs. So it’s data-driven, not guesswork.

Tom: That’s the part I love. It’s like teaching a translator by showing them a conversation in two languages side by side, instead of handing them a dictionary.

Jane: And they claim it works — they got up to ninety-six percent accuracy on real data from a university collaboration system. That’s pretty impressive.

Tom: It is. But I’m curious about how they actually pulled that off. That’s what we’re going to get into next — the guts of the method.

Paper discussion segment 2: Jane: So we’ve set the stage. The paper is “Event Abstraction for Enterprise Collaboration Systems to Support Social Process Mining,” and the problem is that collaboration system logs are too messy for process mining. Now let’s talk about how they actually solved it.

Tom: Right. So the authors — Jonas Blatt, Patrick Delfmann, and Petra Schubert from the University of Koblenz — they took a supervised machine learning approach. And here’s the clever bit: they didn’t try to guess what the high-level activities should be. They actually observed real users clicking around in the system and recorded those clicks as the “high-level” truth.

Jane: So they built an observer tool that tracks what buttons people press. That gives them a clean log of actual user intentions — like “I created a wiki page” or “I uploaded a file.” Meanwhile, the system itself is logging every tiny event in the background.

Tom: And then they compare the two. They take the fine-grained system logs and the observed high-level logs, and they train a model that learns which sequences of low-level events correspond to which high-level activities.

Jane: And the tricky part is that it’s not a simple one-to-one mapping. They list five challenges in the paper. One, a single high-level action can produce multiple low-level events. Two, two different high-level actions can overlap in time. Three, the order of low-level events can vary. Four, the same low-level event can be part of different high-level activities. And five, some low-level events are just noise — they call them “ghost activities.”

Tom: Ghost activities — I love that name. Like when the system logs that someone visited a community page, but that doesn’t really tell you anything about what they were trying to do.

Jane: Exactly. So their algorithm has to handle all five of those at once. And that’s what makes it different from previous work. They actually reviewed twenty-five existing event abstraction approaches and found that none of them could handle all five challenges simultaneously.

Tom: And that’s the gap they’re filling. Their method, ECSEA, uses two maps. One map goes from low-level activities to possible high-level activities. The other map goes from high-level activities to the sequences of low-level events that were observed for them, along with a count of how often each sequence appeared.

Jane: So it’s like building a phrasebook. You see that “wiki.page.created” followed by “wiki.page.follow” usually means someone created a wiki page. And you also learn that “wiki.page.updated” can be part of two different high-level activities depending on context.

Tom: And then when they apply the model to a new low-level log, they use a greedy sliding window approach. They look at the first few events, check which high-level activity they most likely belong to, and then merge them into one event.

Jane: And they use a threshold to decide whether a mapping is good enough, and they calculate an error score based on how different the observed sequence is from the learned ones. That’s how they handle the varying order problem.

Tom: It’s a pretty elegant design. But the real question is — does it actually work? And that’s what we’re going to talk about next, because they ran some pretty thorough evaluations.

Paper discussion segment 3: Tom: So we’ve covered the method. Now let’s get into the results, because that’s where the rubber meets the road. The paper is “Event Abstraction for Enterprise Collaboration Systems to Support Social Process Mining,” and the authors ran two evaluations.

Jane: Right. First, they did a synthetic test. They took a real event log from the BPI Challenge two thousand twenty — that’s a well-known dataset in the process mining community — and they treated it as if it were a high-level log. Then they artificially generated low-level logs from it, with different numbers of low-level activities per high-level activity.

Tom: And they made those synthetic logs match the messy characteristics we talked about — multiple events per activity, overlapping events, varying order, and so on. They generated seventy different low-level logs and trained models on each one.

Jane: And the results were really strong. The accuracy was always above ninety-eight percent, even when they had eight low-level activities for every high-level activity. And they only used about ten percent of the traces for training.

Tom: That’s a big deal. It means you don’t need to observe the system forever. You can observe a small sample, train the model, and then apply it to years of historical data.

Jane: Then they did a real-world test. They used a system called UniConnect — that’s an actual enterprise collaboration system at their university with over three thousand users. They observed it for three months, recorded the high-level clicks, and extracted the corresponding low-level logs.

Tom: And they got ninety-six percent accuracy on that real data. That’s not just a lab experiment — that’s a working system.

Meng: Hey, can I jump in here? I’m wondering about the practical side. If I’m an IT team at a company, what does it take to actually deploy this? Do I need to install that observer tool?

Jane: Great question, Meng. So the observer is a browser-based tool that records clicks. And the key insight is that you only need to run it once per system type. Once you’ve trained the model on one instance of, say, HCL Connections, you can apply it to other instances of the same system — as long as they haven’t been heavily customized.

Meng: So it’s a one-time setup cost. That’s actually pretty reasonable. But what about the accuracy drop from ninety-eight percent to ninety-six percent — is that just noise, or is there something about real data that’s harder?

Tom: That’s a good observation. Real data has more variability. Users don’t follow clean paths. They interrupt themselves, they switch tasks, they do things in weird orders. The synthetic data was generated to be messy, but real users are messier.

Jane: And they also mention that the model can be reused for other instances of the same system type. They actually tested it on another instance called KoCo and it worked. So the investment pays off across multiple deployments.

Meng: That’s the kind of thing that makes this actually viable in a company. You don’t want to redo the whole observation process for every department.

Tom: Exactly. And the authors see this as a preprocessing step for something bigger — social process mining. Once you have clean high-level logs, you can start looking for collaboration patterns, like how teams work on documents together or how knowledge flows through an organization.

Jane: And that’s where the real value is. It’s not just about making pretty process maps. It’s about understanding how people actually collaborate.

Conclusion: Tom: Alright, let’s wrap this up. We’ve been talking about “Event Abstraction for Enterprise Collaboration Systems to Support Social Process Mining” — and honestly, this is one of those papers that solves a real bottleneck.

Jane: Yeah. The problem was that collaboration systems generate logs that are too fine-grained for process mining. You’d get spaghetti models that nobody could read. And the existing event abstraction approaches couldn’t handle all the messiness — overlapping events, varying order, ghost activities, all that.

Tom: And what these authors did was build a method that learns from observed user behavior. They watch people click, they compare that to the system logs, and they train a model that can automatically convert messy low-level logs into clean high-level ones.

Jane: And the results are solid — over ninety-eight percent accuracy on synthetic data and ninety-six percent on real data from a university system. Plus, the model is reusable across instances of the same software.

Meng: And from a practical standpoint, that one-time observation cost makes it feasible for real companies. You don’t need to run the observer forever.

Lu: And I think the bigger picture here is that this opens the door to actually studying collaboration as a process. We’ve been able to study how orders flow through an ERP system for decades. Now we can study how knowledge flows through a team, how ideas get developed on a wiki, how a community forms around a document. That’s a whole new frontier.

Lalam: And if I may add — this could change how organizations understand their own culture. When you can see patterns in collaboration, you can identify which teams are siloed, which workflows are actually working, and where knowledge gets stuck. That’s not just a technical improvement; it’s a way to make organizations more humane and effective.

Tom: Beautifully said. So that’s it for this paper. We’ve covered the problem, the method, the results, and the implications. Next up, we’ve got another paper that’s going to push into a completely different corner of process mining. Stay tuned.

Jane: Thanks for listening, everyone. We’ll see you on the next one.

Jonas Blatt, Patrick Delfmann, Petra Schubert

University of Koblenz

cs.LG, cs.AI

Submitted: 2023-08-09

Updated: 2026-08-18

Comments: 8 pages, 1 figure, 3 tables

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 46/100

Key concepts

Event Abstraction
The process of taking highly detailed, messy records of computer actions (like keystrokes) and grouping them into broader, meaningful actions. For example, turning many individual clicks into a single action like 'wrote an email.'
Enterprise Collaboration Systems
Digital platforms used by teams for sharing files, writing wiki pages, posting blogs, and commenting—essentially the digital office. The paper studied these systems to analyze team workflows.
Social Process Mining
A field that applies process mining techniques not to structured business software (like ERPs), but to collaborative, social software. Its goal is to find patterns in how people actually work together.
ECSEA
The method developed by the authors (Enterprise Collaboration System Event Abstraction). It learns how to group tiny log entries into meaningful high-level activities by observing real user clicks and comparing them to system logs.

Terminology

Summary

Summary

This paper addresses the challenge of applying Process Mining (PM) to Enterprise Collaboration Systems (ECS) by introducing a novel event abstraction approach called ECSEA (ECS Event Abstraction). The authors note that while PM has been successfully applied to process-oriented systems like ERP and workflow management systems, it is less suited for communication- and document-oriented ECS because their event logs are very fine-granular and applying PM to them results in spaghetti models that are overly complex and hard to interpret. The research direction is called Social Process Mining (SPM), which aims to identify collaboration patterns in ECS and focuses on how users of ECS 'move through the system' i.e., which functions they typically execute in which order and how they collaborate on social content.

The paper identifies five specific challenges (C1-C5) with ECS event logs that must be addressed by any event abstraction technique:

  • C1 Multiple LL events: Particular HL events may result in multiple LL events. Thus, a HL activity may be expressed by multiple LL activities. Furthermore, a particular HL activity may be expressed by different sets of LL activities.

  • C2 Overlapping events: Multiple LL events that are related to two (or more) HL events may overlap temporally. Moreover, the interval between the time-stamps of multiple LL events that are triggered by an activity is not always the same.

  • C3 Different LL event ordering: The ordering of the LL events might vary slightly. This is because the underlying system performs tasks independently so that other activities may be logged first.

  • C4 Multiple LL to HL activity mappings: A LL activity may be triggered by multiple HL activities. Thus, it is possible that a LL activity is part of more than one mapping. This characteristic represents an m:n mapping.

  • C5 Ghost activities: There may exist LL events with no related HL activity.

The authors conducted a literature review of 25 publications on event abstraction and found that none of the existing approaches meet all these challenges simultaneously. They created a concept matrix (Table II) showing that while some approaches address parts of the challenges, none work without manual preprocessing based on a-priori domain knowledge (challenge C6). For example, approaches by Baier and Mendling, Senderovich et al., and Mannhardt et al. require predefined mappings, interaction sets, or activity patterns that need domain expertise. The approach by Tax et al. based on Conditional Random Fields is most similar but is not able to detect overlapping HL activities (represented as interleaving LL events) and requires annotated event logs where each event has a label referring to the HL activity, which is not available in the ECS context.

The ECSEA approach is based on supervised machine learning. It trains a model by comparing observed high-level (HL) traces with related low-level (LL) traces. The HL traces are gathered by observing the ECS using a click path observer, which records clicks on particular elements (e.g., buttons), which represent certain HL activities. This observation needs to be done only once per system type because other instances produce similar LLLs, which can be converted into HLLs by the already trained ECSEA model.

The model is formally defined as: "Let m = (llc, hlc) be a model with two maps llc and hlc. The map llc assigns single LL activities to a set of HL activities, i.e., llc: Al 7→ X ⊆ Ah. Further, let SLL = ⟨x1, x2,..., xn⟩xi ∈ Al the universe of sequences of LL activities. The map hlc assigns single HL activities to a set of sequences of LL activities that count their occurrences, i.e., hlc: Ah 7→ 2SLL, N "

The training phase uses a fitting function that takes a LL trace and a related HL trace (with the same case identifier), along with parameters τ (maximal time-span between LL and HL event timestamps) and Γ (a set of attribute names used to group similar LL events). The algorithm creates sequences of LL events that share the same attributes in Γ, then maps each HL event to a sub-sequence of LL events where the time distance is below τ and the HL event has minimal temporal distance. The model is trained iteratively, and hyperparameter optimization is performed using grid search to maximize accuracy, which is calculated using the normalized Damerau–Levenshtein distance between generated and original HL traces.

The application phase uses a greedy algorithm based on sliding windows (Algorithm 1). It processes a LL trace by repeatedly extracting a window of events (with the same Γ attributes and within time-span τ), finding the best mapping using the model (Algorithm 2), and merging the mapped events into a new HL event. The timestamp of the new event is determined by a parameter Φ (timestamp-merge-type) with valid values MIN, MAX, MEAN, and MEDIAN. Ghost events are handled by removing the first event if no progress is made in an iteration. After processing all windows, events with the same activity, same Γ attributes, and timestamp distance below τ are merged.

The evaluation was conducted in two parts. First, with synthetic data using the PermitLog from the BPI Challenge 2020, the authors split each activity into multiple LL activities and generated 70 different LLLs with varying configurations (2, 3, 4, 5, 6, 7, and 8 new LL activities per HL activity). They trained models using 706 of 7065 traces and achieved accuracy is always above 98% and drops slightly the more LL activities are used per HL activity. The timestamp-merge-type Φ influenced accuracy only slightly. Second, with real-world data from UniConnect, an operational large-scale Enterprise Collaboration System with more than 3000 users, the authors observed the HLL for three months using their observer, extracted the corresponding LLL, and trained an ECSEA model. They reached an accuracy up to 96% and successfully applied the model to KoCo, another instance of HCL Connections.

The authors conclude that ECSEA produces accurate HLLs and that the trained model can also be used to abstract LLL from different instances of the same system type (without major customizations), resulting in a high reusability of the trained model in research and practice. They acknowledge the limitation that a HLL must be recorded through observation but note this has to be done only once. Future work includes finding a suitable case identifier beyond the workspace (e.g., the set of social documents that is jointly worked on) and extending the algorithm to create HL events with start and end lifecycle transactions using min and max timestamps, as well as adding weight factors for hlc mappings optimized with a genetic algorithm.

Improvements for AI systems

Based on the scientific paper, here are the specific improvements I can make to an AI system, along with what the improved system can do:


Improvement: Implement a supervised machine-learning model (ECSEA) that learns mappings between low-level (LL) system events and high-level (HL) user activities. The model uses two maps: llc (LL activity → set of possible HL activities) and hlc (HL activity → sequences of LL activities with occurrence counts). Training is done by comparing observed HL traces (from click-path observers) with system-generated LL traces.

What the improved AI system can do:

  • Automatically convert raw, fine-grained event logs (e.g., from Enterprise Collaboration Systems like HCL Connections) into interpretable, high-level activity logs.

  • Handle cases where multiple LL events map to a single HL activity (C1), overlapping temporal events (C2), varying LL event ordering (C3), many-to-many LL-to-HL mappings (C4), and ghost events with no HL meaning (C5).

  • Operate without requiring a-priori domain knowledge or manual preprocessing (C6), unlike existing approaches.

Improvement: Implement the apply function (Algorithm 1) that processes a LL trace using a sliding window. The window is defined by a maximal time-span τ and grouping attributes Γ (e.g., user ID). The algorithm iteratively:

  • Builds a window of LL events with matching attributes and temporal proximity.

  • Uses getBestMapping (Algorithm 2) to find the best HL activity by minimizing a normalized Damerau–Levenshtein distance, weighted by mapping frequency.

  • Merges matched LL events into a single HL event, assigns a timestamp based on a configurable merge type (MIN, MAX, MEAN, MEDIAN), and removes them from the trace.

  • Handles ghost events by removing the first event if no mapping is found in an iteration.

Improvement: Implement a grid-search-based hyperparameter optimization during training. The system iterates over different values for:

  • τ (max time-span between LL and HL events),

  • ϑ (mapping threshold for accepting/rejecting mappings),

  • Φ (timestamp merge type: MIN, MAX, MEAN, MEDIAN).

The model with the highest accuracy on the training set (using normalized Damerau–Levenshtein distance) is selected and then validated on a held-out test set to assess overfitting.

Improvement: Train the ECSEA model once on a controlled observation of an ECS (e.g., three months of click-path data) and reuse it for other instances of the same system type (e.g., different HCL Connections deployments) without re-observation.

Improvement: Integrate ECSEA as the first step in a larger SPM pipeline that includes:

  • Case identification (e.g., using social documents as case IDs),

  • Process discovery,

  • Frequent subgraph mining for pattern detection.

Improvement: The evaluation showed that training on only 10% of the traces (706 out of 7065) yields >98% accuracy. The system can be configured to train on a small, representative sample, reducing the need for extensive observation periods.

The improved AI system can:

  • Transform raw ECS logs into interpretable high-level activity logs with ≥96% accuracy, enabling process mining on collaboration systems.

  • Handle all five identified ECS log challenges (multiple events, overlaps, ordering variability, many-to-many mappings, ghost events) without manual intervention.

  • Reuse trained models across system instances, reducing deployment cost.

  • Optimize its own hyperparameters for maximum accuracy on new data.

  • Support real-time and batch processing of event streams.

  • Enable Social Process Mining to uncover collaboration patterns, which is currently impossible with state-of-the-art PM tools.

These improvements directly address the paper’s goal of making collaboration work in ECS interpretable and analyzable, with immediate applicability to HCL Connections and similar platforms.

Abstract

One aim of Process Mining (PM) is the discovery of process models from event logs of information systems. PM has been successfully applied to process-oriented enterprise systems but is less suited for communication- and document-oriented Enterprise Collaboration Systems (ECS). ECS event logs are very fine-granular and PM applied to their logs results in spaghetti models. A common solution for this is event abstraction, i.e., converting low-level logs into more abstract high-level logs before running discovery algorithms. ECS logs come with special characteristics that have so far not been fully addressed by existing event abstraction approaches. We aim to close this gap with a tailored ECS event abstraction (ECSEA) approach that trains a model by comparing recorded actual user activities (high-level traces) with the system-generated low-level traces (extracted from the ECS). The model allows us to automatically convert future low-level traces into an abstracted high-level log that can be used for PM. Our evaluation shows that the algorithm produces accurate results. ECSEA is a preprocessing method that is essential for the interpretation of collaborative work activity in ECS, which we call Social Process Mining.

Related papers