Workflow Cards: Structured Summaries of Workflow Executions Using Provenance Data
Nicola Giuseppe Marchioro, Gabriele Padovani, Amal Gueroudji, Rafael Ferreira da Silva, Wesley Brewer, Valentine Anantharaj, Sandro Fiore, Renan Souza
University of Trento · Argonne National Laboratory · Oak Ridge National Laboratory
cs.DC, cs.AI
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: Accepted at eScience2026
Code: https://github.com/ORNL/flowcept
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: Workflow Cards: Structured Summaries of Workflow Executions Using Provenance Data This paper introduces Workflow Cards, a structured documentation artifact designed to summarize workflow execution
Terminology
Summary
Workflow Cards: Structured Summaries of Workflow Executions Using Provenance Data
This paper introduces Workflow Cards, a structured documentation artifact designed to summarize workflow execution provenance data in a format that is readable by both humans and large language models (LLMs). The work is motivated by the observation that existing documentation frameworks such as Model Cards and Data Cards focus on static artifacts (datasets and trained models) but overlook the workflow executions that produce, transform, and evaluate them. These executions hold critical details about data preparation, parameter choice, runtime behavior, resource use, and intermediate transformations, precisely where bias, performance variation, and reproducibility gaps tend to originate.
The paper has two main parts. First, it defines a Workflow Card template informed by a representative set of provenance questions that surface from the execution-level data missing from Model and Data Cards. Second, it evaluates how effectively LLMs use Workflow Cards to understand workflow executions compared with querying provenance databases through a schema-based interface.
The contributions of the work are: (i) introducing Workflow Cards as workflow-oriented documentation artifacts for exposing workflow provenance in a human- and LLM-readable format; (ii) organizing common workflow provenance questions from the literature into three categories (input/dataset, workflow, output/model) that guide the Workflow Card design; (iii) implementing an open-source provenance-aware Workflow Card generation approach integrated with the provenance systems yProv4ML and Flowcept; and (iv) evaluating Workflow Cards both as complementary documentation artifacts alongside Model and Data Cards and as an LLM-facing provenance representation, showing on a real HPC ML workflow that Workflow Cards nearly double answer quality for LLM-based provenance question answering compared with schema-based querying.
The Workflow Card template is organized as a top-down summary that begins with a high-level view of the run and progressively descends into per-activity detail. It opens with a short Workflow block that names and describes the run in plain language, followed by a Summary block that records the identity and lifecycle of the execution, including a unique execution identifier, the workflow version, start and end timestamps, total duration, overall status, and the entrypoint repository, branch, and commit that pin the exact code used. An Infrastructure block then captures the computational environment, covering the operating system, hardware, runtime, resource manager, and a snapshot of the software dependencies. The Workflow Overview gathers run-level aggregates such as activity and task counts, governing arguments, aggregate resource usage across CPU, memory, GPU, disk, and network, and a free-text Observations field for anomalies or decisions made mid-run, before the Activities block repeats a compact record for each stage that reports its tasks, timing, status, hosts, and summarized inputs and outputs. A final Significant Workflow Artifacts block lists only the inputs and outputs needed to reproduce or interpret the run, each identified by a short name and a resolvable reference. Every field is marked as Required or Optional, and a tilde denotes a value that was not captured, which keeps the card lightweight while signaling missing information rather than hiding it. Because each card describes a single immutable execution, it deliberately omits metadata that may evolve independently of the run, and it favors aggregate, high-level descriptions over exhaustive traces so that it stays small enough to read directly and to pass to an LLM as context.
The evaluation is conducted across two independent benchmarks sharing the same protocol. Benchmark I assesses whether Workflow Cards provide execution-level information absent from existing Model and Data Cards. Benchmark II evaluates their accessibility advantage for LLMs compared to a lightweight schema-based interface to a provenance database. Two LLMs act as judges throughout: a smaller model, nemotron-nano-3 (30B parameters), and a substantially larger one, gpt-oss-120b (120B parameters). An LLM-as-a-Judge approach is used because it scales evaluation across the full question set and has been shown to correlate strongly with human judgment on answer-rating tasks, and it is additionally validated against an independent human rater in Benchmark II.
For Benchmark I, a set of simulated machine learning fine-tuning workflows was constructed using publicly available Hugging Face repositories. For each workflow configuration, the pretrained Model Card, associated Data Cards, and fine-tuned Model Card were collected. Since public Workflow Cards do not yet exist, executions were simulated and Workflow Cards were generated with the proposed template. Workflow-level fields were synthetically generated using Claude Sonnet 4.6, constrained by the real Model and Data Cards. The evaluation compared provenance question answering under three different information settings: (i) Full Knowledge, in which all documentation artifacts were concatenated and provided simultaneously to the LLM; (ii) Single-card, in which each card was evaluated independently; and (iii) Leave-one-out, where all cards except one were provided to the LLM.
Results from Benchmark I show that the least informative card, the one with the largest average drop from the all-cards reference, is the pretrained Model Card. The Workflow Card, by contrast, matches the full set, reaching 0.619 against the all-cards reference of 0.611 while being, on average, 33.2 kB smaller, showing the same answer quality at a fraction of the context size, roughly 8k fewer tokens at about four characters per token. In the leave-one-out setting, removing the Workflow Card produces the largest drop-off, followed by the Fine-tuned Model Card. Quantitatively, the execution-level information unique to the Workflow Card more than doubles accuracy over the remaining cards: removing it drops the mean score from the all-cards reference of 0.611 to 0.265 (a 2.3× drop), by far the largest effect of any artifact; removing the Data Card or the Pretrained Model Card, by contrast, changes the score only marginally. The pairwise cross-entropy between cards accounts for the small effect of removing the Pretrained Model Card: much of its information is also captured by the Fine-tuned Model Card, but not the reverse, an asymmetry reflected in the cross-entropy between the two cards.
When performance is isolated per question focus, among the Data and Model Cards, the card matching a question's focus performs best (e.g., the Data Card on data questions). However, the Workflow Card alone matches or exceeds the All-Cards reference on every question type (Data, Model, and Workflow), despite using less context. Specifically, for Data questions the Workflow Card scores 0.6884 versus the All-Cards reference of 0.6612; for Model questions it scores 0.6688 versus 0.6035; and for Workflow questions it scores 0.4005 versus 0.3999.
For Benchmark II, a scientific machine learning pipeline for the DLESyM climate forecasting model was used. The workflow spans four steps: (i) pre-processing raw data using the REDI tool, (ii) fine-tuning DLESyM, (iii) forecasting, and (iv) evaluating forecasts via Empirical Orthogonal Function. The workflow was instrumented using both Flowcept and yProv4ML to capture provenance across all steps. The card generator maps captured entities, activities, agents, task timestamps, resource records, and I/O metadata into the template blocks. A 17-question subset targeting mainly workflow provenance was used. Two representations of this same captured provenance were compared, differing only in how it is presented to the LLM: a lightweight schema-based querying approach, where the LLM is only given the schema and structural provenance rather than the raw records, and a Workflow Card generated from the same provenance data and supplied directly as context.
Results from Benchmark II show that on average, the card setting is noticeably better: across both answering models, moving from schema-based querying to Workflow Cards nearly doubles the mean rating (from ≈0.45 to ≈0.87), an improvement on which the LLM-as-a-Judge and human evaluations agree. The spread of ratings is similar for both. Agreement among judges is strongest when cards are the input. For the schema-based approach, the LLM judges score considerably harsher than the human rater, likely because raw schema-and-structure responses are harder for an LLM to interpret than a pre-generated card. No clear self-preference pattern was found in the LLM-as-a-Judge ratings: nemotron-nano-3 rates its own answers lower than gpt-oss-120b's, while gpt-oss-120b grades slightly more generously overall. The human ratings follow the same trend as the model judges, supporting the use of the LLM judges as a calibrated automated evaluation. Quantitatively, the human ratings agree closely with the LLM judges (Pearson r = 0.74 for the mean of the two judges, and 0.66 and 0.77 individually; mean absolute difference 0.13 on the 0 to 1 scale).
The paper concludes that Workflow Cards capture unique, process-level provenance absent from existing artifacts and, on a real HPC workflow, nearly double LLM answer quality on provenance questions compared with schema-based querying. By making execution provenance directly readable to both people and LLMs, this work takes one more step toward the transparency and reproducibility of scientific workflows. The Workflow Card template, benchmark question sets, LLM prompts, and the code to reproduce the benchmarks are publicly available, and the updated versions of Flowcept and yProv4ML featuring Workflow Card generation support are open-source.
Improvements for AI systems
Improvements to AI Systems:
-
Provenance-Aware Context Generation: AI systems can automatically generate structured Workflow Cards from execution logs, provenance databases, or instrumentation hooks (e.g., via Flowcept/yProv4ML), condensing raw traces into human- and LLM-readable summaries. This enables the AI to answer questions about data lineage, parameter choices, resource usage, and intermediate transformations without querying complex schemas.
-
Enhanced Question-Answering over Workflow Executions: By ingesting Workflow Cards as context, AI systems can nearly double answer quality (from ≈0.45 to ≈0.87) on provenance-related queries compared to schema-based interfaces. The improved system can directly respond to questions like
What was the exact commit used?
orWhich GPU resources were consumed during preprocessing?
with higher accuracy and lower latency. -
Complementary Documentation Integration: AI systems can combine Workflow Cards with existing Model Cards and Data Cards to form a complete knowledge base. This allows the AI to reason across artifacts—e.g., explaining how a specific training run's hyperparameters (from Workflow Card) affected model performance (from Model Card) or dataset biases (from Data Card)—with a 2.3× accuracy drop when Workflow Card is missing, highlighting its critical role.
-
Efficient Context Utilization: Workflow Cards are compact (33.2 kB smaller than full documentation sets, 8k fewer tokens), enabling AI systems to process more workflows within context limits. The improved system can handle multiple executions simultaneously, compare runs, or provide historical analysis without truncation or performance degradation.
-
Automated Reproducibility Audits: AI systems can use Workflow Cards to verify reproducibility by checking if all required fields (e.g., commit hash, environment, inputs/outputs) are present and consistent. The system can flag missing provenance (marked with tilde) and suggest corrective actions, improving trust in scientific workflows.
-
Cross-Question-Type Robustness: The improved AI system can answer data, model, and workflow questions with equal or better accuracy than using all documentation combined (e.g., 0.6884 vs. 0.6612 for data questions), making it a universal interface for execution-level insights without needing to switch between different card types.
-
Scalable LLM-as-a-Judge Evaluation: AI systems can adopt the paper's evaluation protocol (LLM-as-a-Judge with cross-model validation) to self-assess answer quality on provenance tasks, using the calibrated correlation with human raters (Pearson r = 0.74) to automate quality control in production pipelines.
-
Adaptive Documentation Generation: AI systems can dynamically generate Workflow Cards at different granularities (high-level summary vs. per-activity detail) based on user queries or task complexity, optimizing for both readability and information density, as demonstrated by the template's hierarchical structure.
What the Improved AI System Can Do:
-
Answer complex provenance questions (e.g.,
Which step caused the memory spike?
orWhat changed between run A and B?
) with near-human accuracy. -
Generate reproducible execution summaries on-the-fly for any instrumented workflow.
-
Provide explainable, evidence-based answers by referencing specific card fields (e.g.,
The run failed at step 3 due to GPU OOM, as shown in the Activities block
). -
Support multi-run comparisons and trend analysis by ingesting multiple Workflow Cards simultaneously.
-
Operate in resource-constrained environments (e.g., edge devices) by using compact cards instead of full provenance databases.
Abstract
Model Cards and Data Cards have demonstrated the value of structured, human-readable documentation for machine learning artifacts, capturing their context, parameters, limitations, and intended use. However, these practices remain focused on static artifacts (the datasets and trained models themselves) while overlooking the workflow executions that produce, transform, and evaluate them. Such executions hold critical details about data preparation, parameter choice, runtime behavior, resource use, and intermediate transformations, precisely where bias, performance variation, and reproducibility gaps tend to originate. To close this gap, we introduce Workflow Cards: structured summaries that condense the machine-readable provenance data of a workflow execution into a form both humans and large language models (LLMs) can read and analyze. This paper has two main parts. First, it defines a Workflow Card template informed by a representative set of provenance questions that surface from the execution-level data missing from Model and Data Cards. Second, it evaluates how effectively LLMs use Workflow Cards to understand workflow executions compared with querying provenance databases through a schema-based interface. Results show that Workflow Cards provide execution-level information absent from existing card types, such as Model Cards and Data Cards, thereby filling an important documentation gap; and that Workflow Cards nearly double answer quality compared with schema-based querying, consistently across LLM-as-a-Judge and human assessments.
Sources
- Navigating Dataset Documentations in AI: A Large-Scale Analysis of Dataset Cards on Hugging Face
- What's documented in AI? Systematic Analysis of 32K AI Model Cards
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing