VirnyFlow: A Design Space for Responsible Model Development

arXiv:2506.01584 · cs.LG, cs.AI, cs.CY · Submitted 2026-08-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "VirnyFlow: Optimizing ML Pipelines for Accuracy, Fairness, and Stability at Scale".

Jane: The paper was written by Denys Herasymuk, Anastasiia Mozghova, Nazar Protsiv, Vladyslav Sydorak and Julia Stoyanovich from Ukrainian Catholic University and New York University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Alright, welcome back to the show, everyone. We are digging into a brand new paper today, and it's called "VirnyFlow: A Design Space for Responsible Model Development." Jane, I have to say, the title alone got me excited.

Jane: It got me excited too, Tom, because it's not just another AutoML paper. It's about building machine learning models in a way that actually cares about the real world, not just raw accuracy.

Tom: Exactly. And the authors here are from Ukrainian Catholic University and NYU — Denys Herasymuk, Nazar Protsiv, and Julia Stoyanovich. That's a solid team.

Jane: Right, and Julia Stoyanovich has been doing really important work on responsible data management, so this fits right into her wheelhouse. The name "VirnyFlow" — "Virny" actually means "faithful" or "true" in Ukrainian, which is a lovely touch.

Tom: Oh, that's a great catch. So the whole idea is that when you're building a model, you're not just trying to maximize one number like accuracy. You're juggling fairness, stability, maybe even things like how consistent the model is when you retrain it.

Jane: And that's the "responsible" part. The paper argues that most AutoML tools treat optimization like a black box — you give it a dataset, it spits out a "best" model, done. But that ignores the context. Who is this model affecting? What groups might it harm?

Tom: Right, and they give this great example in the intro. There's a data scientist named Ann working on public health insurance eligibility. She needs to balance accuracy with fairness across sex and race, and even intersectional groups like Black women specifically.

Jane: And that's not something you can just automate away. The paper says fairness can't be fully automated because deciding what's fair depends on the socio-technical setting. A human has to make those judgment calls.

Tom: So VirnyFlow isn't trying to replace the data scientist. It's giving them a flexible playground to experiment, define their own objectives, and iterate. That's a fundamentally different philosophy from the "just give me the best model" approach.

Jane: I love that framing. It's not about removing the human from the loop — it's about giving the human better tools to make informed decisions.

Tom: And that's what we're going to unpack today. But first, let me just say — this paper is not shy about its ambitions. They're calling it the "first design space for responsible model development." That's a bold claim.

Jane: It is, but I think they back it up. We'll get into the actual system design next, but the core idea is that you can define your own evaluation protocol — your own metrics, your own weights, your own sensitive attributes — and then the system optimizes across the entire ML pipeline.

Tom: So not just hyperparameter tuning, but also data preprocessing, fairness interventions, model selection, all of it, jointly.

Jane: Exactly. And that's a huge deal because errors introduced early in the pipeline can propagate downstream. If you're not cleaning your data responsibly, no amount of tuning will fix the bias.

Tom: Alright, I'm hooked. Let's get into how this actually works under the hood.

Summary: Tom: So we're back, and we're still on "VirnyFlow: A Design Space for Responsible Model Development." Jane, you gave us the big picture — now let's talk about what the system actually does.

Jane: Right. So the paper describes a five-step optimization process. First, you define your search space — that's all the possible pipeline components, like which imputation method to use, which fairness intervention, which model.

Tom: And then the system has to pick which pipelines to try. That's where it gets clever.

Jane: Very clever. They use something called a multi-armed bandit approach. Think of it like a slot machine where each arm is a different logical pipeline — a combination of components. The system keeps track of which arms have performed well historically and prioritizes those, but it also explores new ones so it doesn't get stuck.

Tom: And there's a scoring model that balances mean performance against variance, with a risk factor you can tune. If you're feeling adventurous, you set the risk factor high and it'll try more uncertain pipelines.

Jane: Exactly. But the real magic is in step three — physical pipeline selection. Once a logical pipeline is chosen, they use multi-objective Bayesian optimization to instantiate it with specific hyperparameters.

Tom: And this is where the multi-objective part shines. Instead of just optimizing F1 score, you can optimize F1, fairness metrics, and stability all at once. The system builds a Pareto front of trade-offs.

Jane: Right, so you're not getting one "best" model. You're getting a set of models that represent different trade-offs, and you as the human get to decide which one fits your context.

Tom: And then there's the pruning strategy. They adapted something called Adaptive Pipeline Selection from Alpine Meadow. The idea is you don't train every pipeline on the full dataset — you start with a fraction, like fifty percent, and if the pipeline is clearly performing badly, you kill it early.

Jane: That saves a ton of compute. And they extended it to handle multiple objectives by computing a weighted sum of errors across all your objectives.

Tom: Right, so it's not just about accuracy — it's about whether the pipeline is making progress on fairness and stability too.

Jane: And here's the thing I really appreciate — the whole system is built for interactivity. You can see results as they come in, adjust your objectives mid-run, and the system adapts.

Tom: They also built it on a distributed architecture with Kafka for messaging and MongoDB for storage. So it scales across multiple nodes, which we'll get to in the experiments.

Jane: But the key takeaway for me is that this isn't a black box. You're not just handing over your problem and getting a model back. You're actively shaping the optimization process.

Tom: And that's what makes it "responsible" — the human stays in the loop, making value judgments about what trade-offs are acceptable.

Jane: Exactly. Now, let's talk about whether it actually works. The experiments are pretty impressive.

Improvements: Tom: We're back on "VirnyFlow: A Design Space for Responsible Model Development," and now we get to the fun part — the results. Jane, what did they find?

Jane: So they ran experiments on five real-world datasets — diabetes, employment, public coverage, heart disease, and a larger employment dataset. And they compared VirnyFlow against two state-of-the-art AutoML systems: Alpine Meadow and auto-sklearn.

Tom: And the results are pretty striking. On the performance side, VirnyFlow consistently beat both baselines when you look at the average score across multiple metrics.

Jane: Right, and that's the key — they're not just comparing F1 scores. They're comparing a composite of accuracy, fairness, and stability. On the diabetes dataset, VirnyFlow got the best F1 and the best label stability. On the employment dataset, it got the best F1 and the best fairness metric for intersectional groups.

Tom: And on the public coverage dataset, it took a slight hit on F1 — about zero point zero three seven — but it massively improved fairness, especially for race. The SRD metric improved by zero point one one, which is huge.

Jane: That's the trade-off story. You can't have everything, but VirnyFlow lets you see those trade-offs clearly and choose what matters for your problem.

Tom: And then there's the scalability study. This is where it gets really interesting.

Jane: Oh, absolutely. They tested with up to one hundred twenty-eight workers across four nodes. Alpine Meadow and auto-sklearn are limited to a single node, so they cap out at thirty-two workers.

Tom: VirnyFlow achieved a speedup of seven point zero seven times on the heart dataset and five point zero three times on the large employment dataset with one hundred twenty-eight workers. That's a massive improvement.

Jane: And even at thirty-two workers, VirnyFlow was competitive with Alpine Meadow and beat auto-sklearn, despite not having the same low-level optimizations.

Tom: But there's a caveat — with fewer than eight workers, VirnyFlow is actually less efficient because of the overhead from Kafka and the distributed architecture.

Jane: Right, so if you're just running on a single machine with a few cores, it might not be the best choice. But if you have access to a cluster, it scales beautifully.

Tom: They also did a sensitivity analysis on configuration settings. They found that increasing the number of physical pipeline candidates per selection reduces runtime but can hurt performance. And they found that using multiple training set fractions for pruning doesn't always help.

Jane: So there's no free lunch, but the defaults they propose seem to be a solid starting point.

Tom: And here's what I find most exciting — they're not just optimizing for accuracy and fairness. They're also optimizing for model stability, which is something almost no other AutoML system does.

Jane: That's a big deal. In high-stakes domains like healthcare, an unstable model can give inconsistent predictions, which is dangerous. VirnyFlow lets you tune for that explicitly.

Tom: So the improvements over existing systems are clear — better multi-objective optimization, better scalability, and support for metrics that actually matter in the real world.

Jane: Now let's think about what this means for the broader field.

Conclusion: Tom: Alright, we're wrapping up our discussion of "VirnyFlow: A Design Space for Responsible Model Development." Jane, what's the big takeaway?

Jane: The big takeaway is that responsible AI isn't just about adding a fairness constraint at the end. It's about building systems that let humans define what "responsible" means in their context, and then optimizing across the entire pipeline to achieve that.

Tom: And VirnyFlow does that by giving users a flexible evaluation protocol — they can define their own metrics, their own weights, their own sensitive attributes, even intersectional groups.

Jane: Right, and it's not just about fairness. It's about stability, uncertainty, and other dimensions of model performance that are often ignored.

Tom: The system also scales — it's not just a toy. With distributed execution, it can handle large datasets and many workers.

Jane: And the comparison against Alpine Meadow and auto-sklearn shows that it's not just a nice idea — it actually outperforms existing systems on both performance and scalability.

Tom: But there are limitations. The paper focuses on tabular data and fixed-structure pipelines. It doesn't cover neural architecture search or unsupervised learning.

Jane: Right, and they acknowledge that. But they also lay out clear future directions — better scoring methods, more advanced visualization interfaces, and further improvements to pruning.

Tom: I think the most impactful thing about this paper is the philosophy shift. It's saying that AutoML shouldn't be a black box that spits out a model. It should be a collaborative tool that helps humans make better decisions.

Jane: And that's a message that resonates beyond just the technical community. It's about accountability, transparency, and putting human judgment at the center of AI development.

Tom: So as we say goodbye to this paper, I want to thank the authors — Denys Herasymuk, Nazar Protsiv, and Julia Stoyanovich — for this contribution.

Jane: And we're looking forward to seeing where this line of research goes. If you're working on responsible AI or AutoML, this is definitely a paper worth reading.

Tom: Alright, that's it for "VirnyFlow: A Design Space for Responsible Model Development." Next up, we've got another exciting paper to dig into. Stay tuned, folks.

Jane: Thanks for listening, everyone. See you on the next one.

Denys Herasymuk, Anastasiia Mozghova, Nazar Protsiv, Vladyslav Sydorak, Julia Stoyanovich

Ukrainian Catholic University · New York University

cs.LG, cs.AI, cs.CY

Submitted: 2026-08-15

Updated: 2026-08-18

Code: https://github.com/denysgerasymuk799/virny-flow

Project page: https://dataresponsibly.github.io/Virny/glossary/disparity_performance_

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 70/100

The gist: VirnyFlow is presented as "the first design space for responsible model development, designed to assist data scientists in building ML pipelines that are tailored to the specific context of their

Key concepts

VirnyFlow
A system described as the 'first design space for responsible model development.' It allows users to define multiple objectives—such as fairness and stability—and jointly optimize an ML pipeline across all components, rather than just maximizing accuracy.
Responsible Model Development
A philosophy that moves beyond treating optimization like a black box. It emphasizes building ML models by considering the real-world context, ensuring the model accounts for potential harms to different groups (like sex or race), and keeping human judgment in the loop.
Multi-objective Bayesian Optimization
A technique used by VirnyFlow that optimizes multiple metrics simultaneously (e.g., F1 score, fairness metric, stability). Instead of one 'best' model, it generates a set of trade-off models represented by a Pareto front.
Multi-armed Bandit Approach
A method used to select which ML pipelines to test. It balances prioritizing pipelines that have performed well historically with exploring new, uncertain combinations of components.

Terminology

Summary

VirnyFlow is presented as the first design space for responsible model development, designed to assist data scientists in building ML pipelines that are tailored to the specific context of their problem. The paper states that "Unlike conventional AutoML frameworks, VirnyFlow enables users to define customized optimization criteria, perform comprehensive experimentation across pipeline stages, and iteratively refine models in alignment with real-world constraints. The system integrates evaluation protocol definition, multi-objective Bayesian optimization, cost-aware multi-armed bandits, query optimization, and distributed parallelism into a unified architecture."

The paper motivates the work by noting that Developing machine learning (ML) models responsibly requires a deep understanding of real-world problems, which are inherently multi-objective, and that Responsible model development extends beyond optimizing for accuracy, requiring an evaluation protocol tailored to the specific context of use and guided by human expertise. The authors argue that constructing a well-suited ML pipeline demands extensive experimentation, iterative refinements, and significant computational resources, and that "Ideally, systems designed to support model developers should follow a human-centric approach, offering a diverse set of pipeline tuning criteria, enabling multi-stage pipeline optimization, and providing an efficient and flexible design space for comprehensive experimentation."

The paper identifies limitations in existing AutoML systems: most AutoML systems define optimization criteria in isolation from the problem context, and Unless fairness is an explicit objective, AutoML may deepen existing disparities. Yet fairness itself cannot be fully automated; deciding what is fair depends on the socio-technical setting and requires human judgment. The paper also notes that Other performance dimensions such as model stability also play a critical role, and that Optimization processes that ignore stability risk producing models that are unreliable in practice, even if they appear accurate in evaluation settings. The authors state that no existing system fully supports a context-sensitive, iterative ML pipeline development guided by human domain expertise.

The paper's main contributions are listed as: System architecture. We present a novel architecture tailored to responsible ML pipeline development. "Evaluation protocol integration. We define and embed a context-sensitive evaluation protocol into the architecture to support multi-stage, multi-objective optimization across the ML lifecycle. This includes tuning for fairness and stability alongside accuracy, with optimization criteria defined over flexible data subsets (e.g., demographic groups and intersections)." "Unified optimization framework. We combine evaluation protocol specification, multi-objective Bayesian optimization, cost-aware bandits, query optimization, and distributed parallelism into a cohesive design space for iterative experimentation." Empirical validation. We show that VirnyFlow outperforms state-of-the-art AutoML systems in both optimization flexibility and computational scalability on five real-world benchmarks.

The paper scopes its work: "we focus on settings characterized by moderately sized tabular datasets and diverse pipeline variants, rather than large-scale datasets or distributed training of complex models. We emphasize pipeline execution optimizations, excluding visual integration, user feedback, or interface design. Additionally, we restrict our consideration to traditional supervised ML pipelines with fixed structures, omitting joint data cleaning and training, neural architecture search, unsupervised learning, and automated data acquisition or preprocessing."

The system design is described as a five-step optimization process: (1) Search space construction, (2) Logical pipeline selection, (3) Physical pipeline selection, (4) Pipeline evaluation, and (5) Iterative refinement. The evaluation protocol is built on top of Virny, a Python library designed for in-depth model performance profiling across multiple dimensions, including accuracy, stability, uncertainty, and fairness, which "is compatible with most tabular ML pipelines and provides ten fairness metrics, including widely used ones like Equalized Odds, as well as newer stability-based and uncertainty-based metrics such as Label Stability Difference."

For logical pipeline selection, the paper states: To achieve this goal, we integrate query optimization concepts from Alpine Meadow into VirnyFlow, re-implementing them from scratch and adapting them for multi-objective optimization. The selection strategy is "formulated as a three-step Multi-Armed Bandit problem: (1 - Select) an arm (i.e., logical pipeline) to run randomly but proportionally to the score. (2 - Store) execution history in the database. (3 - Adjust) scores, repeat from step (1). The scoring model is defined as: s = ∑ni=1 wi · mui + (theta/c) · ∑ni=1 wi · deltai, where n is the number of optimization objectives, wi is the weight assigned to each objective, mui and deltai are the mean and standard deviation of the logical pipeline plan quality across multiple objectives, and c is the cost, or execution time, for a logical pipeline based on past history. The parameter theta acts as a risk factor that determines how much variance is tolerated when selecting a pipeline."

For physical pipeline selection, "Each physical pipeline is generated using multi-objective Bayesian optimization (BO) to tune the pipeline across multiple stages, including data cleaning and the use of fairness-enhancing interventions, and multiple objectives, including predictive accuracy, fairness, and stability. The system uses OpenBox, a framework that offers a standardized set of single- and multi-objective BO optimizers (e.g., EI, EHVI, MESMO), including support for constraints and parallelization. The paper notes that each time the BO-advisor instantiates a physical pipeline, it jointly tunes multiple stages of the logical pipeline to align them with the optimization criterion, leveraging cross-stage interactions to improve overall performance."

For pipeline evaluation and interactivity, the paper states: "To enable incremental computation and early termination of unpromising pipelines, we adopt the Adaptive Pipeline Selection (APS) algorithm from Alpine Meadow, a bandit-based pruning strategy that detects poorly performing pipelines without utilizing the entire training set. We extend APS... to support multiple objectives. The halting criterion is based on the idea that if the partial training error of a pipeline exceeds the best test error observed so far, the pipeline is terminated. The paper states that Interactivity is embedded into both the scoring model and the pipeline pruning logic, ensuring that promising results are presented to users earlier."

For distributed execution, VirnyFlow combines distributed execution, fine-grained parallelism, asynchronous communication, and asynchronous programming. The architecture consists of Task Manager, Workers, Distributed Queue, and an external database (MongoDB). The paper notes that the combination of fine-grained parallelism, where pipelines are executed as independent tasks, and asynchronous communication via Distributed Queue ensures efficient resource utilization. The system also incorporates fault tolerance by storing tasks in the external database first, allowing a user can restart VirnyFlow and resume execution from the last saved state.

The experiments answer three research questions. For RQ1 (functionality), the paper demonstrates its ability to optimize ML pipelines based on multiple objectives across three datasets (diabetes, folk-emp, folk-pubcov), showing that VirnyFlow effectively optimizes ML pipelines under diverse multi-objective criteria, including fairness across binary and intersectional groups and model stability. For RQ2.1 (performance), comparing against Alpine Meadow and auto-sklearn, VirnyFlow consistently outperforms both Alpine Meadow and auto-sklearn in terms of the average score across all three datasets. For RQ2.2 (scalability), "VirnyFlow consistently outperforms both Alpine Meadow and auto-sklearn in runtime, reaching speedups of up to 5.03 on folk-emp-big and 7.07 on heart with 128 workers, substantially surpassing Alpine Meadow at its 32-worker maximum. For RQ3 (sensitivity), the paper shows a clear trade-off: increasing the number of physical pipeline candidates (k) reduces runtime but generally degrades performance, and notes that using multiple training set fractions for pruning does not consistently reduce runtime and may even increase it."

The paper concludes: "This paper introduces VirnyFlow, the first design space for responsible model development, aimed at helping data scientists build ML pipelines that are customized to the specific context of their problems. By integrating a context-sensitive evaluation protocol, we enable multi-stage, multi-objective optimization that goes beyond traditional performance metrics. Our approach unifies diverse techniques, including multi-objective Bayesian optimization, cost-aware multi-armed bandits, query optimization, and distributed parallelism, into an interactive, flexible, and efficient design space for experimentation. Extensive empirical evaluation demonstrates that VirnyFlow achieves superior performance and scalability compared to existing state-of-the-art AutoML solutions, introducing a novel perspective on responsible ML systems and establishing a strong foundation for future research."

Future work includes advanced scoring methods for multi-objective optimization, incorporating decision-maker preferences, and refining pruning mechanisms for multi-objective optimization and further improving resource efficiency.

Improvements for AI systems

Based on the paper, here are specific improvements I can implement in an AI system, along with what the improved system can do:


1. Multi-Objective Optimization with Context-Sensitive Evaluation Protocol

Improvement: Replace single-metric optimization (e.g., accuracy or F1 only) with a weighted multi-objective optimization framework that includes fairness metrics (TPRD, TNRD, FNRD, SRD), stability metrics (Label Stability), and accuracy, all defined over flexible subgroups (including intersectional groups like sex & race).

What the improved AI system can do:

  • Accept a user-defined evaluation protocol specifying which metrics to optimize, their weights, and the demographic groups to evaluate.

  • Simultaneously optimize for accuracy, fairness, and stability, rather than treating them as post-hoc checks.

  • Handle binary and intersectional protected attributes (e.g., sex, race, sex&race) in the optimization loop.

  • Produce a Pareto front of trade-off solutions, allowing the user to select a model that best fits their socio-technical context.

2. Multi-Stage Pipeline Optimization (Beyond Model Tuning)

3. Cost-Aware Multi-Armed Bandit for Logical Pipeline Selection

4. Adaptive Pipeline Selection (APS) with Multi-Objective Pruning

5. Distributed Parallelism with Fault Tolerance and Asynchronous Communication

6. Multi-Objective Bayesian Optimization (MOBO) with User-Defined Weights

7. Integrated Stability Optimization (Label Stability)

8. Interactive Experiment Management with Query Optimization

9. Extensible Architecture for Custom Metrics and Models

10. Sensitivity-Aware Configuration Defaults

Summary of Capabilities of the Improved AI System:

The improved AI system is a context-aware, multi-objective, multi-stage ML pipeline optimizer that:

  • Lets users define what “good” means (accuracy, fairness, stability, or any combination, over any demographic groups).

  • Optimizes the entire pipeline (data cleaning → fairness intervention → model → hyperparameters) jointly.

  • Scales to large datasets and distributed clusters with fault tolerance.

  • Provides transparent trade-offs and interactive feedback, so users can make informed, responsible decisions.

  • Is extensible to new metrics, models, and fairness definitions, making it a research-friendly platform for responsible ML development.

Abstract

Developing machine learning (ML) models requires a deep understanding of real-world problems, which are inherently multi-objective. In this paper, we present VirnyFlow, the first design space for responsible model development, designed to assist data scientists in building ML pipelines that are tailored to the specific context of their problem. Unlike conventional AutoML frameworks, VirnyFlow enables users to define customized optimization criteria, perform comprehensive experimentation across pipeline stages, and iteratively refine models in alignment with real-world constraints. Our system integrates evaluation protocol definition, multi-objective Bayesian optimization, cost-aware multi-armed bandits, query optimization, and distributed parallelism into a unified architecture. We show that VirnyFlow significantly outperforms state-of-the-art AutoML systems in both optimization quality and scalability across five real-world benchmarks, offering a flexible, efficient, and responsible alternative to black-box automation in ML development.

Sources

Related papers