Designing for Ethical AI: HCI Feature Considerations to Improve Fairness and User Experience in AutoML use for Human Resources
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Designing for Ethical AI: HCI Feature Considerations to Improve Fairness and User Experience in AutoML use for Human Resources".
Jane: The paper was written by Sundaraparipurnan Narayanan from University of Bordeaux and IAE.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the show, everybody. Today we are digging into a paper that's been making the rounds, and it's called "Designing for Ethical AI: HCI Feature Considerations to Improve Fairness and User Experience in AutoML use for Human Resources." Jane, that title is a mouthful, but it's about something that affects almost everyone.
Jane: It really does, Tom. And I love that we're talking about this because it's about the tools companies use to hire people. We're not just talking about a robot reading resumes. We're talking about the software that decides who even gets a first interview, and whether that software is fair.
Tom: Exactly. And the paper is essentially a doctoral thesis from the University of Bordeaux. The author, Sundaraparipurnan Narayanan, is looking at Automated Machine Learning, or AutoML. These are tools that let non-experts build AI models without being a data scientist.
Jane: So, imagine an HR manager who wants to predict which employees might quit, or wants to screen a pile of resumes. They can use this AutoML tool. But the big question the paper asks is, do these tools help that HR manager be fair, or do they accidentally bake in the same biases we've had for decades?
Tom: And that's the crucial part. The paper isn't just saying "AI is biased." It's saying the design of the tool itself, the buttons, the dashboards, the warnings it does or doesn't show, is what makes the difference. It's about the user experience of fairness.
Jane: Right. If the tool shows you a model's accuracy but hides the fact that it's rejecting qualified women at a higher rate, the HR manager might think they're doing a great job. The paper calls this a "black box" problem, and it's a huge deal for anyone who's ever applied for a job.
Tom: So we're going to break down how these tools are failing, what the researchers found when they tested eight of them, and what they recommend. This is a big one, Jane. It's about whether the future of work is going to be more equitable or just more efficiently biased.
Jane: And that's the hook. Let's get into the summary of the paper next, because the findings are pretty stark.
Summary: Tom: So, Jane, we've set the stage. Now let's talk about what this paper actually did. It's not just a think piece. They ran a full evaluation on eight different AutoML tools, both the kind you click around in and the kind you code with.
Jane: And what they found is that most of these tools are failing on fairness. The paper says it plainly: fairness features are often missing, or they're so buried in the interface that a normal person would never find them. Out of all the features they looked for, only about twenty-nine percent actually existed in the tools.
Tom: That's a wild number. So, the majority of the time, the tools aren't even giving users the option to check for bias. And when they do, the paper shows it's often just a metric on a screen with no explanation of what to do about it.
Jane: Exactly. They tested these tools on real HR datasets, like the IBM employee attrition dataset and a recruitment dataset from Utrecht. And the results were all over the place. One tool, PyCaret, heavily favored male candidates in one test. Another, DataRobot, completely excluded candidates who didn't fit into a binary gender category.
Tom: And that's the scary part. These aren't abstract numbers. That means a qualified person was rejected because of a flaw in the software, and the person using the tool might have had no idea. The paper calls this "automation bias," where users just trust the machine's output.
Jane: Right. The study also found that the tools with a graphical interface, like Dataiku and DataRobot, were much better at supporting fairness than the code-based libraries like AutoGluon or FLAML. But even the best ones had huge gaps.
Tom: So the summary is that the technology is promising, but the design is letting everyone down. It's putting the burden on the user to be a fairness expert, which most HR managers aren't.
Jane: And that's the core problem the paper is trying to solve. It's not about making the algorithms smarter; it's about making the tools more responsible. We'll talk about their specific suggestions next, because they have a whole framework for it.
Improvements: Tom: So we know the problem. What's the fix? The paper doesn't just complain; it lays out a whole framework for how to design these tools better. Jane, what's the big idea?
Jane: The big idea is that fairness has to be built into the design, not bolted on as an afterthought. They call it a "Human-Computer Interaction" framework, and it has five parts. Think of it like building a car with safety features as standard, not as an optional extra you have to pay for.
Tom: So what are those five parts? Give us the quick tour.
Jane: First, there's "Contracts and User Development." That means the tool needs to clearly tell you what it can and can't do, and it should train you on how to spot bias. Second, the "User Interface" itself needs to make fairness metrics visible and easy to understand, like a dashboard with a green light for fair and a red light for biased.
Tom: So, not just a number like "zero point two" that means nothing to a normal person, but a clear signal.
Jane: Exactly. Third is "Information Architecture," which is about organizing the information so you can actually find the fairness settings. Fourth is "Human Augmentation," which is about letting the user step in and override the machine when something looks wrong.
Tom: And the fifth?
Jane: The fifth is "Care and Responsibility." That's about accountability. The tool should have audit trails so you can see what decisions were made and why, and it should have a system for reporting incidents when bias is found.
Tom: And the paper even suggests things like "nudges." So, if the tool detects that a model is treating a group unfairly, it could pop up a warning saying, "Hey, this model is rejecting eighty percent of female applicants. Do you want to fix this?" That's a practical feature.
Jane: Right. And the research shows that these features aren't just about being ethical. They're about building trust. If a company is going to buy this software, they need to know it won't get them sued. So, fairness becomes a selling point.
Tom: So it's a business case, not just a moral one. Let's dig into the first page of the paper next, because it sets up this whole problem in a really compelling way.
First Page: Tom: So we've talked about the findings and the recommendations. Let's go back to the very beginning of "Designing for Ethical AI" and look at the first page, because it frames the whole problem so well. Jane, what stood out to you?
Jane: The first page talks about Industry four point zero and how AI is being used to make hiring "smarter." It mentions these tools that can parse resumes and rank candidates, and it sounds great on the surface. But then it immediately warns that these models are trained on data from past human decisions, which were often biased.
Tom: So the machine learns from our mistakes. It learns that men were hired more in the past, so it continues to hire more men. The paper calls this a "disruptive innovation," but it's disruptive in a bad way if it just automates discrimination.
Jane: Exactly. And the page makes a really important point about risk. It says that using AI carries risks that could affect individuals, groups, and entire organizations. It's not just about a bad hire; it's about perpetuating discrimination and harming people's lives.
Tom: And it mentions the business side. It says companies could face legal repercussions and liabilities for using AI that causes harm. We've seen this happen with Amazon's recruiting tool that was scrapped for being sexist, and with the HireVue facial analysis controversy.
Jane: Right. And that's the key takeaway from that first page. It sets up the central tension of the entire thesis: AutoML promises to democratize AI and make it accessible, but if we don't design it carefully, it will just democratize bias.
Tom: It's a powerful framing. It's not saying "don't use AI." It's saying "if you're going to use it, you have a responsibility to make it fair, and the tools need to help you do that."
Jane: And that responsibility falls on the companies making the tools. The paper argues that fairness is no longer optional. It's a core feature for product-market fit, especially in a regulated area like HR. It's about building trust so that people will actually want to use the product.
Tom: So, from the very first page, it's clear this is a serious, well-researched call to action. Let's wrap this up in our conclusion and think about what this means for the future.
Conclusion: Tom: Alright, Jane, we've covered a lot of ground on "Designing for Ethical AI: HCI Feature Considerations to Improve Fairness and User Experience in AutoML use for Human Resources." Let's bring it home. What's the one thing listeners should remember?
Jane: The one thing is that fairness in AI isn't a technical problem; it's a design problem. The paper shows that the tools we have are powerful, but they're not built to help people be fair. They're built to be fast and accurate, and fairness gets left behind.
Tom: And that has real-world consequences. We saw the data. Some tools under-hire qualified women. Others completely ignore non-binary candidates. This isn't hypothetical. It's happening right now in hiring processes around the world.
Jane: But the good news is that the paper offers a roadmap. By focusing on Human-Computer Interaction, by making fairness visible, by letting users intervene, and by building in accountability, we can create tools that actually help us build a more equitable workforce.
Tom: So, it's a hopeful message, but a demanding one. It demands that software developers take responsibility, and it demands that companies using these tools ask the right questions.
Jane: And it demands that we, as a society, pay attention. Because the algorithms making decisions about our lives are only going to become more common. We need to make sure they're making those decisions fairly.
Tom: Well said, Jane. That's a wrap on this paper. It's been a fantastic discussion, and I think we've only scratched the surface. Thanks to everyone for tuning in, and we'll be back soon with the next paper to dissect. Until then, keep questioning the black boxes.
Jane: Bye, everyone!
Sundaraparipurnan Narayanan
University of Bordeaux · IAE
cs.HC, cs.AI
Submitted: 2026-05-28
Updated: 2026-08-11
Comments: 295 pages; Doctoral Thesis;28 figures; 42 tables; 248 references
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 56/100
The gist: Summary This thesis, titled "Designing for Ethical AI: HCI Feature Considerations to Improve Fairness and User Experience in AutoML use for Human Resources," investigates the problem of fairness in
Key concepts
- AutoML
- Automated Machine Learning refers to tools that allow non-experts to build AI models without needing to be data scientists. The paper examines how the design of these tools affects fairness when used in Human Resources applications.
- Black Box Problem
- This occurs when an AI tool shows a result, like a model's accuracy, but hides critical information about its internal workings or biases. This makes it difficult for users to understand why certain decisions are being made, such as rejecting qualified candidates.
- Human-Computer Interaction (HCI) Framework
- This is the five-part design framework proposed to make AI tools fair. It focuses on building fairness into the design from the start, including making metrics visible on dashboards, allowing users to override biased decisions, and creating audit trails for accountability.
Terminology
Summary
Summary
This thesis, titled Designing for Ethical AI: HCI Feature Considerations to Improve Fairness and User Experience in AutoML use for Human Resources,
investigates the problem of fairness in Automated Machine Learning (AutoML) tools when used for hiring and other human resources (HR) functions. The research integrates perspectives from regulations, business, and human-computer interaction (HCI) to argue that fairness is not only an ethical concern but a key product feature that enhances usability, builds trust, and facilitates enterprise adoption.
The study is motivated by the potential impact of algorithmic bias on people, which can arise from models developed using AutoML by non-expert users. It addresses a gap in existing research, which has largely focused on the technical performance of AutoML tools rather than their ethical and social implications, particularly regarding regulatory compliance and fairness. The research problem is framed around the risk that AutoML tools, which are designed to simplify model building, can produce biased or discriminatory models, exposing both the companies that create them and those that use them to legal, reputational, and financial risks.
The central research questions are:
-
(RQ1) What fairness problems and solutions are there in popular AutoML tools for hiring data?
-
(RQ2) How do interface design, transparency, and feedback help or hurt fairness?
-
(RQ3) What are the missing pieces in human involvement (HITL), fairness visuals, and reporting?
-
(RQ4) What should AutoML providers focus on in product design to make fairness features easy to use and widely adopted by businesses?
The research is grounded in several foundational theories, including the Technology Acceptance Model (TAM), Innovation Diffusion Theory (IDT), Human-Centered AI (HCAI), Cognitive Load Theory (CLT), and Affordance Theory. TAM is used to understand how fairness features influence adoption through perceived usefulness and perceived ease of use. IDT explains how innovation features like transparency and auditability influence the diffusion of new tools. HCAI, CLT, and Affordance Theory are used to evaluate AutoML interfaces, emphasizing the need for transparent, controllable, and reliable tools that reduce cognitive load and make fairness features perceptible and actionable for non-experts.
The study employs a mixed-methods approach. The qualitative component includes structured HCI audits, heuristic evaluations, and cognitive walkthroughs that simulate how non-expert business users attempt to find bias, apply solutions, and create audit reports. This is supplemented by analyzing user contracts and documents for fairness disclosures. The quantitative component tests eight well-known AutoML tools using curated HR datasets with sensitive information. The resulting models are assessed using legal fairness metrics such as Demographic Parity Difference and Equalized Odds to determine how well AutoML outputs align with fairness expectations in real-world hiring scenarios.
The evaluation framework is structured around five dimensions: (1) Contracts and User Development, focusing on clear disclosures and user training; (2) User Interface and Experience Design, emphasizing intuitive interfaces and visual explanations; (3) Information Architecture, ensuring information is logically organized; (4) Human Augmentation, supporting collaboration through explainability and feedback loops; and (5) Care and Responsibility, addressing ethical and safety considerations through governance tools and transparency measures.
The selection of tools and datasets was systematic. For code-based libraries, only those with a strong GitHub presence (more than 3,500 stars and at least 50 contributors) were included, leading to the selection of AutoGluon, FLAML, PyCaret, and H2O AutoML (Python). For GUI-based tools, the focus was on ease of use, business adoption, and availability of free or research versions, resulting in the selection of Dataiku, DataRobot, H2O AutoML Studio, and Alteryx RapidMiner. Six publicly available, tabular HR datasets with demographic and hiring-related information were selected for benchmarking.
The evaluation findings reveal consistent problems with fairness transparency, user control, and bias mitigation across the reviewed tools. Key issues include a heavy reliance on users to detect and address bias without sufficient system support, interfaces that reinforce automation bias, limited user control and lack of human-in-the-loop features, and a lack of visual tools for comparing model performance across demographic groups. The quantitative analysis showed that none of the tools produced models that were fair across all datasets, and each model had partial or heavily biased results depending on the metric. For instance, in the Utrecht Recruitment dataset, DataRobot completely excluded “Other” gender candidates, and PyCaret heavily favored males. In the Employee Promotion dataset, RapidMiner demonstrated high bias by under-promoting males.
The thesis introduces a new framework for evaluating HCI focused on fairness, which is used to critique current tools and demonstrate how to build fairness into product development. The practical recommendations emphasize making fairness a core part of AutoML tools by offering clear, interactive ways to see bias, improving feedback to support human oversight, and building accountability into both the interface and documentation. These recommendations are organized across the five HCI dimensions and include clearer user contracts, more intuitive interface designs with fairness visualizations, better information architecture, stronger human augmentation features like feedback loops, and more robust accountability mechanisms.
The research makes both theoretical and practical contributions. Theoretically, it reframes fairness in AutoML as a strategic product feature that supports usability, trust, and market adoption, shifting the focus from model accuracy to user empowerment and responsible AI. Practically, it provides an actionable evaluation framework, benchmarks eight major AutoML tools, and offers design recommendations that reduce cognitive load, mitigate automation bias, and improve fairness transparency. The novelty of the work lies in addressing the ethical and social dimensions of AutoML, an area often overlooked in favor of technical performance. The thesis concludes by outlining future research paths, including expanding bias mitigation to other sectors, exploring fairness across the AI lifecycle, and improving transparency in fairness-performance trade-offs through interactive and user-friendly design features.
Improvements for AI systems
Based on the thesis, here are the specific improvements that can be made to AI systems, particularly Automated Machine Learning (AutoML) tools used in Human Resources (HR) and hiring:
1. Proactive Fairness Diagnostics and Guardrails
-
Improvement: Integrate automated bias detection and mitigation directly into the model-building pipeline, rather than relying on users to identify issues post-hoc.
-
What the improved system can do: Automatically flag proxy variables (e.g., zip code correlating with race) and multicollinearity issues during data preprocessing. It will trigger alerts when fairness metrics (e.g., Demographic Parity Difference, Equalized Odds) exceed legal thresholds (e.g., the EEOC's 80% rule) and block or warn against model deployment if a protected group is disproportionately disadvantaged (e.g., TPR = 0 for a specific gender).
2. Fairness-Aware Visualization and Trade-off Analysis
-
Improvement: Implement interactive dashboards that visualize fairness metrics (TPR, FPR, DPD, PRP) across demographic groups, alongside performance metrics.
-
What the improved system can do: Provide users with a
bias comparison grid
and sliders to adjust fairness thresholds in real-time, showing the immediate trade-off between accuracy and fairness. This allows non-expert users to visually understand the impact of their decisions and select models that meet both performance and ethical requirements.
3. Human-in-the-Loop (HITL) with Meaningful Override
-
Improvement: Design interfaces that support human oversight and intervention, not just passive review. This includes explicit override capabilities for fairness-related decisions.
-
What the improved system can do: Allow users to set custom fairness constraints, apply specific bias mitigation techniques (e.g., reweighting), and manually override automated model choices that might lead to discriminatory outcomes. The system will log these interventions for auditability, ensuring traceability and accountability.
4. Enhanced Transparency and Explainability for Non-Experts
-
Improvement: Move beyond simple feature importance to provide contextual, plain-language explanations of why bias exists and how it was mitigated.
-
What the improved system can do: Generate
fairness reports
that explain the causes of bias (e.g.,The model relies heavily on 'Years of Experience,' which correlates with age, leading to a higher false positive rate for younger candidates
). It will also surface group-wise SHAP values to show how specific features impact different demographic groups, making the system's logic understandable to HR professionals without data science expertise.
5. Structured Feedback Loops and Iterative Refinement
-
Improvement: Create formalized mechanisms for users to provide feedback on fairness outcomes and have that feedback directly influence model retraining.
-
What the improved system can do: Provide a feedback form within the interface where users can flag a model's output as biased. The system will then use this feedback to adjust fairness parameters, re-run the AutoML process, and visually show the user how their input changed the model's fairness metrics and performance.
6. Contractual and User Development Clarity
-
Improvement: Embed clear disclosures about the system's limitations and acceptable use cases directly into the user interface and documentation.
-
What the improved system can do: Present a
Data Card
andModel Card
that explicitly state the tool's inability to guarantee fairness, the specific types of bias it can detect, and the contexts where it should not be used (e.g., making final hiring decisions without human review). This sets appropriate expectations and reduces the risk of misuse by non-expert users.
Abstract
This thesis examines the fairness of Automated Machine Learning (AutoML) tools in human resource hiring systems through the combined lenses of regulation, business strategy, and Human-Computer Interaction (HCI). It argues that fairness is no longer merely an ethical concern but a critical determinant of usability, trust, legal compliance, and organizational adoption. While AutoML platforms improve efficiency by simplifying model selection and deployment, they also risk perpetuating discriminatory outcomes when trained on biased historical hiring data. Existing platforms prioritize technical performance over fairness, leaving non-expert business users unable to detect or mitigate bias effectively. The study investigates fairness gaps in AutoML tools through four research questions focused on fairness mechanisms, interface transparency, human oversight, and product design priorities. Drawing on frameworks such as the Technology Acceptance Model, Innovation Diffusion Theory, Human-Centered AI, Cognitive Load Theory, and Affordance Theory, the research evaluates both usability and fairness alignment. Using qualitative HCI audits and quantitative testing of eight AutoML platforms on HR datasets, the findings reveal widespread deficiencies in transparency, user control, and bias mitigation support. The thesis proposes a five-dimensional HCI-based fairness evaluation framework and recommends embedding fairness directly into AutoML product design to improve accountability, adoption, and ethical sustainability in AI-driven hiring systems.
Sources
- A Multivocal Literature Review on the Benefits and Limitations of Automated Machine Learning Tools
- To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making
- From Soft Classifiers to Hard Decisions: How fair can we be?
- Techniques for Automated Machine Learning
- A Comprehensive Empirical Study of Bias Mitigation Methods for Machine Learning Classifiers
- Fair prediction with disparate impact: A study of bias in recidivism prediction instruments
- A Bandit-Based Algorithm for Fairness-Aware Hyperparameter Optimization
- FairFed: Enabling Group Fairness in Federated Learning
- Towards Fair and Explainable AI using a Human-Centered AI Approach
- An Open Source AutoML Benchmark
- UniAutoML: A Human-Centered Framework for Unified Discriminative and Generative AutoML with Large Language Models
- Can AutoML outperform humans? An evaluation on popular OpenML datasets using AutoML Benchmark
- Retiring $\Delta$DP: New Distribution-Level Metrics for Demographic Parity
- Equality of Opportunity in Supervised Learning
- Evaluating Fairness Metrics in the Presence of Dataset Bias
- Towards Evaluating Exploratory Model Building Process with AutoML Systems
- Visualization for Human-Centered AI Tools
- Automated Machine Learning, Bounded Rationality, and Rational Metareasoning
- The Roles and Modes of Human Interactions with Automated Machine Learning Systems
- What Can AutoML Do For Continual Learning?
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support