Variable Selection in the Context of AI Fairness

summary

Video file (mp4)

The gist

(1) discussing current efforts to understand fairness beyond mathematical and statistical analysis, (2) bridging theory and practice by discussing fairness auditing and compliance with international

This episode discusses

The paper

Variable Selection in the Context of AI Fairness · Read on arXiv

Ivan Luciano Danesi, Chiara Frigerio, Fabio Maccaferri, Giorgio Alessandro Motta, Pietro Zecca

UniCredit S.p.A. · Università Cattolica del Sacro Cuore · Cetif

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Variable Selection in the Context of AI Fairness".

Jane: The paper was written by Ivan Luciano Danesi, Chiara Frigerio, Fabio Maccaferri, Giorgio Alessandro Motta and Pietro Zecca from UniCredit S.p.A. and Università Cattolica del Sacro Cuore and Cetif.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Alright, welcome back to the show, everyone. We've got a fresh one from the arXiv this week, and it's called "Variable Selection in the Context of AI Fairness." Jane, I have to say, the title alone got me hooked, because it's pointing at something that feels really counterintuitive.

Jane: Oh, absolutely, Tom. And I think that's exactly why we need to talk about it. The paper is by a team from UniCredit and Università Cattolica in Milan, and they're basically asking a question that flips a lot of common wisdom on its head. We usually think that if we want a fair AI, we should just remove the sensitive stuff, like race or gender, from the data.

Tom: Right, just don't look at it, and then it can't be biased, right? That's the old "fairness through unawareness" idea.

Jane: Exactly. But this paper argues that this approach, while it sounds clean, can actually make things worse. They're saying that by throwing away those variables, you lose information that the model needs to make accurate predictions, and that loss of accuracy can actually hurt the very groups you were trying to protect.

Tom: So it's like trying to fix a leaky pipe by just not looking at the water meter? You're not fixing the problem, you're just hiding it.

Jane: That's a great way to put it. They even walk through some real-world examples, like in credit scoring. You can remove race from the model, but the postal code or the type of spending habits can act as a proxy for it anyway. So the bias just sneaks back in through the back door.

Tom: And that's the core tension, right? We want fairness, but we also want accuracy. The paper suggests these aren't necessarily in conflict, but the way we've been choosing variables is forcing us to pick one over the other.

Jane: Precisely. And they're not just philosophizing about it. They actually build a mathematical framework to show that keeping all the variables, including the sensitive ones, gives you a better starting point. You can then audit the model for fairness after the fact, rather than just hoping it's fair because you hid the data.

Tom: So it's a more honest approach. You keep everything on the table, you see what the model does, and then you correct it. That sounds like a much more mature way to handle this than just pretending the problem doesn't exist.

Jane: It really is. And it aligns with what regulators in Europe are starting to demand. The EU AI Act doesn't just say "don't be biased," it says you need to actively manage and document how you're addressing bias. You can't just say you removed the variable and call it a day.

Tom: Okay, so we've got the big idea. But I'm curious about the actual math they use to back this up. That's what we're digging into next.

Summary: Tom: So, Jane, we've established that this paper, "Variable Selection in the Context of AI Fairness," is arguing against the "just remove the sensitive data" approach. But what's the actual meat of their argument? What are they showing us mathematically?

Jane: Well, they set up a scenario where you have a full dataset, let's call it the complete set of variables. Then you have a subset, which is what you get when you start deleting things, like the sensitive attributes. They show that if you train a model on the full set, you get the best possible error rate, the lowest prediction error.

Tom: Because you have all the information, so the model can make the most informed guesses.

Jane: Exactly. Now, if you train on the subset, your error rate goes up. It's suboptimal, mathematically speaking. But here's the kicker: they then look at fairness. They define a function that measures fairness under different definitions, like demographic parity or equal opportunity.

Tom: And what do they find when they compare the fairness of the full model versus the reduced model?

Jane: They lay out three possible outcomes. The reduced model could be less fair, equally fair, or more fair than the full model. But they argue that in two out of those three cases, the full model is either better or just as good. And even in the third case, where the reduced model is somehow fairer, they show you can use the information from the full model to fix that specific unfairness.

Tom: So they're saying that starting with everything is almost always the right move, and if it's not, you can still fix it.

Jane: Right. And they even have a clever trick for that third case. If the reduced model is fairer for a specific group, you can create a hybrid model. You use the full model for most people, but for that specific group, you use the reduced model's output. It's a surgical fix.

Tom: That's really smart. It's like having a general practitioner and a specialist. You use the general one for most things, but you call in the specialist when you need that specific expertise.

Jane: That's a perfect analogy. And the key point is that you can only do that if you still have access to the sensitive variables. If you've already thrown them away, you can't identify who needs the specialist treatment. You're flying blind.

Tom: So the paper is essentially saying that variable exclusion is a blunt instrument that takes away your ability to do nuanced fairness work later.

Jane: Exactly. And it also connects back to explainability. If you don't have the variables, you can't explain why a decision was made for a specific person. You lose the ability to audit and understand the model.

Tom: So it's not just about accuracy, it's about the whole toolkit for responsible AI. I'm starting to see why this is such a big deal. But I'm wondering, what does this mean for the people actually building these systems? What's the practical advice here?

Improvements: Tom: So we've got the theory, and it sounds great on paper. But Jane, what does this actually mean for a team of engineers or data scientists working on a real product? How does this paper change what they should do on Monday morning?

Jane: I think the biggest takeaway is a shift in workflow. Instead of starting with a list of variables to exclude, you start with the full dataset. You train your model on everything, you optimize for accuracy, and then you run a rigorous fairness audit on the output.

Tom: So the variable selection happens after you see the results, not before.

Jane: Exactly. And this is where I think we need to bring in some other voices. Lu, you've been working on fairness in machine learning for a while now. Does this match what you see in practice?

Lu: It does, and I think it's a much more robust approach. The problem with preemptive exclusion is that you're making a decision based on assumptions about correlation that you haven't actually tested in your specific model. By keeping everything in, you let the data speak for itself. You can then use statistical methods, like the error optimization they mention, to see which variables actually matter.

Meng: But from an engineering standpoint, that sounds expensive. Training on a full dataset with all those variables takes more compute, more time, and more memory. Is that a realistic ask for every team?

Jane: That's a fair point, Meng. And the paper does acknowledge that you need to be smart about this. They're not saying you should keep every single variable forever. They're saying you should start with everything to find the optimal model, and then use objective criteria to trim it down.

Lu: And the key is that the trimming should be based on statistical significance and error contribution, not on a gut feeling that a variable is "too sensitive." If a variable doesn't help the model, it will drop out naturally. If it does help, then removing it is going to hurt accuracy, and you need to know that.

Meng: So it's like a two-stage process. First, you find the best model with everything. Then, you look at the fairness metrics and decide if you need to make any trade-offs. That's a lot more deliberate than just scrubbing the data upfront.

Jane: Exactly. And the paper even provides a mathematical guarantee that this approach, where you start full and then adjust, will keep your error rate bounded. You won't end up with a wildly inaccurate model just to gain a tiny bit of fairness.

Lalam: I would add that this approach also improves the cultural and ethical posture of an organization. By keeping the sensitive variables in the loop, you are signaling a commitment to understanding the impact of your model, rather than hiding from it. It fosters a culture of accountability, which is exactly what the EU AI Act is pushing for.

Tom: So it's not just a technical fix, it's a cultural shift. You're building a system that is transparent about its biases and actively working to correct them.

Jane: And that's the real improvement the paper suggests. It moves us from a place of ignorance and hope to a place of knowledge and control.

Conclusion: Tom: Well, we've covered a lot of ground on "Variable Selection in the Context of AI Fairness." Jane, can you help us wrap this up and say goodbye to this paper?

Jane: I'd love to. The core message is that trying to achieve fairness by deleting sensitive variables is a flawed strategy. It hurts accuracy, it doesn't actually remove bias because proxies sneak in, and it destroys your ability to explain and audit your model.

Tom: And the alternative they propose is to start with everything, optimize for accuracy, and then use a fairness audit to make targeted corrections.

Jane: Exactly. It's a more honest and more effective path. It acknowledges that fairness is a complex, context-dependent issue that can't be solved with a simple deletion. It requires careful analysis and, as Lu and Lalam pointed out, a cultural commitment to accountability.

Tom: And it lines up perfectly with the regulatory direction in Europe, where they're demanding that companies actively manage bias rather than just claim it doesn't exist.

Jane: Right. This paper gives us a mathematical foundation for doing that. It's a bridge between the abstract ethical goal of fairness and the concrete reality of building a model that works in the real world.

Tom: So we're saying goodbye to this paper with a sense of hope. It gives us a practical, rigorous way to build AI that we can trust.

Jane: Absolutely. And it sets the stage for a much more interesting conversation about how we define and measure fairness in the first place. But that's a topic for another paper.

Tom: Indeed it is. Thanks for joining us, everyone. We'll be back next week with another deep dive into the latest research.

Jane: Take care, and keep questioning the data.

More episodes

← Home