Automating the Detection of Requirement Dependencies Using Large Language Models

summary

Video file (mp4)

The gist

The paper introduces LEREDD (LLM-Enabled REquirement Dependency Detection), an automated approach for detecting various types of direct dependencies between Natural Language (NL) requirement pairs.

In short

The episode discusses a paper automating requirement dependency detection using Large Language Models (LLMs). The system, LEREDD, uses Retrieval-Augmented Generation and In-Context Learning to classify dependencies between software requirements with high accuracy. The authors demonstrate its effectiveness across automotive systems and show it outperforms traditional methods.

Key concepts

Requirement Dependency
This refers to a relationship between two requirements in a system, such as one requirement depending on another. For example, a roof requirement depends on the wall requirement being present first.
LEREDD
This is the system introduced in the paper, standing for LLM-Enabled Requirement Dependency Detection. Its basic function is to take a list of requirements and tell you which pairs are dependent and classify the type of dependency between them.
Retrieval-Augmented Generation (RAG)
RAG provides context to an LLM by feeding it relevant sections from a software specification before it makes a judgment. This allows the model to understand specific domain details, such as identifying what a subsystem like 'BCS' refers to in the context of braking control.
In-Context Learning (ICL)
ICL involves showing an LLM examples within the prompt for each dependency type. This teaches the model by example, allowing it to better distinguish between different dependency classifications like 'Requires' or 'Conflicts' based on provided instances.

Terminology used across episodes

This episode discusses

The paper

Automating the Detection of Requirement Dependencies Using Large Language Models · Read on arXiv

Ikram Darif, Feifei Niu, Manel Abdellatif, Lionel C. Briand, Ramesh S., Arun Adiththan

University of Ottawa · École de technologie supérieure · University of Limerick · General Motors

Requirements are inherently interconnected through various types of dependencies. Identifying these dependencies is essential, as they underpin critical decisions and influence a range of activities throughout software development. However, this task is challenging, particularly in modern software systems, given the high volume of complex, coupled requirements. These challenges are further exacerbated by the ambiguity of Natural Language (NL) requirements and their constant change. Consequently, requirement dependency detection is often overlooked or performed manually. Large Language Models (LLMs) exhibit strong capabilities in NL processing, presenting a promising avenue for requirement-related tasks. While they have shown to enhance various requirements engineering tasks, their effectiveness in identifying requirement dependencies remains unexplored. In this paper, we introduce LEREDD, an LLM-based approach for automated detection of requirement dependencies that leverages Retrieval-Augmented Generation (RAG) and In-Context Learning (ICL). It is designed to identify diverse dependency types directly from NL requirements. We empirically evaluate LEREDD against two state-of-the-art baselines. The results show that LEREDD provides highly accurate classification of dependent and non-dependent requirements, achieving an accuracy of 0.93, and an F1 score of 0.84, with the latter averaging 0.96 for non-dependent cases. LEREDD outperforms zero-shot LLMs and baselines, particularly in detecting fine-grained dependency types, where it yields average relative gains of 94.87% and 105.41% in F1 scores for the Requires dependency over the baselines. We also provide an annotated dataset of requirement dependencies encompassing 813 requirement pairs across three distinct systems to support reproducibility and future research.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Automating the Detection of Requirement Dependencies Using Large Language Models".

Jane: The paper was written by Ikram Darif, Feifei Niu, Manel Abdellatif, Lionel C. Briand, Ramesh S. et al. from University of Ottawa and École de technologie supérieure and University of Limerick and General Motors.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title and Authors: Tom: Welcome back to the show, everybody. Today we're digging into a paper that's got a title that sounds like a mouthful but is actually about something really practical: "Automating the Detection of Requirement Dependencies Using Large Language Models."

Jane: And Tom, I gotta say, when I first read that title, I thought, okay, this is about software requirements, which sounds dry, but the problem they're tackling is one of those hidden time-sinks in engineering that nobody talks about.

Tom: Right, and the authors are a mix of academic and industry folks — Ikram Darif, Feifei Niu, Manel Abdellatif, Lionel Briand, and then Ramesh S and Arun Adiththan from General Motors. That's a big deal, right?

Jane: Huge. When you see GM on the author list, you know this isn't just theoretical. They're dealing with real automotive systems where requirements are complex and interconnected.

Tom: So Jane, for our listeners who aren't software engineers — what's a requirement dependency? Why should I care?

Jane: Think of it like building a house. You have a requirement that says "the house needs a roof" and another that says "the house needs walls." The roof depends on the walls being there first. If you change the walls, you've got to know the roof is affected. In software, especially in cars, these dependencies are everywhere.

Tom: And the problem is, in a modern car, you might have hundreds of requirements. Manually checking every pair to see if they depend on each other? That's thousands of pairs. It's exhausting and error-prone.

Jane: Exactly. And the paper points out that when people skip this step, you get project delays, rework, even safety issues. In a car, a missed dependency could mean a safety feature doesn't work because something else changed.

Tom: So they're using large language models — the same kind of tech behind chatbots — to automate this detection. And the authors are claiming some pretty impressive results. We'll get into the numbers soon.

Jane: The exciting part is the team structure. You've got academic rigor from the university folks and practical urgency from GM. That combination usually means the research is actually going to be useful, not just published and forgotten.

Tom: And that's what we're going to explore today — how they built this system, what it achieves, and whether it can actually help engineers in the real world. Stick around.

Jane: We're just getting started with "Automating the Detection of Requirement Dependencies Using Large Language Models," and there's a lot to unpack.

Summary: Tom: So we've got the title and the team. Now let's talk about what this paper actually does. Jane, can you break down the core idea for us?

Jane: Sure. The paper introduces a system called LEREDD — that's LLM-Enabled REquirement Dependency Detection. The basic idea is: give it a list of requirements from a software spec, and it tells you which pairs are dependent on each other and what kind of dependency it is.

Tom: And they're not just saying "these two are related." They're classifying into specific types — like "Requires," "Implements," "Conflicts," "Contradicts," "Details," "Is similar," and "Is a variant."

Jane: Right. So "Requires" means one thing can't happen without the other. "Conflicts" means fulfilling one restricts the other. These distinctions matter because they affect how engineers plan and manage changes.

Tom: And the clever part is how they get the LLM to do this well. They use two tricks: Retrieval-Augmented Generation, or RAG, and In-Context Learning, or ICL. Can you explain those in plain terms?

Jane: Think of RAG as giving the model a cheat sheet. Before asking it to judge a pair of requirements, you feed it relevant sections from the software spec — the domain context. So if the requirements mention a "BCS" subsystem, the model knows from the context that BCS is the braking control system.

Tom: And ICL?

Jane: That's showing the model examples. For each dependency type, they retrieve a few similar requirement pairs from a labeled dataset and show the model "here's what a 'Requires' dependency looks like, here's what 'No dependency' looks like." It's like teaching by example.

Tom: So the model gets both the context and examples before making its call. And they tested this across three automotive systems — Traffic Jam Assist, Automated Parking Assist, and Adaptive Driving Beam. They annotated over eight hundred requirement pairs manually.

Jane: That's a lot of careful human work. Two annotators independently labeled these pairs, then resolved disagreements. That gives us a solid ground truth to measure against.

Tom: And the results? They're claiming an average accuracy of about ninety-three percent and an F1 score of eighty-four percent. But the really impressive number is for detecting "No dependency" — that's ninety-six percent F1.

Jane: That's huge because in real systems, most requirement pairs are NOT dependent. If you can reliably filter those out, you save engineers an enormous amount of time. They only need to focus on the small fraction that actually have dependencies.

Tom: And they compared against two baseline methods — a traditional TF-IDF retrieval approach and a fine-tuned BERT model. LEREDD beat both, especially on the fine-grained dependency types.

Jane: The gains on "Requires" were dramatic — like ninety-five percent to one hundred five percent relative improvement in F1 score over the baselines. That's not a small bump; that's a game-changer.

Tom: So the summary is: they built a system that uses LLMs with context and examples to accurately detect requirement dependencies, and it outperforms existing methods. We're going to dig into the methodology and what makes it work next.

Improvements: Tom: We've covered the basics. Now let's talk about what makes LEREDD better than just asking a chatbot "are these two requirements related?" Jane, what's the real innovation here?

Jane: The key improvement is that they didn't just use an LLM in a naive way. They systematically tested different prompting strategies. And they found that zero-shot — just asking the model cold — performs poorly on fine-grained dependencies. The model tends to over-predict dependencies and struggles with rare types.

Tom: So they added the two-layer knowledge retrieval. And they didn't just pick arbitrary settings — they ran two hundred sixteen experiments to tune the few-shot parameters. That's thorough.

Jane: Very thorough. They tested different embedding models — SBERT versus BGE-M3. They tested different similarity metrics — cosine versus Euclidean. They tested different ways to aggregate similarity between requirement pairs. And they tested different numbers of examples, from one to nine.

Tom: And the winner was SBERT with Euclidean distance, using a max-similarity aggregation, and four examples per dependency type. Why does that combination work?

Jane: The max aggregation is interesting. Instead of averaging similarity across all four requirement combinations, they take the best match for each requirement in the target pair. That makes sense because a single highly relevant example is more informative than several mediocre ones.

Tom: And for RAG, they found that ten chunks of five hundred characters each was optimal. Not the whole document — that adds noise. Not too few — that lacks context. There's a sweet spot.

Jane: Exactly. And the improvements were consistent. Adding few-shot examples boosted F1 scores significantly. Adding RAG on top of that gave another boost. The combined approach achieved relative gains of about eighty-two percent for "Implements" dependencies compared to zero-shot.

Tom: So the improvement isn't just about using an LLM — it's about using it intelligently. Giving it the right context, the right examples, and the right amount of both.

Jane: And there's another important finding. They tested cross-dataset — meaning they trained on one system and tested on another. That's the realistic scenario because in practice, you don't have labeled data for every new system you work on.

Tom: And how did it hold up?

Jane: Really well. The accuracy dropped only slightly — about one point six percent on average. The fine-tuned BERT baseline, by contrast, degraded significantly. That shows LEREDD is robust and generalizes across different systems.

Tom: That's the practical win. Engineers can use this on a new project without needing to manually annotate thousands of requirement pairs first.

Jane: And that's the kind of improvement that moves this from a research curiosity to a deployable tool. We'll talk about the bigger implications next.

Conclusion: Tom: We've covered a lot of ground on "Automating the Detection of Requirement Dependencies Using Large Language Models." Let's wrap up with the big picture. Jane, what's the takeaway for our listeners?

Jane: The takeaway is that LLMs, when used with the right techniques — retrieval-augmented generation and in-context learning — can reliably automate a task that's been a bottleneck in software engineering for decades. Detecting requirement dependencies isn't glamorous, but it's essential.

Tom: And the numbers back it up. Ninety-six percent F1 on detecting non-dependencies means you can trust the system to filter out the noise. That frees engineers to focus on the actual dependencies that need human judgment.

Jane: And the cross-dataset robustness means this isn't a one-trick pony. It works across different automotive systems, which suggests it could generalize to other domains too — aerospace, medical devices, any complex system with lots of requirements.

Tom: The authors also released their annotated dataset — eight hundred thirteen requirement pairs across three systems. That's a gift to the research community. It gives others a benchmark to test against.

Jane: And that's how science progresses. Someone else can build on this, improve it, extend it to indirect dependencies, which the paper mentions as future work.

Tom: There's also the human element. The paper notes that manual annotation is costly and error-prone. Tools like LEREDD don't replace engineers — they augment them. They handle the tedious screening so humans can focus on the nuanced cases.

Jane: And in safety-critical domains like automotive, catching a missed dependency early can prevent a recall or worse. The impact here goes beyond productivity — it's about safety and reliability.

Tom: So as we say goodbye to this paper, I want to thank the authors for a rigorous, practical contribution. And Jane, what's next on our reading list?

Jane: We've got another paper lined up that builds on some of these ideas — I'll tease it as "requirement evolution and impact analysis." Should be a good follow-up.

Tom: Sounds great. Thanks for joining us, everyone. We'll see you on the next episode.

Jane: Take care, and keep reading.

More episodes

← Home