Automating the Detection of Requirement Dependencies Using Large Language Models

arXiv:2602.22456 · cs.SE, cs.AI · Submitted 2026-08-14 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Automating the Detection of Requirement Dependencies Using Large Language Models".

Jane: The paper was written by Ikram Darif, Feifei Niu, Manel Abdellatif, Lionel C. Briand, Ramesh S. et al. from University of Ottawa and École de technologie supérieure and University of Limerick and General Motors.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title and Authors: Tom: Welcome back to the show, everybody. Today we're digging into a paper that's got a title that sounds like a mouthful but is actually about something really practical: "Automating the Detection of Requirement Dependencies Using Large Language Models."

Jane: And Tom, I gotta say, when I first read that title, I thought, okay, this is about software requirements, which sounds dry, but the problem they're tackling is one of those hidden time-sinks in engineering that nobody talks about.

Tom: Right, and the authors are a mix of academic and industry folks — Ikram Darif, Feifei Niu, Manel Abdellatif, Lionel Briand, and then Ramesh S and Arun Adiththan from General Motors. That's a big deal, right?

Jane: Huge. When you see GM on the author list, you know this isn't just theoretical. They're dealing with real automotive systems where requirements are complex and interconnected.

Tom: So Jane, for our listeners who aren't software engineers — what's a requirement dependency? Why should I care?

Jane: Think of it like building a house. You have a requirement that says "the house needs a roof" and another that says "the house needs walls." The roof depends on the walls being there first. If you change the walls, you've got to know the roof is affected. In software, especially in cars, these dependencies are everywhere.

Tom: And the problem is, in a modern car, you might have hundreds of requirements. Manually checking every pair to see if they depend on each other? That's thousands of pairs. It's exhausting and error-prone.

Jane: Exactly. And the paper points out that when people skip this step, you get project delays, rework, even safety issues. In a car, a missed dependency could mean a safety feature doesn't work because something else changed.

Tom: So they're using large language models — the same kind of tech behind chatbots — to automate this detection. And the authors are claiming some pretty impressive results. We'll get into the numbers soon.

Jane: The exciting part is the team structure. You've got academic rigor from the university folks and practical urgency from GM. That combination usually means the research is actually going to be useful, not just published and forgotten.

Tom: And that's what we're going to explore today — how they built this system, what it achieves, and whether it can actually help engineers in the real world. Stick around.

Jane: We're just getting started with "Automating the Detection of Requirement Dependencies Using Large Language Models," and there's a lot to unpack.

Summary: Tom: So we've got the title and the team. Now let's talk about what this paper actually does. Jane, can you break down the core idea for us?

Jane: Sure. The paper introduces a system called LEREDD — that's LLM-Enabled REquirement Dependency Detection. The basic idea is: give it a list of requirements from a software spec, and it tells you which pairs are dependent on each other and what kind of dependency it is.

Tom: And they're not just saying "these two are related." They're classifying into specific types — like "Requires," "Implements," "Conflicts," "Contradicts," "Details," "Is similar," and "Is a variant."

Jane: Right. So "Requires" means one thing can't happen without the other. "Conflicts" means fulfilling one restricts the other. These distinctions matter because they affect how engineers plan and manage changes.

Tom: And the clever part is how they get the LLM to do this well. They use two tricks: Retrieval-Augmented Generation, or RAG, and In-Context Learning, or ICL. Can you explain those in plain terms?

Jane: Think of RAG as giving the model a cheat sheet. Before asking it to judge a pair of requirements, you feed it relevant sections from the software spec — the domain context. So if the requirements mention a "BCS" subsystem, the model knows from the context that BCS is the braking control system.

Tom: And ICL?

Jane: That's showing the model examples. For each dependency type, they retrieve a few similar requirement pairs from a labeled dataset and show the model "here's what a 'Requires' dependency looks like, here's what 'No dependency' looks like." It's like teaching by example.

Tom: So the model gets both the context and examples before making its call. And they tested this across three automotive systems — Traffic Jam Assist, Automated Parking Assist, and Adaptive Driving Beam. They annotated over eight hundred requirement pairs manually.

Jane: That's a lot of careful human work. Two annotators independently labeled these pairs, then resolved disagreements. That gives us a solid ground truth to measure against.

Tom: And the results? They're claiming an average accuracy of about ninety-three percent and an F1 score of eighty-four percent. But the really impressive number is for detecting "No dependency" — that's ninety-six percent F1.

Jane: That's huge because in real systems, most requirement pairs are NOT dependent. If you can reliably filter those out, you save engineers an enormous amount of time. They only need to focus on the small fraction that actually have dependencies.

Tom: And they compared against two baseline methods — a traditional TF-IDF retrieval approach and a fine-tuned BERT model. LEREDD beat both, especially on the fine-grained dependency types.

Jane: The gains on "Requires" were dramatic — like ninety-five percent to one hundred five percent relative improvement in F1 score over the baselines. That's not a small bump; that's a game-changer.

Tom: So the summary is: they built a system that uses LLMs with context and examples to accurately detect requirement dependencies, and it outperforms existing methods. We're going to dig into the methodology and what makes it work next.

Improvements: Tom: We've covered the basics. Now let's talk about what makes LEREDD better than just asking a chatbot "are these two requirements related?" Jane, what's the real innovation here?

Jane: The key improvement is that they didn't just use an LLM in a naive way. They systematically tested different prompting strategies. And they found that zero-shot — just asking the model cold — performs poorly on fine-grained dependencies. The model tends to over-predict dependencies and struggles with rare types.

Tom: So they added the two-layer knowledge retrieval. And they didn't just pick arbitrary settings — they ran two hundred sixteen experiments to tune the few-shot parameters. That's thorough.

Jane: Very thorough. They tested different embedding models — SBERT versus BGE-M3. They tested different similarity metrics — cosine versus Euclidean. They tested different ways to aggregate similarity between requirement pairs. And they tested different numbers of examples, from one to nine.

Tom: And the winner was SBERT with Euclidean distance, using a max-similarity aggregation, and four examples per dependency type. Why does that combination work?

Jane: The max aggregation is interesting. Instead of averaging similarity across all four requirement combinations, they take the best match for each requirement in the target pair. That makes sense because a single highly relevant example is more informative than several mediocre ones.

Tom: And for RAG, they found that ten chunks of five hundred characters each was optimal. Not the whole document — that adds noise. Not too few — that lacks context. There's a sweet spot.

Jane: Exactly. And the improvements were consistent. Adding few-shot examples boosted F1 scores significantly. Adding RAG on top of that gave another boost. The combined approach achieved relative gains of about eighty-two percent for "Implements" dependencies compared to zero-shot.

Tom: So the improvement isn't just about using an LLM — it's about using it intelligently. Giving it the right context, the right examples, and the right amount of both.

Jane: And there's another important finding. They tested cross-dataset — meaning they trained on one system and tested on another. That's the realistic scenario because in practice, you don't have labeled data for every new system you work on.

Tom: And how did it hold up?

Jane: Really well. The accuracy dropped only slightly — about one point six percent on average. The fine-tuned BERT baseline, by contrast, degraded significantly. That shows LEREDD is robust and generalizes across different systems.

Tom: That's the practical win. Engineers can use this on a new project without needing to manually annotate thousands of requirement pairs first.

Jane: And that's the kind of improvement that moves this from a research curiosity to a deployable tool. We'll talk about the bigger implications next.

Conclusion: Tom: We've covered a lot of ground on "Automating the Detection of Requirement Dependencies Using Large Language Models." Let's wrap up with the big picture. Jane, what's the takeaway for our listeners?

Jane: The takeaway is that LLMs, when used with the right techniques — retrieval-augmented generation and in-context learning — can reliably automate a task that's been a bottleneck in software engineering for decades. Detecting requirement dependencies isn't glamorous, but it's essential.

Tom: And the numbers back it up. Ninety-six percent F1 on detecting non-dependencies means you can trust the system to filter out the noise. That frees engineers to focus on the actual dependencies that need human judgment.

Jane: And the cross-dataset robustness means this isn't a one-trick pony. It works across different automotive systems, which suggests it could generalize to other domains too — aerospace, medical devices, any complex system with lots of requirements.

Tom: The authors also released their annotated dataset — eight hundred thirteen requirement pairs across three systems. That's a gift to the research community. It gives others a benchmark to test against.

Jane: And that's how science progresses. Someone else can build on this, improve it, extend it to indirect dependencies, which the paper mentions as future work.

Tom: There's also the human element. The paper notes that manual annotation is costly and error-prone. Tools like LEREDD don't replace engineers — they augment them. They handle the tedious screening so humans can focus on the nuanced cases.

Jane: And in safety-critical domains like automotive, catching a missed dependency early can prevent a recall or worse. The impact here goes beyond productivity — it's about safety and reliability.

Tom: So as we say goodbye to this paper, I want to thank the authors for a rigorous, practical contribution. And Jane, what's next on our reading list?

Jane: We've got another paper lined up that builds on some of these ideas — I'll tease it as "requirement evolution and impact analysis." Should be a good follow-up.

Tom: Sounds great. Thanks for joining us, everyone. We'll see you on the next episode.

Jane: Take care, and keep reading.

Ikram Darif, Feifei Niu, Manel Abdellatif, Lionel C. Briand, Ramesh S., Arun Adiththan

University of Ottawa · École de technologie supérieure · University of Limerick · General Motors

cs.SE, cs.AI

Submitted: 2026-08-14

Updated: 2026-08-18

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 69/100

The gist: The paper introduces LEREDD (LLM-Enabled REquirement Dependency Detection), an automated approach for detecting various types of direct dependencies between Natural Language (NL) requirement pairs.

Key concepts

Requirement Dependency
This refers to a relationship between two requirements in a system, such as one requirement depending on another. For example, a roof requirement depends on the wall requirement being present first.
LEREDD
This is the system introduced in the paper, standing for LLM-Enabled Requirement Dependency Detection. Its basic function is to take a list of requirements and tell you which pairs are dependent and classify the type of dependency between them.
Retrieval-Augmented Generation (RAG)
RAG provides context to an LLM by feeding it relevant sections from a software specification before it makes a judgment. This allows the model to understand specific domain details, such as identifying what a subsystem like 'BCS' refers to in the context of braking control.
In-Context Learning (ICL)
ICL involves showing an LLM examples within the prompt for each dependency type. This teaches the model by example, allowing it to better distinguish between different dependency classifications like 'Requires' or 'Conflicts' based on provided instances.

Terminology

Summary

The paper introduces LEREDD (LLM-Enabled REquirement Dependency Detection), an automated approach for detecting various types of direct dependencies between Natural Language (NL) requirement pairs. The approach leverages Retrieval-Augmented Generation (RAG) to extract domain-specific context from the Software Requirement Specification (SRS) document, and utilizes In-Context Learning (ICL) to dynamically retrieve relevant examples for each dependency type and the no-dependency case.

The paper states that Modern software systems are becoming increasingly complex, driven by the growing and advanced demands of stakeholders and users. Requirements are not independent entities. They are inherently interconnected through dependencies of varying natures. Identifying these dependencies is essential because they underpin critical software development decisions and influence activities throughout the software development life cycle. However, detecting requirement dependencies remains error-prone, time-consuming, and cognitively challenging, largely due to the ambiguity introduced by the predominant use of Natural Language (NL) to author requirements, as well as to the reliance on manual effort for detection.

The framework processes a set of n NL requirements extracted from an SRS document, generating a list of n(n−1)/2 unique requirement pairs. These pairs are processed by a two-phase pipeline:

Knowledge Retrieval Phase applies a dual-strategy process:

  1. Contextual retrieval employs RAG to extract domain-specific information from the SRS document, using a fixed-size chunking strategy where the k=10 most semantically similar chunks are retrieved.

  2. Dynamic examples retrieval identifies and retrieves semantically similar requirement pairs for all dependency types from an annotated dataset, using the SBERT model for embeddings and a maximum-similarity aggregation formula.

Dependency Inference Phase uses GPT-4.1 (identified as the best-performing model) with a structured prompt containing: formal definitions of dependency types, retrieved examples, domain-specific context, and specific instructions. For each requirement pair, LEREDD generates three outputs: (1) a prediction classifying the dependency type or indicating no dependency, (2) a confidence score on a 5-point Likert scale, and (3) a rationale justifying the prediction.

The approach supports seven dependency types: Requires, Implements, Conflicts, Contradicts, Details, Is similar, and Is a variant, derived from a state-of-the-art classification, with Implements introduced for the industrial partner's needs.

The authors collected SRS documents from Michigan State University's Requirements Engineering course website, selecting three automotive systems: Traffic Jam Assist (TJA), Automated Parking Assist (APA), and Adaptive Driving Beam (ADB). The curated dataset contains 40, 25, and 50 requirement statements for ADB, TJA, and APA respectively. Two annotators independently annotated requirement pairs, achieving a Cohen's kappa of 0.43 (moderate agreement), with all disagreements resolved by consensus. The final annotated dataset comprises 813 requirement pairs across the three systems, with the distribution: No Dependency (642), Requires (95), Implements (30), Details (21), Conflicts (18), and Is similar (7).

RQ1: Which SOTA LLM is more effective? Four LLMs were evaluated (GPT-4.1, Llama 3.1, Gemma 20B, Mistral 7B) in zero-shot settings. GPT-4.1 consistently outperformed others, achieving the highest F1 scores across all systems (ADB: 0.40, TJA: 0.29, APA: 0.47), with an average of 0.39. The paper notes: "Zero-shot models typically over-predict dependencies. Zero-shot GPT-4.1 can reliably identify non-dependent requirement pairs, achieving an average F1 score of 0.87 for the No dependency class across the three systems. However, the model struggles with fine-grained dependency types."

RQ2: What is the best prompting strategy? A comprehensive suite of 216 experiments was conducted. The optimal few-shot configuration uses SBERT, Euclidean distance, maximum similarity aggregation, and 4 examples per dependency type, combined with a RAG configuration of ten 500-character chunks. This setup achieves average F1 scores of 0.95, 0.70, and 0.70 for No dependency, Requires, and Implements, yielding relative gains of 5.56%, 34.62%, and 81.82% respectively compared to zero-shot GPT-4.1.

RQ3: Intra-dataset comparison with baselines. LEREDD was compared against two baselines: TF-IDF & LSA (retrieval-based) and fine-tuned BERT (ML-based). LEREDD consistently achieved the best accuracy and F1 scores across all systems, reaching 0.97 and 0.92 on ADB, 0.92 and 0.76 on TJA, and 0.89 and 0.85 on APA. For the Requires dependency, LEREDD improved F1 scores from 0.22 and 0.38 to 0.92 on ADB, yielding relative gains of 318% and 142% over the baselines. The paper reports: "It yields relative gains of 33.33% over TF-IDF & LSA and 61.54% over fine-tuned BERT across all dependency types and systems."

RQ4: Cross-dataset comparison with baselines. Six experiments covering all training-testing combinations were conducted. LEREDD achieved an average accuracy of 0.915 and an average F1 score of 0.76 across all experiments. For the Requires dependency, average F1 scores increased from 0.33 and 0.22 to 0.70, yielding relative gains of 112.12% and 218.18% respectively. For No dependency, average F1 scores improved from 0.86 and 0.47 to 0.95, yielding relative gains of 10.47% and 102.13%. The paper notes: Its accuracy is comparable to that of the intra-dataset evaluation, with the average decreasing by only 1.61%.

The paper reports that LEREDD achieves an average accuracy of 92.66% and an F1 score of 84.33%, across all dependency classes and the evaluated systems, with the No dependency class achieving an average F1-score of 96%. The authors emphasize: By effectively filtering these cases, LEREDD can substantially reduce the analysts' overhead for dependency analysis.

Regarding computational cost, "TF-IDF & LSA required an average of 2.48 seconds, while fine-tuned BERT and LEREDD required an average of 4 minutes 3 seconds and 1 minute 48 seconds, respectively."

The paper concludes: "Zero-shot LLMs are insufficient for accurately identifying fine-grained dependency types and often yield low F1 scores even in No dependency cases, underscoring the need for task-specific guidance. Third, much better results can be obtained with a small set of carefully selected, relevant examples. Fourth, regarding RAG, retrieval precision plays a more critical role than retrieval volume."

The authors state: "As future work, we will extend LEREDD to capture indirect and implicit requirement dependencies. In addition, we will investigate the utility of predicted dependencies for impact analysis when requirements evolve over time."

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems and what the improved system can do:

  • Integrate a two-stage retrieval pipeline: (a) contextual retrieval using RAG with fixed-size chunking (500-character chunks, top-10 retrieval), and (b) dynamic example retrieval using SBERT embeddings with Euclidean distance and max-similarity aggregation.

  • Use the formula Score max avg = (max(sim(R1,Ra), sim(R1,Rb)) + max(sim(R2,Ra), sim(R2,Rb))) / 2 to select the top-4 most similar examples per dependency type.

What the improved system can do:

Automatically detect seven types of requirement dependencies (Requires, Implements, Conflicts, Contradicts, Details, Is similar, Is a variant) from natural language requirements with an average F1 score of 0.84 and accuracy of 0.93, outperforming zero-shot LLMs by 81.82% relative gain on Implements dependency.

  • Add a post-inference layer that forces the LLM to output three structured fields: Dependency Status, Rationale, and Confidence Score (0–5 Likert scale), using exact label formatting.

  • Implement temperature=0 for deterministic outputs and a confidence threshold to flag low-confidence predictions for human review.

  • Implement a dynamic example retrieval mechanism that does not rely on the target system's annotated data, using examples from a different system (e.g., ADB examples for TJA testing).

  • Add a fallback mechanism: if no examples are available for a dependency type, use the zero-shot mode with RAG context only.

  • Add a post-processing layer that adjusts predictions using the known class distribution (e.g., No dependency = 79%, Requires = 11.7%) to correct over-prediction of dominant classes.

  • Implement a two-stage decision: first classify as dependent/non-dependent, then fine-tune the dependency type.

  • Use SBERT (all-mpnet-base-v2) with Euclidean distance and max-similarity aggregation to select 4 examples per dependency type dynamically, rather than static examples.

  • Implement a saturation detection mechanism: stop adding examples when F1 score improvement is < 1% (optimal range: 3–5 examples).

  • Implement a chunk-size and count tuner: use 500-character chunks with 200-character overlap, retrieve top-10 chunks, and exclude the full document to avoid noise.

  • Add a relevance filter: only include chunks with cosine similarity > 0.5 to the target pair.

  • Include a module that uses the improved system to automatically generate annotated requirement pairs with dependency labels, rationales, and confidence scores, following the paper's annotation guidelines (Cohen's kappa = 0.43, consensus resolution).

The improved AI system can:

  • Process natural language requirement documents and output structured dependency classifications with explanations and confidence scores.

  • Operate effectively in both intra-dataset and cross-dataset settings, making it deployable for new systems without retraining.

  • Reduce manual dependency analysis effort by up to 96% for non-dependent pairs while maintaining high accuracy on fine-grained dependencies.

  • Provide a reproducible, open-source pipeline (with the paper's replication package) for benchmarking and future research.

Abstract

Requirements are inherently interconnected through various types of dependencies. Identifying these dependencies is essential, as they underpin critical decisions and influence a range of activities throughout software development. However, this task is challenging, particularly in modern software systems, given the high volume of complex, coupled requirements. These challenges are further exacerbated by the ambiguity of Natural Language (NL) requirements and their constant change. Consequently, requirement dependency detection is often overlooked or performed manually. Large Language Models (LLMs) exhibit strong capabilities in NL processing, presenting a promising avenue for requirement-related tasks. While they have shown to enhance various requirements engineering tasks, their effectiveness in identifying requirement dependencies remains unexplored. In this paper, we introduce LEREDD, an LLM-based approach for automated detection of requirement dependencies that leverages Retrieval-Augmented Generation (RAG) and In-Context Learning (ICL). It is designed to identify diverse dependency types directly from NL requirements. We empirically evaluate LEREDD against two state-of-the-art baselines. The results show that LEREDD provides highly accurate classification of dependent and non-dependent requirements, achieving an accuracy of 0.93, and an F1 score of 0.84, with the latter averaging 0.96 for non-dependent cases. LEREDD outperforms zero-shot LLMs and baselines, particularly in detecting fine-grained dependency types, where it yields average relative gains of 94.87% and 105.41% in F1 scores for the Requires dependency over the baselines. We also provide an annotated dataset of requirement dependencies encompassing 813 requirement pairs across three distinct systems to support reproducibility and future research.

Sources

Related papers