Measuring Curriculum Alignment across Topical Coverage, Competency, and Cognitive Depth: A Longitudinal Framework Applied to CS2013 and CS2023

arXiv:2606.19469 · cs.AI, cs.SE · Submitted 2026-06-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Measuring Curriculum Alignment across Topical Coverage, Competency, and Cognitive Depth: A Longitudinal Framework Applied to CS2013 and CS2023".

Jane: The paper was written by Sherzod Turaev, Mary John, Saja Aldabet, Mamoun Awad, Nazar Zaki et al. from United Arab Emirates University and Abu Dhabi Polytechnic.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: We are kicking things off today with a paper that has quite a mouthful of a title: "Measuring Curriculum Alignment across Topical Coverage, Competency, and Cognitive Depth: A Longitudinal Framework Applied to CS2013 and CS2023."

Jane: It certainly is long, Tom, but it really boils down to asking if what students learn in their computer science degrees actually matches up with global industry standards.

Tom: Exactly, and Sherzod Turaev and his team from United Arab Emirates University are doing the heavy lifting to find out.

Jane: They were joined by Mary John from Abu Dhabi Polytechnic and several other researchers to tackle this massive mapping problem.

Tom: I find it fascinating that they aren't just looking at a single snapshot of a program, but are actually comparing how things change over a whole decade.

Lu: That longitudinal perspective is so creative because it allows us to see if a curriculum is truly evolving or if it's just staying static while the world moves on.

Meng: I do wonder how they managed to handle such an enormous amount of data without the whole thing turning into a mess of spreadsheets.

Jane: They actually organized everything into structured corpora, which helped them map thousands of different topics and learning outcomes systematically.

Meng: That would definitely save university administrators a lot of headaches compared to trying to do all that manual tagging by hand!

Lu: And since they are comparing two different versions of the guidelines, it's almost like they're providing a time machine for curriculum design.

Lalam: This kind of work ensures that our educational foundations remain strong even as the landscape of technology shifts beneath us.

Tom: We will get into the specific nuts and bolts of how they actually built this measuring system in our next segment.

Summary: Tom: Now that we have set the stage, let's talk about how this team actually executed their plan to audit these degrees.

Jane: They developed what they call a "human-in-the-loop" pipeline, which essentially uses AI to suggest potential matches and then lets humans confirm them.

Tom: I read that they didn't just pick one AI model and assume it was the best one for the job.

Jane: No, they actually benchmarked seven different retrievers to see which one performed best on this specific type of text.

Lu: It was a really interesting finding because a smaller, more specialized model actually outperformed some of the much larger ones since it was better suited for short phrases.

Meng: I am curious about how they ensured the human part of that loop was actually consistent and didn't just rely on one person's opinion.

Jane: They had a second rater check a sample of the work and used Cohen’s kappa to measure how much they agreed, which gave them very high reliability scores.

Meng: So we aren't just blindly trusting an algorithm to decide what a student knows; there is a real layer of human verification involved.

Lu: I love that combination of automated speed and human wisdom; it makes the whole process feel much more robust and trustworthy.

Lalam: This creates a verifiable way to certify that our educational standards are actually being met in the classroom every single year.

Tom: It is a very sophisticated approach, but let's see what happens when they apply it to real-world curriculum data in a moment.

Improvements: Tom: We just heard about the methodology, but the actual results reveal some pretty significant gaps in how these programs are taught.

Jane: One of the most striking things they uncovered was something called an "articulation gap."

Tom: That basically means a program might be teaching a topic, but it's failing to actually state that skill as a formal learning outcome.

Jane: Right, so you might be learning the material in class, but it isn't officially promised in your course outcomes for employers to see.

Tom: And then there is the "depth gap," which is even more concerning when you look at newer standards like CS2023.

Jane: Under the old two thousand thirteen rules, programs were mostly hitting the right level, but under the new requirements, that dropped to about seventy-six percent.

Lu: I can see this being used to create much more precise learning paths that target those specific areas where depth is lacking.

Meng: I noticed they also found some critical omissions too, like parallel and distributed computing having zero percent coverage in one instance.

Jane: That is a huge red flag for any student wanting to work in modern, high-performance software engineering!

Meng: If universities can use this tool to see exactly where their curriculum is thin, they can actually fix it much faster.

Lu: It really moves curriculum design away from guesswork and toward a much more scientific approach.

Lalam: By addressing these gaps, we are essentially raising the floor for what an educated professional looks like in our society.

Tom: --- CONCLUSION ---

Conclusion: Tom: We have reached the end of our discussion on "Measuring Curriculum Alignment across Topical Coverage, Competency, and Cognitive Depth: A Longitudinal Framework Applied to CS2013 and CS2023."

Jane: It's a powerful study that reminds us that simply covering a topic isn't enough; you have to teach it at the right level.

Tom: And you have to be able to clearly state those competencies so students and employers know exactly what they can do.

Lu: I am so excited about the potential for this to make education much more dynamic and responsive!

Meng: It is a very practical tool that could help engineers and educators alike build better programs.

Lalam: This framework helps ensure our educational culture stays in perfect step with our technological progress.

Tom: Thanks for joining us today, everyone, we will see you next time!

United Arab Emirates University · Abu Dhabi Polytechnic

cs.AI, cs.SE

Submitted: 2026-06-17

Updated: 2026-09-15

Comments: 27 pages, 5 figures, 9 tables

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 83/100

The gist: This paper introduces a "reproducible, human-in-the-loop pipeline" designed to measure how undergraduate computer science programs align with international curricular guidelines.

Key concepts

Articulation gap
An articulation gap occurs when a program teaches a topic but fails to state that skill as a formal learning outcome. This means students learn the material, but it isn't officially documented in course outcomes for employers to recognize their specific competencies.
Depth gap
A depth gap happens when a curriculum covers a topic but fails to teach it at the required cognitive level. Under newer CS2023 standards, researchers found that programs met the necessary depth only about seventy-six percent of the time.
Human-in-the-loop pipeline
This methodology uses AI to suggest potential matches between data points and then relies on humans to confirm them. This approach combines the speed of automated algorithms with human wisdom to ensure the results are robust, accurate, and trustworthy.

Terminology

Summary

This paper introduces a reproducible, human-in-the-loop pipeline designed to measure how undergraduate computer science programs align with international curricular guidelines. It addresses the critical lack of reliable ways to assess not only topical coverage but also whether competencies are articulated and taught at the recommended cognitive depth. By applying this framework longitudinally to both CS2013 and CS2023, the study provides a way to distinguish between persistent structural gaps and changes that reflect the evolution of disciplinary standards.

The Three-Lens Framework

The researchers move beyond simple topical overlap by measuring alignment through three nested lenses of increasing stringency. This formal framework allows for a more nuanced diagnosis of curricular alignment:

  1. Knowledge: Measures topical coverage to determine if a unit is substantively taught.

  2. Skill: Evaluates competency articulation to see if the content is articulated as a stated outcome in course learning outcomes.

  3. Cognitive Depth: Assesses whether articulated competencies are delivered at the recommended cognitive level using a five-level scale derived from Bloom’s revised taxonomy.

This multi-dimensional approach enables the identification of an articulation gap—where content is taught but not stated as an outcome—and a depth gap, where competencies are present but taught below the required level.

The Retrieve-then-Confirm Pipeline

To ensure accuracy, the study employs a staged methodology that separates high-recall candidate generation from high-precision confirmation. The process begins with semantic retrieval, where seven different retrievers were benchmarked against pooled relevance judgments to find the most effective configuration. A reciprocal-rank-fusion ensemble was identified as the strongest retriever for proposing candidate matches between program text and guideline knowledge units.

Following retrieval, human experts perform the expensive, high-precision act of confirming these candidates under explicit definitions of coverage. To ensure scientific rigor, the resulting maps are validated by an independent second rater to establish substantial agreement, with Cohen’s kappa scores of 0.64 for CS2023 and 0.69 for CS2013. This design ensures that automation accelerates the work without being trusted to constitute the result.

Longitudinal Findings and Gap Diagnosis

By mapping a single accredited program against both CS2013 and CS2023, the study provides a longitudinal comparison that reveals how coverage shifts when guidelines are restructured. The program demonstrated a near-constant currency in coverage, encompassing roughly 50% of both sets of knowledge units. Crucially, the framework distinguishes between:

  • Persistent gaps: Omissions like parallel and distributed computing and foundations of programming languages that were present in both standards.

  • Standard-driven changes: Improvements in areas like mathematical and statistical foundations, which reflect the new CS2023 structure rather than changes in the program itself.

The analysis also includes a gap diagnosis against ABET accreditation criteria, separating deliberate specialization from accreditation-relevant omission. A major finding was that while competency presence remained high, the program's depth-adequacy rate dropped significantly under CS2023 compared to CS2013. The authors conclude this reflects the newer guideline’s raised expectations rather than a deficiency in the program.

Improvements for AI systems

1. Multi-Dimensional Alignment Engine

  • Capabilities: Instead of performing simple semantic similarity for topical overlap, this system quantifies alignment through three nested lenses: topical coverage (presence of subject matter), competency articulation (whether the skill is explicitly stated in learning outcomes), and cognitive depth (whether the outcome reaches the required level on a domain-specific Bloom’s Taxonomy scale). It can specifically diagnose articulation gaps (content taught but not promised as an outcome) and depth gaps (promised skills that fall below the required cognitive threshold, such as teaching at an Understand level when Analyze is required).

2. Task-Specific Retriever Benchmarking Module

  • Capabilities: This module prevents the deployment of suboptimal embedding models in RAG (Retrieval-Augmented Generation) pipelines by automatically benchmarking various embedding families (e.g., sentence-level vs. long-context models) against a small, human-verified gold standard relevance set specific to the task. It can detect when high-parameter, long-context models underperform on short-text/sentence-level semantic retrieval compared to lightweight, specialized sentence models, ensuring optimal retrieval precision before full-scale deployment.

3. Longitudinal Gap Diagnostic System

  • Capabilities: This system performs comparative analysis of a dataset or curriculum against multiple generations of a regulatory or disciplinary standard. It can distinguish between structural gaps (omissions that persist across different versions of a standard, indicating long-term neglect) and standard-driven changes (gaps that appear only in the newest version, indicating new requirements). This allows for prioritized decision-making in compliance and development.

4. Cognitive Verb-to-Level Validator

  • Capabilities: This component automatically audits instructional or professional outcomes by mapping leading action verbs to a structured, domain-specific taxonomy (e.g., Bloom’s for Computing). It can flag specific outcomes that fail to meet the required cognitive ambition of a standard, allowing for automated suggestions to refine Understand or Apply level outcomes into Analyze, Evaluate, or Create level outcomes.

Abstract

Undergraduate computer science is governed by international curricular guidelines revised about once a decade, yet programs lack a reliable way to measure how completely they cover the current guideline and how coverage shifts when it changes. Existing analyses rely on topic models or manual tagging, seldom report reliability, do not benchmark the matching method, and examine topical overlap at a single point in time. We address these gaps with a staged pipeline that separates candidate generation from confirmation, applied to one accredited Bachelor of Science in Computer Science against Computer Science Curricula 2013 (CS2013) and 2023 (CS2023). Semantic retrieval proposes candidate course-to-knowledge-unit matches, a large language model confirms each against an explicit coverage rule, and an independent expert validates the resulting map. Benchmarking seven retrievers against pooled relevance judgments, we find that no automatic configuration reaches acceptable precision and recall, peaking at an F1 of 0.55 and inflating apparent coverage once tuned for recall, establishing retrieval as a candidate generator, not a measurement. Each map was validated by two independent experts and reconciled to a consensus, with substantial first-pass agreement (Cohen's kappa 0.64 and 0.69); the reported coverage is the lenient end of a sensitivity band whose strict end lies about seven points lower. Coverage of CS2023 is 48.4 percent of knowledge units, 59.4 percent by recommended hours, and about 28 percent of topics, and sixty-nine percent of covered units rest on a single course. The program articulates most competencies it covers yet meets the recommended cognitive depth far less often under CS2023 than under CS2013, a gap that survives a sensitivity analysis of the mapping, while structural gaps stay separable from artifacts of the standard's evolution. The instrument is reusable and released.

Sources

Related papers