RAIL: An Automatic Classifier of the Artificial Intelligence Readiness Level

arXiv:2608.13428 · cs.AI · Submitted 2026-08-13 · Read on arXiv

Juan Irving Vasquez, Juan Terven, Laura-Ivoone Garay-Jimenez

Instituto Politécnico Nacional

cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: Under review at journal

Code: https://github.com/irvingvasquez/RAIL

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 50/100

The gist: This paper addresses the challenge of assessing the maturity of artificial intelligence technologies, which is essential for investment decisions, project management, and policy monitoring.

Terminology

Summary

This paper addresses the challenge of assessing the maturity of artificial intelligence technologies, which is essential for investment decisions, project management, and policy monitoring. The authors note that available readiness frameworks are heterogeneous and difficult to apply automatically: "the adaptation of Technology Readiness Levels to AI lacks AI-specific gating criteria, the Machine Learning Technology Readiness Levels presuppose access to internal process artifacts, and AI/data readiness dimension models employ scales that resist direct comparison."

The paper makes two contributions. First, it unifies three frameworks into the Unified AI Readiness Level (AIRL), "a nine-level ordinal scale built on an environmental evidence ladder and complemented by dimensional caps (covering specification, data existence, data quality, data legality, expert knowledge, and algorithmic maturity) together with a generality-anchoring rule and explicit assignment disciplines, so that a readiness level becomes decidable from a natural-language description of the work alone." Second, it proposes RAIL (Readiness Assessment via Independent LLM-experts), "a panel-of-experts classifier that operationalizes the scale: one evidence agent and six independent dimension agents, each a large language model with a narrowly scoped mandate, deliver verdicts that a deterministic minimum rule aggregates and a chief expert reviews under asymmetric authority, confirming or lowering the panel's recommendation but never raising it above the caps."

The method was tested on several research works, showing consistency and avoiding overestimation from monolithic LLM classifiers. Code is available at https://github.com/irvingvasquez/RAIL.

The paper begins by noting that "Artificial intelligence has moved from laboratory research to a strategic asset in industry, government, and science, and with this transition the question of maturity has become as consequential as the question of performance." Investors, project managers, and policy bodies all need to place AI work on a common maturity scale. The Technology Readiness Levels (TRL) introduced by NASA has long served this role for conventional engineered systems, but "the TRL scale, however, was conceived for deterministic hardware and software, and a growing body of literature shows that it transfers poorly to artificial intelligence: AI systems are data-dependent, stochastic, and sensitive to distribution shift, so that a system validated in the laboratory may degrade unpredictably once deployed."

Several AI-specific readiness frameworks have been proposed, including the European Commission's AI Watch contextualization of TRLs for AI, the Machine Learning Technology Readiness Levels (MLTRL), and dimensional models decomposing readiness into AI and data specific facets. However, "these frameworks are heterogeneous and difficult to apply automatically: the AI adaptation of the TRL lacks AI-specific gating criteria, MLTRL presupposes access to internal process artifacts that are rarely visible in a textual description of a work, and the dimensional models employ scales of differing lengths that resist direct comparison."

Regarding automatic estimation, the authors note that "Bibliometric and clustering methods estimate maturity only at the level of entire technology fields and at coarse granularity; supervised ensembles depend on small, domain-specific labeled datasets; and recent single-LLM assistants, while promising, inherit the biases and hallucination risks of a single model and still deviate from expert judgment in a non-negligible fraction of cases. The literature exhibits a twofold gap: AI-specific frameworks lack automated, reproducible estimation mechanisms, and existing estimators fail to reproduce the property that makes manual assessment reliable."

The paper's contributions are: (1) the Unified AI Readiness Level (AIRL) integrating the EU interpretation of TRL, MLTRL, and the dimensional model of Eljasik-Swoboda et al. into a single framework; and (2) RAIL, a panel-of-experts classifier. By construction, the resulting classifier is sound with respect to the framework, conservative under ambiguity, neutral under silence (lack of evidence), and auditable.

The TRL scale was introduced by NASA as a nine-level metric ranging from the observation of basic principles (TRL 1) to a system proven in an operational environment (TRL 9). Several adaptations have been proposed:

  • Eljasik-Swoboda et al. extended the READINESSnavigator tool with six AI- and data-specific readiness dimensions (algorithmic, data-quality, data-legal, etc.)

  • Martínez-Plumed et al. (AI Watch) contextualized the nine TRLs for AI and introduced bidimensional readiness-versus-generality charts

  • Lavin et al. formalized MLTRL, a ten-level (0–9) systems-engineering framework featuring non-monotonic 'switchbacks,' standardized TRL Cards, and gated multidisciplinary reviews

  • Browne et al. proposed the AI Readiness Level (AIRL) framework gating TRL progression behind minimum thresholds in five dimensions: alignment, justified confidence, governance, human readiness, and data readiness

  • Müller et al. proposed the Use Case-Centered AI Readiness Level (UCAIRL), an eleven-level scale

  • Other adaptations include a unified nine-level TRL ladder for clinical AI, TRL-based mapping of AI maturity in maternal health, an organizational AI-readiness index for the public sector, and an enterprise-level AIRL scale

Early work used soft-computing techniques (neural networks, genetic algorithms, fuzzy logic). Chukhray et al. trained a stacking ensemble on 56 university R&D projects. Bibliometric approaches by Dastoor et al. fitted S-curves to publications, patents, and grants. Jain and Kumar clustered 136 technology trends with unsupervised methods. Most recently, Betancourt et al. leveraged LLMs, fine-tuning LLaMA 2 and GPT-3.5-Turbo on approximately 2,500 TRL-specific samples; on eight real prototypes, the assistant matched expert assessments exactly in 50% of cases and within one level in a further 37.5%.

The authors identify key limitations: bibliometric methods operate at coarse granularity, supervised ensembles depend on small labeled datasets, and single-LLM assistants inherit the biases and hallucination risks of a single model. Moreover, virtually all estimation work targets the classical TRL scale, and does not yet address the richer, multidimensional AI-readiness frameworks.

The AIRL integrates three frameworks: the EU adaptation of TRL by Martínez-Plumed et al., MLTRL by Lavin et al., and the dimensional model of Eljasik-Swoboda et al. "None of the three frameworks alone suffices for the classification task addressed in this work: the TRL scale lacks AI-specific gating criteria, ML-TRL presupposes access to internal process artifacts that are rarely visible in a textual description of a work, and the dimensional model of Eljasik-Swoboda et al. employs heterogeneous scales of five to nine levels that resist direct comparison."

The AIRL scale comprises nine ordinal levels whose primary discriminating variable is the evaluation environment in which evidence of functioning has been produced. The levels are:

  • AIRL 1: basic principles have been formulated but no experiment has been executed

  • AIRL 2: a concrete application concept exists and that exploratory experiments have been performed on sample, toy, or synthetic data

  • AIRL 3: the experimental proof of principle: the approach is validated in a testbed, typically against public benchmarks or simulated data, with defined metrics and baseline comparisons

  • AIRL 4: the constituent components (model, data pipeline, and interfaces) have been integrated and shown to work together in a controlled environment

  • AIRL 5: the technology has been validated on real, representative data of the target use case in a relevant environment, and when the evaluation includes application-oriented measures in addition to conventional ML metrics

  • AIRL 6: the demonstration of the technology as a capability... the model no longer operates in isolation but as a module of a larger workflow, demonstrated in a relevant environment before stakeholders beyond the research team

  • AIRL 7: an actual system prototype functioning in the operational environment, typically as a pilot, field trial, beta program, or shadow deployment on live data

  • AIRL 8: "the system complete and qualified: verified against the full set of requirements, subjected to deployment-oriented testing regimes such as canary or shadow tests, and, in regulated domains, certified by the competent authority"

  • AIRL 9: reserved for systems proven in sustained operation, requiring monitoring for data drift, concept drift, and performance degradation, together with defined retraining or feedback processes

The AIRL incorporates six dimensions as caps: specification, data existence, data quality, data legality, expert knowledge, and algorithmic maturity. "When a description provides explicit evidence that one of these dimensions is deficient... the assigned level is bounded above by the level compatible with that deficiency, irrespective of the sophistication of the reported experiments. The final level is the minimum of the environmentally evidenced level and the caps induced by explicitly evidenced dimensional gaps. Silence regarding a dimension is treated as neutral; only affirmative evidence of a gap triggers a cap."

The scale internalizes the readiness-generality tradeoff as an anchoring rule: Each description is classified at the level of generality that it itself claims. Three further disciplines govern assignment: (1) evidence takes precedence over intention, (2) ambiguity is resolved conservatively toward the lower of two adjacent candidate levels, and (3) maturity is not inherited — incorporating a pretrained component does not transfer readiness, and reworked systems are classified at the current state of the reworked component.

The authors note that a single LLM attempting to assign AIRL in one pass conflates three judgments that the framework deliberately keeps distinct and exhibits two failure modes: hallucinating gaps for dimensions about which descriptions are merely silent, and letting strong experimental evidence argue past explicitly stated deficiencies. The RAIL architecture therefore uses a panel of experts.

The panel comprises eight assessing agents and one deterministic operator in three stages: (1) one evidence expert and six dimension experts examine the description independently and in parallel; (2) a deterministic aggregator combines verdicts through the minimum rule; (3) a chief expert receives the complete panel report and pronounces the final classification under asymmetric authority.

The evidence expert is responsible solely for the environmental ladder. It identifies the most advanced evaluation environment with demonstrated evidence and returns a provisional level in 1,...,9 with decisive evidence. It also applies the assignment disciplines: compositional minimum over essential components (R1), exclusion of aspirational claims (R2), non-inheritance of maturity (R5), generality anchoring (R6), and treatment of reworked systems (R7). It emits two Boolean flags: whether the application belongs to a regulated domain and whether regulatory certification is explicitly stated.

Each of six dimension experts handles exactly one dimension: specification (D1), data existence (D2), data quality (D3), data legality (D4), expert knowledge (D5), and algorithmic maturity (D6). Each outputs a cap ci ∈ 9, 6, 4, where 9 means no cap, and 6 and 4 denote two severities of explicitly evidenced deficiency. Two disciplines govern: neutrality of silence (caps require affirmative textual evidence) and evidentiary traceability (any cap must be accompanied by a verbatim quotation). A fail-neutral policy interprets malformed responses as ci = 9.

The aggregated recommendation is computed as:

l* = min(l, c1,..., c6, ρ)

where ρ = 7 if r = 1 and z = 0 (regulated domain without certification), otherwise ρ = 9. This is the direct formalization of the minimum principle inherited from Eljasik-Swoboda et al. and of the system-level composition rule of Lavin et al. The rule is exact, auditable, and immune to persuasion.

The chief expert reviews the full deliberation for quality control and produces the final justification. Its authority is asymmetric: the final classification is AIRL = min(v, l*), where v is the chief's proposed level. The presiding agent may confirm the panel's recommendation or lower it, but may never raise it above the deterministic bound. Dissent is recorded verbatim in the classification output.

The classifier is: sound (no output can exceed the bound implied by evidenced environmental level and dimensional gaps), conservative (ambiguity resolved downward), neutral under silence (caps require affirmative quoted evidence), auditable (every classification carries provisional level, binding caps with quotations, applied rules, and dissent), and modular (agents are independent and may be revised without retraining the rest).

The evaluation used a combinatorial experiment with protocols (TRL, AI TRL, AIRL) and classification strategies (Monolithic LLM, RAIL). The corpus consisted of ten master's and doctoral theses from CIDETEC. For each thesis, the title, abstract, experimental summary, and conclusions were extracted. Experiments used Qwen3:32B through Ollama on an RTX 4090.

The four configurations produced markedly different distributions:

  • TRL (Mono): mean 4.7 (range 4–6)

  • AI TRL (Mono): mean 6.2 (range 4–7, with 7 as modal value assigned to six of ten theses)

  • AIRL (Mono): mean 4.7 (range 3–5, with 5 assigned to eight documents)

  • AIRL (RAIL): mean 4.5 (range 3–5)

Three regularities emerged. First, the AI-adapted TRL under monolithic prompting is systematically the most generous configuration: it exceeds the original TRL on nine of ten documents, by +1.5 levels on average. For the Vasquez 2009 thesis, the classifier assigned TRL 7 despite acknowledging the test was limited to a single object, introducing uncertainty about broader generality — demonstrating the argue-past failure mode.

Second, the AIRL protocol under a monolithic classifier exhibits the opposite pathology in attenuated form: compression toward the center of the scale. Eight of ten documents received level 5, and the dimensional apparatus almost never fires. The panel lowered two labels (from 5 to 4) where explicit dimensional gaps existed. For the Alvarez 2023 tomato detection thesis, RAIL's D3 expert located and quoted el conjunto de validación real carecía de balance en sus clases (explicitly stated class imbalance), capping the label at 4.

Third, RAIL is empirically conservative with respect to its monolithic counterpart: it never exceeds the monolithic AIRL label on any document.

The simulation-only thesis (Gante 2023) was the cleanest case: evidence expert placed it at level 3, all dimension experts returned neutral caps, chief expert confirmed with empty dissent. The cloud-deployed segmentation thesis (Brito 2024) exercised the opposite path: evidence expert proposed level 7, but the deterministic layer capped the label at 4, with the chief expert's dissent recording the tension. Across the audited classifications the chief expert confirmed the deterministic recommendation in every case and never lowered it.

Decomposed deliberation multiplies inference. The monolithic baseline classifies a document in approximately 22 seconds, while full RAIL classification requires approximately 131 seconds per document — a factor of six. For batch production of audited training labels, this cost is justified; for interactive use, the monolithic AIRL with deterministic post-check is a defensible low-cost option.

"Ten documents from a single institution, in a single genre (graduate theses in robotics and applied computer vision, in Spanish and English), classified by a single model at a single seed, support an analysis of agreement structure and mechanism, not a claim of accuracy." The summarization of experimental sections by a separate LLM introduces further mediation.

The paper introduced the Unified AI Readiness Level (AIRL) and the RAIL panel-of-experts classifier. The experimental evaluation confirmed the failure modes motivating the architecture: monolithic classifiers with AI-adapted TRL systematically inflated maturity, while monolithic classifiers with the full AIRL rulebook compressed outputs toward the center and rarely activated dimensional caps. In contrast, the proposed RAIL never exceeded its monolithic counterpart, and its independent dimension experts recovered explicitly stated gap that the same underlying model overlooked when judging holistically.

The dissent mechanism preserved informative disagreements. Benefits come at roughly six times the computational cost of single inference, a price that is justified for the batch production of audited training labels.

Future work will proceed along three lines: the construction of a larger, multi-institutional corpus with expert-labeled readiness levels to enable proper accuracy and inter-rater agreement studies.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems:

Improvement: Replace monolithic LLM classifiers with a panel of specialized agents, each with a narrowly scoped mandate (evidence assessment, dimensional gating, final review).

Capability: The improved system can classify complex, multi-faceted artifacts (e.g., research papers, project reports) with higher accuracy and less bias, because it separates distinct judgment types (environmental evidence vs. dimensional gaps) that a single model tends to conflate.

Improvement: Implement a two-tier decision structure: independent agents produce verdicts → deterministic minimum-rule aggregation → a chief agent that can only confirm or lower, never raise.

Improvement: Train/instruct agents to treat missing information as neutral (no cap) rather than as evidence of deficiency, requiring affirmative textual evidence for any negative judgment.

Improvement: Classify based on the most advanced demonstrated evaluation environment (not claimed intent), and anchor to the level of generality the text itself claims.

Improvement: For each of six dimensions (specification, data existence, data quality, data legality, expert knowledge, algorithmic maturity), require the agent to output a cap level (9/6/4) only when accompanied by a verbatim quotation from the source.

Improvement: Implement a policy where malformed or non-responsive agent outputs default to the most permissive value (no cap).

Improvement: Add deterministic rules (e.g., if regulated domain and no certification stated, cap at level 7) that operate independently of LLM judgment.

Improvement: Offer two operational modes: full panel (6x cost, audited, for batch label production) and monolithic-with-deterministic-check (low-cost, for interactive use).


What the improved AI system can do overall: It can assess the maturity of AI research/technology descriptions from natural language alone, producing conservative, auditable, and explainable readiness levels (1–9) that avoid the systematic inflation and central-tendency biases of single-model classifiers, while remaining robust to incomplete input and partial component failures.

Sources

Related papers