The Systems Engineering Approach in Times of Large Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "The Systems Engineering Approach in Times of Large Language Models".
Jane: The paper was written by Christian Cabrera, Viviana Bastidas, Jennifer Schooling and Neil D. Lawrence from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Summary of Challenges: Tom: So, the authors start by laying out all these challenges, and they're pretty severe when applying LLMs to real-world problems.
Jane: They point out that because LLMs are probabilistic models, they can't be entirely reliable; they can even hallucinate information.
Meng: That lack of certainty is a huge operational problem for any system that relies on those outputs.
Lu: It’s not just the technical failure, though; the paper emphasizes that these issues fundamentally challenge the alignment and reliability of the whole socio-technical system.
Tom: And we've got challenges like black-boxes, which make accountability nearly impossible because users don't understand how decisions are made.
Jane: That's what they mean by "intellectual debt," when you can’t trace why a component produced a certain output.
Meng: From an implementation standpoint, that lack understanding means debugging becomes incredibly difficult.
Lalam: The paper also highlights sustainability concerns, pointing out the immense carbon footprint required to train these large models.
Lu: It’s a huge environmental cost we're ignoring in the pursuit of rapid AI advancement, and that needs to be addressed early in the maintainability phase.
Systems Engineering Approaches: Tom: Now, moving into the core findings of how people are solving these problems, the paper shows that current research prioritizes alignment and reliability immensely.
Jane: About eighty-seven percent of the papers reviewed focused on making sure those systems behave as expected.
Meng: That makes sense; you have to make sure a system does what its requirements state before worrying about anything else.
Lu: The authors found that the most used principles are Systems Views, Top-Down methods, and the Problem-Solving Cycle.
Tom: Why do you think those three specific approaches are so dominant?
Jane: They seem to give us a way to look at the problem from different perspectives while also ensuring we aren't just solving one side of the problem.
Meng: It’s about decomposing the large issue into manageable parts, which is what that Top-Down principle allows us to do in practice.
Lalam: I see how System Views helps define a path for accountability, especially when we look at things like healthcare systems where multiple actors are involved.
Lu: The paper also looked at how researchers are tackling interpretability and accountability using different lenses.
Jane: Although less common than reliability work, those solutions often rely heavily on the Systems View principle as well.
Tom: And they found that for maintainability and sustainability, some approaches lean toward cost-benefit analysis or creating self-maintaining systems.
Meng: Self-maintenance is critical if we want to reduce that operational overhead and ensure longevity in the field.
Gaps and Future Work: Tom: The paper gives us a lot of concrete examples, but it also points out some real gaps in the current work.
Jane: They suggest that relying only on traditional top-down methods isn't enough because AI is so complex and fluid.
Lu: We need to be much more creative with how we define requirements if we want to move beyond just having stakeholders inside one organization.
Meng: The practical challenge is operationalizing high-level concepts like "sustainability" into the actual code or metrics.
Lalam: How do you translate a concept like "truthfulness" into a measurable engineering requirement? That seems hard to pin down.
Lu: It requires dynamic tools and flexible architectures to manage the interaction between actors and data sources throughout the system’s lifecycle, not just static plans.
Tom: And there's also the issue of how fast technology is changing; everything is evolving so quickly.
Jane: The current methodologies are often static and rely on prior knowledge, which isn't helpful when new learning models emerge every few months.
Meng: We need something that can dynamically adapt to a shifting technological landscape, not just a fixed process model.
Lalam: This feels like the next frontier is moving from static design toward dynamic adaptation across different levels of the system.
Conclusion: Tom: So, bringing it all back together, "Systems engineering for artificial intelligence-based systems: A review in time" really gives us a roadmap.
Jane: It shows that by prioritizing context and addressing things like alignment and reliability, we have a solid foundation for developing AI systems.
Meng: It confirms that the Systems Engineering approach is essential for building dependable AI solutions in critical industries.
Lu: The paper lays out how to build an ecosystem of architectural patterns and design artifacts needed for sustainable deployment.
Lalam: My takeaway is that this provides the vision we need, driving our work toward a culture where AI is integrated thoughtfully and responsibly.
Tom: It’s been a fantastic discussion on how to make sure these powerful LLMs work together with sound engineering principles.
Jane: We'll be looking forward to seeing how these ideas help in the next segment of our show.
Christian Cabrera, Viviana Bastidas, Jennifer Schooling, Neil D. Lawrence
cs.AI, cs.CY, cs.SE
Submitted: 2024-11-13
Updated: 2026-08-20
Comments: This paper has been accepted for the upcoming 58th Hawaii International Conference on System Sciences (HICSS-58)
Code: https://github.com/cabrerac/semi-automatic-literature-survey
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: The full text of "Llinas, J., Fouad, H., & Mittu, R.
Key concepts
- Black-boxes
- This refers to AI systems where users cannot understand how decisions are made. This lack of transparency makes accountability nearly impossible because the process that produced a specific output is unknown.
- Intellectual Debt
- This term describes the difficulty in tracing why a component of an AI system produced a certain output. When this debt exists, debugging and understanding the system's failures becomes incredibly difficult.
- Systems Views
- A core principle highlighted in the paper, Systems Views allow for looking at a problem from different perspectives. This helps ensure that solutions are comprehensive and do not only address one isolated aspect of the overall system.
- Top-Down methods
- This is a dominant principle used in solving complex AI problems. It allows for decomposing a large, overarching issue into smaller, more manageable parts for practical implementation and development.
Terminology
Summary
The full text of "Llinas, J., Fouad, H., & Mittu, R. (2021). Systems engineering for artificial intelligence-based systems: A review in time" is required to extract the summary. The provided input contains only the citation metadata and does not include the body or abstract of the scientific paper.
Improvements for AI systems
(Initial Disclaimer: Due to the breadth of the provided literature—encompassing systems engineering, ethics, safety assurance, and model architectures—the improvements require a holistic overhaul across the entire AI lifecycle (MLOps). The following enhancements detail a comprehensive framework for building a robust, trustworthy, and certifiable AI system.)
The system must move beyond treating ML models as black boxes and instead be designed as a modular, composable system.
-
Improvement: Implement a formal System Architecture Decomposition layer. The AI function must be broken down into discrete, verifiable modules (e.g., Data Ingestion Module to Feature Extraction Module to Core Prediction Module to Decision Synthesis Module).
-
Technical Enhancement: Enforce Formal Verification at Module Interfaces. Every input/output contract between these modules must be rigorously defined and verified against formal methods (e.g., temporal logic or pre/post-conditions) to guarantee type safety, range constraints, and logical consistency regardless of the underlying model's complexity.
-
Impact: Eliminates cascading failures and ensures that a failure in one module does not propagate unpredictable errors throughout the entire system.
The system must be designed to detect and mitigate operational drift, data corruption, or adversarial attacks in real-time.
-
Improvement: Integrate a Runtime Monitoring and Anomaly Detection Layer. This layer operates independently of the core AI prediction engine.
-
Technical Enhancement: Implement three simultaneous monitoring checks:
-
Input Drift Detection: Continuously compare incoming live data distributions (P live) against the established training distribution (P train). If statistical divergence (e.g., using Maximum Mean Discrepancy, MMD) exceeds a predefined threshold, the system must flag an alert and revert to a defined safe state (e.g., defaulting to human-in-the-loop review).
-
Adversarial Robustness Testing: Incorporate real-time checks against known adversarial perturbation vectors (-norms) during inference.
-
Output Plausibility Checking: Implement domain-specific guardrails (e.g., if the system predicts a negative patient outcome, but all input vital signs are optimal, the output is flagged as implausible).
- Impact: Guarantees that the system operates within its certified performance envelope and prevents catastrophic failures due to real-world data shifts or malicious attacks.
The system must provide transparent justifications for every high-stakes decision, satisfying regulatory requirements and building user trust.
-
Improvement: Implement a mandatory Multi-Layered Explanation Generation Module. The system cannot simply output a score; it must output an explainable decision package.
-
Technical Enhancement: This module must generate three distinct levels of explanation:
-
Local Explanation (Feature Attribution): For any given prediction, calculate and display the precise contribution of every input feature (e.g., using SHAP values or LIME) to the final score, allowing the user to see why a specific decision was made.
-
Global Explanation (Model Behavior): Provide high-level documentation detailing how model changes or data shifts affect overall system performance, satisfying auditors and regulators.
-
Counterfactual Explanation: For any negative outcome (e.g.,
denied loan
), the system must generate the minimal change required in the input features to achieve a desired positive outcome (e.g.,If your income were 15% higher, the loan would have been approved
).
- Impact: Provides full traceability and accountability, transforming opaque predictions into auditable, actionable intelligence.
The improved system is not merely an algorithm; it is a Certified Decision Support Platform that can:
-
Operate with Guaranteed Boundaries: It autonomously determines if the input data falls outside its certified operational domain (P live not equal to P train). If so, it immediately ceases automated action and requires human review, preventing catastrophic misclassification based on novel or corrupted inputs.
-
Justify Every Decision: It provides a three-tiered explanation package (Feature Attribution, Global Context, and Counterfactuals) for every output. This means that instead of receiving a diagnosis or recommendation, the user receives:
Diagnosis X is recommended because [Feature A] contributed 40% and [Feature B] contributed 35%. If Feature C were improved by Y amount, the diagnosis would shift to Z.
-
Support Continuous Regulatory Compliance: Because its architecture is modular and its performance metrics are continuously monitored against defined safety thresholds, the system provides an immutable audit log that proves when, how, and why it made a decision—a critical capability for high-stakes regulated industries (e.g., medicine, finance, autonomous transport).
Sources
- Machine Learning Systems: A Survey from a Data-Oriented Perspective
- Generative AI and Process Systems Engineering: The Next Frontier
- Five Ps: Leverage Zones Towards Responsible AI
- Requirements are All You Need: The Final Frontier for End-User Software Engineering
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection