Machine Learning Systems: A Survey from a Data-Oriented Perspective

arXiv:2302.04810 · cs.SE, cs.AI, cs.LG · Submitted 2025-07-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Machine Learning Systems: A Survey from a Data-Oriented Perspective".

Jane: The paper was written by Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux et al. from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: We’ve seen that many systems are still stuck in traditional service-oriented architectures, but the authors of "Machine Learning Systems: A Survey from a Data-Oriented Perspective" have some very clear recommendations for overcoming those hurdles.

Jane: They aren't just suggesting one thing; they're offering several practical strategies to move towards a more data-centric design. The advice is essentially about shifting focus from how services talk to what the data itself does.

Lu: I particularly like the suggestions around using dataflow and publish/subscribe patterns, which naturally create those loose couplings that make systems more flexible and scalable.

Meng: And it’s not just for big systems; even if your application is smaller, adopting these principles helps manage resources efficiently, which is a huge win for any startup looking to keep costs down.

Lalam: I think the advice also emphasizes building trust into the architecture by making data management visible and auditable, which builds a more reliable digital culture.

Tom: The paper identifies that "data coupling" is a key mechanism here, meaning components interact by reading from and writing to shared data mediums instead of direct calls. How does that change the way we build things?

Jane: It replaces those traditional API calls with something much looser, which means when components fail or when the system needs to handle massive amounts of information, it doesn's not everything collapsing into one rigid point.

Lu: That loose coupling is what allows for asynchronous interactions by design, making sure that a delay in one process doesn't halt the entire workflow.

Meng: When we have real-time systems, this is vital; you can’t afford to wait for a synchronous call if the data flow needs to keep moving continuously.

Lalam: It ensures our digital processes are more resilient and that our software supports continuous streams of information rather than just waiting for a single request.

Tom: We're really looking at how this shift from using data mediums is being implemented in practice, which will lead us right into the next section where we analyze the specific sub-principles.

Summary: Tom: The survey's summary of "Machine Learning Systems: A Survey from a Data-Oriented Perspective" gives us a clear picture of what works and what doesn't, especially when looking at those three core principles. It’s important to understand the strengths and weaknesses here.

Jane: It shows that while data-driven systems are unavoidable—they all use ML—only about seventy-five percent partially or fully adopt a shared data model, which is a big step but not perfect.

Lu: The findings on decentralization are equally telling; most systems still rely on centralized storage, which speaks to the current inertia in the industry toward familiar cloud setups.

Meng: But even though that centralization persists, the authors show how local data chunks can be used to improve things like privacy and fault tolerance in certain scenarios.

Lalam: This summary really highlights that our systems need to be more adaptable, and that trusting the data flow itself is a much better way to build robust software.

Tom: Let’s look at "Openness" as described in the paper—how do we move from static systems to open ones?

Jane: The paper notes that almost seventy percent of reviewed systems use autonomous entities, which is great for handling unpredictable environments where human intervention isn's possible.

Lu: And even better, about a quarter of those papers use asynchronous communication protocols like publish/subscribe, which makes the system much more dynamic and capable of handling large volumes.

Meng: This move to open systems is critical when we think about things like IoT devices joining the network or adapting to changing conditions in manufacturing plants.

Lalam: The goal isn't just for the components to exist; they have to be able to interact freely and autonomously, ensuring our digital infrastructure supports growth.

Tom: We've covered the core findings and how they lead into practical application; now we move toward a final wrap-up of what this all means for the world.

Conclusion: Tom: So, we’ve explored "Machine Learning Systems: A Survey from a Data-Oriented Perspective," and it's clear that while data-driven systems are the norm, fully embracing a "Data-Oriented" architecture remains quite rare.

Jane: We're seeing that the industry is still leaning heavily on centralized cloud solutions, which makes adopting decentralized architectures much harder to implement right now.

Lu: But I think this research shows us a pathway to create systems that are truly fault-tolerant and maintain their integrity over time, even in complex real-world applications.

Meng: It’s a call for developers to start thinking about data as the central component of their design, not just another service in the stack.

Lalam: We need to use this framework to build software that is not only efficient but also trustworthy and inherently adaptable for a global audience.

Tom: This survey provides a solid foundation for future work, highlighting where we are and what challenges remain.

Jane: I hope this discussion helps listeners understand the importance of moving towards data-centric designs in the real-world applications of AI.

Lu: We need to keep pushing our boundaries with these principles, making sure that the data flow is a critical part of our engineering DNA.

Meng: It' time for us to apply these architectural lessons across more practical projects and start seeing more decentralized solutions in production.

Lalam: Let's hope this "Machine Learning Systems: A Survey from a Data-Oriented Perspective" gives us the clarity we need to build better, smarter, digital futures.

Conclusion: Tom: So, we’ve spent time looking at this survey, and it’s clear that while data-driven systems are unavoidable in modern applications, fully embracing a Data-Oriented Architecture remains quite rare across the board.

Jane: I agree; it really highlights how much of the industry is still leaning on traditional service models despite the challenges those creates with privacy and latency.

Meng: And as a practical matter, it shows that unless we address those issues, building truly resilient ML-based systems will be incredibly difficult for startups trying to scale efficiently.

Lu: It’s fascinating how much more potential there's in designing systems around the flow of data rather than just looking at the current architectures.

Lalam: I think this entire discussion is about building a digital culture where trust and transparency aren't just features, but fundamental requirements for how we operate.

Tom: That ties back to the core findings of "Machine Learning Systems: A Survey from a Data-Oriented Perspective," which shows that data coupling is key to making things work.

Jane: It’s encouraging to see that even though adoption is low, there are paths forward toward decentralization and flexibility for real-time needs.

Meng: I hope the advice in this paper becomes a starting point for more tool development, because right now it's mostly theoretical guidance.

Lu: I wonder how far we can push these decentralized concepts when we consider combining them with complex edge computing environments?

Lalam: The vision is clear; we' are moving toward systems that not only work but also autonomously adapt to ensure better service for everyone.

Tom: It’s a powerful roadmap for the future, and I think it gives us a lot to think about as we move into our next topic.

Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar et al.

cs.SE, cs.AI, cs.LG

Submitted: 2025-07-16

Updated: 2026-08-20

Comments: Under review CSUR

DOI: 10.1145/3769292

Code: https://github.com/cabrerac/semi-automatic-literature-survey

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 84/100

The gist: The survey, "Machine Learning Systems: A Survey from a Data-Oriented Perspective," details the evolving architectural patterns necessary for deploying robust machine learning models in complex,

Key concepts

Data Coupling
This is a key mechanism where system components interact by reading from and writing to shared data storage mediums instead of making direct API calls. This loose coupling ensures that if one component fails, the entire system does not collapse when handling massive amounts of information.
Data-Oriented Architecture
This design approach shifts the focus from how services communicate to what the data itself does. It prioritizes data as the central component, aiming to build more robust and adaptable software by trusting the flow of information within a system.
Publish/Subscribe Patterns
These patterns allow for asynchronous communication where components interact without direct calls. One entity publishes data that other interested components subscribe to receive it, creating loose couplings that make systems more flexible and scalable.

Terminology

Summary

The survey, Machine Learning Systems: A Survey from a Data-Oriented Perspective, details the evolving architectural patterns necessary for deploying robust machine learning models in complex, real-world environments. The central thesis revolves around shifting focus from purely algorithmic performance to the underlying data flow and system reliability required for operational ML systems.

A significant portion of the survey addresses architectural paradigms, emphasizing that modern systems must adopt a data-centric view rather than solely a service-oriented one. This is highlighted by concepts such as The Data Dichotomy: Rethinking the Way We Treat Data and Services [131], advocating for architectures that prioritize data lineage and flow. Specifically, the paper explores dataflow management in the Internet of Things: Sensing, control, and security [142], illustrating how data must be managed across distributed edge environments. Furthermore, it reviews methodologies for ensuring accountability through Decision provenance: Harnessing data flow for accountable systems [129].

The survey provides a comprehensive look at implementing ML within dynamic computing infrastructures. It discusses the need for adaptive and distributed solutions, referencing Dynamic Service Placement in Multi-Access Edge Computing [133] and the principles of Software engineering of self-adaptive systems [143]. For system design, it reviews established patterns, including Architectural patterns for microservices: A systematic mapping study [134], and the development of A data-oriented architecture for loosely coupled real-time information systems [139].

In terms of ML implementation and reliability, the paper covers both advanced techniques and systematic review areas. It reviews Machine learning architecture and design patterns [141] and provides an overview of automated methods, such as Automated machine learning: Review of the state-of-the-art [140]. Practical applications are detailed across various domains, including healthcare, exemplified by An Early Warning System for Hemodialysis Complications Utilizing Transfer Learning from HD IoT Dataset [128], and industrial settings with methods like Machine learning based acoustic defect detection in factory automation [148].

The paper also dedicates attention to critical non-functional requirements. Security and privacy are paramount concerns, necessitating specialized techniques such as those outlined in Privacy heroes need data disguises [136] and scalable detection mechanisms like those used for Harnessing the Nature of Spam in Scalable Online Social Spam Detection [145].

In conclusion, the survey posits that successful ML systems require a holistic integration of robust dataflow management, adaptive distributed architectures (such as those found in Fog or Edge Computing environments), and rigorous consideration of provenance, privacy, and reliability to transition from theoretical models to dependable operational infrastructure.

Improvements for AI systems

Based on the principles outlined in this survey, we implement a shift from traditional Service-Oriented Architectures (SOA) and Microservices to a Data-Oriented Architecture (DOA). This approach fundamentally changes how components interact, moving away from ephemeral API calls to persistent data flow.

A. Implementation of Data Coupling (Data as First Class Citizen):

  • Mechanism: Instead of components interacting via synchronous Remote Procedure Calls (RPC) or REST APIs, we establish a Shared Data Model. Components act as producers and consumers, reading from and writing to this model.

  • Specific Technologies: We utilize durable data mediums such as Apache Kafka streams, RabbitMQ message queues, or structured storage like HDF5/HBase (for batch processing) to serve as the shared state.

  • Resulting Engineering Advantage: This creates a permanent, auditable record of the system's entire operational history (the Data Flow Design).

B. Prioritization of Decentralization and Local-First Processing:

  • Mechanism: We replace centralized cloud dependence with distributed, autonomous entities that store local data chunks (Local First) and interact via peer-to-peer protocols.

  • Specific Technologies: Utilizing distributed computing frameworks like Apache Spark (for parallel processing) or Distributed Hash Tables (DHT) to manage data partitioning across multiple nodes.

  • Resulting Engineering Advantage: This mitigates the impact of physical distance, drastically reducing end-to-end latency and eliminating single points of failure.

C. Adoption of Asynchronous Openness:

  • Mechanism: Components are designed as asynchronous entities. They do not wait for synchronous responses but subscribe to relevant data producers via a common Message Exchange Protocol.

  • Specific Technologies: Implementing protocols like MQTT or leveraging the publish/subscribe patterns inherent in platforms such as Kafka.

  • Resulting Engineering Advantage: This achieves maximum decoupling, ensuring that a failure in one component does not halt the entire system, significantly increasing fault tolerance and scalability.


By adopting these DOA principles, the resulting AI system gains specific functionalities that were impossible or highly inefficient under traditional architectures:

1. Real-Time Predictive Maintenance and Fault Detection:

  • The system can continuously monitor its own operational state by querying the shared data model. It detects subtle shifts in data distribution (e.g, sensor drift, hardware degradation) and autonomously triggers adaptive maintenance routines before failures occur.

2. Edge-Based Privacy and Security Compliance:

  • Since local data chunks are processed locally (Decentralization), sensitive user or industrial data never needs to be aggregated into a centralized cloud repository. This ensures strict data ownership and compliance with privacy regulations (e.g., GDPR), as the processing happens near the data source.

3. Autonomous Self-Adaptation:

  • The system can analyze its historical state (via Data Coupling) to identify performance bottlenecks or model drift. It can then trigger re-training or hyperparameter adjustments in a non-blocking, asynchronous manner, maintaining optimal performance without human intervention.

4. Scalable Big Data Processing:

  • The system handles massive, continuous data streams (Big Data/Data as First Class Citizen) by partitioning the load across distributed nodes (Decentralization), allowing it to scale horizontally to meet unpredictable real-world demands without requiring complex re-architecture or service orchestration.

Sources

Related papers