Position: It is Time to Virtualize Foundation Models with a Self-evolving Operating System Layer
cs.AI, cs.LG, cs.MA, cs.OS
Submitted: 2026-09-16
Updated: 2026-09-16
Journal ref: ICML 2026 Position Paper Track
Code: https://github.com/tursodatabase/agentfs
Project page: https://nexi-lab.github.io/nexus
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: AI applications have shifted from single, monolithic foundation models (FM) to compound agentic systems.
Terminology
Abstract
AI applications have shifted from single, monolithic foundation models (FM) to compound agentic systems. Yet today's stacks remain fragmented: even as protocols (e.g., MCP, A2A) ease tool/agent connectivity, each framework embeds an implicit runtime for state, memory, budgets, and guardrails, making behavior non-portable and governance brittle. It mirrors computing before operating systems, when every program re-implemented basic services. This position paper argues that the field now needs a Foundation Model Operating System (FMOS) -- a system layer that virtualizes FM interactions analogous to how virtual machines abstract physical hardware, giving applications the illusion of dedicated, trustworthy FM instances with effectively unbounded capabilities. Internally, the FMOS orchestrates knowledge across memory tiers, model selection and resource allocation, and verification and policy enforcement. Like the human brain switching between fast intuition and slow deliberation, the FMOS learns when to intervene and when to let inference proceed directly and continuously adapting its policies based on operational experience.
Sources
- Representation Engineering for Large-Language Models: Survey and Research Challenges
- Reasoning Language Models: A Blueprint
- InferCept: Efficient Intercept Support for Augmented Large Language Model Inference
- DBOS: A Proposal for a Data-Centric Operating System
- L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
- Training Large Language Models to Reason in a Continuous Latent Space
- Forecast2Anomaly (F2A): Adapting Multivariate Time Series Foundation Models for Anomaly Prediction
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
- EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems
- What Challenges Do Developers Face in AI Agent Systems? An Empirical Study on Stack Overflow & GitHub Issues
- When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs
- Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems
- A Blueprint Architecture of Compound AI Systems for Enterprise
- AIOS: LLM Agent Operating System
- MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling
- s1: Simple test-time scaling
- Large Language Model Routing with Benchmark Datasets
- MemGPT: Towards LLMs as Operating Systems
- vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection