A Taxonomy of Architecture Options for Foundation Model-based Agents: Analysis and Decision Model
cs.SE, cs.AI
Submitted: 2024-08-06
Updated: 2026-09-09
Comments: Accepted
Code: https://github.com/yoheinakajima/babyagi
License: http://creativecommons.org/licenses/by/4.0/
The gist: The rapid advancement of AI technology has led to widespread applications of agent systems across various domains.
Terminology
Abstract
The rapid advancement of AI technology has led to widespread applications of agent systems across various domains. However, the need for detailed architecture design poses significant challenges in designing and operating these systems. This paper introduces a taxonomy focused on the architectures of foundation-model-based agents, addressing critical aspects such as functional capabilities and non-functional qualities. We also discuss the operations involved in both design-time and run-time phases, providing a comprehensive view of architectural design and operational characteristics. By unifying and detailing these classifications, our taxonomy aims to improve the design of foundation-model-based agents. Additionally, the paper establishes a decision model that guides critical design and runtime decisions, offering a structured approach to enhance the development of foundation-model-based agents. Our contributions include providing a structured architecture design option and guiding the development process of foundation-model-based agents, thereby addressing current fragmentation in the field.
Sources
- The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey
- GPT-4 Technical Report
- Dense Passage Retrieval for Open-Domain Question Answering
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
- PaLM 2 Technical Report
- LLaMA: Open and Efficient Foundation Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Automating Customer Service using LangChain: Building custom open-source GPT Chatbot for organizations
- Enhancing Trust in LLM-Based AI Automation Agents: New Considerations and Future Challenges
- The Prompt Report: A Systematic Survey of Prompt Engineering Techniques
- Large Language Model based Multi-Agents: A Survey of Progress and Challenges
- The Rise and Potential of Large Language Model Based Agents: A Survey
- A Review of Prominent Paradigms for LLM-Based Agents: Tool Use (Including RAG), Planning, and Feedback Learning
- Towards autonomous system: flexible modular production system enhanced with large language model agents
- Unleashing the Emergent Cognitive Synergy in Large Language Models: A Task-Solving Agent through Multi-Persona Self-Collaboration
- BloombergGPT: A Large Language Model for Finance
- GAIA: a benchmark for General AI Assistants
- OS-Copilot: Towards Generalist Computer Agents with Self-Improvement
- pFedMoE: Data-Level Personalization with Mixture of Experts for Model-Heterogeneous Personalized Federated Learning
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties