Octopus v2: On-device language model for super agent
cs.CL
Submitted: 2024-04-02
Updated: 2026-09-11
Code: https://github.com/ggerganov/llama.cpp
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Language models have shown effectiveness in a variety of software applications, particularly in tasks related to automatic workflow.
Terminology
Abstract
Language models have shown effectiveness in a variety of software applications, particularly in tasks related to automatic workflow. These models possess the crucial ability to call functions, which is essential in creating AI agents. Despite the high performance of large-scale language models in cloud environments, they are often associated with concerns over privacy and cost. Current on-device models for function calling face issues with latency and accuracy. Our research presents a new method that empowers an on-device model with 2 billion parameters to surpass the performance of GPT-4 in both accuracy and latency, and decrease the context length by 95%. When compared to Llama-7B with a RAG-based function calling mechanism, our method enhances latency by 35-fold. This method reduces the latency to levels deemed suitable for deployment across a variety of edge devices in production environments, aligning with the performance requisites for real-world applications.
Sources
- GPT-4 Technical Report
- AnyTool: Self-Reflective, Hierarchical Agents for Large-Scale API Calls
- MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
- LoRA: Low-Rank Adaptation of Large Language Models
- Active Retrieval Augmented Generation
- MIMIC-IT: Multi-Modal In-Context Instruction Tuning
- A Survey on Retrieval-Augmented Text Generation
- TaskMatrix.AI: Completing Tasks by Connecting Foundation Models with Millions of APIs
- AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
- Prompt Injection attack against LLM-integrated Applications
- MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
- Generation-Augmented Retrieval for Open-domain Question Answering
- Efficient Estimation of Word Representations in Vector Space
- ART: Automatic multi-step reasoning and tool-use for large language models
- Gorilla: Large Language Model Connected with Massive APIs
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
- TPTU: Large Language Model-based AI Agents for Task Planning and Tool Usage
- Small LLMs Are Weak Tool Learners: A Multi-LLM Agent
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Accelerating LLM Inference with Staged Speculative Decoding
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering