Meta-TTL: Meta-Learning Self-Improvement Policies for Language Agents
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Meta-TTL: Meta-Learning Self-Improvement Policies for Language Agents".
Tom: Test-Time Learning (TTL) enables language agents to iteratively refine their performance through repeated interactions at inference time, and this paper introduces META-TTL,
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, we're diving into a really interesting paper today called "Meta-TTL: Meta-Learning Self-Improvement Policies for Language Agents." It sounds super technical, but the main idea is about how AI agents can get better just by interacting with things repeatedly during testing.
Jane: That’s right, Tom, and it suggests we move away from having humans manually set these adaptation rules because the system itself can learn what works best through experience.
Lu: I think this points toward a much more flexible kind of AI where the learning mechanism isn't hardcoded but is discovered through interaction with task environments.
Meng: From an engineering standpoint, it’s exciting to think about systems that can dynamically adjust their behavior mid-run instead of just following a pre-set script.
Lalam: I see this as a potential culture shift; instead of relying on static instructions, we could build agents that organically refine their operational style based on real-world performance.
The paper's summary: Tom: So, what does this paper actually propose? Basically, it introduces META-TTL as a framework that treats test-time learning as a bi-level optimization problem to discover the best adaptation policies for these language agents.
Jane: It sets up an inner TTL loop where the agent tries out different adaptation policies across episodes and measures how well those policies help the agent fix errors from one episode to the next.
Lu: The outer loop is what's really clever; it uses evolutionary search over a diverse set of training tasks to optimize that adaptation policy, meaning it learns from a distribution of environments rather than just one.
Meng: So, instead of fixing one rule for how an agent should adapt, this system tries to find the optimal rule across many different scenarios.
Lalam: That idea of learning the adaptation policy itself is really compelling because it means the AI develops its own strategy for improvement over time, which feels much more dynamic than a fixed setup.
The paper's improvements: Tom: The authors suggest that by learning this adaptation policy directly from task environments, we can achieve sustained test-time improvement that existing methods struggle with.
Jane: They argue that hand-engineering these policies based on human intuition is inefficient because the optimal adaptation strategies are specific to the tasks themselves.
Lu: The framework involves two main policies: an actor policy for immediate action and an adaptation policy that updates the actor based on accumulated experience, which is learned across tasks.
Meng: It shifts the focus from just optimizing the agent's initial behavior to optimizing *how* it should learn and adapt its behavior over time.
Lalam: I find the idea of a meta-prompt being the learnable component particularly interesting; it allows us to define *how* an agent should diagnose failures and guide its next steps, which is a sophisticated level of control.
Conclusion: Tom: Wrapping up this discussion on "Meta-TTL: Meta-Learning Self-Improvement Policies for Language Agents," the core implication is that we can move toward agents that continuously refine their own methods during inference time using learned adaptation policies.
Jane: This means we stop relying on static, pre-programmed rules and start having AI systems that iteratively improve their performance as they interact with novel situations.
Lu: It suggests a way for meta-learning to extract genuinely transferable knowledge from task distributions so the agent can adapt efficiently when facing something new.
Meng: Practically speaking, this means we could build agents that are much more resilient and capable of handling out-of-distribution tasks because they've learned robust strategies rather than just memorizing specific solutions.
Lalam: I think the biggest impact here is on developing AI that exhibits sustained self-improvement, which fundamentally alters how we design and deploy complex language models in real-world applications.
National University of Singapore
cs.LG, cs.AI
Submitted: 2026-04-01
Updated: 2026-09-28
Code: https://github.com/zzzlou/meta-ttl
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 89/100
The gist: Test-Time Learning (TTL) enables language agents to iteratively refine their performance through repeated interactions at inference time, and this paper introduces META-TTL, a framework that
Key concepts
- Test-Time Learning (TTL)
- A method where an AI agent refines its behavior repeatedly during actual use (inference) by interacting with the environment. The agent adapts its strategy based on past attempts to correct errors in real-time, rather than learning beforehand.
- Bi-level Optimization
- The framework uses two nested loops: an inner loop for the agent to perform TTL episodes and an outer loop for meta-training. The outer loop searches through many tasks to find the best 'meta-prompt' that guides the inner TTL process toward better adaptation policies.
- Meta-Prompt ($ε$)
- This is the learnable component that defines how an agent adapts. It acts as a high-level instruction set specifying what past experience to focus on, how to diagnose failures, and what kind of guidance to generate for future actions.
- Actor Policy vs. Adaptation Policy
- The actor policy decides immediate actions within a single game episode. The adaptation policy operates at a higher level; after an episode ends, it analyzes the experience and produces an updated actor policy for the next attempt, learning how to improve over time.
Terminology
Summary
Test-Time Learning (TTL) enables language agents to iteratively refine their performance through repeated interactions at inference time, and this paper introduces META-TTL, a framework that formulates discovery of effective adaptation policies as a bi-level optimization problem. This approach learns an adaptation policy from task environments by optimizing it for downstream improvement at test time, suggesting that optimal adaptation policies should be learned from task environments rather than being hand-engineered based on human intuition.
How it works
The framework is structured around a bi-level optimization problem consisting of an inner loop and an outer meta-training loop. The inner TTL loop executes the standard TTL process where an LLM agent interacts with the environment over a series of episodes and adapts based on prior attempts, measuring how well a candidate adaptation policy helps the agent correct errors across sequential episodes. Guided by this performance, the outer loop employs evolutionary search over a diverse distribution of training tasks to continually optimize the adaptation policy.
Learned Adaptation Policies for Language Agents
The core mechanism involves two distinct policies: an actor policy that determines behavior within a single episode, and an adaptation policy that updates the actor policy based on accumulated experience. The goal is to learn this adaptation policy, which maps past experience to future behavioral improvement, rather than relying on fixed or hand-crafted rules. This learning is instantiated in prompt space, where the actor's behavior is mediated through system prompt rewriting rather than parameter updates. Specifically:
-
The actor policy determines behavior within an episode:
The actor policy π determines behavior within a single episode, selecting actions given the current observation.
-
The adaptation policy operates at a higher level:
the adaptation policy f operates at a higher level: after each episode, it observes the accumulated experience and produces an updated actor policy for the next attempt.
-
The learnable component is the meta-prompt ϕ, which fully specifies the adaptation policy:
The meta-prompt ϕ fully specifies the adaptation policy: it determines what aspects of past experience the meta-agent attends to, how it diagnoses failures, and what form of guidance it produces.
Reflective Meta-Training
The objective of meta-training is to find a meta-prompt ϕ∗ that maximizes expected TTL performance on the training tasks: ϕ∗ = argmax ϕ Eg∼Dtrain W-AUC(ξϕg).
This is achieved through reflective prompt evolution, similar to GEPA, where candidate meta-prompts are proposed through reflection and selected by session-level W-AUC. The process involves:
-
Proposal and Local Validation: A parent meta-prompt is sampled, run on a training task (TTL session), and a candidate is proposed based on reflection. This candidate undergoes local validation on the same task to ensure improvement before proceeding to global validation.
-
Global Validation: Candidates that pass local validation are evaluated across all validation tasks in Dval, and the pool of experts is updated if a new best score is achieved for any task.
-
Expert Selection: After training, a single optimized meta-prompt ϕ∗ is selected from the expert pool by choosing the expert with the highest average validation score, often using per-game z-score normalization to select more uniformly strong candidates.
Emergent Adaptation Policies
The optimized meta-prompt (ϕ∗) exhibits several qualitatively distinct features that emerge through evolutionary optimization:
-
Mandatory structured output: The meta-prompt specifies six required output sections, including
diagnosis of what happened,
durable game facts,
and arecommended route with save points.
This structure forces the meta-agent to separate diagnosis, fact extraction, planning, and scripting. -
Explicit credit assignment protocol: The meta-prompt requires itemizing which actions scored points and how to reproduce them, which actions caused death or threats, which wasted turns (dead ends), and which blocked progress (locked doors).
-
Grounded fact accumulation: A
Game facts to remember
section must record map links, required triggers, working command syntax, and non-working verbs the parser rejected. This is constrained to be evidenced by the most recent episode log to prevent hallucination. -
Exploration management: The meta-prompt enforces a disciplined exploration policy:
at most one new experiment per episode, always under a save/restore point,
with an explicit fallback if two attempts at the same approach fail. -
Concrete action scripts: The meta-prompt requires a
15–25 command opening script that reproduces known scoring actions quickly before attempting new objectives.
-
Conditional fact banks: Game-specific knowledge is included but activated only when the game identity is confirmed from the episode log via a
CRITICAL ADAPTATION RULE
ensuring irrelevant fact banks are ignored.
Evaluation
META-TTL was evaluated on Jericho, WebArena-Lite, and τ2-bench across both in-distribution (ID) and out-of-distribution (OOD) settings.
Improvements for AI systems
Here are specific, actionable improvements for AI systems based on the META-TTL framework:
-
Improve on-the-fly performance in novel or unfamiliar environments by implementing a bi-level meta-learning loop that learns an adaptation policy (encoded as a natural language meta-prompt) rather than relying on fixed, hand-crafted heuristics.
-
Enable language agents to achieve sustained, iterative self-improvement across sequential episodes during inference time by dynamically updating their behavioral adaptation policy based on accumulated experience from previous failures and successes.
-
Develop a
Meta-Agent
capable of acting as a learned learning algorithm that maps past trajectories (experience history) to future behavioral improvement by optimizing its system prompt against a distribution of training tasks. -
Create agents that demonstrate strong generalization capabilities to out-of-distribution (OOD) tasks, where the learned adaptation policy encodes transferable strategies that allow them to solve unseen environments effectively, as evidenced by gains on benchmarks like Jericho and τ2-bench OOD settings.
-
Implement a robust
Expert Selection
mechanism using per-game z-score normalization during meta-training to ensure the agent learns task-agnostic adaptation strategies rather than overfitting to easily improved or single tasks, leading to more reliable generalization. -
Produce highly structured and actionable feedback by forcing the adaptation policy (meta-prompt) to follow a mandatory output format that explicitly includes diagnosis, durable fact extraction, next-episode priorities, and concrete command scripts tailored for the current game context.
-
Enhance agent planning by requiring the learned policy to enforce disciplined exploration management (e.g., one new branch per episode under save/restore points) and explicit credit assignment protocols that itemize which specific actions caused progress versus those that were wasted turns or dead ends.
-
Create an optimization strategy where the adaptation mechanism itself is treated as a learnable object, allowing researchers to compare the efficacy of optimizing the adaptation policy versus optimizing the primary task-solving actor, demonstrating superior gains in W-AUC and overall performance across diverse benchmarks.
Abstract
Test-Time Learning (TTL) enables language agents to iteratively refine their performance through repeated interactions with the environment at inference time. At the core of TTL is a self-improvement policy that updates the actor policy based on experience from previous episodes, thereby improving future behavior. Existing methods rely on hand-crafted self-improvement rather than optimizing them for downstream improvement. We argue that optimal self-improvement policies should be learned from task environments, not hand-engineered based on human intuition. To achieve this, we introduce Meta-TTL, a framework that formulates the discovery of effective self-improvement policies as a bi-level optimization problem. Within this framework, the inner loop executes the standard TTL process, measuring how effectively a candidate self-improvement policy helps an agent correct errors across sequential episodes. Guided by the agent's performance, the outer loop performs reflective meta-training across diverse training tasks, using a balanced improvement score (BIS) to balance task contributions during candidate selection. We evaluate Meta-TTL on Jericho, WebArena-Lite, and-bench across both in-distribution (ID) and out-of-distribution (OOD) settings. Meta-TTL consistently outperforms existing baselines, improving TTL over the strongest baseline by up to 23% on ID tasks and 27% on OOD tasks. These results suggest that the optimized self-improvement policy encodes transferable meta-strategies that generalize beyond the training task distribution.
Sources
- Self-Improving LLM Agents at Test-Time
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Language Models are Few-Shot Learners
- Learning to Self-Evolve
- A Survey on In-context Learning
- A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
- Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
- Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
- A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence
- EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers
- MetaReflection: Learning Instructions for Language Agents using Past Reflections
- Interactive Fiction Games: A Colossal Adventure
- EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems
- Adaptation of Agentic AI: A Survey of Post-Training, Memory, and Skills
- Meta-RL Induces Exploration in Language Agents
- EvoX: Meta-Evolution for Automated Discovery
- Self-Refine: Iterative Refinement with Self-Feedback
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Reflexion: Language Agents with Verbal Reinforcement Learning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks