Grounding LTL Tasks in Sub-Symbolic RL Environments for Zero-Shot Generalization

arXiv:2602.09761 · cs.LG, cs.AI · Submitted 2026-08-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Grounding LTL Tasks in Sub-Symbolic RL Environments for Zero-Shot Generalization".

Jane: The paper was written by Matteo Pannacci, Andrea Fanti, Elena Umili and Roberto Capobianco from Sapienza University of Rome and Sony AI.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the channel, folks. Today we're digging into a fresh arXiv paper that's got a mouthful of a title — "Grounding LTL Tasks in Sub-Symbolic RL Environments for Zero-Shot Generalization." Jane, what's your first read on that?

Jane: Tom, I love this one because it's tackling a problem that sounds abstract but is actually super practical. The title is basically saying: we want a robot or an agent to follow instructions written in a formal logic called LTL, but it only sees raw images — no one tells it what those images mean.

Tom: Right, so LTL is Linear Temporal Logic — it's a way to write instructions that involve time, like "first get the key, then open the door, and never touch the lava." And the paper's team — Matteo Pannacci, Andrea Fanti, Elena Umili, and Roberto Capobianco — they're asking: can an agent learn to follow those instructions when it has to figure out what "key" and "door" look like all by itself?

Jane: Exactly. And that's the "grounding" part. The agent has to connect the symbols in the formula — key, door, lava — to the actual pixels it sees. Most previous work just assumed you had a perfect mapping already. This paper says, no, let's learn that mapping at the same time we learn the policy.

Tom: And the "zero-shot generalization" part is the kicker. They don't just want the agent to nail the training tasks. They want it to handle brand-new instructions it's never seen before — longer ones, more complex ones — without any retraining.

Jane: Which is a huge deal, because that's what makes an agent actually useful in the real world. You don't want to retrain your robot every time you give it a slightly different chore list.

Tom: So the title is really three big promises in one: formal instructions, learned perception, and generalization to unseen tasks. That's ambitious.

Jane: It is. And the authors actually deliver on it, which we'll get into. But first, let's just sit with how much is being asked here. Most RL agents need a reward signal that tells them exactly what's good. Here, the reward is sparse — you only get a +one when the whole task is done, and a-one if you fail. That's a hard setting.

Tom: And they're doing it in environments where the map is randomized every episode, so the agent can't even memorize where things are. It has to actually understand what it's looking at.

Jane: That's the right kind of hard. Let's talk about how they pulled it off.

Summary: Tom: So we've got the title unpacked. Now let's get into what the paper actually does. Jane, can you walk us through the core idea?

Jane: Sure. The key insight is that they use something called Neural Reward Machines — NRMs for short. Think of it like a state machine that tracks how far along you are in the task. The agent gets a reward only when it completes the whole thing, but the NRM lets the agent learn which observations correspond to which symbols by backpropagating through that structure.

Tom: So instead of needing someone to label every image, the structure of the task itself provides the supervision. That's clever.

Jane: Exactly. The agent sees an image, guesses which symbol is present, and that guess moves the state machine forward or not. When the reward comes, the agent can trace back which guesses were right and which were wrong. It's like learning by doing, with the task logic as the teacher.

Tom: And they combine that with an existing method called LTL2Action, which handles the policy side — figuring out what action to take given the current formula and the current observation. But LTL2Action assumed you knew the labeling function. This paper removes that assumption.

Jane: Right. They also pretrain the part of the network that understands the LTL formulas themselves — the syntax, the temporal operators, all that. They do this on a toy environment called LTLBootcamp, where the agent just has to produce the right sequence of symbols. That pretraining gives the agent a head start on understanding what the formulas mean, before it ever sees the real environment.

Tom: And the results? They test on two environments — a Minecraft-like grid world and a continuous 2D world called FlatWorld. In the Minecraft-like one, their method gets essentially the same success rate as if the agent had been given the true symbol grounding. That's the upper bound, and they're right there.

Jane: Yeah, the numbers are striking. On partially-ordered tasks, their method hits a one point zero average total return, while the previous state-of-the-art baseline — that's Kuo et al. from two thousand twenty — gets around zero point nine five. And on global avoidance tasks, the baseline basically collapses to near zero, while their method stays around zero point nine nine.

Tom: So the baseline just fails on the harder task class, and this method doesn't even break a sweat.

Jane: And the grounder — the thing that maps images to symbols — converges to near-perfect accuracy within about a million frames, which is only five percent of the full training run. So the perception problem gets solved fast, and then the policy can focus on the actual task.

Tom: That's a beautiful result. The hard part wasn't perception — it was figuring out how to learn perception from sparse rewards, and they cracked it.

Jane: They did. And the fact that they can generalize to longer and more complex formulas — up to twelve conjunctions when they only trained on four — that's the zero-shot promise delivered.

Improvements: Tom: So we've covered what they did and the results. Now let's talk about what this actually improves over the existing landscape. Jane, what's the biggest leap here?

Jane: The biggest leap is dropping the assumption that you have a perfect symbol grounding. In the real world, you almost never have that. You have a camera, you have pixels, you have noise. This paper says: we can learn the mapping and the policy together, with the same experience, and still match the performance of someone who cheated and knew the ground truth.

Tom: And that's not just a small convenience — it changes what kinds of environments we can deploy these agents in.

Jane: Exactly. The previous state-of-the-art method for this sub-symbolic setting, Kuo et al., encodes the LTL formula directly into the network structure. It's clever, but it doesn't generalize well. When the tasks get longer or more complex, it falls apart. This method, by separating the formula understanding from the perception, is much more robust.

Tom: Let's bring in Lu and Meng for this. Lu, you're the researcher — what excites you most about this approach?

Lu: I think the most exciting thing is the semi-supervised nature of the grounder training. They're using the structure of multiple tasks to disambiguate what each symbol means. A single task might be ambiguous — you could map the same observations to different symbols and still get the same reward. But when you have many tasks, the ambiguity gets resolved. That's a really elegant use of multi-task learning.

Meng: And from an engineering standpoint, I love that they only store episodes with at least one non-zero reward for grounder training. That's a huge efficiency win. You're not wasting compute on episodes where nothing happened. You're selectively using the most informative data.

Jane: Right, and they also mention that they don't need dense rewards anymore. Previous work with Neural Reward Machines relied on potential-based shaping to get enough signal. This paper shows that the multi-task structure alone provides enough supervision — you just need enough tasks.

Lu: That's a significant theoretical point. It means the reasoning shortcuts problem — where the model finds an unintended mapping that still satisfies the structure — gets mitigated by task diversity. The more tasks you have, the harder it is to cheat.

Meng: And they tested that. The grounder accuracy goes above ninety-five percent in both environments. That's not a fluke; that's a robust learning signal.

Tom: So the improvement isn't just "we made it work a little better." It's "we removed a core assumption and the method still performs at the level of the privileged baseline."

Jane: Exactly. And that's what makes this a genuine step forward for the field, not just an incremental tweak.

Conclusion: Tom: Alright, let's wrap up our discussion of "Grounding LTL Tasks in Sub-Symbolic RL Environments for Zero-Shot Generalization." Jane, what's the one-sentence takeaway you want listeners to remember?

Jane: The takeaway is that you can teach an agent to follow formal temporal instructions from raw images alone, without any hand-labeled symbols, and it can then handle instructions it's never seen before — as long as you give it enough diverse tasks to learn from.

Tom: And that's a big deal because it moves us closer to agents that can be given arbitrary instructions in the real world — robots in warehouses, assistants in homes — without needing someone to manually wire up every concept.

Lu: I'd add that the method is modular. The grounder, the LTL encoder, the policy — they're separate pieces. That means you could swap in a different policy algorithm, or a different grounder architecture, and the overall framework still holds.

Meng: And from a practical standpoint, the pretraining on LTLBootcamp is cheap and environment-agnostic. You do it once, and it transfers to any environment with the same formula distribution. That's a real engineering win.

Jane: There are still limitations, of course. They didn't crack global avoidance tasks in the continuous FlatWorld environment — all methods failed there. And the symbols are all about location — being at a certain cell or zone. Real-world symbols might be about objects, relationships, actions.

Tom: But the foundation is solid. And the authors themselves point to future work: curriculum learning, reward shaping informed by task progression, and even avoiding the automata construction altogether using fuzzy logic conformance checking.

Lu: That last one is exciting. If you can train the grounder without explicitly building the automata, you could scale to much more complex formulas.

Tom: Well, that's a great note to end on. We'll be keeping an eye on this line of work. Thanks to everyone who tuned in — we'll see you next time with another paper from the arXiv.

Jane: Until then, keep learning, and remember: even a robot needs to know what a door looks like before it can walk through it. Bye, everyone.

Matteo Pannacci, Andrea Fanti, Elena Umili, Roberto Capobianco

Sapienza University of Rome · Sony AI

cs.LG, cs.AI

Submitted: 2026-08-16

Updated: 2026-08-18

Comments: Preprint currently under review

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 69/100

Key concepts

LTL (Linear Temporal Logic)
LTL is a formal logic used to write instructions for an agent. It describes sequences of actions over time. For example, it can specify that an agent must first retrieve a key and then open a door, or ensure it never touches lava.
Grounding
Grounding is the process where the an agent learns to connect abstract symbols (like 'key' or 'door') to their corresponding visual representation (pixels). The system learns this mapping itself, rather than assuming it was pre-labeled by a human.
Zero-Shot Generalization
This is the ability for an AI agent to successfully execute instructions that were not part of its original training set. The system handles complex, unseen commands without needing a new training phase.

Terminology

Summary

Summary

This paper addresses the problem of training a Reinforcement Learning (RL) agent to follow multiple temporally-extended instructions expressed in Linear Temporal Logic (LTL) in sub-symbolic environments, where the mapping between raw observations and the symbols appearing in the formulae (the symbol grounding) is unknown. The authors state: We drop this unrealistic assumption by jointly training a multi-task policy and a symbol grounder with the same experience.

The proposed method builds on Neural Reward Machines (NRMs) to jointly learn a policy and a symbol grounder. The paper notes: NRMs frame the problem as a Semi-Supervised Symbol Grounding (SSSG) task to provide indirect supervision to the symbol grounder based on the structure of the LTL formulae. The system is structured as four main modules: a grounder module mapping raw observations to symbols (Lθ: S → ∆(P)), an environment feature module extracting features from observations (f img θ: S → Rn), an LTL module extracting feature vectors for the original formula and all progression steps (fθLT L: LLTL (P) → Rm), and an RL module implementing the chosen RL algorithm. The pipeline closely resembles that of LTL2Action [Vaezipoor et al., 2021], with one key difference: in our system, all modules are trainable, whereas in LTL2Action the grounder module is known a priori.

The grounder is trained using NRMs, which are described as "a probabilistic extension of Reward Machines (RMs) [Toro Icarte et al., 2022] that model non-Markovian rewards via an automata-based structure, while incorporating uncertainty in transitions, rewards, and symbol grounding." The training uses multiple NRMs—one per training task—that share the same grounder but differ in their Moore Machine. Grounder training consists of two phases: data collection and update. During collection, RL interactions are stored as state and reward sequences, and the grounder is trained by minimizing the cross-entropy between predicted rewards and ground-truth rewards. The paper notes: "While [Umili et al., 2024] addresses this issue using dense, potential-based rewards, in our training setup we successfully eliminate the reliance on dense rewards by exploiting the structure induced by multiple training tasks."

The LTL module uses a Relational Graph Convolutional Network (R-GCN) built from the formula's Abstract Syntax Tree (AST). The authors tested three training schemes and found that using only the bootcamp (LTL-Bootcamp pretraining with a frozen LTL module) was most effective in our sub-symbolic setup, because a fixed LTL module stabilizes learning in the presence of the uncertain grounder.

Experiments were conducted on two environments: a Minecraft-like discrete grid-world environment (7×7 grid with 5 atomic propositions, randomized maps each episode, RGB image observations of size 56×56) and the continuous FlatWorld environment (2D plane with continuous movement actions, colored circular zones as propositions, randomized placement each episode). Tasks included Partially-Ordered Tasks (conjunctions of sequences that can be executed in parallel) and Global Avoidance Tasks (requiring avoidance of a certain atom for the entire task duration).

Results on the Minecraft-like environment show the method shows both a faster and more stable convergence and significantly higher performances with respect to the baseline from [Kuo et al., 2020], both in terms of success rate and steps needed to solve each formula. For Partially-Ordered Tasks, the method achieved an average total return of 1.000 (base), 0.999 (+dep.), and 1.000 (+conj.), compared to the baseline's 0.951, 0.430, and 0.842, and the upper bound's 1.000, 0.999, and 1.000. For Global Avoidance Tasks, the method achieved 0.990, 0.767, and 0.962, compared to the baseline's 0.031, -0.016, and-0.012, and the upper bound's 0.992, 0.745, and 0.964. The grounder module consistently converges within 1 million frames (5% of a full training run).

On FlatWorld, for Partially-Ordered Tasks, the method achieved average total returns of 0.999 (base), 0.991 (+dep.), and 0.930 (+conj.), compared to the baseline's 0.903, 0.805, and 0.138, and the upper bound's 0.999, 0.998, and 0.983. For Global Avoidance Tasks, all methods fail to learn the global avoidance tasks, with the method achieving-0.008, -0.207, and-0.185, compared to the baseline's-0.012, -0.250, and-0.299, and the upper bound's 0.002, -0.175, and-0.152. Despite the policy failure, the grounder still surpasses 95% accuracy.

The main contributions listed are: "a method to jointly learn a multi-task policy and a symbol grounder by exploiting the indirect supervision provided by the environment; an extension of LTL2Action which lifts the assumption of knowing the labeling function; an empirical evaluation of the proposed method against the state-of-the-art baseline and the upper bound of knowing the true labeling function."

The paper concludes: "Our method shows little to no performance loss with respect to the upper bound of having access to the true symbol grounder, both on training formulae and on more complex unseen formulae. Moreover, it significantly improves over the previous state-of-the-art." Future directions mentioned include robustness of the symbol grounder to raw observations, reward shaping informed by task progression, curriculum learning, and using fuzzy logic conformance checking to train the grounder without explicitly constructing automata.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system, along with what the improved system can now do.


Improvement: I will replace the standard assumption of a known labeling function with a trainable neural grounder module that maps raw observations (e.g., images) to probability distributions over atomic propositions. This grounder is trained jointly with the policy, using the same experience buffer, via Neural Reward Machines (NRMs) that provide indirect supervision from sparse rewards.

What the improved AI can do: The agent can now operate in sub-symbolic environments (e.g., raw camera feeds) without any pre-defined mapping from pixels to symbols. It learns to recognize objects, locations, or states (e.g., at the door, holding the egg) purely from task rewards and the structure of LTL formulae, even when the environment layout is randomized each episode.

Improvement: I will use a frozen, pre-trained LTL encoder (based on a Relational Graph Convolutional Network over the formula's Abstract Syntax Tree) that is trained on a simple LTLBootcamp sequence-generation task. This encoder is then combined with the grounder and policy, allowing the system to condition on the progressed formula (via LTL progression) without fine-tuning the encoder during RL.

Improvement: I will implement the paper's key insight: training the grounder on multiple tasks simultaneously provides enough learning signal to avoid the reasoning shortcut problem (where the grounder learns an incorrect but reward-satisfying mapping). I will also filter training episodes to only store those with at least one non-zero reward or where LTL progression reaches a terminal state (⊤ or ⊥), improving data efficiency.

Improvement: I will replace the grounder's categorical output (argmax) with a probabilistic grounding, and use the NRM's differentiable transition to propagate gradients through time. The policy is conditioned on the concatenation of the environment feature vector and the LTL feature vector (from the progressed formula), enabling Markovian state representation from non-Markovian histories.

Improvement: I will adopt the paper's architecture: a shared grounder, a CNN-based environment encoder, a frozen GNN-based LTL encoder, and a PPO-based RL module. I will also use the paper's hyperparameter settings (e.g., learning rate 0.0003, GAE-λ 0.95, PPO clip 0.2) and the specific training procedure (bootcamp pretraining, then joint training with frozen LTL encoder).

Improvement: I will implement the grounder as a CNN that processes 56×56 RGB images (Minecraft) or continuous 2D coordinates (FlatWorld), and I will randomize the environment layout (item positions, zone placements) at every episode. The grounder is trained to be invariant to layout changes, focusing on visual features (icons, colors) rather than fixed positions.

Improvement: I will freeze the LTL encoder after pretraining, preventing backpropagation from disrupting the learned semantics of LTL operators. This stabilizes training in the presence of an initially inaccurate grounder, as shown in the paper's ablation study (Figure 6).

Improvement: I will implement the paper's approach for Global Avoidance tasks (e.g., avoid lava until you reach the door), which uses the co-safe LTL fragment with a single avoidance proposition. I will also note the paper's finding that this task class is harder and may not converge in continuous environments.

The improved AI system can:

  • Learn to follow multiple, complex, temporally-extended instructions (LTL) directly from raw sensory input (images or continuous states).

  • Automatically discover the meaning of symbols (e.g., egg, lava, door) without any human-provided labels.

  • Generalize to new, longer, and more complex instructions that were never seen during training.

  • Operate in both discrete grid-worlds and continuous 2D navigation environments with randomized layouts.

  • Achieve performance comparable to systems that have access to the true symbol grounding, while significantly outperforming the only existing baseline for sub-symbolic multi-task RL.

Abstract

In this work we address the problem of training a Reinforcement Learning agent to follow multiple temporally-extended instructions expressed in Linear Temporal Logic in sub-symbolic environments. Previous multi-task work has mostly relied on knowledge of the mapping between raw observations and symbols appearing in the formulae. We drop this unrealistic assumption by jointly training a multi-task policy and a symbol grounder with the same experience. The symbol grounder is trained only from raw observations and sparse rewards via Neural Reward Machines in a semi-supervised fashion. Experiments on vision-based environments show that our method achieves performance comparable to using the true symbol grounding and significantly outperforms state-of-the-art methods for sub-symbolic environments.

Sources

Related papers