Skill-SLM: Agent Skill-driven Small Language Models for Reliable Robot Operation

summary

Video file (mp4)

The gist

The gist Skill-SLM proposes a framework that reformulates SLM-driven robot operation as a task-decomposition and skill-composition problem, which substantially outperforms distillation-oriented

In short

Skill-SLM reframes robot operation as decomposing tasks into subtasks and composing them using learned skills. It builds a skill library via a grammar, uses two agents for decomposition and orchestration, and employs progressive orchestration to ensure coherent, reliable robot actions. This method outperforms distillation-based approaches on unseen tasks.

Key concepts

Skill Library Construction
This involves creating a structured database of robot skills. A novel context-free grammar is used with an LLM to discover necessary operational skills from task instructions. These discovered skills are then grounded in real robot APIs and documented with detailed instructions for the SLM.
Decomposition and Orchestration Agents
Skill-SLM uses two specialized agents: one decomposes a complex task instruction into smaller subtasks and selects the right skills from the library. A second agent takes these selected skills and sequences them to generate executable robot operation code segments, forming the final plan.
Progressive Skill Orchestration
This strategy improves reliability by generating robot code step-by-step for each subtask. The generation of the current step's code is conditioned on the results of all previous steps, allowing the SLM to maintain context and reference intermediate results, ensuring smooth transitions between skills.

Terminology used across episodes

This episode discusses

The paper

Skill-SLM: Agent Skill-driven Small Language Models for Reliable Robot Operation · Read on arXiv

Wenhao Wang, Yanyan Li, Jiawei Yuan

Department of Computer & Information Science, University of Massachusetts Dartmouth · Department of Computer Science & Engineering, California State University San Marcos

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Skill-SLM: Agent Skill-driven Small Language Models for Reliable Robot Operation".

Dev: The gist Skill-SLM proposes a framework that reformulates SLM-driven robot operation as a task-decomposition and skill-composition problem, which substantially outperforms distillation-oriented baselines,

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: So we're looking at Skill-SLM today, the paper is titled "Skill-SLM: Agent Skill-driven Small Language Models for Reliable Robot Operation," and it’s written by Wenhao Wang and Yanyan Li. It boils down to taking robot operation away from just teaching models through examples, which is what a lot of the current work has been doing.

Dev: Exactly, this isn't about just making bigger language models; it's about how we structure the whole process for an SLM to actually do something reliable on a physical robot. It’s moving from just pattern matching to actual planning and execution based on defined skills.

Taro: What I find interesting is that they are framing this as a task decomposition and skill composition problem, which means instead of one giant model trying to figure out everything at once, you break the job down into smaller, manageable steps first.

Rosa: That’s right, it’s about breaking things down. They propose a framework where an SLM takes a big instruction and breaks it into subtasks and then picks skills from a library to do those subtasks, and finally an orchestration agent puts those skills together into actual robot code.

Dev: It seems like the main idea is creating these two specialized agents, one for breaking things down and one for putting the pieces together, which addresses how you actually get from a natural language command to physical movement.

Taro: It’s interesting because it tackles generalization too; they mention this approach is supposed to help with unseen tasks that require generalizing capabilities beyond what was seen in the training data.

Rosa: Right, so we're talking about moving toward a system that can handle more varied robot jobs without needing a completely new massive dataset every single time.

The paper's summary: Rosa: The paper dives into how they built this skill library first, which is the foundation of the whole thing. They propose a novel robot operational skill aware context-free grammar, or CFG, to automatically extract all the skills a robot needs for specific tasks.

Dev: So they use an LLM, specifically configured with GPT-five point five as mentioned in their work, to look at a set of task instructions—about fifty of them—and generate a CFG from that corpus. This CFG then identifies things like object selection, movements, and inspections.

Taro: That part where they use the LLM to induce the CFG is key because it’s how they get an initial inventory of what skills are even possible for a robot to do in this context.

Rosa: Once those skills are found, they build the skill library by grounding each one in actual robot APIs with a human expert. Each skill gets a structured document, which includes instructions on purpose, inputs, preconditions, and the implementation procedure using those APIs.

Dev: The document structure they use is pretty specific: it’s defined as d(s) = m(s), b(s), where the body part b(s) gives the SLM all the necessary details to implement that specific skill.

Taro: And then you have these two agents, right? A decomposition agent that breaks tasks down using that library, and an orchestration agent that takes those chosen skills and turns them into executable robot operation code segments.

Rosa: Exactly. The whole point of this summary is showing how they structure the entire workflow as a chain: instruction to decomposition to selection to orchestration to execution.

The paper's improvements: Dev: One major improvement they focus on is this progressive skill orchestration strategy, which they introduce for reliability. Instead of just picking one skill and running it, the SLM generates the code segment for the current subtask while conditioning that generation on what happened in all previous steps.

Rosa: That means if you’re doing a sequence of moves, the model can refer back to variables or targets that were set earlier in the plan. It helps preserve those cross-skill dependencies during execution, which is crucial when skills rely on each other.

Taro: I see how that addresses a problem where simple one-shot orchestration might get an intermediate step right but then lose track of the overall plan coherence later on.

Dev: It’s about keeping the context alive as you build the final sequence of operations, which improves the reliability of what ends up being executed on the robot.

Rosa: And they synthesized training data using LLM teachers to create two separate datasets: one for decomposition and one for orchestration. This distillation process is used to teach SLMs how to perform this compositional reasoning effectively.

Taro: So, they are essentially using the powerful reasoning of an LLM teacher to distill its skill-based planning knowledge into the smaller SLMs, which is a smart way to bridge that gap in capability.

Conclusion: Rosa: So, looking at everything we've covered in Skill-SLM, it’s clear they’ve moved away from just hoping an SLM can learn by seeing examples. They’ve proposed a concrete structure where the SLM decomposes a task into subtasks, selects skills from a library, and then orchestrates them progressively for reliable execution.

Dev: The results show that this framework achieves the highest average TSR across all three task sets they tested. And importantly, on new-combination tasks or when new skills are introduced without extra distillation, it maintains performance close to what you’d see on seen-pattern tasks.

Taro: That suggests their compositional reasoning isn't just stuck in the training set; it can actually be extended through this skill augmentation, which is a pretty significant finding for generalization.

Rosa: It really shows that for onboard robot operation, having a framework focused on task decomposition and skill composition can lead to much more reliable and generalizable systems across different robot platforms.

Dev: They also highlighted some limitations. They pointed out that direct skill extraction, where you just pull skills straight from the task corpus, doesn't do well on new-capability tasks, meaning we can’t just rely on the model finding everything automatically.

Taro: And they also found that removing the distillation process causes a big drop in performance, showing that just having prompts and skill documents isn't enough for base models to get this skill-driven reasoning right by themselves.

Rosa: So, in summary, Skill-SLM provides a robust approach to task planning and robot operations by creating a formal skill library, using agent decomposition and orchestration with progressive context awareness for the final code generation.

Dev: It’s a promising direction for reliable onboard SLM-driven robot operation because it provides that structure where we know exactly what's supposed to happen at every step.

More episodes

← Home