LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents".
Jane: The paper was written by Aofan Yu, Chenyu Zhou, Tianyi Xu, Zihan Guo, Rong Shan et al. from Shanghai Jiao Tong University and Sun Yat-sen University and Shanghai Innovation Institute and OPPO Research Institute.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back, everyone. Today we're looking at a paper that's got a title that really makes you stop and think: "LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents." Jane, I have to say, even the title is a mouthful, but it's pointing at something pretty fundamental.
Jane: It really is, Tom. And I think the simplest way to break it down is this: right now, when we want an AI agent to know how to do something, like clean a house or search the web, we write out instructions and paste them into the prompt every single time. That's the "in-context" part. This paper says, what if instead we could bake those instructions directly into the model's weights, like a little plug-in module?
Tom: Right, and that's the "in-weight" part. So instead of reading a manual every time, the model just has the knowledge installed. And the "Latent" part is because the skill isn't visible as text anymore—it's hidden in these mathematical adjustments to the model.
Jane: Exactly. And the authors are from a bunch of places, Shanghai Jiao Tong University, Sun Yat-Sen University, and OPPO Research. They're calling these adjustments LoRA adapters. LoRA is a technique that lets you tweak a big model without retraining the whole thing—it's like adding a small, specialized circuit to a big computer.
Tom: And the clever bit is how they make those circuits. They train a separate "hypernetwork" that reads the skill text and, in one go, produces the adapter. So you write a skill once, it gets compiled into this plug-in, and then you can just attach it to the model whenever you need it.
Jane: Which solves a real headache. If you have a hundred skills, you don't want to stuff all that text into every single prompt. That's expensive and it can confuse the model. This way, you just load the right plug-in.
Tom: And it's modular. You can swap skills in and out without retraining the main model. That's a huge deal for practical use. I'm curious what Lu thinks about this shift from reading instructions to having them installed.
Lu: I think it's a genuinely different way of thinking about knowledge. We're so used to prompting as the only interface. This paper suggests that the model's weights themselves can be a kind of storage medium, and that has implications for how we think about model memory and even model security.
Jane: Security is a good point. If the skill isn't in the prompt, it's not as exposed to prompt injection attacks. We'll get into that more later. For now, I'm just excited about the idea of a model that doesn't need to read the manual every single time.
Tom: Same here. And the results they show are pretty striking. We're going to dig into those numbers next.
Summary: Tom: So we've set the stage with the basic idea of "LatentSkill." Now let's talk about what they actually found. Jane, the numbers on their main benchmarks are pretty wild.
Jane: They are. They tested on ALFWorld, which is a text-based home robot simulator, and Search-QA, which is about answering questions using web searches. On ALFWorld, their method, LatentSkill, improved the success rate by over twenty-one points on the "seen" tasks and over thirteen points on the "unseen" tasks compared to just putting the skill text in the prompt.
Tom: And that's not even the headline. The headline is that they did this while using sixty-four percent fewer tokens on ALFWorld and seventy-two percent fewer on Search-QA. So they're getting better results with way less text going into the model.
Lu: That's the efficiency argument, and it's compelling. But what I find more interesting is that the performance gain isn't just about efficiency. The model is actually behaving better. It's completing tasks in fewer steps. It's not just a compressed prompt; it's a better representation of the skill.
Meng: From an engineering standpoint, that token reduction is huge. Prefill time, which is the time it takes to process the input, scales with the number of tokens. If you're running an agent that makes hundreds of decisions, cutting the input size by two-thirds is a massive speedup and cost saving. It makes these agents much more viable in production.
Jane: And they didn't just test on one thing. They tested on two very different benchmarks. ALFWorld is about embodied interaction, like picking up objects and using lamps. Search-QA is about reasoning over retrieved documents. The fact that it works on both suggests the approach is pretty general.
Tom: Right, and they also compared against a bunch of strong baselines, like Reflexion and AdaPlanner. LatentSkill beat them all on average. It wasn't even close on some tasks.
Lu: The unseen split is the real test, though. That's where the model has to handle tasks it hasn't seen during training. And it still improved by thirteen points. That tells me the hypernetwork isn't just memorizing the training skills; it's learning a general mapping from text to behavior.
Meng: I want to know about the practical side. How big is this hypernetwork? Is it feasible to run this alongside the main model?
Jane: The paper doesn't give exact parameter counts for the hypernetwork, but it's a Transformer-based model, and it's trained once. At inference time, you just run it once per skill to generate the adapter, and then it's just a LoRA on the main model. The overhead is minimal.
Tom: And that's the beauty of it. You compile the skill once, cache the adapter, and then it's just a matter of loading it. We'll talk more about what that means for control and combining skills in a bit.
Lu: I think the most exciting implication is that we might be moving toward a world where skills are like software libraries you can install, rather than documents you have to read. That's a fundamental shift in how we interface with AI.
Improvements: Tom: We've talked about the performance and efficiency. But the paper goes deeper. They show that these generated skill adapters have some really interesting properties. Jane, you want to take the controllability one?
Jane: Sure. So they found you can scale the effect of a skill by just multiplying the adapter weights by a number, which they call alpha. If alpha is zero, the skill has no effect. If it's one, it's at full strength. And they found that performance follows an inverted-U curve. Too little alpha and the skill doesn't help. Too much, and it actually hurts.
Meng: That's a really useful knob for an engineer. You can tune the strength of a skill per task. They showed that harder tasks, like the Pick2 task which has a low baseline, benefit from a higher alpha. So you could have a system that automatically adjusts the injection strength based on the task difficulty.
Lu: And it's not just a binary on-off switch. It's continuous control. That's something you can't do with text in a prompt. You can't have "half a skill" in text. But you can have a half-strength LoRA. That opens up a lot of possibilities for fine-grained control.
Tom: And then there's the composition part, which I think is the most mind-bending. They showed you can combine skills by just adding their adapters together in weight space. But there's a catch.
Jane: Right, the catch is that you have to decompose the skills into aligned components first. They tested this on ALFWorld by combining a "Look" skill with a "Pick" skill. If you just add the whole adapters together, it doesn't work well. But if you break each skill into shared components, like "general behavior" and "mistake avoidance," and task-specific components, and only add the task-specific ones, it works beautifully.
Meng: That makes sense from a signal processing perspective. If you add two signals that share a common component, you're amplifying that common component twice. By separating out the shared parts and adding them once, you avoid that distortion.
Lu: Exactly. And this is where it gets really interesting for the future. Imagine a skill marketplace where you can buy a "search" component and a "reasoning" component and just add them together to create a new, more powerful skill. That's the kind of modularity this enables.
Tom: And they showed that this composition actually works. On the Look task, the component merging approach got eighty-four point six percent success on seen episodes, compared to sixty-one point five percent for just the Look skill alone. So it's not just preserving capability; it's genuinely adding complementary behavior.
Jane: And they also tested robustness. They perturbed the skill text, like paraphrasing it or adding noise, and the performance held up much better than when the skill was in the prompt. And they tested against prompt injection attacks, where a malicious instruction is added to the prompt. LatentSkill was much more resistant because the skill isn't in the prompt to be hijacked.
Lu: That's a significant security advantage. Skills as weights are not readable text, so they're not exposed to the same attack surface. It's not perfect, but it's a meaningful improvement.
Meng: So the improvements here are threefold: you get continuous control over skill strength, you get a principled way to compose skills, and you get better robustness. That's a pretty complete package.
Conclusion: Tom: Well, we've covered a lot of ground on "LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents." Let's try to wrap it all up.
Jane: I think the core message is that skills don't have to live in the prompt. By moving them into the model's weights as LoRA adapters, you get a representation that's more efficient, more controllable, and more robust.
Tom: And the results back that up. Better performance on ALFWorld and Search-QA, with a fraction of the token overhead. And the ability to scale and compose skills is a game-changer for building modular agent systems.
Lu: For me, the most profound implication is that this gives us a new substrate for knowledge. We're no longer limited to text as the only way to encode procedural knowledge. Weight space has structure, and we're just starting to learn how to use it.
Meng: From my side, the practical impact is clear. This makes agent systems cheaper to run and easier to maintain. You can update a skill by just recompiling its adapter, without retraining the whole model. That's a huge operational win.
Jane: And let's not forget the security angle. Keeping skills out of the prompt reduces their exposure to injection attacks. That's a real benefit for anyone deploying agents in an untrusted environment.
Tom: So, "LatentSkill" is a paper that takes a familiar problem—how to give agents skills—and offers a fresh solution that's both practical and theoretically interesting. It's not just a compression trick; it's a different way of thinking about what a skill is.
Lu: And it opens up so many questions. How does this scale to hundreds of skills? Can we do skill arithmetic to create novel behaviors? How does this interact with continual learning? This is just the beginning.
Jane: We'll be watching this line of research closely. Thanks for joining us, and we'll see you next time on the arXiv channel.
Tom: Take care, everyone.
Shanghai Jiao Tong University · Sun Yat-sen University · Shanghai Innovation Institute · OPPO Research Institute
cs.CL, cs.AI
Submitted: 2026-06-04
Updated: 2026-08-26
Comments: 16 pages, 4 figures
Code: https://github.com/yuaofan0-oss/LatentSkill
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 85/100
The gist: LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents The paper introduces LatentSkill, a framework that "converts textual agent skills into LoRA adapters through
Key concepts
- In-Context Skills
- The traditional method where instructions or skills are written out as text and pasted into the prompt every time an AI agent needs to perform a task. This can be expensive and cumbersome for many skills.
- In-Weight Latent Skills
- A novel approach where skill knowledge is baked directly into the model's weights, like a specialized plug-in module. This makes the skill mathematically hidden rather than visible as text in the prompt.
- LoRA Adapters
- A technique used to tweak a large AI model without needing to retrain the entire thing. It functions by adding small, specialized circuits or adapters that can be attached to modify the model's behavior for specific skills.
- Hypernetwork
- A separate network trained by the authors that reads skill text and, in one step, generates the necessary LoRA adapter. This allows a single skill to be compiled into a plug-in module efficiently.
Terminology
Summary
LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents
The paper introduces LatentSkill, a framework that converts textual agent skills into LoRA adapters through hypernetwork-based adapter generation.
Instead of delivering skills through the context window, LatentSkill represents them in weight space.
Given a skill description, a trained hypernetwork generates a skill-specific LoRA adapter in a single forward pass, which is then mounted on the backbone LLM during inference.
The original skill text is no longer included in the prompt, reducing both context cost and exposure as readable text.
The authors state their contributions as: "(1) We propose LatentSkill, a framework that converts textual agent skills into modular LoRA adapters through a hypernetwork. (2) We show that LatentSkill improves over direct in-context skill prompting on ALFWorld and Search-QA while reducing the context overhead introduced by skill text. (3) We analyze the generated skill weights and show that they exhibit domain-level structure, can be controlled through LoRA scaling, and can be composed when skills are decomposed at the right granularity."
Let Mθ denote a frozen backbone LLM, and s be a textual skill document. A skill compiler Gϕ maps the skill document to a set of LoRA updates: ∆s = Gϕ(s). With ∆s mounted, the model predicts from the task history alone: pθ,ϕ(yt ht, s) = pθ⊕α∆s(yt ht), where ⊕ denotes LoRA-based parameter augmentation and α controls injection strength. For each target module m, the generated update follows the standard LoRA form ∆Ws(m) = Bs(m)As(m), and mounting adds a scaled low-rank update: W′ = W + (α/r)BsAs.
The compiler is pretrained on a corpus of textual skill documents (approximately 171K deduplicated skill documents crawled from GitHub, totaling roughly 300M tokens). Two document-level pretraining tasks are used: a reconstruction task where the compiler reads the complete skill document s, and the adapted backbone receives a reconstruction instruction as input and is trained to reproduce the original document s,
and a completion task where we construct a truncated prefix s̃ by randomly removing the latter part of the document; the compiler reads s̃, and the adapted backbone is trained to complete the full skill document.
Only compiler parameters are updated.
The compiler is fine-tuned with teacher agent trajectories. For each pair (si, τi), the compiler generates one latent skill... The same adapter is mounted throughout the entire trajectory.
The objective encourages the adapter to capture skill-level, trajectory-consistent policy information rather than per-step adaptations.
At inference, skill compilation is separated from agent execution.
Each skill is compiled once and stored in an adapter cache. For a single skill, the cached adapter is mounted with injection coefficient αk. For multiple skills, adapters are composed in weight space: ∆K = Σ αk C[k]. The framework also supports component-level composition where a skill can be decomposed into semantic components... each component can be compiled independently.
Backbone: Qwen3-8B (frozen). The skill compiler is a Transformer-based hypernetwork.
Benchmarks:
-
ALFWorld: text-based interactive environment with six task categories (Pick, Look, Clean, Heat, Cool, Pick2), evaluated on seen (140 episodes) and unseen (134 episodes) splits.
-
Search-QA: seven search-augmented QA datasets (NQ, TriviaQA, PopQA, HotpotQA, 2WikiMultihopQA, MuSiQue, Bamboogle), with training data from NQ and HotpotQA and the remaining five as out-of-domain evaluation sets.
Baselines: Vanilla, Few-shot, Reflexion, AdaPlanner (ALFWorld); Vanilla, CoT, Few-shot, R1-Instruct, RAG (Search-QA); and In-context Skill which uses the same skill as LatentSkill but places it in the prompt rather than encoding it as LoRA weights.
On ALFWorld, LatentSkill reaches 74.3% and 69.4% average success on both the seen and unseen splits, improving over In-Context Skill by 21.4 and 13.4 points.
On Search-QA, it achieves the highest average EM score of 35.6.
Gains are especially pronounced on multi-step tasks: on unseen Pick2, LatentSkill reaches 70.6%, surpassing the second-best method by 41.2 points.
"On ALFWorld, LatentSkill reduces prefill overhead by 64.1% relative to In-Context Skill, while improving average success by 21.4 and 13.4 points on the seen and unseen splits, respectively. On Search-QA, it reduces context overhead by 72.2% and improves average EM by 3.0 points."
LatentSkill also shortens trajectories: On the seen split, it achieves the fewest average steps per episode, reducing the trajectory length from 35.0 steps for the Vanilla backbone to 28.4 steps.
The authors apply Multidimensional Scaling (MDS) to LoRA weights. In-domain skills form clear domain-level clusters: the 5 ALFWorld skills and 3 Search skills are separated in weight space, with an inter-cluster distance of 0.0887 and higher within-domain than cross-domain similarity (0.982 vs. 0.910).
After SFT, the inter-cluster distance decreases to 0.0704, a 20.6% reduction, while both within-domain and cross-domain similarities increase,
suggesting SFT introduces shared agent-level behavioral patterns while preserving skill-specific structure.
This organization generalizes to out-of-distribution skills (Code, Finance, Writing from GitHub): OOD skills from Code, Finance, and Writing form separated clusters, with within-domain similarities of 0.783, 0.9664, and 0.9681, respectively, each exceeding the corresponding cross-domain similarities.
Sweeping α ∈ 0, 0.1, 0.2, 0.3, 0.5, 0.6, 0.8, 1.0, 1.2 on ALFWorld, "task performance follows an inverted-U curve. On the seen split, average success rises from 43.57% at α=0 to 74.29% at α=0.6, but drops to 22.86% at α=1.2. The unseen split shows the same pattern, peaking at 70.90% when α=0.5 and falling to 8.21% at α=1.2. The authors note
tasks with weaker backbone baselines often require stronger skill injection and that
the optimal α varies across tasks."
Using Look as target skill and Pick as auxiliary skill, five configurations are evaluated: Look-Only, Pick-Only, Direct Merging (adding complete skill LoRAs), Text Merging (concatenating skill texts before compilation), and Component Merging (separately compiling aligned components and combining LoRAs). Component Merging achieves the best performance on both splits, reaching 84.6% on seen episodes and 77.8% on unseen episodes,
adding successful episodes while losing none of Look-Only's original successes.
Direct Merging and Text Merging fail to improve over Look-Only on unseen split. The authors conclude: skill composition in LoRA space requires semantic alignment between text decomposition and weight addition.
Four perturbations are applied to skill text: Paraphrase, Plaintext (removing Markdown), Reorder (shuffling bullets), and Noise (injecting irrelevant sentences). LatentSkill maintains a consistent advantage over In-context Skill under all perturbations. On ALFWorld, the margin ranges from 17.2 to 24.3 points.
Notably, removing Markdown formatting causes no degradation on ALFWorld, suggesting that the generated LoRA captures skill semantics rather than relying on surface formatting.
Under prompt-level attacks, Under Hijack, In-context Skill drops from 52.9% to 8.57% on ALFWorld, while the weight-space variant retains 38.6%.
Under Extract, In-context Skill is vulnerable because the skill text is directly present in the prompt, whereas weight-space storage reduces direct plaintext exposure.
Analysis of 7 LoRA injection positions in Qwen3-8B shows attn o and mlp down exhibit substantially higher discriminability gaps than the remaining five positions (pretrain: 0.056/0.105; SFT: 0.050/0.094), while attn q/k/v show gaps close to zero.
The full:o+d configuration uses only 2 out of 7 injection positions yet retains 93.3% of the full configuration's performance on the seen split (59.3 vs. 63.6).
The stable rank of all skill LoRAs ranges from approximately 2.35–2.40 (pretrain) to 2.17–2.23 (SFT), while a randomly initialized LoRA of the same shape yields a stable rank of 837.87, a difference of roughly 380×.
Additionally, the top 2 singular directions alone capture approximately 67% of the total energy, and the top 5 directions capture approximately 93%.
After SFT, the stable rank of all skills decreases uniformly by approximately 0.17,
indicating SFT systematically concentrates skill knowledge into fewer singular directions.
The authors acknowledge: This work evaluates LatentSkill on two agent benchmarks, ALFWorld and Search-QA... they do not exhaust the full diversity of agent deployment scenarios.
Additionally, "all experiments use Qwen3-8B as the frozen backbone LLM with a fixed LoRA configuration. The behavior of the skill compiler and the properties of the generated latent skills may vary with different model families, model scales, or adapter configurations."
"LatentSkill converts textual agent skills into modular LoRA adapters through a pretrained hypernetwork, moving reusable procedural knowledge from context space into weight space. Across ALFWorld and Search-QA, this design improves over direct in-context skill prompting while substantially reducing the repeated prefill overhead introduced by skill text... the generated skill LoRAs form a structured semantic geometry, can be controlled through the injection coefficient, and can be composed in parameter space when skill components are properly aligned."
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system, along with the resulting capabilities.
-
Improvement: Replace the system's reliance on injecting long textual instructions (skills) into the prompt at every decision step. Instead, I will integrate a Transformer-based hypernetwork (the
skill compiler
) that takes a skill document as input and outputs a set of LoRA (Low-Rank Adaptation) weight matrices in a single forward pass. These LoRA adapters are then mounted onto the frozen backbone LLM. -
What the improved system can do:
-
Eliminate per-step skill token overhead: The system no longer needs to re-read the same skill text hundreds of times during a long-horizon task. This reduces prefill token usage by up to 72.2% (as demonstrated on Search-QA) and speeds up inference.
-
Maintain modularity: Skills can be loaded, unloaded, or swapped dynamically without retraining the backbone. This allows for a plug-and-play skill library where new skills are added by simply compiling their text into a LoRA, and old ones are removed by discarding the adapter.
-
Reduce prompt injection attack surface: Since the skill is stored in weight space, it is not exposed as plaintext in the prompt. This makes the system significantly more robust to
Hijack
andExtract
attacks, retaining 38.6% success on ALFWorld under a hijack attack where the in-context baseline drops to 8.57%. -
Improvement: Introduce a scalar injection coefficient
αthat linearly scales the generated LoRA weights before they are added to the backbone. This provides a continuous control knob for the influence of a skill, moving beyond the binary choice ofskill present
orskill absent.
-
What the improved system can do:
-
Adapt to task difficulty: The system can automatically tune
αper task or per episode. For example, on the ALFWorld unseen split, the system can increaseαfor harder tasks (like Pick2, which has a low baseline) to boost success from 23.53% to 88.24%, while using a lowerαfor easier tasks to avoid over-injection and performance collapse. -
Prevent catastrophic over-injection: The system can monitor performance and scale back
αif it detects that the skill is disrupting the backbone's general capabilities (e.g., performance drops from 74.29% to 22.86% whenαis doubled from 0.6 to 1.2). This allows for a safety mechanism to maintain stable performance. -
Improvement: Implement a composition mechanism that operates on semantically decomposed skill components rather than whole skills. Instead of naively adding two full LoRA adapters (which over-amplifies shared components), the system will:
-
Decompose a skill into aligned components (e.g.,
general behavior,
mistake avoidance,
task-specific action
). -
Compile each component into a separate LoRA.
-
Combine the LoRAs by adding them, but only including shared components once.
-
What the improved system can do:
-
Combine complementary skills without interference: The system can successfully merge a
Look
skill with aPick
skill. This achieves 84.6% success on seen episodes and 77.8% on unseen episodes, a significant improvement over the 61.5% and 72.2% baselines. It preserves the target skill's capability while successfully adding the complementary search behavior from the auxiliary skill. -
Avoid the
double-counting
failure mode: By preventing shared components from being amplified twice, the system avoids the behavioral disruption seen in direct merging, where the model perceives a target but fails to execute the pick-up action. -
Improvement: Based on the paper's finding that skill-specific knowledge is concentrated in the
attn oandmlp downmodules (with discriminability gaps of 0.056 and 0.105 vs. near-zero for others), I will configure the hypernetwork to generate LoRAs only for these two positions, rather than all 7 possible positions in a transformer block. -
What the improved system can do:
-
Reduce parameter overhead by 70%: The system only needs to generate and store LoRAs for 2 out of 7 modules, significantly reducing memory and storage requirements for the skill library.
-
Improve out-of-distribution generalization: This targeted injection not only retains 93.3% of the full configuration's performance on the seen split but actually improves performance on the unseen split (63.4% vs. 61.2%). This is because the other 5 positions contribute mostly noise that hinders generalization to new environments.
-
Improvement: Integrate a diagnostic tool that analyzes the generated LoRA weights using metrics like stable rank and singular value energy distribution. This allows for real-time quality assessment of the compiled skills.
-
What the improved system can do:
-
Detect poorly encoded skills: The system can flag a skill if its LoRA weights have an unusually high stable rank (indicating a noisy, non-compressed encoding) or if the top singular values do not capture a sufficient proportion of energy (e.g., <90% for top-5).
-
Monitor training progress: During the SFT phase, the system can track the stable rank decrease (from 2.35 to 2.18) as a metric to confirm that the hypernetwork is successfully compressing skill knowledge into a more concentrated and efficient representation, ensuring the training is on track.
Sources
- Text-to-LoRA: Instant Transformer Adaption
- Doc-to-LoRA: Learning to Instantly Internalize Contexts
- FireAct: Toward Language Agent Fine-tuning
- Generative Adapter: Contextualizing Language Models in Parameters with A Single Forward Pass
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- SkillRet: A Large-Scale Benchmark for Skill Retrieval in LLM Agents
- In-context Autoencoder for Context Compression in a Large Language Model
- SkillProbe: Security Auditing for Emerging Agent Skill Marketplaces via Multi-Agent Collaboration
- HyperNetworks
- LoRA: Low-Rank Adaptation of Large Language Models
- LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- Towards Secure Agent Skills: Architecture, Threat Taxonomy, and Security Analysis
- SkillOS: Learning Skill Curation for Self-Evolving Agents
- SkillMAS: Skill Co-Evolution with LLM-based Multi-Agent System
- SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass
- ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
- Prompt Injection attack against LLM-integrated Applications
- AdaPlanner: Adaptive Planning from Feedback with Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering