Ready2Blend: From Natural-Language Instructions to Composable Alignment Prompts

arXiv:2609.39365 · cs.CL, cs.AI · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Ready2Blend: From Natural-Language Instructions to Composable Alignment Prompts".

Tom: Continual alignment requires Large Language Models (LLMs) to adapt to new requirements without forgetting previously acquired behaviors, and this paper introduces Ready2Blend,

Jane: First, who's behind it and why it matters.

Title and authors: Tom: Let's talk about the title and who wrote this paper; it’s "Ready2Blend: From Natural-Language Instructions to Composable Alignment Prompts," which tells us they are focusing on taking those flexible instructions and making them composable.

Jane: The authors are Jeesu Jung, Hwan Jang, Juseon Do, Jeonghwan Choi, Jinho Choo, Sungwoo Nam, Seungki Hong and Hwanjun Song from Korea Advanced Institute of Science and Technology.

Lu: Their work on this paper is particularly interesting because they address the fundamental tension between natural language flexibility and the need for stable knowledge retention in continual alignment.

Meng: I’m curious how much of this relies on the structure of that fixed-length prompt they generate; if that prompt structure is rigid, it limits how much new information we can inject later.

Lalam: What excites me most about their title is the concept of composable prompts, which suggests a way to build complex behaviors by combining simpler learned pieces rather than having one monolithic instruction.

The paper's summary: Tom: So, what does the core idea of "Ready2Blend: From Natural-Language Instructions to Composable Alignment Prompts" actually boil down to for us? Essentially, they introduce a method where they map each new alignment requirement onto a fixed-length prompt and store these prompts in a bank.

Jane: To put it simply, instead of retraining the entire Large Language Model when we want it to follow a new rule, they learn that rule as an alignment prompt over the model’s frozen structure, allowing them to keep all their old knowledge intact.

Lu: The mechanism involves using something called AlignFormer to translate the natural language requirement into this fixed-length alignment prompt which is then stored in a modular bank.

Meng: So, they are essentially treating each requirement like a separate instruction file that gets fed into the model, rather than merging it directly into the main instruction set. That sounds like a way to manage complexity better during fine-tuning phases.

Lalam: I think what's key here is that these prompts can be blended together later at inference time using weighted vector arithmetic, which means you can select and combine different requirements based on how important they are for the specific task at hand.

The paper's improvements: Tom: The paper points out a few major improvements in their methodology, specifically focusing on composability regularization and inference-time blending as the way they achieve their results.

Jane: They use Composability Regularization to make sure that the geometry of these learned prompts actually makes sense when you try to combine them later; it enforces consistency based on how the original text requirements related to each other.

Lu: That geometric transfer aspect is quite sophisticated; it means they aren't just learning random vectors, but they are anchoring the prompt space based on the semantic structure of the original natural language.

Meng: From a practical standpoint, that regularization step seems important because if those prompts don't align geometrically, blending them would result in nonsense behavior for the AI.

Lalam: And then at inference time, they use weighted vector arithmetic to blend these prompts together; this is what enables that inference-time personalization you can control with user preferences by adjusting the weights.

Conclusion: Tom: So, to wrap up on "Ready2Blend: From Natural-Language Instructions to Composable Alignment Prompts," we've seen how they use AlignFormer and Composability Regularization to create a system where alignment requirements are learned independently but can be combined arithmetically at inference.

Jane: The main implication is that we can achieve strong performance on continual alignment tasks, reaching ninety-three point one–ninety-eight point five percent of joint training results, while saving about two point six to four point three times the training time compared to methods like CPPO and LifeAlign.

Lu: What I find most significant is that it proves you can decouple alignment from backbone training entirely by learning prompts over a frozen LLM, which opens up new avenues for building very specialized AI behaviors without needing constant parameter updates.

Meng: The efficiency gain is compelling; achieving those results in less than half the training time makes this approach much more viable for environments where rapid adaptation is necessary.

Lalam: And the ability to do inference-time personalization through weighted prompt blending means we are moving toward AI systems that can adapt their style or adherence to instructions instantly based on context, which is a major step for practical deployment.

Jeesu Jung, Hwan Jang, Juseon Do, Jeonghwan Choi, Jinho Choo, Sungwoo Nam, Seungki Hong

Korea Advanced Institute of Science and Technology

cs.CL, cs.AI

Submitted: 2026-09-30

Updated: 2026-09-30

Importance score: 91/100

The gist: Continual alignment requires Large Language Models (LLMs) to adapt to new requirements without forgetting previously acquired behaviors, and this paper introduces Ready2Blend, a method that combines

Key concepts

Composable Alignment
This approach learns each new alignment requirement separately as an independent prompt over a static model backbone. The key idea is that these individually learned prompts are structured so they can be combined mathematically during runtime, allowing users to mix and match preferences without retraining the main LLM.
AlignFormer
AlignFormer is a mechanism used to translate human-readable text requirements into a standardized, fixed-length alignment prompt. It uses a frozen sentence encoder and a stack of Transformer blocks that iteratively refine the requirement's meaning into specific slots before projecting it into the LLM's embedding space.
Composability Regularization
This technique ensures that the prompts learned for different requirements have compatible geometric relationships in the model's space. It enforces consistency rules between prompts and their original text requirements, guaranteeing that combining these learned prompts through arithmetic operations yields meaningful results.

Terminology

Summary

Continual alignment requires Large Language Models (LLMs) to adapt to new requirements without forgetting previously acquired behaviors, and this paper introduces Ready2Blend, a method that combines natural language flexibility with learned alignment to achieve this goal. The gist is: Ready2Blend learns each alignment requirement as a prompt over a frozen backbone and blends them at inference.

The Core Problem and Formulation

Lifelong alignment involves continually incorporating evolving requirements—whether across tasks or preferences—while preserving prior behaviors, which is difficult because post-training methods often require repeated parameter updates that risk catastrophic forgetting. The paper shifts the focus from updating the model's backbone to updating prompt space. This formulation is termed composable alignment, where each requirement is learned independently as a prompt over a frozen backbone, and these prompts are structured so that any subset can be jointly controlled at inference.

AlignFormer: Requirement-to-Prompt Translation

Ready2Blend utilizes AlignFormer, a variant of Q-Former, to map each natural language requirement to a fixed-length alignment prompt. This process involves several steps:

  1. The textual requirement R is first mapped by a frozen pre-trained sentence encoder (SentEncoder) into an embedding space, yielding e.

  2. A stack of Transformer decoder blocks (Decoder), conditioned on randomly initialized learnable query tokens Z(0), attends to e via cross-attention and self-attention, iteratively refining the representation to absorb the semantics of R into k slots, resulting in Z(L).

  3. This final representation is projected by Proj2 into the LLM embedding space (size h) to yield the alignment prompt P = Proj2(Z(L)).

Composability Regularization: Structuring Prompt Geometry

Since independently learned prompts may occupy incompatible regions of the representation space, Ready2Blend employs Composability Regularization to ensure meaningful arithmetic combination. This regularization transfers the semantic geometry of a pre-trained sentence encoder into the prompt space by enforcing two consistency conditions on the learned prompts (Eq. 3):

  1. Point-wise Consistency: Each alignment prompt is placed consistently with its textual requirement R.

  2. Pair-wise Consistency: The relation between prompts Pi and Pj mirrors that between their requirements Ri and Rj, anchoring the geometry to the stored prompts rather than a re-encoding by the updated AlignFormer.

Inference-Time Blending and Personalization

At inference time, alignment is achieved by blending selected prompts from the Alignment Prompt Bank (B) using weighted vector arithmetic (Eq. 4):

)&Blend(S) = P; R∈S wR PR, where wR ≥ 0, P; R∈S wR = 1. The resulting mixed prompt preserves the length and scale of a single prompt regardless of the number of selected requirements. This allows for weighted personalization, where user preferences translate directly into blending weights (wR), and composition is order-free because blending is commutative. The default composition operator used is arithmetic averaging, which maintains a fixed prompt length regardless of the number of composed requirements, avoiding additional inference-time token costs. The method requires only a few prompt tokens (k=4) at inference time. Ready2Blend matches strong post-training methods in final quality with competitive retention and significantly less training time. It is the only frozen-backbone method to reach this level.

Performance and Efficiency

Ready2Blend demonstrates superior performance across two continual alignment settings: task-incremental alignment and preference-incremental alignment. In both settings, it matches strong post-training methods, reaching 93.1–98.5% of a joint-training reference with competitive retention (BWT). Furthermore, Ready2Blend is the only frozen-backbone method to achieve this level of performance while requiring only a few prompt tokens and up to 4.3× less training time compared to CPPO and LifeAlign. The modular design enables weighted personalization and order-free composition without retraining, proving that lifelong alignment can accumulate composable prompts around one that never changes. The use of a frozen backbone means adding a new requirement neither updates the backbone nor edits a shared instruction, ensuring each earlier requirement keeps its own learned prompt intact. This robustness is further confirmed by its stability across different alignment orders and LLM judge choices.

Key Contributions Summary

"We formulate composable alignment, where alignment requirements are learned independently as they arrive, rather than through joint multi-objective training, yet remain jointly controllable at inference without repeatedly updating the large backbone model."

We introduce Ready2Blend, combining AlignFormer and Composability Regularization to learn independently trainable yet arithmetically composable alignment prompts.

The method successfully decouples alignment from backbone training by learning prompts over a frozen LLM, achieving high final performance with low training and inference overhead. It also enables inference-time personalization through weighted prompt blending.

Improvements for AI systems

Here are the specific improvements that can be made to AI systems based on the Ready2Blend paper, focusing on capabilities derived from its architecture and training methodology:

  1. Enhance Continual Learning Capability without Catastrophic Forgetting:

  2. Enable Dynamic, Inference-Time Personalization:

  3. Achieve Robust Alignment Composition via Semantic Geometry Transfer:

  4. Improve Training Efficiency for Lifelong Alignment Tasks:

  5. Create Order-Invariant Multi-Requirement Steering Mechanisms:


  1. Enhance Continual Learning Capability without Catastrophic Forgetting:

The system can maintain a history of diverse alignment requirements (tasks or preferences) encountered over time without needing to retrain the entire foundational LLM backbone for every new requirement. This is achieved by learning lightweight, requirement-specific prompts rather than updating model weights.

  1. Enable Dynamic, Inference-Time Personalization:

The system can adapt its output behavior instantaneously based on a specific user's needs (e.g., preference for conciseness vs. abstractiveness) by dynamically blending and reweighting the learned prompts at inference time. This allows for highly granular, user-specific control without needing separate fine-tuned models for every user profile.

  1. Achieve Robust Alignment Composition via Semantic Geometry Transfer:

The system can combine multiple, independently learned alignment requirements (e.g., be truthful AND be concise) into a single, coherent instruction prompt during inference. This composition is semantically meaningful because the prompts are regularized to share a common geometric structure derived from the original natural language definitions, ensuring that blending results in a predictable and accurate combined behavior rather than semantic cancellation or distortion.

  1. Improve Training Efficiency for Lifelong Alignment Tasks:

The system can acquire a wide variety of alignment requirements with significantly reduced training time compared to traditional post-training methods (CPPO, LifeAlign) by only optimizing the lightweight prompt parameters, while keeping the massive LLM backbone frozen. This makes lifelong alignment feasible in resource-constrained or rapidly evolving deployment scenarios.

  1. Create Order-Invariant Multi-Requirement Steering Mechanisms:

The system can steer the frozen LLM toward a desired set of behaviors regardless of the sequence in which those requirements were learned or applied. By using arithmetic blending (as opposed to sequential model updates), the final behavior is decoupled from the training order, making it more robust and reliable in long-running, multi-stage alignment processes.

Sources

Related papers