First Things First: Teaching LLM-Based Agents to Prioritize Must-Haves before Nice-to-Haves
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "First Things First".
Jane: Recent progress in multimodal large language models (MLLMs) has fueled enthusiasm for their potential as autonomous agents, but current systems struggle when faced with complex,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: We're starting with the title and authors of "First Things First: Teaching LLM-Based Agents to Prioritize Must-Haves before Nice-to-Haves." It seems like the authors are focusing on how models handle these layered requirements in a way that standard instruction following doesn't.
Jane: The authors are Tianjie Ju, Xinyue Xu, Wanxuan Sun, Lingxiao Diao, Gongshen Liu, and Zhuosheng Zhang. I think their focus is really on creating a benchmark to test this specific type of reasoning capability in multimodal models.
Lu: They’re introducing something called FTF-BENCH to evaluate these three distinct scenarios: must-have requirements that lead to a unique solution, multiple answers where nice-to-have requirements dictate the ranking, and unanswerable cases where the agent should just abstain.
Meng: That structure makes sense because real user requests are rarely just one simple instruction; they usually have layers of constraints that need careful sorting. I wonder how comprehensive this benchmark is across different domains like e-commerce or booking services.
Lalam: It’s smart that they’ve designed it to reflect realistic scenarios from those specific areas, which means the results will be much more relevant when we apply these models to actual customer service tasks.
The paper's summary: Tom: Now, let's talk about what the paper actually summarizes regarding its core findings and what this means for us. Essentially, they are showing that current MLLMs often fail because they don’t know how to weigh hard constraints against soft preferences correctly.
Jane: They summarize that the existing benchmarks mostly test generic instruction following, but this work focuses on measuring whether models can identify and prioritize those must-have requirements before optimizing for anything else.
Lu: The paper summarizes the introduction of First Things First Reinforcement Learning, which they designed to explicitly optimize reasoning over these multi-priority user requirements using a specific reward framework.
Meng: That reward framework sounds like it's trying to force the model to be careful about both being right and being accurate regarding what's mandatory versus optional. It’s an attempt at making the model self-aware of its prioritization task.
Lalam: The summary mentions that this approach involves training with a multiobjective reward that checks output format, answer correctness, and requirement classification into must-have or nice-to-have categories. That seems like a really robust way to guide the learning process.
The paper's improvements: Tom: The paper details the specific improvements they propose for MLLMs by moving beyond standard instruction tuning. They introduce FTFRL, which is this reinforcement learning approach designed to enhance requirement-aware reasoning.
Jane: Instead of just relying on standard training methods, they suggest a reward system composed of four key dimensions: a Format Reward for schema adherence, an Accuracy Reward using an Ranswer function for correctness, a Requirement Reward to assess correct classification into must-have or nice-to-have, and a KL Constraint to stabilize the policy updates.
Lu: That Requirement Reward is what I find particularly interesting because it directly penalizes both false positives—when the model treats something optional as mandatory—and false negatives, where it ignores an essential requirement.
Meng: If the model can learn to correctly classify those requirements using that F1 score metric, it means we’re not just teaching it to produce a good answer; we’re teaching it *how* to think about what kind of answer is acceptable in the first place.
Lalam: That focus on penalizing both types of errors gives us a much better signal for training, because getting either wrong is bad, but misclassifying the requirement type is also a mistake that needs to be corrected.
Conclusion: Tom: So, to wrap up this discussion on "First Things First: Teaching LLM-Based Agents to Prioritize Must-Haves before Nice-to-Haves," the main implication is that requirement-aware reasoning capability is a simple yet effective way to boost the generalization of MLLM agents across various complex tasks.
Jane: They conclude that this approach helps improve the performance of models in scenarios where they have to resolve trade-offs or even abstain from generating a response when no solution exists.
Lu: I think what stands out is that they found that even without explicit training on other reasoning benchmarks like LogicVista or MathVision, these models showed consistent improvements after being trained on FTF-BENCH.
Meng: That transferability across different reasoning tasks suggests that learning to prioritize requirements doesn't just help with service scenarios; it actually strengthens the model's fundamental ability to do complex logical and mathematical deduction in multimodal contexts.
Lalam: It’s a really powerful concept, and I think the authors make a very clear case that this method provides a straightforward path toward making MLLM agents more reliable in production environments.
Tom: Exactly, so we're looking at this paper as showing how to move MLLMs from just being pattern matchers to actual reasoning partners capable of handling layered complexity. That's where we leave it for today.
School of Computer Science, Shanghai Jiao Tong University · ByteDance
cs.CV
Submitted: 2026-09-04
Updated: 2026-09-30
Comments: Accepted at EMNLP 2026 (Findings)
Code: https://github.com/claire62/FTF-RL
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 88/100
The gist: Recent progress in multimodal large language models (MLLMs) has fueled enthusiasm for their potential as autonomous agents, but current systems struggle when faced with complex, structured user
Key concepts
- First Things First Benchmark (FTF-BENCH)
- This is a benchmark containing 3,649 real-world problems from areas like e-commerce and booking. These problems are specifically designed to test an agent's ability to handle three scenarios: must-have requirements that have a single correct answer, multiple answers where nice-to-have items dictate ranking, and cases where the agent must abstain from answering.
- First Things First Reinforcement Learning (FTFRL)
- FTFRL is a reinforcement learning method that trains language models using a 'multiobjective reward.' This reward system checks four things: the output format, answer correctness, whether requirements were correctly classified as must-have or nice-to-have, and stability constraints. This explicitly teaches the model to prioritize essential needs.
- Requirement Reward (Rrequirement)
- This component of the reward framework measures how accurately a model identifies and categorizes user requirements into 'must-have' or 'nice-to-have' groups using an F1 score. This is vital because it penalizes both ignoring essential constraints (false negatives) and incorrectly adding extra constraints (false positives).
- Upper vs. Direct Setting
- These settings describe how the model receives input. The 'Direct' setting means the model reads the original, colloquial request. The 'Upper' setting means the model receives structured gold labels detailing all requirements beforehand. Results show that models perform better in the 'Upper' setting, proving that parsing and prioritizing requirements from natural language is a key source of error.
Terminology
Summary
Recent progress in multimodal large language models (MLLMs) has fueled enthusiasm for their potential as autonomous agents, but current systems struggle when faced with complex, structured user requirements that involve both mandatory and optional constraints. This work addresses this critical gap by examining how MLLMs reason under three distinct requirement scenarios: must-have requirements that uniquely determine a solution, multiple answers where nice-to-have requirements dictate prioritization, and cases where no solution exists requiring abstention. The paper introduces First Things First Reinforcement Learning (FTFRL) to explicitly optimize reasoning over these multi-priority user requirements, demonstrating that enhancing requirement-aware reasoning capability provides a simple yet effective pathway to improve the generalization of MLLM agents across popular logical and mathematical reasoning tasks.
The Problem: Catastrophic Failures in Requirement Prioritization
Existing MLLMs frequently exhibit catastrophic failures
in real-world service scenarios because they lack a mechanism to prioritize hard requirements over soft requirements. The paper constructs the First Things First Benchmark (FTF-BENCH), which contains 3,649 carefully constructed problems reflecting realistic service scenarios across e-commerce, booking, and map/ride-hailing. These tasks are categorized into three settings: (i) must-have requirements with Single Answer (a unique valid candidate satisfies all must-haves); (ii) Multiple Answers where candidates are ranked via nice-to-have requirements; and (iii) Unanswerable, where the agent should abstain from generating a response.
Evaluation shows that current MLLMs often misinterpret task requirements, violate must-have requirements, and produce invalid solutions,
especially in the multiple-answer and unanswerable settings.
The Proposed Solution: First Things First Reinforcement Learning (FTFRL)
To address this gap, the authors propose FTFRL, a reinforcement learning approach designed to enhance requirement-aware reasoning. This method moves beyond standard SFT or RLHF by training MLLMs with a multiobjective reward that checks well-formed outputs, answer correctness, and requirement classification into must-have and nice-to-have.
The reward framework is composed of four key dimensions:
-
Format Reward: Verifies adherence to an XML-style schema for requirement classification, intermediate reasoning steps, and final answers.
-
Accuracy Reward: Measures whether the final prediction matches the ground truth answer using a MLLM-based
Judger
function (Ranswer). -
Requirement Reward: Assesses correct identification and classification of requirements into must-have and nice-to-have categories using Macro-averaged F1 score (Rrequirement). This component is crucial because it
penalizes both false positives (e.g., over-constraining) and false negatives (e.g., ignoring essential requirements).
-
KL Constraint: Used to keep the updated policy close to a reference policy model to stabilize updates.
Evaluation and Results on FTF-BENCH
The authors systematically evaluate state-of-the-art MLLMs, including proprietary models like Gemini 2.5 Pro and GPT-5, against FTF-BENCH under two settings: Direct (reading the original colloquial request) and Upper (receiving structured gold requirement labels). The results confirm that Upper exceeds Direct across most scenarios,
proving that the primary source of error is not visual perception alone but the failure to correctly parse and prioritize requirements from natural language. Specifically, FTF-RL yields substantial improvements on Qwen2.5-VL models, with gains exceeding 26% in the multiple-answer scenario for Qwen2.5-VL-7B after reinforcement learning.
Generalization Across Reasoning Benchmarks
The paper investigates the transferability of requirement-aware reasoning by comparing MLLMs trained on FTF-BENCH with their baselines on other logic and math reasoning benchmarks, including LogicVista, MathVision, and InfoQA. The findings reveal that even without any explicit training on these reasoning benchmarks, the MLLMs exhibit consistent improvements after reinforcement learning
on most such tasks. This suggests that requirement-aware reasoning not only strengthens the understanding of complex user intent but also stimulates the general reasoning ability of MLLMs,
indicating that learning to prioritize requirements transfers to broader reasoning skills.
Conclusion and Future Directions
The study concludes that enhancing requirement-aware reasoning capability provides a simple yet effective pathway to improve generalization of MLLM agents.
The authors note several key observations, including:
-
The most severe errors occur in the multiple-answer and unanswerable settings where models must resolve trade-offs or abstain.
-
Scaling effects are not monotonic; smaller models may lack the capacity to distinguish requirements, while larger models can exhibit an opposite failure mode by
overfitting to the surface form of user prompts.
-
The multi-objective RL framework is crucial, as removing any reward component leads to a consistent performance drop across all datasets.
Improvements for AI systems
Here are the specific improvements that can be made to existing Multimodal Large Language Models (MLLMs) by implementing the proposed First Things First Reinforcement Learning (FTF-RL) framework, along with what these improved systems can achieve:
-
The core improvement is a shift from generic instruction following to a structured, requirement-aware reasoning pipeline. Existing MLLMs often fail because they treat
must-have
(hard constraints) andnice-to-have
(soft preferences) requirements with equal weight, leading to catastrophic failures where they either violate essential conditions or overfit irrelevant details. -
The improved system can reliably distinguish between necessary conditions and optional preferences in complex, real-world scenarios (e-commerce, booking, map navigation).
-
Specifically, the improved MLLM will be capable of:
Discussing complex trade-offs under strict constraints (e.g., Book the cheapest hotel with two bedrooms AND a sea view if possible
). The system will correctly prioritize the mandatory requirements (price and capacity) before optimizing for the optional requirement (sea view), ensuring feasibility is never compromised.
-
The system will exhibit robust error handling in
Unanswerable
scenarios. Instead of hallucinating a plausible but incorrect solution when no candidate meets all hard criteria, the improved agent will reliably abstain from generating an answer, thus preventing costly errors in autonomous decision-making systems. -
The improved MLLM will learn to generate verifiable intermediate reasoning steps (as enforced by the XML format reward) that explicitly separate requirement classification from final decision-making. This allows for
interpretability,
enabling researchers and users to diagnose exactly where a failure occurred—whether it was a misunderstanding of the input image, misclassification of a constraint, or flawed trade-off selection. -
The system will demonstrate enhanced generalization across diverse reasoning tasks (LogicVista, MathVision, InfoQA) even when trained primarily on service scenarios. This suggests that learning to prioritize requirements provides a fundamental boost to the model's ability to perform complex logical and mathematical deduction in multimodal contexts, rather than just rote instruction following.
-
The system will be trained with a multi-objective reward function that simultaneously optimizes for:
Discussing correct output format (XML schema adherence).
Achieving high answer accuracy against ground truth.
Accurately classifying requirements into must-have
vs. nice-to-have.
-
By optimizing the requirement classification reward (using Macro-F1 score), the system will learn to penalize both false positives (treating a nice-to-have as mandatory) and false negatives (ignoring a must-have), resulting in balanced optimization that faithfully represents user intent.
-
The improved agent will exhibit better scaling behavior. The research indicates that while very large models can overfit to surface forms, the FTF-RL training process helps scale models more effectively by ensuring they maintain a balanced handling of requirement prioritization across different model sizes (e.g., the 32B model showing superior balance compared to the 72B model).
-
The system will be capable of generating highly structured, verifiable outputs (JSON format with explicit reasoning tags), which is crucial for integrating MLLMs into production agentic workflows where output parsing and downstream logic are critical.
Sources
- Qwen2.5-VL Technical Report
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System
- Evaluation of Instruction-Following Ability for Large Language Models on Story-Ending Generation
- MPCC: A Novel Benchmark for Multimodal Planning with Complex Constraints in Multimodal Large Language Models
- LearnAct: Few-Shot Mobile GUI Agent with a Unified Demonstration Benchmark
- TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
- SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
- MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning
- Instruction Tuning with GPT-4
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- A Fano-Style Accuracy Upper Bound for LLM Single-Pass Reasoning in Multi-Hop QA
- LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts
- R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
- CFBench: A Comprehensive Constraints-Following Benchmark for LLMs
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models