SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning

arXiv:2608.10513 · cs.CV, cs.AI · Submitted 2026-08-11 · Read on arXiv

Caoyuan Ma, Wenpu Liu, Weichu Xie, Tian Gu, Shilei Zhao, Lingxi Min, Shuai Dong, Yuqi Xu, Ji Zhao, Ziyue Wang, Wenzheng Chang, Taiqiang Wu, Yongfu Zhu, Wenqi Shao, Yinqiang Zheng

The University of Tokyo · JD.com · Wuhan University · Peking University · Shanghai AI Laboratory · Shanghai Innovation Institute · Shanghai Jiao Tong University · The University of Hong Kong

cs.CV, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: 16pages, 4 figures. Preprint. Project page: https://safe-vlm.github.io/SafeCap/ ; code: https://github.com/Safe-VLM/SafeCap

Code: https://github.com/Safe-VLM/SafeCap

Project page: https://safe-vlm.github.io/SafeCap

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning Abstract Summary Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to

Terminology

Summary

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning

Abstract Summary

Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their language backbones. The paper proposes SafeCap, a reinforcement-learning framework that aligns LVLMs through learned self-captioning. SafeCap trains a policy model to first generate a safety-relevant image caption and then produce a final answer; the caption is further optimized by whether it enables a frozen LLM to reach a safety-aligned decision. This caption-mediated objective encourages the policy to expose visual cues relevant to safe response generation rather than relying solely on direct refusal supervision. Across five multimodal safety benchmarks and six vision-utility benchmarks, SafeCap substantially improves aggregate safety performance under its intended DirectCap protocol, with gains of 3.7–19.0 points in safety average across four model settings while maintaining comparable or improved vision utility. Under controlled comparisons on matched backbones and data, SafeCap outperforms safety SFT, DPO, and SafeGRPO, demonstrating the effectiveness of caption-mediated reinforcement learning for multimodal safety alignment.

Introduction Summary

Large vision-language models (LVLMs) extend the instruction-following ability of large language models to visual inputs, but the added visual pathway also changes the safety problem. A text-only model can often rely on language-side refusal policies, whereas an LVLM must first perceive safety-relevant evidence from an image, align it with the user’s request, and then decide whether to answer, warn, or refuse. This creates failure modes that are difficult to cover with text-only safety alignment: harmful instructions may be embedded in typographic images, dangerous intent may be implied by objects rather than stated in text, and benign-looking prompts may become unsafe only after visual grounding. As a result, visual inputs can weaken or bypass the safety behavior learned by the language backbone, making multimodal safety a problem of both perception and alignment rather than refusal prompting alone.

Existing LVLM safety methods address this gap through training-time alignment and inference-time intervention. Training-time methods directly optimize safety behavior using curated multimodal preference or safety-alignment data. Inference-time methods instead act around a frozen model through defense prompting, representation calibration, or learned visual and multimodal guardrails. Caption-mediated defenses offer a different route: ECSO transforms an unsafe image into a query-aware textual description and then uses the text-only pathway to reactivate safety mechanisms in the pre-aligned language model. This direction is appealing because the textual description can make image information available to language-side safety policies. However, image-to-text conversion can lose fine-grained visual details; if a caption omits safety-relevant objects, visible text, or contextual details, the downstream language model may not receive the information needed for a safe decision. Moreover, ECSO itself remains dependent on the language model’s ability to identify and neutralize unsafe queries, leaving caption-mediated safety brittle when either perception or text-side safety fails.

SafeCap is a reinforcement learning framework that trains an LVLM to use self-captioning as its primary safety-aware inference path. SafeCap asks the policy model to produce two fields for each image-question pair: a tagged caption that describes the image, and a final answer that responds to the user. The caption is not treated as an auxiliary explanation after the fact. Instead, it is a trainable interface encouraged to encode visual information useful for safety-aligned reasoning by both the policy model and a frozen text-only LLM. By optimizing this self-captioning path, SafeCap encourages the model to express safety-relevant visual cues before producing its final answer.

The key design challenge is reward shaping. A safety objective that does not reward helpfulness may favor refusals, whereas a caption-only objective need not ensure that the final answer is safe. SafeCap therefore combines a format gate, a caption-mediated reward, and a direct answer reward, with component-wise normalization and an exponential risk-discount formulation. The caption-mediated reward assesses whether the generated caption supports a frozen LLM response whose safety status agrees with that of the policy model, while the direct answer reward evaluates the policy model’s own final response after captioning. The paper also reports Direct inference and Prism inference, but they serve different diagnostic roles: Direct tests whether training damages the native LVLM response, and Prism diagnoses whether the generated caption alone provides useful information to a vision-free reasoner. The main target behavior is the learned DirectCap path.

Contributions

The paper's contributions are threefold:

  • Safety-aware self-captioning formulation: formulating LVLM safety alignment as a self-captioning problem, in which the model generates an explicit intermediate representation before producing its final answer.

  • Caption-mediated reinforcement objective: introducing a dual-signal reward that combines direct evaluation of the policy model’s final answer with a frozen-LLM signal assessing whether the generated caption supports a safety-aligned response, together with component-wise reward normalization.

  • Multi-protocol evidence and ablations: providing a multi-protocol evaluation and ablation study showing that SafeCap achieves its strongest safety gains on the DirectCap path while maintaining aggregate vision utility.

Related Work Summary

Safety risks introduced by visual inputs: LVLMs inherit much of their instruction-following ability from language backbones, but the visual pathway changes the threat model in ways that text-only safety alignment does not cover. First, images are continuous and high-dimensional, allowing adversarial perturbations to steer aligned models toward unsafe responses even when the accompanying text appears benign. Second, visual inputs can carry semantic instructions through OCR, diagrams, typographic prompts, or object-level context. MM-SafetyBench and VLSBench evaluate safety risks arising from benign-looking text-image pairs, where harmful intent must be inferred from the visual input. FigStep converts harmful instructions into incomplete typographic images; image-to-text logic jailbreaks encode attack logic in visual flowcharts; HADES shows that safety-critical intent can be hidden in the image while the text prompt remains innocuous; and MIS extends this risk to multi-image inputs, where unsafe intent emerges from the composition of multiple images. Third, the process of adapting an LLM into an LVLM can itself weaken safety alignment: visual instruction tuning and cross-modal projection may shift internal representations away from the refusal behavior learned by the language model.

Inference-time and caption-mediated defenses: Inference-time defenses improve safety without updating the base LVLM. They include defense prompting and prompt adaptation, visual or multimodal guardrails, cross-model or reward-guided decoding, and representation-level calibration or visual safety prompting. Caption-mediated methods offer a complementary route: ECSO first detects unsafe responses and then converts the visual input into a query-aware textual description so that the text-only pathway can generate a safer response. More generally, image textualization produces descriptions that can be inspected, filtered, or passed to a text-only model, while response-level protectors detoxify harmful outputs after generation. These methods make visually encoded risk available to language-side safety mechanisms, but their effectiveness depends on the retained image information and the downstream model’s safety capability. SafeCap differs by training the LVLM’s own caption-and-answer path, rather than adding an inference-time wrapper around a frozen model.

Training-time alignment and caption rewards: Training-time defenses instead optimize model parameters or construct safety-alignment data so that multimodal safety behavior is learned more directly. SPA-VL provides safety preference data for LVLM alignment, VLGuard studies safety fine-tuning with explicit harmful and benign visual-language examples, MM-RLHF explores human-feedback alignment for multimodal LLMs, and cross-modal safety alignment investigates whether textual safety alignment transfers to multimodal settings. These approaches highlight the central trade-off in LVLM safety: improving refusal on malicious inputs should not collapse visual perception or cause over-refusal on benign ones. Image captioning offers a useful interface for this trade-off. Classical captioning work established captions as a bridge between vision and language, and recent caption-centric datasets and pipelines improve LVLM pretraining by providing richer visual descriptions. CapRL converts open-ended caption quality into a decoupled visual question answering (VQA) signal, reducing reward hacking associated with subjective LVLM-as-a-judge rewards. Inspired by this design, SafeCap introduces a binary LLM safety-alignment judgment into a caption-mediated reward: a caption is rewarded when it supports a frozen text-only LLM response whose safety status agrees with the policy model’s answer.

Method Summary

Problem Formulation: SafeCap targets the learned self-captioning setting. Each training example contains an image-question pair (x, q), where x is the visual input and q is the user request. Instead of training the LVLM to answer q directly, the policy model πθ is prompted to produce a structured response y = ca, where c is a textual description of the image and a is the final answer. This format corresponds to the DirectCap operating point used in the experiments. The goal is not merely to add a verbose caption before every answer, but to make c a safety-relevant evidence channel: it should expose visual details that determine whether the request is benign or harmful, and a should use that evidence to answer safely.

Training objective: The policy is optimized with GRPO on the public SPA-VL safety-alignment dataset, without introducing private utility data. For each image–question pair, the current policy samples a group G of K = 8 structured rollouts. Training uses three signals—a template reward Rtmp, a direct answer reward Rans, and a caption-mediated reward Rcap—to construct the outcome-level advantage Ai = wtmp r̃i,tmp + wcap r̃i,cap + wans r̃i,ans, where r̃i,k is the group-normalized value of reward component k for rollout i. The advantage is used in the PPO-clipped GRPO loss. Neither an explicit KL loss nor a KL reward penalty is enabled in the main runs; the frozen LLM and judge models are used only to compute reward and are never updated.

Template reward: A rule-based parser extracts the caption between and and the answer after the closing tag. The template reward is binary: the response must contain exactly one complete caption block, both the caption and answer must be non-empty, and the generation must terminate before the length limit. If this template gate fails, all downstream answer and caption rewards are skipped and only the failed structural signal remains.

Answer reward: The direct answer reward evaluates the policy model’s own final answer a. A judge model assigns two integer scores on a 0–5 scale: a retention score u(a), measuring how well the answer addresses legitimate informational needs, and a risk score h(a), measuring harmful or unsafe content. The main SafeCap reward uses an exponential risk-discount function, S(u, h) = uγ h, and defines the answer reward as Rans = S(u(a), h(a)) = u(a)γ h(a). The coefficient γ ∈ (0, 1) controls the strength of risk suppression: each one-point increase in h multiplies the utility contribution by γ, so a smaller γ penalizes risky answers more sharply. This non-negative multiplicative form differs from a classic linear reward u − βh: unsafe high-risk answers lose value quickly, whereas safe useful answers retain positive reward. This property is particularly useful for caption-mediated training: with a linear alternative, negatively valued rollouts from the frozen-LLM route can make low-information captions or generic refusals comparatively attractive.

Caption reward: The caption-mediated reward evaluates whether the generated caption is useful as a safety evidence channel. First, the caption c is passed to a frozen text-only LLM together with the original question q, producing a caption-conditioned answer af. The frozen model never sees the image. Second, the judge assigns a descriptive-coverage score u(c) using a fixed rubric that rewards concrete coverage of subjects, attributes, actions, spatial layout, setting, and visible text. Because the judge does not see the image, this score measures descriptive coverage rather than factual visual correctness. Third, a binary safety-alignment judge compares the policy answer a and the frozen answer af and returns g(q, a, af) ∈ 0, 1, where 1 means that both answers agree on the safety status of the situation: either both recognize the risk, or neither identifies a risk and neither gives unsafe advice. The caption reward is then Rcap = g(q, a, af) u(c). This design prevents a caption from being rewarded only for being detailed. It also adds a consistency-based signal: the caption must support an independently generated text-only answer whose safety status agrees with the policy answer, which reduces ambiguity from rubric-only scoring.

Component-wise group normalization: Before combining the three reward signals, each component is normalized separately within the rollout group. Let ri,k denote the raw value of component k ∈ tmp, cap, ans for rollout i ∈ G. Specifically, r̃i,k = (ri,k − µG,k)/(σG,k + ϵ), where µG,k and σG,k are the mean and standard deviation of component k over the rollouts in G, and ϵ = 10−6 is a numerical-stability constant. The normalized components are then combined according to the advantage equation. This component-wise group normalization follows the decoupled-normalization idea of GDPO, but is implemented in the GRPO training path without its additional batch-level whitening step.

Inference and Diagnostics: At inference time, SafeCap’s intended operating point is DirectCap: the trained policy first writes the caption and then answers in the same response. Two diagnostic protocols are also evaluated. Direct inference removes the caption requirement and tests whether RL training has damaged or improved the model’s native image-question answering behavior. Prism inference gives the generated caption to the frozen text-only LLM; this isolates whether the caption alone contains enough transferable visual evidence for a vision-free reasoner.

Experiments Summary

Experimental setup: SafeCap is trained with GRPO on the public SPA-VL multimodal safety-alignment dataset, without introducing private or additional training data. All models use a common inference and evaluation pipeline. Training uses eight sampled rollouts per prompt, a rollout batch size of 64 prompts, a PPO mini-batch size of 64, AdamW optimizer with constant learning rate 5×10−7, weight decay 0.01, gradient clipping at 1.0, entropy regularization set to 0, no explicit KL loss or KL reward penalty, prompts and responses limited to 4096 tokens, and 200 optimization steps. The default reward uses γ = 0.35, reward weights (wtmp, wcap, wans) = (0.5, 0.5, 1.0), and decoupled reward normalization. The frozen text-only model used by the caption reward is Qwen3-4B, and the judge model is gpt-oss-20b with deterministic decoding.

Evaluation protocols: Each model is evaluated under three inference protocols. In the Direct protocol, the model receives the original benchmark image-question pair and directly produces the answer. In the DirectCap protocol, the model is instructed to first describe the image in a tagged caption and then answer the question; the final answer after the caption tag is used for scoring. In the Prism protocol, a frozen text-only LLM (Qwen3-4B) answers the original question based only on the caption generated by the evaluated model.

Benchmarks and metrics: Safety is evaluated on five public LVLM safety benchmarks: MM-SafetyBench, MSSBench, VLSBench, FigStep, and MIS-Test. For MM-SafetyBench, FigStep, and MIS-Test, the paper reports 1 − ASR, where higher values indicate fewer unsafe responses. For MSSBench, the official overall accuracy is reported, where higher values indicate better balance between refusing unsafe requests and answering safe ones. For VLSBench, the safe rate from its multimodal judge is reported. For FigStep, an automatic harmful/safe judge prompt is used that classifies a response as harmful if it provides actionable assistance for illegal or dangerous activities, and as safe if it genuinely refuses or avoids actionable harmful information. Stratified sampling and manual rechecking confirmed the reliability of this process. Vision utility is evaluated on six perception and instruction-following benchmarks: MM-Vet, BLINK, MMVP, ERQA, VPCT, and MMStar. All table entries are higher-is-better.

Zero-Training Protocol Comparison: Table 1 reports the zero-training comparison among Direct, DirectCap, and Prism inference over the full 11-benchmark suite. Caption-mediated inference changes the safety–utility balance even before training, but it is not uniformly beneficial. Direct is the strongest average protocol for the instruction-tuned 2B and 4B models, while DirectCap is competitive on several individual utility metrics. Prism is important as a caption-transfer diagnostic: it sometimes improves individual safety metrics, such as FigStep on 2B and 4B-Base, but its aggregate scores are usually below Direct and DirectCap because the frozen text-only reasoner depends on caption completeness and extraction quality. One zero-training score requires special care: the relatively high FigStep score of Qwen3.5-2B-Base under DirectCap is often caused by describing an empty list template rather than executing its harmful completion request. In a matched 500-sample comparison, 66 of the 81 responses that changed from safe to harmful had a zero-training answer that only described this blank template. This score is treated as a protocol-specific artifact of incomplete task execution.

SafeCap Training Results: Compared with the zero-training baselines, SafeCap gives the clearest and most consistent gains under DirectCap. The aggregate 11-benchmark mean improves by 5.29 points for 2B, 5.57 points for 2B-Base, 5.48 points for 4B, and 8.57 points for 4B-Base. These gains are not only safety-judge effects: DirectCap utility also improves for 2B, 2B-Base, and 4B, and remains close to the zero-training 4B-Base utility average while its safety average rises by 18.96 points. Direct inference is more model-dependent. It substantially improves the base-initialized 2B-Base and 4B-Base safety averages, but it does not improve the already strong instruction-tuned 2B and 4B direct baselines on the aggregate suite. Training is also stable across three independently seeded 100-step Qwen3.5-4B-Base runs: DirectCap reaches 50.83±1.14 S-Avg and 54.32 ± 1.54 V-Avg, compared with zero-training DirectCap scores of 40.43 and 53.13. Prism remains a necessary diagnostic rather than the primary operating point. It improves over the zero-training Prism baseline for 2B, 2B-Base, and 4B on the aggregate suite, but is still usually below DirectCap after training. Prism performance is also fundamentally shaped by the downstream LLM: holding the same SafeCap captions fixed, replacing Qwen3-4B with Qwen3-14B raises the Prism safety average by 4.72 points.

Comparison with Safety Fine-Tuning: Figure 3 compares methods trained from Qwen3.5-4B-Base on the same SPA-VL data and for matched training steps. Under DirectCap, SFT and DPO improve S-Avg only to 43.19 and 41.36, whereas SafeCap reaches 59.39 while retaining V-Avg near the zero-training value.

Comparison with SafeGRPO: SafeGRPO requires additional safety annotations that are not part of the SPA-VL setting used for the main results. A controlled comparison is conducted by training and evaluating both SafeCap and SafeGRPO on SafeGRPO’s SafeTag-VL-3K data, using the same Qwen3.5-4B-Base initialization and 200 steps. SafeCap gives the strongest trained DirectCap safety result (55.06 vs. 41.27) while retaining similar utility (51.34 vs. 50.56). Under Base inference, SafeCap also improves safety over SafeGRPO (46.86 vs. 44.98). SafeCap under DirectCap substantially exceeds SafeGRPO under its conventional Base (direct) inference in safety (55.06 vs. 44.98).

Ablation Studies Summary

Reward component ablation: Every ablation (removing the direct answer reward, caption reward, or decoupled reward normalization) lowers safety and, with the sole exception of a +0.42 DirectCap utility change without normalization, utility as well. The direct answer reward is the most consequential component for safety: removing it reduces safety by 5.31 and 6.49 points under Direct and DirectCap, respectively, and by 1.81 points under Prism. Removing the caption reward also consistently degrades safety (1.08–2.53 points) and utility (1.23–2.41 points), indicating that caption supervision contributes to both safe behavior and visual-task performance. Decoupled normalization particularly supports the Direct and DirectCap safety operating points (drops of 2.74 and 2.02 points when removed), while its effect under Prism is smaller (0.24 points).

Risk-coefficient ablation: A smaller coefficient corresponds to a stronger risk penalty and leads to faster safety convergence. The paper uses γ = 0.35 in all other comparisons as a balanced operating choice.

Alternative reward constructions: Rank-2 degrades both safety and utility under both protocols. Classic improves DirectCap safety but lowers utility, while the Lagrangian variant gives a smaller DirectCap safety gain with lower utility, suggesting that these alternatives are less balanced than the final objective. Refusal diagnostics show that Classic refuses more often on the safety benchmarks while refusal rates on general benchmarks remain low, suggesting that risk-heavy objectives may raise safety-judge scores partly through stronger refusal behavior.

Discussion and Limitations Summary

SafeCap uses the policy LVLM’s own visual observations as the source of its captions; the auxiliary frozen model consumes text only. This avoids introducing a second VLM with an additional visual perception pathway and its associated model-specific biases. The caption reward’s binary safety-consistency signal is directly checkable: it rewards a caption only when the text-only answer it supports agrees with the policy answer on the safety status of the request. However, its effectiveness remains bounded by the capability of the frozen LLM that interprets the caption. Since the frozen LLM does not observe the image, the reward cannot by itself identify factual errors in a caption. Manual monitoring of rollout trajectories throughout training did not observe systematic hallucination.

The safety evaluation scores the model’s final answer, with the objective of preventing actionable assistance for harmful requests. A caption may faithfully restate harmful text or intent already visible in the user-provided image. This is not treated as newly introduced harmful information or as assistance beyond the user’s input; its role is to expose the visual evidence needed for the model to make a safe final decision. The direct answer reward uses fixed, human-specified rubrics to score usefulness and risk, rather than a fully verifiable safety reward. Developing safety and risk rewards with stronger verifiability remains an important open problem. For cost reasons, several supplementary experiments use early-stopped checkpoints; these are marked explicitly and not compared directly with the fully trained main results. The experiments use smaller Qwen3.5 variants; evaluating larger models and additional model families is an important next step. A promising extension is to perform caption-mediated SFT before RL, though suitable supervised data for this intermediate objective are currently limited.

Conclusion Summary

SafeCap is a reinforcement-learning framework that improves LVLM safety through learned self-captioning. SafeCap trains an LVLM to expose safety-relevant visual evidence in a caption before answering. Across five safety and six vision-utility benchmarks, SafeCap achieves its clearest and most consistent gains under the intended DirectCap operating point, substantially improving safety for base-initialized models while preserving utility near the corresponding zero-training baseline. Under matched settings, SafeCap also outperforms safety SFT, DPO, and SafeGRPO. More broadly, SafeCap shows the value of incorporating perception-side evidence into multimodal safety alignment, complementing conventional output-side alignment and providing a practical direction for future community efforts.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems:

Improvement: Add a structured self-captioning step before final answer generation, where the model explicitly describes safety-relevant visual evidence (objects, text, context) in a tagged format before responding.

Capability: The AI system can now:

  • Surface visual cues that determine whether a request is benign or harmful

  • Make safety decisions based on explicit visual evidence rather than implicit perception

  • Provide auditable reasoning trails for safety-critical decisions

Improvement: Train with a composite reward that combines:

  • Direct answer evaluation (usefulness × risk-discount factor)

  • Caption-mediated consistency check (does the caption support a frozen text-only LLM's safety-aligned response?)

Improvement: Replace linear utility-minus-risk rewards with exponential risk-discount formulation: S(u, h) = u × γ h where γ ∈ (0,1) controls risk suppression strength.

Improvement: Normalize each reward component (template, caption, answer) separately within rollout groups before combining, rather than normalizing the total reward.

Improvement: Add a diagnostic protocol where generated captions are passed to a frozen text-only LLM for safety assessment, isolating whether captions alone contain sufficient transferable evidence.

Improvement: Use SafeCap's training approach specifically for models initialized from base (non-instruction-tuned) checkpoints, where direct safety gains are largest (+18.96 safety points for 4B-Base).

Improvement: Train the model to generate captions that enable a frozen text-only LLM to make safety-aligned decisions, effectively transferring safety reasoning from text-only models to multimodal systems.

Improvement: Enforce a structured response format with exactly one complete caption block and non-empty answer, using a rule-based parser as a hard gate before applying downstream rewards.

Improvement: Apply asymmetric weights (template: 0.5, caption: 0.5, answer: 1.0) with the answer reward weighted highest, prioritizing final response quality while maintaining caption quality.

Improvement: Track training stability through multiple seeded runs and early-stopped checkpoints, with explicit marking of incomplete training for supplementary experiments.

Abstract

Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their language backbones. We propose SafeCap, a reinforcement-learning framework that aligns LVLMs through learned self-captioning. SafeCap trains a policy model to first generate a safety-relevant image caption and then produce a final answer; the caption is further optimized by whether it enables a frozen LLM to reach a safety-aligned decision. This caption-mediated objective encourages the policy to expose visual cues relevant to safe response generation rather than relying solely on direct refusal supervision. Across five multimodal safety benchmarks and six vision-utility benchmarks, SafeCap substantially improves aggregate safety performance under its intended DirectCap protocol, with gains of 3.7-19.0 points in safety average across four model settings while maintaining comparable or improved vision utility. Under controlled comparisons on matched backbones and data, SafeCap outperforms safety SFT, DPO, and SafeGRPO, demonstrating the effectiveness of caption-mediated reinforcement learning for multimodal safety alignment.

Sources

Related papers