Context-Dependent Affordance Reports in Vision-Language Models
cs.CL, cs.AI, cs.LG
Submitted: 2026-02-14
Updated: 2026-09-12
Comments: 16 pages, 2 figures. Substantial revision: corrects empty-response handling, withdraws percentage-of-meaning and latent-manifold interpretations, and adds matched-question controls on two model configurations. Code, data and source: https://doi.org/10.5281/zenodo.22721059
Code: https://github.com/studiofarzulla/semantic-vision
License: http://creativecommons.org/licenses/by/4.0/
The gist: Vision-language models produce different object and use descriptions under different persona prompts, but low overlap alone does not identify an affordance effect.
Terminology
Abstract
Vision-language models produce different object and use descriptions under different persona prompts, but low overlap alone does not identify an affordance effect. We audit an earlier seven-prompt study and add matched-question controls. In the historical Qwen pilot, 363 of 3,213 parsed responses contain empty object lists. These affect 2,037 of 9,244 comparisons, with the implementation assigning zero lexical overlap to every affected pair. Conditioning on nonempty reports raises pooled word Jaccard from 0.095 to 0.121 and sentence cosine from 0.415 to 0.511. A previously named chef-specific Tucker factor loses its concentrated loading under missing-cell and complete-nonempty analyses. We withdraw the functional-manifold interpretation and the conversion of similarity scores into percentages of meaning. A new experiment uses 48 images absent from the original pilot, four personas, a shared three-object task, two wordings, and two requested seeds in each of two model configurations. Qwen3.5-9B's persona-minus-wording cosine-distance contrast is 0.0146 (95% CI [-0.0003, 0.0295]; 37 complete images). Ollama llava:13b's persona-minus-wording cosine-distance contrast is-0.0266 (95% CI [-0.0412, -0.0121]; 32 complete images). Persona-associated variation does not uniformly exceed wording or sampling variation. The study provides a reproducible analysis of context-conditioned reports while separating response availability, content similarity, and the limits of inference from linguistic outputs.
Sources
- Leverage Task Context for Object Affordance Ranking
- RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation
- Qwen3-VL Technical Report
- Self-Explainable Affordance Learning with Embodied Caption
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering