Teaching Vision-Language Models to Use the Scale They Are Given: Label-Free Equivariance Training for Metric Physical Reasoning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Teaching Vision-Language Models to Use the Scale They Are Given".
Jane: Metric questions about monocular video carry a symmetry that costs nothing to state or to check: the answer is proportional to the reference the model is given.
Tom: First, who's behind it and why it matters.
Paper summary: Jane: To put it simply, they are using this exact symmetry to supervise the model, which means you don't need any answers written down for training; you just use the inherent physics of how measurements relate to each other to guide the learning process.
Lu: The implication here is that we can inject formal physical constraints directly into the training loop for vision-language models, effectively building symmetry into a white-box model structure, which is something I’ve been thinking about a lot in my own research.
Meng: Practically speaking, this suggests that instead of needing massive datasets annotated with precise measurements, we could train models using only the inherent physical consistency of the visual data and the reference scales provided by the input.
Lalam: For me, this means that future AI systems won't just be good at recognizing objects; they will be inherently better at reasoning about how those objects interact physically in a consistent way, which could really improve how we build culture around these kinds of intelligent systems.
Tom: So, to wrap up the main point of "Teaching Vision-Language Models to Use the Scale They Are Given: Label-Free Equivariance Training for Metric Physical Reasoning," it’s about using this exact symmetry to give vision models better metric grounding without needing any ground truth data or annotations.
Jane: It really boils down to teaching the model the fundamental physical rule that dictates how answers must change when you adjust the scale of the input measurements.
Lu: This work opens up a new avenue for imposing dimensional analysis constraints in machine learning, which is a powerful idea for making models more physically coherent across different scales.
Meng: It’s a way to test if a model is really grasping physical reality, not just pattern matching on specific examples, which I think has huge implications for how we validate these systems.
Lalam: If we can make models respect this exact relationship between scale and answer, it means they will handle new physical situations much more reliably than before.
Conclusion: Tom: So, we've been diving deep into this paper about EquiSD and how it fixes that measurement problem in vision models. Now we're getting to the wrap-up, and I want to talk about what this whole concept actually means for us listening right now.
Jane: Exactly! We’ve been looking at the technical details of how they use that homogeneity relation to train the model without needing any ground-truth answers, so let's look at those titles and authors. This paper is titled "Teaching Vision-Language Models to Use the Scale They Are Given: Label-Free Equivariance Training for Metric Physical Reasoning," and it was authored by a team of researchers who are clearly really pushing the boundaries of what we can teach these models about the physical world.
Lu: I think what’s fascinating is that they're not just tweaking existing methods; they've identified this exact physical symmetry as a fundamental constraint, which opens up entirely new ways to constrain model learning processes.
Meng: From an engineering standpoint, it’s exciting because if we can train models using only their own outputs and the inherent structure of the data, we drastically reduce our reliance on expensive human annotation efforts for metric tasks.
Lalam: For me, this means that in the future, vision systems won't just be good at spotting things; they'll possess an innate understanding of scale and proportion because they've been trained on that physical consistency.
Tom: That’s a big shift! So, the core idea is taking this exact physical relationship—where scaling measurements by a factor alpha forces the answer to scale by exactly alpha —and using it as a label-free training signal.
Jane: It really is simple to grasp if you think about it like this: we're giving the AI a rule about how its answers must behave when you stretch or shrink the world it sees, and that rule is derived purely from geometry.
Lu: And the way they formulate that E-step to read off the scale-free ratio implicitly in the model’s own answer is really clever; it’s a self-referential constraint on its reasoning process.
Meng: I wonder how robust this constraint holds up when we move from simulated environments, like those videos they tested against, to messy real-world footage with varying lighting and camera angles.
Lalam: If this method successfully transfers from synthetic data to actual QuantiPhy videos, it means the physical understanding we build in the lab can actually apply to unpredictable real-world situations.
Tom: That transfer capability is what really makes this compelling; they showed significant performance gains across different scales on unseen videos, which is a huge win for generalization.
Jane: It shows that by focusing on these fundamental physical symmetries instead of just memorizing specific examples, we can build models that are inherently more reliable when faced with new visual information.
Lu: The implication is that we could start imposing physical laws directly onto the architecture of vision-language models, moving them closer to true physical reasoning systems.
Meng: So, the real impact here is a potential reduction in the engineering overhead needed to make these complex AI systems capable of performing accurate metric tasks reliably.
Lalam: If this work helps build more physically consistent AI, it could fundamentally alter how we design and interact with intelligent systems across many different fields.
Kaizhen Tan, Yang Feng, Heqing Du, Siru Tao, Xin Xu, Hanzhe Hong
New York University · Columbia University · Carnegie Mellon University
cs.CV
Submitted: 2026-09-01
Updated: 2026-10-04
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 87/100
The gist: Metric questions about monocular video carry a symmetry that costs nothing to state or to check: the answer is proportional to the reference the model is given.
Key concepts
- Homogeneity Relation
- This is an exact physical constraint stating that if every quantity in a world space is scaled by a factor alpha, the correct answer must also scale by exactly alpha. This relationship arises from projective geometry and serves as the key to fixing how models handle metric reasoning.
- EquiSD Training Procedure
- This is a label-free training method that makes the homogeneity relation concrete. It involves two steps: first, finding a scale-free ratio implicit in the model's own answer, and second, fine-tuning the model on targets derived from this projection. This process enforces equivariance without needing ground-truth answers.
- Response Slope (s)
- This metric measures how much a model uses the supplied scale information. A slope of 1 means the model uses the reference scale perfectly, while a slope of 0 indicates the answer is determined only by what it remembers about objects on screen. This analysis shows current models only use scale information partially.
- Metric Grounding
- This refers to a model's ability to correctly convert visual measurements into physical units. The paper shows that by enforcing the homogeneity relation, models become much better at grounding metric questions because they learn how scaling affects the final answer, leading to improved accuracy across various scales.
Terminology
Summary
Metric questions about monocular video carry a symmetry that costs nothing to state or to check: the answer is proportional to the reference the model is given. This work introduces EquiSD, a label-free training procedure that exploits this exact physical symmetry—the homogeneity relation—to improve metric grounding in vision-language models without requiring any ground-truth annotations.
The Core Problem and Constraint
Metric questions about video require vision-language models to use supplied real-world references to convert visual measurements into physical units, but current models only use this scale information partially. The paper identifies the homogeneity relation as an exact, annotation-free constraint on metric physical reasoning,
stating that for a question where every world-space quantity is rescaled by a common factor α, the correct answer must change by exactly that factor: y(V, Sαq) = α y(V, q).
This constraint follows from projective geometry and is the key to fixing the deficit in metric grounding.
The EquiSD Training Procedure
EquiSD makes this exact symmetry concrete in two steps. First, it reads off the scale-free ratio implicit in the model’s own answer, which is the projection of that answer onto the family of functions [the symmetry] allows.
Second, it fine-tunes the model on that projection. The targets are derived from this process: The projected targets encode the equivariance relation without introducing any magnitude the model did not produce itself,
and training is done on exactly those targets
with a loss restricted to answer tokens only. This procedure requires no ground-truth answers and only one model query per training video.
Quantifying Model Deficit
The paper quantifies how much of the supplied scale a model uses by probing it with the rescaled prompt, presenting it with the rewriting Sαq. The response slope is measured using Eq. (2), where s = d log ˆy(α) / d log α,
which equals 1 for a model that uses the reference as given and 0 for one whose answer is set entirely by what it remembers about the objects on screen. This analysis shows that current models use this scale information only partially,
with accuracy staying concentrated near the familiar scale of the depicted objects.
Performance Gains and Generalization
Training with EquiSD significantly improves model performance across scales. On held-out simulated videos, EquiSD increases a 3B model’s median response slope from 0.66 to 0.94 and improves mean relative accuracy by 9.2 points across scales. The learned relation generalizes to unseen world scales and transfers without adaptation to real QuantiPhy videos,
where accuracy increases by 6.4 points, showing that the label-free signal retains 93% of what exact simulator answers buy where on simulated video it retained 66%.
Ablations and Further Insights
The research separates the constraint from self-distillation by testing five arms in Table 8. It demonstrates that fine-tuning on its own outputs does not reproduce the equivariance gain, whereas the EquiSD procedure does. Furthermore, comparing a model against its own un-intervened answer shows that the models hold the relevant physical fact and stop applying it once the answer has to be a measurement.
The study concludes that Physics supplies further constraints, among them conservation laws and invariance under a change of frame, each constraining how outputs relate rather than what they are.
Future Directions
The constraint is valid only within specific settings: a time base fixed by the video, worldspace anchors that are lengths or length rates, and a target of dimension L T −k.
Reaching beyond these boundaries requires a different relation rather than simply using a wider scaling factor α grid. The work also releases IVP-Sim, 1,000 rendered rigid-body videos with exact states and single-parameter interventions whose log-log sensitivity exponents are checked against their closed forms.
The gist
EquiSD, a label-free training procedure that exploits the homogeneity relation—the exact physical symmetry where scaling every world-space quantity by α results in the correct answer scaling by exactly α—improves metric grounding in vision-language models without requiring any ground-truth annotations. This method is effective across various model sizes and successfully transfers from synthetic data to real videos, recovering most of the performance gains offered by simulator supervision.
How it works
-
The homogeneity relation is identified as an exact constraint:
y(V, Sαq) = α y(V, q).
-
EquiSD performs an E-step: It queries the base model at a log-symmetric grid A and calculates the ratio that best explains the set:
log Rˆ = median [log ˆy(α) − log α − log ρ].
Improvements for AI systems
Here are specific improvements for AI systems based on this research:
-
Enhance Metric Grounding in Vision-Language Models (VLMs): Implement a training procedure, such as EquiSD, that exploits the inherent homogeneity relation of physical quantities (Eq. 1: the correct answer scales by exactly the same factor α when all world-space quantities are rescaled by α).
-
Enable Scale-Invariant Physical Reasoning: Train models using label-free supervision derived from their own predictions projected onto this scale-equivariant family of functions, rather than relying on expensive, hand-annotated metric ground truth for every query.
-
Improve Robustness to World Scale Changes: The improved system will maintain high accuracy (e.g., MRA > 90%) when presented with world scales significantly larger or smaller than those in the training data (e.g., 4 orders of magnitude), a capability that current models lack, leading to consistent performance across vast physical scenarios.
-
Increase Accuracy on Real-World Data: By leveraging EquiSD and transferring knowledge from synthetic simulations (IVP-Sim) to real-world videos, the system's accuracy on unconstrained real video tasks will increase significantly (e.g., gaining 6.4 MRA against untrained models).
-
Develop State-of-the-Art Physical Reasoning Benchmarks: Create new benchmarks that specifically test a model's reliance on metric grounding versus its general physical mechanism knowledge, using the scale-free control to rigorously separate these two deficits.
-
Improve Counterfactual and Intervention Capabilities: The system can better predict outcomes after interventions (e.g., changing initial conditions) by using the
relative arm
method in a test-time correction setting, allowing it to determine how much an outcome changes based on physical parameters without needing absolute unit grounding at inference time. -
Develop Self-Correcting Inferences: Implement an inference mechanism that can correct frozen models by identifying the specific scale at which their prediction matches their own ungrounded response, thereby recovering the exact ratio even without external ground truth.
Sources
- Qwen2.5-VL Technical Report
- LLMPhy: Parameter-Identifiable Physical Reasoning Combining Large Language Models and Physics Engines
- Test-Time Consistency in Vision Language Models
- Filtered-CoPhy: Unsupervised Learning of Counterfactual Physics in Pixel Space
- $\Delta$ynamics: Language-Based Representation for Inferring Rigid-Body Dynamics From Videos
- QuantiPhy: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language Models
- SIMPACT: Simulation-Enabled Action Planning using Vision-Language Models
- Ill-Posed by Design: Probing Evidence Use in VLMs
- CRIPP-VQA: Counterfactual Reasoning about Implicit Physical Properties via Video Question Answering
- Causal Scaffolding for Physical Reasoning: A Benchmark for Causally-Informed Physical World Understanding in VLMs
- Dimensionless machine learning: Imposing exact units equivariance
- PhysMind: From Video to Executable Worlds for Training-Free Physical Reasoning
- ContPhy: Continuum Physical Concept Learning and Reasoning from Videos
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models