Modeling The Object Representations Underlying Human Physical Reasoning

arXiv:2602.12486 · cs.CV, cs.AI · Submitted 2026-02-12 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Modeling The Object Representations Underlying Human Physical Reasoning".

Jane: Humans appear to represent objects for intuitive physics with coarse,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So what’s the main point of "Modeling The Object Representations Underlying Human Physical Reasoning"? Essentially, the authors are investigating if vision models develop those human-like bodies—those coarse shapes with smooth concavities—and under what conditions they emerge.

Jane: They propose using a time-to-collision behavioral paradigm to compare model predictions against actual human movement, using a specific alignment metric to measure the difference between how models and humans treat concave versus convex regions.

Lu: The core finding they highlight is that alignment with human behavior follows an inverse U-shaped curve depending on the model's size, training time, and effective capacity. Small or briefly trained models tend to under-segment into blobs, while very large models over-segment with fine boundary wiggles.

Meng: That inverse curve tells us that there's a sweet spot, an intermediate granularity where the model's concavity effects most closely mimic human concavity effects. It sounds like the researchers are showing that this isn't just about accuracy but about capturing a specific type of geometric feature needed for physics intuition.

Lalam: If we can map these computational conditions—like intermediate pruning or mid-training checkpoints—to achieving better representations, it gives us a tangible way to tune our models toward more physically plausible outputs.

Conclusion: Tom: Looking at the authors and the title of "Modeling The Object Representations Underlying Human Physical Reasoning," I think they’ve managed to lay out a really clear path connecting computational mechanics to human perception of objects.

Jane: It seems they are suggesting that these intuitive, coarse bodies aren't some innate bias but rather an emergent property that arises from the constraints placed on the model during training, like having limited capacity or time.

Lu: That idea is compelling because it suggests that if we can understand these resource constraints, we might be able to design models specifically to capture physics-efficient representations for things like object layout and physical affordances.

Meng: Practically speaking, the implication is that instead of just pushing for higher pixel accuracy in segmentation, we might actually benefit from techniques like intermediate pruning or specific training schedules to get those human-like coarse bodies.

Lalam: For our culture here at the startup, this means we can start exploring these checkpoints and pruning strategies not just for speed, but for building representations that inherently support more intuitive physical reasoning in our AI systems.

Tom: It really boils down to this: finding that ideal body granularity isn't a magic setting; it’s a specific regime reached by combining mid-training checkpoints, intermediate pruning strengths, and intermediate-sized architectures.

Jane: So the authors are arguing that not too fine and not too coarse is actually what appears optimal for physical reasoning in these representations.

Lu: That consistent filling-in effect they observe strongly suggests that similar coarse encodings could appear in biology as an efficient compromise under limits on brain size, neural capacity, and metabolic cost.

Meng: I see the practical implication being that we need to stop optimizing purely for boundary wiggles and start looking at how different levels of complexity affect the physical plausibility of the predicted object layout.

Lalam: And from my perspective, this work gives us a framework to investigate how our vision models can be shaped to better reflect the kinds of approximations humans use when they're doing physics reasoning.

Harvard University · Weizmann Institute of Science · Bocconi University

cs.CV, cs.AI

Submitted: 2026-02-12

Updated: 2026-10-01

Importance score: 87/100

The gist: Humans appear to represent objects for intuitive physics with coarse, volumetric “bodies” that smooth concavities – trading fine visual details for efficient physical predictions – yet their

Key concepts

Concavity Effect
This measures how a model's prediction of a shape changes when comparing concave (indented) versus convex (outward curving) regions. Humans show a distinct preference for how they represent these differences, which models aim to mimic.
Ideal Body Granularity
This is the optimal level of detail in an object's shape—neither too rough nor too detailed. The study found this intermediate granularity best aligns with human physical reasoning, suggesting it is a naturally efficient compromise for physical prediction.
Resource Constraints
The paper argues that human-like bodies are not hard-coded biases but rather the result of limitations. When models have limited training time or capacity, they simplify shapes to save resources, leading to smooth, coarse representations that mimic how humans intuitively process physics.

Terminology

Summary

Humans appear to represent objects for intuitive physics with coarse, volumetric “bodies” that smooth concavities – trading fine visual details for efficient physical predictions – yet their internal structure is largely unknown. This research investigates whether vision segmentation models acquire these human-like bodies and under what conditions they emerge, suggesting that such representations arise from resource constraints rather than bespoke biases.

The gist

Across all manipulations (model size, training time, and effective capacity), alignment with human behavior follows an inverse U-shaped curve: small/briefly trained/pruned models under-segment into blobs; large/fully trained models over-segment with boundary wiggles; and an intermediate “ideal body granularity” best matches humans.

How it works

The study introduces a comparison pipeline and alignment metric using a time-to-collision (TTC) behavioral paradigm to compare vision model predictions against human behavior. The core of the experiment involves varying three key factors: model size, training time, and effective capacity via unstructured pruning. The pipeline maps the vision model’s output to human behavior by simulating future motion based on inferred object masks and true kinematics.

The comparison is quantified using a metric that measures the difference between models' concavity effects and human concavity effects:

  1. The study defines the condition-averaged human TTC as:

TTChuman(τ, g) = 1 / N (τ, g) X n∈N(τ,g) TTC(n)human.

  1. It then calculates the human concavity effect at ground-truth time τ as:

∆human(τ) = TTChuman(τ, concave)−TTChuman(τ, convex).

  1. An analogous quantity is computed for the model: ∆model(τ) = TTCmodel(τ, concave) − TTCmodel(τ, convex).

  2. The primary alignment metric is the absolute difference between these effects: E(τ) = ∆model(τ) − ∆human(τ). Lower E¯ indicates that the model’s concavity–convex difference more closely matches humans’ concave–convex differences.

Key Findings Across Manipulations

The analysis systematically varied three factors to observe how human-like bodies emerge:

  1. Training Time: Moderately approximate body representation emerges from medium training time, and aligns best with human behavior. The error E¯ is high at the very beginning of training, decreases to a minimum at an intermediate number of updates, before rising again as the model continues to optimize for pixel-accurate segmentation.

  2. Pruning of Neurons: Magnitude-based pruning showed that Intermediate pruning levels yield the lowest error, indicating that human-like concave–convex differences arise at an intermediate effective capacity. With no pruning, models produce overly fine, fragmented masks; with very strong pruning, they collapse to coarse blobs and lose the concavity effect.

  3. Model Size: Intermediate-sized models (e.g., B1–B2) achieve the lowest error, resulting in a concave–convex TTC shift that most closely matches human behavior. Small models under-segment objects into overly coarse blobs, while the largest models oversegment with pixel-perfect masks.

The Ideal Body Granularity Regime

The results consistently point to an ideal body granularity regime where object bodies are neither too coarse nor too over-detailed, and the concavity effect best matches human behavior. This regime is reached by different routes: mid-training checkpoints, intermediate pruning strengths, and mid-sized architectures. This suggests that similar coarse encodings in biology may arise as an efficient compromise under constraints on brain size, neural capacity, and metabolic cost, rather than a speciesspecific characteristic. The consistent “fillingin” effect observed in the visual intuition suggests that models simplify concave regions into smoother, coarser body representations.

Computational Account

The paper concludes that human-like approximate object representations emerge from generic resource constraints rather than specialized inductive biases. "Shorter training, smaller networks, and limited representational capacity emphasize low-frequency geometry and smooth over concavities; prolonged optimization and greater capacity sharpen boundaries and drift toward recognition-oriented detail. This supports the view that not too fine, not too coarse" appears optimal for physical reasoning: bodies that capture mass layout and physical affordances without overfitting boundary wiggles. This suggests simple knobs—early checkpoints, modest architectures, or light pruning—to elicit physics-efficient representations from standard models.

Contributions

The study contributes a general pipeline for comparing segmentation models with human physical judgments and identifies computational conditions under which humanlike body representations emerge. It provides a mechanistic account of why humans may simplify object geometry during intuitive physical reasoning by focusing on the differential treatment of concave vs. convex regions under matched physical conditions, rather than generic segmentation coarseness.

Improvements for AI systems

Here are specific, actionable improvements for AI systems based on the findings of this research:


)Based on the paper Human-Like Coarse Object Representations in Vision Models, here are specific improvements for developing more physically intuitive and resource-efficient AI systems:

  1. Improvements in Segmentation Architectures (Focusing on Ideal Body Granularity):

  2. Improved Robustness to Resource Constraints (Focusing on Pruning):

  3. Optimized Training Regimes (Focusing on Training Time):

  4. Development of Physics-Aware Predictive Models:

  5. Architectural Refinement: Implement a Granularity-Aware Segmentation Head

  6. Improved Robustness to Resource Constraints: Implement Adaptive Pruning Strategies

  7. Optimized Training Regimes: Employ Dynamic Checkpointing for Coarse Representation Capture

  8. Development of Physics-Aware Predictive Models: Integrate Concavity/Convexity Feature Extraction into Downstream Tasks

The resulting improved AI systems can perform the following specific tasks:

  1. A system that predicts physical interactions (e.g., collision times, object trajectories) with higher accuracy in dynamic, real-world scenarios by using segmentation masks that capture the essential body structure—smooth approximations of mass distribution—rather than overly fine pixel details.

  2. A model capable of operating efficiently on edge devices or in low-resource environments (e.g., mobile robotics) without sacrificing crucial physical reasoning capabilities, as the system will be tuned to use only an intermediate level of segmentation granularity that balances recognition detail and physical affordances.

  3. An AI agent that exhibits human-like predictive behavior during interaction tasks, such as navigating crowded spaces or predicting object motion in a simulation, because its internal representation of objects aligns with how humans simplify geometry for intuitive physics.

Sources

Related papers