The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "The Percept-V Challenge".
Jane: Multimodal Large Language Models (MLLMs) are being tested on simple visual perception problems to determine if they match human capabilities, and this research introduces Percept-V,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So Jane, we've just finished looking at the core of this paper, "The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?". The main point is that cognitive science research looks at visual perception as a sign of intelligence early in development. This study introduces Percept-V, which is a dataset made up of six thousand program-generated images spread across thirty domains designed to test seven specific skills from the TVPS-four framework <ref:2508.21143#pg0>. The big question they're asking is whether Multimodal Large Language Models can match human performance on these foundational visual perception tasks.
Jane: That sounds really fascinating, Tom; so the researchers are focusing specifically on those basic visual skills instead of just complex reasoning or world knowledge that many other benchmarks test. They are using a framework called TVPS-four to categorize perception into things like visual discrimination and form constancy, which helps them structure the testing around real cognitive science <ref:2508.21143#pg0,visual discrimination and form constancy>. The claim is that even though current Multimodal Large Language Models can handle really complicated tasks, there's limited research specifically looking at how they do on these simpler perception challenges.
Lu: I think the real intrigue here lies in designing a dataset where the reasoning and knowledge needed to solve the problems are kept minimal, which sets up a direct test of pure perception ability without relying on deep prior understanding. The fact that they structured it so that forty-three percent of domains require multiple skills makes it a much tougher assessment for these models because they have to apply different perceptual skills at the same time. It opens up a lot of creative avenues for testing how these models integrate visual understanding with sequential and spatial memory, which is where I see huge potential.
Meng: From an engineering standpoint, I'm interested in how they managed to create six thousand uncontaminated images across thirty domains that are still simple enough to test perception without needing heavy reasoning capabilities <ref:2508.21143#pg0>. The practical impact for us is seeing where the current perception bottlenecks are before we try to build the next generation of multimodal systems. We need this data to understand the baseline capabilities they're currently hitting versus what humans can actually do in these specific areas.
Lalam: I think Percept-V is really important because if we can isolate and measure where an AI fundamentally struggles with basic visual understanding, we can pinpoint exactly where our culture and training need to focus their efforts for improvement. The fact that the models are expected to solve these domains easily because they are simple suggests a huge gap between their current capabilities and human performance that needs careful attention.
Paper summary: Tom: Exactly, Lalam; it’s about finding those specific perceptual weak points rather than just chasing more complex reasoning abilities. Jane, how does this dataset help us understand the scope of what these Multimodal Large Language Models actually know about visual inputs?
Jane: It helps by systematically testing a wide range of visual skills, from simple things like counting missing objects to recognizing form constancy in mirror images. The paper shows that these domains are constructed to be relatively straightforward, meaning the required problem-solving skills are kept low, so we're really testing pure perception and memory recall. It gives us a clear map of which specific perceptual abilities they might lack when faced with visual complexity.
Lu: And the structure itself is telling; by having these domains test different skills in concert, it moves beyond just checking off individual boxes; it tests how well the AI can apply those distinct skills together, which is a much more realistic scenario for real-world visual tasks. That combination of testing skills in concert is what I think reveals the deeper limitations.
Meng: If we look at the structure, we see that even when problems get bigger—when there are more objects or complexity increases—the performance drops fast across all evaluated models. That rapid degradation suggests that the underlying visual processing mechanisms they use aren't robust enough to scale up with visual complexity without significant improvement in their core perception skills.
Lalam: So, it seems the main message of "The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?" is that despite their advanced reasoning, current Multimodal Large Language Models still show substantial gaps when it comes to basic visual perception compared to human ability. This isn't about them failing at complex math; it's about the foundational way they process and understand visual input itself.
Tom: Right, so we’re talking about a significant performance drop as the problem size increases, which means even when the tasks are simple, increasing their visual complexity quickly reveals where they fall short compared to humans. Jane, can you explain what that comparison against human performance looks like in simple terms for our listeners?
Jane: Well, the paper points out that human performance is often more than forty percent higher than the closest Multimodal Large Language Model when tackling these specific perception tasks. For example, in one skill area called Visual Spatial Relationship, the best Multimodal Large Language Model only achieved about seventy-three point zero eight percent accuracy, while humans showed accuracy around ninety-six point four two percent in that same skill.
Lu: That gap is quite large when you put it into perspective; it shows that the models aren't just slightly off on a few details; they are missing a whole level of nuanced visual understanding required for those spatial relationships or figure-ground distinctions. It really highlights how far we still are from true perceptual parity.
Paper summary: Meng: From an engineering standpoint, that forty percent difference means that any system relying heavily on these perception skills without significant architectural changes will be inherently limited in its performance ceiling when dealing with visual complexity. We can't just tweak the reasoning layers if the input understanding itself is weak at the base level.
Lalam: That speaks directly to where we need to focus our development efforts; it tells us that improving their basic visual processing ability is a necessary prerequisite before we expect them to handle more advanced applications successfully. It’s about building a solid foundation in vision first.
Tom: So, summarizing the core message of "The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?", it's that this paper introduces Percept-V to specifically test if Multimodal Large Language Models can match human capabilities on simple visual perception problems, and the results show they exhibit weak performance compared to very high human performance when handling visual complexity.
Jane: And what this means for us is that these foundational perceptual skills are where we see the most significant limitations in current AI systems when compared to humans, even when those tasks aren't overly difficult in terms of explicit reasoning. It sets a clear benchmark for understanding where the actual hurdles lie for multimodal learning.
Lu: I see immense potential here for new research directions; by isolating these seven skills and seeing how models struggle with them individually and in combination, we can develop targeted training strategies that address these specific perceptual deficits rather than just broad model tuning. It’s a fantastic roadmap for future work.
Meng: For practical implementation, this suggests that any future vision-based AI pipeline needs a dedicated module focused purely on robust visual discrimination and form constancy before it even tries to engage in higher-level reasoning tasks. We need reliable inputs first.
Lalam: I think the implication is that we need to treat basic perception skills as a core component of building multimodal intelligence, not just an afterthought or something we hope the model learns implicitly. Addressing these perceptual gaps systematically could really elevate the capability of our systems across all domains and make them much more reliable in visual tasks.
Tom: So, to wrap up this discussion on "The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?", it boils down to this paper showing that MLLMs are currently behind humans on simple visual perception tasks because of systematic limitations in handling visual complexity, and the proposed solution is a focused effort on improving those basic skills.
Jane: And that sets up a lot of exciting work for the future; we’re looking at how to bridge that gap between current model performance and human perceptual capabilities using this kind of rigorous testing framework. It really frames what needs to be prioritized in developing next-generation vision models.
Conclusion: Tom: So, we've been diving deep into Percept-V, and now we're getting to the wrap-up of this paper titled "The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?"
Jane: That sounds like a good spot to really nail down what this research actually means for us as listeners. The authors are showing that these multimodal models aren't quite matching human performance on basic visual perception tasks, even when the problems aren't overly complicated in terms of reasoning.
Lu: I think the real takeaway here is how foundational these visual skills are; they’re like the bedrock upon which more complex AI understanding has to be built. This paper lays out a very clear map of where the current systems fall short in simple visual tasks.
Meng: From an engineering standpoint, it means we can't just focus on adding more abstract reasoning layers if the underlying perception is shaky; we have to fix that base layer first for any real-world deployment.
Lalam: I see this as a signal that improving basic visual comprehension is a crucial step toward making AI truly capable and reliable in everyday applications, which could significantly improve how we interact with technology.
Tom: Exactly, Lalam; it’s not about chasing the hardest problems first, it's about making sure the fundamental visual understanding is solid across all models. Jane, can you tell our audience what this gap looks like in simple terms?
Jane: Well, think of it like this: human perception is about way more than just looking at an image; it involves recalling details and understanding spatial relationships in a way that current AI struggles to replicate accurately on these tests. The paper shows that human performance often exceeds the best models by quite a bit, especially when dealing with how objects relate to each other in space or how they maintain their identity when things change around them.
Lu: And the structure of Percept-V is clever because it tests those skills together, which reveals where the models fail to apply those skills in concert. This isn't just about one skill being weak; it's about integrating them poorly.
Meng: If we look at the results, performance degrades very quickly as the visual complexity increases; that tells me that scaling up the number of objects or intricate patterns is a major hurdle for these models right now.
Lalam: That rapid degradation suggests we need to focus our development efforts on making those models more robust in handling increasing levels of visual detail, which could have a big impact on how we build future applications.
Tom: So, the title itself sets the stage perfectly: it’s a challenge because these models aren't cracking simple perception problems as well as we thought they could. And that opens up some massive avenues for what comes next in AI development.
Jane: It really frames what needs to be prioritized; we need to focus on building more reliable visual processing capabilities before we can expect these multimodal systems to handle truly complex, real-world scenarios with the same level of accuracy as humans.
Samrajnee Ghosh, Naman Agarwal, Hemanshu Garg, Chinmay Mittal, Parag Singla, Mausam
IIT Delhi
cs.CL, cs.CV
Submitted: 2025-08-28
Updated: 2026-10-02
Importance score: 77/100
The gist: Multimodal Large Language Models (MLLMs) are being tested on simple visual perception problems to determine if they match human capabilities, and this research introduces Percept-V, a dataset
Key concepts
- Percept-V
- An extensive dataset of 6000 program-generated images across 30 domains. It is designed to isolate foundational visual skills by testing tasks that require minimal specialized knowledge, drawing from cognitive science.
- TVPS-4 Framework
- A framework categorizing human perception into seven distinct skills, such as visual discrimination and form constancy. This framework helps researchers systematically map the specific visual abilities being tested in each domain of the Percept-V dataset.
- Visual Complexity
- The inherent difficulty of a visual task, often measured by the number of objects present in an image. The research found that MLLMs consistently degrade in performance as this complexity increases, highlighting their inability to handle intricate visual scenes.
- Skill Application
- The ability of MLLMs to apply multiple visual skills simultaneously when solving a problem. Since 43% of Percept-V domains require multiple skills, the study assesses how well models can integrate different perceptual abilities together.
Terminology
Summary
Multimodal Large Language Models (MLLMs) are being tested on simple visual perception problems to determine if they match human capabilities, and this research introduces Percept-V, a dataset designed to isolate these foundational visual skills. The gist is that SoTA proprietary and open-source MLLMs show weak performance compared to very high human performance on Percept-V, revealing systematic limitations in handling visual complexity.
The Dataset: Percept-V
Percept-V is an extensive dataset comprising 6000 program-generated, uncontaminated images divided into 30 domains. Each domain tests one or more skills from the TVPS-4 framework, which categorizes human perception into seven skills such as visual discrimination and form constancy. The construction draws from cognitive science literature to ensure the tasks require minimal problem-solving skills and no specialized domain knowledge.
The dataset is structured to systematically assess performance across varying levels of difficulty. Each domain contains 200 images with varying problem sizes, operationalized through the number of objects present, which captures its inherent complexity. The instances vary in sizes from 1 to 20 for each domain, resulting in a total of 6000 instances across the 30 domains.
Perceptual Skills Tested (TVPS-4 Framework)
The dataset maps each domain to a combination of TVPS-4 skill categories. The seven skills evaluated are:
-
Visual Discrimination
-
Visual Memory
-
Visual Sequential Memory
-
Visual Spatial Relationship
-
Visual Figure Ground
-
Visual Form Constancy (FC)
-
Visual Closure
The paper notes that skills associated with visual memory and visual sequential memory test the ability to recall a detail once it has been presented and then removed from the field of view.
Furthermore, 43% of the domains in Percept-V require multiple skills, allowing for an assessment of MLLMs' ability when applying skills in concert.
Experimental Setup and Models
The experiments compare four state-of-the-art proprietary MLLMs (GPT-5-mini, GPT-4o, o4-mini, and Gemini 2.5 Flash) against two open source models (Qwen 2.5 VL Instruct and DeepSeek VL2 Tiny). The models tested include both reasoning/thinking models and those equipped with ‘think with image’ capabilities. All experiments were run primarily in a zero-shot mode to control costs, though one-shot prompting was also explored.
Key Findings on Performance Gaps
The results reveal substantial gaps between MLLM and human performance on these simple visual perception tasks.
Across all evaluated models, performance goes down rather fast
as the number of objects in the image increases.
Across all evaluated models, we observe consistent performance degradation as problem size increases, suggesting systematic limitations in handling visual complexity.
The analysis shows that MLLMs are relatively consistent across domains for some skills (e.g., ‘inside circles’ has best overall performance), but they exhibit broad deficits in visual understanding.
For instance, the domain ‘colors present’ showed the worst overall performance, suggesting MLLMs get confused by related colors like turquoise and blue.
Performance vs. Size and Human Comparison
Performance drops significantly with increasing problem size across all skills. While thinking models (o4-mini and Gemini) solve smaller problems with near-perfect accuracy,
this performance drops dramatically with increasing size.
GPT-5-mini, o4-mini, and Gemini are among the best performing models but are relatively robust to size change compared to other LLMs.
The human study confirmed these deficits: human performance is more than 40% points higher than the closest MLLM.
For example, in Visual Spatial Relationship, the best MLLM performance (Gemini) was around 73.08% compared to about 96.42% accuracy shown by humans in the same skill. The overall average performance across all skills for all models was only 55.22%.
Format Following and Conclusion
Most proprietary models hardly falter in following the answer formats specified in the question,
with GPT-4o showing 0% format errors.
However, accuracy issues stem primarily from a lack of perceptual skills. The conclusion is that MLLMs demonstrate significant gaps between themselves and humans on basic perception tasks, suggesting they are rather behind, especially on basic perception skills.
Percept-V is proposed as an important step towards understanding the visual perception limitations that must be addressed in the future to develop more robust MLLMs.
The gist
SoTA proprietary and open-source MLLMs show weak performance compared to very high human performance on Percept-V, revealing systematic limitations in handling visual complexity.
How it works
Improvements for AI systems
Based on the analysis of the provided scientific paper, here are specific improvements for AI systems and what those improved systems can achieve:
-
Improve foundational visual perception capabilities by developing models specifically trained and benchmarked on simple, uncontaminated perception tasks (like those in Percept-V).
-
Develop multimodal LLMs that exhibit superior generalization across problem sizes within a single domain.
-
Enhance the robustness of MLLMs against increasing perceptual complexity (i.e., as the number of objects or distractors increases).
-
Improve model performance on skills that are inherently difficult for current MLLMs, specifically those involving complex spatial relationships or
inside-out
reasoning (e.g., Visual Closure). -
Develop specialized modules within MLLMs dedicated to handling basic visual perception tasks without relying solely on high-level reasoning or specialized domain knowledge.
The improved AI system, leveraging these improvements, can perform the following:
-
Perform accurate, low-level visual discrimination (e.g., distinguishing between circles and triangles of different sizes/colors) with high fidelity across various image complexities.
-
Accurately solve tasks involving spatial reasoning (e.g., counting objects relative to a background, identifying locations within a grid, or determining object positions based on geometric relationships).
-
Maintain consistent performance when the visual complexity of an input image increases (i.e., they do not suffer catastrophic performance drops when more objects are present).
-
Effectively solve tasks requiring pattern recognition and memory recall from simple visual sequences (e.g., recalling the order of colors in concentric circles or matching outlines).
-
Serve as a reliable foundation for applications requiring high-precision visual understanding, such as automated quality control in manufacturing, basic object recognition systems, or enhanced computer vision pipelines where foundational perception is critical before complex reasoning is applied.
Sources
- What MLLMs Learn about When they Learn about Multimodal Reasoning
- BLINK: Multimodal Large Language Models Can See but Not Perceive
- Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
- Do You See Me : A Multidimensional Benchmark for Evaluating Visual Perception in Multimodal LLMs
- Exploring Multimodal Perception in Large Language Models Through Perceptual Strength Ratings
- MMRo: Are Multimodal LLMs Eligible as the Brain for In-Home Robotics?
- Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- A Brief Report on LawGPT 1.0: A Virtual Legal Assistant Based on GPT-3
- NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers
- SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models
- CMMMU: A Chinese Massive Multi-discipline Multimodal Understanding Benchmark
- Exploring Perceptual Limitation of Multimodal Large Language Models
- Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering