Samrajnee Ghosh, Ashish Goswami, Naman Agarwal, Hemanshu Garg, Chinmay Mittal, Mausam, Parag Singla
Indian Institute of Technology, Delhi · COLM 2026
Percept-V is a benchmark of 6,000 program-generated, uncontaminated images across 30 domains, each grounded in one or more of the seven visual perception skills of the TVPS-4 cognitive framework. The tasks require minimal reasoning and no specialised knowledge, isolating basic perception. Every domain spans 20 problem sizes, so accuracy can be measured as a function of visual complexity.
Humans score 95.4%. The best of eight evaluated MLLMs reaches 66.9%, and every model degrades sharply as the number of objects grows.
The dataset, the per-domain image generation scripts and the evaluation harness will be released here.
@article{ghosh2025percept,
title={The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?},
author={Ghosh, Samrajnee and Agarwal, Naman and Garg, Hemanshu and Mittal, Chinmay and Singla, Parag and others},
journal={arXiv preprint arXiv:2508.21143},
year={2025}
}