Polaris-Bench: Official Benchmark and Evaluation Suite
Xia Hu1,
Zhenrui Yue1,
Brian Potetz1,
Howard Zhou1,
Leonidas Guibas1,2,
Chun-Ta Lu3,
Zhicheng Wang1
1Google DeepMind Β 2Stanford University Β 3Google Research
Current Multimodal Large Language Models (MLLMs) achieve strong performance on visual reasoning benchmarks, but do these scores reflect genuine visual perception? In this work, we identify a pervasive vulnerability: the Cartesian Shortcut.
Many prominent visual reasoning benchmarks are built on orthogonal, grid-based layouts that can be readily discretized into explicit textual coordinates. We find that frontier models frequently exploit this property: across over 3,800 questions from 9 prominent visual reasoning benchmarks, they explicitly invoke textual coordinates (e.g., "row 2", "(x,y)") in over 56% of intermediate Chain-of-Thought reasoning, offloading reasoning from visual perception to text-based deduction. This shortcut confounds the evaluation of visual reasoning on grid-based testbeds.
To re-evaluate visual reasoning when this shortcut is unavailable, we introduce Polaris-Bench, which re-formulates 53 visual reasoning tasks across 5 cognitive categories in Polar coordinate space, each paired with a Cartesian counterpart that preserves task semantics and serves as a controlled reference under consistent logical constraints. The Polar layout disrupts the orthogonal structure that models exploit.
Across 14 state-of-the-art MLLMs, frontier models achieving 69β83% on Cartesian layouts collapse to 31β39% on Polar equivalents, while humans drop only 5.7 points (94.5% β 88.8%). Thinking gains largely vanish on Polar layouts, prompting interventions (conversion hints, 5-shot in-context examples) fail to close the gap, and comparable drops arise on other non-orthogonal layouts (hexagonal tilings, wave and arc deformations). These findings show that current MLLMsβ visual reasoning performance is strongly coupled to orthogonal grid structure.
All models evaluated under high reasoning mode. Sorted by Polar accuracy (P). Full per-category results on the project page.
| # | Model | Type | Cartesian (%) | Polar (%) | Drop (Ξ) |
|---|---|---|---|---|---|
| π€ | Human | Baseline | 94.5 | 88.8 | -5.7 |
| 1 | GPT-5.2 | Closed | 77.4 | 39.2 | -38.2 |
| 2 | Gemini-3.1-Pro | Closed | 82.6 | 35.9 | -46.7 |
| 3 | Qwen3.5-397B-A17B | Open | 72.9 | 35.0 | -37.9 |
| 4 | Gemini-3-Flash | Closed | 71.0 | 33.8 | -37.2 |
| 5 | Kimi-k2.5 | Open | 69.0 | 31.1 | -37.8 |
| 6 | Gemma-4-31B | Open | 60.5 | 31.0 | -29.5 |
| 7 | Claude-Sonnet-4.6 | Closed | 44.4 | 25.9 | -18.5 |
| 8 | Gemini-2.5-Pro | Closed | 38.4 | 25.3 | -13.2 |
| 9 | Gemini-3.1-Flash-Lite | Closed | 46.8 | 24.6 | -22.2 |
| 10 | Gemma-4-26B | Open | 47.2 | 22.9 | -24.4 |
| 11 | Grok-4-Fast-Reasoning | Closed | 31.0 | 22.3 | -8.7 |
| 12 | Grok-4-0709 | Closed | 33.0 | 21.8 | -11.2 |
| 13 | Gemini-2.5-Flash | Closed | 32.2 | 21.1 | -11.2 |
| 14 | Mistral-Small-2503 | Open | 19.4 | 19.0 | -0.4 |
| - | Random Baseline | Baseline | 15.8 | 15.8 | 0.0 |
The complete Polaris-Bench evaluation dataset (all 10,800 multimodal problem instances across 53 tasks and 4 coordinate systems) is hosted on the Hugging Face Hub:
https://huggingface.co/datasets/google/polaris-bench
You can load the full dataset directly using the Hugging Face datasets library:
from datasets import load_dataset
# Load the full 10,800 evaluation problem instances
dataset = load_dataset("google/polaris-bench", split="test")
# Inspect an example
sample = dataset[0]
print(f"Task: {sample['task']} | Coordinate System: {sample['question_type']}")
print(f"Question: {sample['question']}")
print(f"Answer: {sample['answer']}")
# sample["image"] is automatically decoded as a PIL Image objectOffline sample tasks: If you want to explore the benchmark locally without downloading the full Hugging Face image dataset, a standalone set of 20 representative paired Cartesian-Polar tasks is included directly in this repository under
examples/sample_tasks/.
polaris-bench/
βββ evaluation/ # Core evaluation module and benchmark scoring
β βββ __init__.py # Package exports (PolarisDataLoader, Evaluator)
β βββ data_loader.py # Dataset loader (supports Hugging Face, local JSON, and sample tasks)
β βββ evaluate.py # Benchmark evaluation script (accuracy by coordinate system and category)
βββ examples/ # Quick start guide and local offline sample instances
β βββ quick_start.py # Standalone demo script (loads data, paired tasks, and runs mock evaluation)
β βββ sample_tasks/ # 20 representative paired tasks for local offline inspection
β βββ sample_tasks.json# Paired Cartesian vs. Polar questions and ground truth
β βββ README.md # Documentation for sample tasks
β βββ images/ # PNG images organized by task
βββ docs/ # Project webpage source (served via GitHub Pages)
β βββ index.html # Interactive project page with leaderboard and task viewer
β βββ static/ # Paper figures, teasers, and website assets
βββ tests/ # Unit test suite
β βββ test_benchmark.py # 9 automated tests for data loading, matching, and metrics
βββ pyproject.toml # Python package build configuration
βββ requirements.txt # Minimal dependencies (datasets, Pillow, etc.)
βββ LICENSE # Apache 2.0 license
βββ README.md # Project documentation and getting started guide
git clone https://github.com/google-deepmind/polaris-bench.git
cd polaris-bench
pip install -r requirements.txtYou can inspect the benchmark using either the HuggingFace Hub dataset or the local offline sample data:
from evaluation.data_loader import PolarisDataLoader
# 1. Load from local file or HuggingFace Hub ('google/polaris-bench')
loader = PolarisDataLoader("examples/sample_tasks/sample_tasks.json")
# 2. Filter by task and coordinate system
cart_examples = loader.get_dataset(task="sudoku", question_type="cartesian")
print(f"Loaded {len(cart_examples)} Cartesian Sudoku tasks.")
# 3. Retrieve paired instances (Cartesian vs. Polar)
pairs = loader.get_paired_dataset(task="sudoku")
cart_sample, polar_sample = pairs[0]
print("Cartesian Image:", cart_sample["image_path"])
print("Polar Image: ", polar_sample["image_path"])
print("Ground Truth: ", cart_sample["answer"])Offline sample data: A standalone set of 20 representative paired tasks is provided in
examples/sample_tasks/for quick local inspection without downloading the full dataset.
Evaluating a model on Polaris-Bench consists of two simple steps: model inference and benchmark scoring.
Query your multimodal model (e.g. Gemini, GPT, Claude, or local open-weights) with each task's image and question prompt. Collect the outputs into a predictions JSON file:
import json
from evaluation.data_loader import PolarisDataLoader
loader = PolarisDataLoader("examples/sample_tasks/sample_tasks.json")
# Run inference with your model and save predictions
predictions = {}
for example in loader:
key = f"{example['task']}::{example['index']}::{example['question_type']}"
# Call your model inference function here:
# predictions[key] = your_model.generate(image=example['image'], prompt=example['question'])
predictions[key] = "A"
with open("predictions.json", "w") as f:
json.dump(predictions, f, indent=2)Supported Prediction Formats:
- 3-part key (
task::index::question_type):
{
"sudoku::example_001::cartesian": "D",
"sudoku::example_001::polar": "B",
"maze::example_042::polar": "A"
}- Nested by coordinate system:
{
"cartesian": {
"sudoku::example_001": "D"
},
"polar": {
"sudoku::example_001": "B"
}
}- List of prediction records:
[
{
"task": "sudoku",
"index": "example_001",
"question_type": "polar",
"prediction": "B"
}
]Run the evaluation script to compute accuracy metrics and the Cartesian-to-Polar drop:
python -m evaluation.evaluate \
--predictions predictions.json \
--ground-truth polaris_bench.jsonTo evaluate only a specific coordinate system, pass --question-type:
python -m evaluation.evaluate \
--predictions predictions.json \
--question-type polarThe script displays a formatted terminal summary table and exports a detailed JSON report (evaluation_report.json), illustrated below on GPT-5.2 evaluation results:
==================================================================================
POLARIS-BENCH EVALUATION REPORT
==================================================================================
Total Evaluated: 10600 | Overall Accuracy: 58.3%
Cartesian Acc: 77.4% | Polar Acc: 39.2% | Drop (Cartesian - Polar): 38.2 pt
----------------------------------------------------------------------------------
Category | Cartesian (%) | Polar (%) | Drop (pt)
----------------------------------------------------------------------------------
Algorithmic Logic & Simulation | 70.7 | 46.7 | 24.0
Combinatorics & Probability | 84.8 | 30.7 | 54.1
Navigation & Routing | 82.2 | 42.8 | 39.4
Spatial Transformation & Geometry | 76.5 | 41.6 | 34.9
Visual Pattern Matching | 70.0 | 34.5 | 35.5
==================================================================================
Each record contains 6 fields (+ image in the HuggingFace Parquet version):
| Field | Type | Description |
|---|---|---|
task |
string | Task name (e.g., sudoku, maze, shape_fitting) |
index |
string | Sample index (e.g., example_001) |
question_type |
string | cartesian, polar, hexagonal, or octagonal |
image |
PIL.Image | Evaluation image (RGBA, ~3000Γ3000 to 4000Γ3600) |
image_path |
string | Relative path (e.g., images/sudoku/sudoku_example_001_cartesian.png) |
question |
string | Evaluation question text |
answer |
string | Ground-truth answer (option labels, digits, coordinates, strings, or lists) |
53 tasks organized into 5 cognitive categories:
| Category | Tasks | Count |
|---|---|---|
| Visual Pattern Matching | pattern_completion, shape_fitting, layer_completion, fragment_matching, template_matching, jigsaw_matching, shape_completion, odd_piece_out, anomaly_detection, letter_collection, pattern_prediction | 11 |
| Spatial Transformation & Geometry | rotation_matching, mirror_reflection, grid_rotation, rotation_center, pivot_rotation, impossible_shape, grid_folding, four_color, area_counting, pipe_lengths, uncut_cells, area_balancing, curve_length | 13 |
| Navigation & Routing | maze, shortest_path, longest_path, bounded_path_finding, wrapping_path_finding, wrapping_navigation, egocentric_navigation, absolute_navigation, wall_follower, rule_based_navigation, monotonic_path, turn_counting, word_search, largest_number_path | 14 |
| Combinatorics & Probability | path_counting, bounded_diagonal_paths, bounded_knight_paths, checkpoint_paths, lattice_paths, wrapping_diagonal_paths, knight_paths, edge_counting, random_walk | 9 |
| Algorithmic Logic & Simulation | n_queens, sudoku, minimum_flips, maximum_collection, collision_detection, bouncing_point | 6 |
@misc{hu2026cartesianshortcutreevaluatevision,
title = {The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space},
author = {Xia Hu and Zhenrui Yue and Brian Potetz and Howard Zhou and Leonidas Guibas and Chun-Ta Lu and Zhicheng Wang},
year = {2026},
eprint = {2605.09883},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2605.09883}
}Copyright 2026 Google LLC
All software is licensed under the Apache License, Version 2.0 (Apache 2.0); you may not use this file except in compliance with the Apache 2.0 license. You may obtain a copy of the Apache 2.0 license at: https://www.apache.org/licenses/LICENSE-2.0
All other materials are licensed under the Creative Commons Attribution 4.0 International License (CC-BY). You may obtain a copy of the CC-BY license at: https://creativecommons.org/licenses/by/4.0/legalcode
Some data was created with inspiration from:
- Babyvision, which is available at https://github.com/UniPat-AI/BabyVision under the Creative Commons Attribution 4.0 International License (CC-BY). You may obtain a copy of the CC-BY license at: https://creativecommons.org/licenses/by/4.0/legalcode.
- EMMA-Bench, which is available at https://github.com/EMMA-Bench/EMMA.
- MathVista: https://github.com/lupantech/MathVista (released under the Creative Commons Attribution-ShareAlike 4.0 International License, CC BY-SA 4.0).
- MEGABench: https://github.com/TIGER-AI-Lab/MEGA-Bench.
Unless required by applicable law or agreed to in writing, all software and materials distributed here under the Apache 2.0 or CC-BY licenses are distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the licenses for the specific language governing permissions and limitations under those licenses.
This is not an official Google product.

