Skip to content

Latest commit

Β 

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space

Polaris-Bench: Official Benchmark and Evaluation Suite

arXiv HuggingFace Project Page License Data License

Xia Hu1, Zhenrui Yue1, Brian Potetz1, Howard Zhou1, Leonidas Guibas1,2, Chun-Ta Lu3, Zhicheng Wang1
1Google DeepMind Β  2Stanford University Β  3Google Research


Overview

Current Multimodal Large Language Models (MLLMs) achieve strong performance on visual reasoning benchmarks, but do these scores reflect genuine visual perception? In this work, we identify a pervasive vulnerability: the Cartesian Shortcut.

The Cartesian Shortcut

The Cartesian Shortcut: Cartesian vs. Polar Visual Reasoning

Many prominent visual reasoning benchmarks are built on orthogonal, grid-based layouts that can be readily discretized into explicit textual coordinates. We find that frontier models frequently exploit this property: across over 3,800 questions from 9 prominent visual reasoning benchmarks, they explicitly invoke textual coordinates (e.g., "row 2", "(x,y)") in over 56% of intermediate Chain-of-Thought reasoning, offloading reasoning from visual perception to text-based deduction. This shortcut confounds the evaluation of visual reasoning on grid-based testbeds.

Polaris-Bench

To re-evaluate visual reasoning when this shortcut is unavailable, we introduce Polaris-Bench, which re-formulates 53 visual reasoning tasks across 5 cognitive categories in Polar coordinate space, each paired with a Cartesian counterpart that preserves task semantics and serves as a controlled reference under consistent logical constraints. The Polar layout disrupts the orthogonal structure that models exploit.

Across 14 state-of-the-art MLLMs, frontier models achieving 69–83% on Cartesian layouts collapse to 31–39% on Polar equivalents, while humans drop only 5.7 points (94.5% β†’ 88.8%). Thinking gains largely vanish on Polar layouts, prompting interventions (conversion hints, 5-shot in-context examples) fail to close the gap, and comparable drops arise on other non-orthogonal layouts (hexagonal tilings, wave and arc deformations). These findings show that current MLLMs’ visual reasoning performance is strongly coupled to orthogonal grid structure.

Representative task pairs in Polaris-Bench across five cognitive categories

Leaderboard

All models evaluated under high reasoning mode. Sorted by Polar accuracy (P). Full per-category results on the project page.

# Model Type Cartesian (%) Polar (%) Drop (Ξ”)
πŸ‘€ Human Baseline 94.5 88.8 -5.7
1 GPT-5.2 Closed 77.4 39.2 -38.2
2 Gemini-3.1-Pro Closed 82.6 35.9 -46.7
3 Qwen3.5-397B-A17B Open 72.9 35.0 -37.9
4 Gemini-3-Flash Closed 71.0 33.8 -37.2
5 Kimi-k2.5 Open 69.0 31.1 -37.8
6 Gemma-4-31B Open 60.5 31.0 -29.5
7 Claude-Sonnet-4.6 Closed 44.4 25.9 -18.5
8 Gemini-2.5-Pro Closed 38.4 25.3 -13.2
9 Gemini-3.1-Flash-Lite Closed 46.8 24.6 -22.2
10 Gemma-4-26B Open 47.2 22.9 -24.4
11 Grok-4-Fast-Reasoning Closed 31.0 22.3 -8.7
12 Grok-4-0709 Closed 33.0 21.8 -11.2
13 Gemini-2.5-Flash Closed 32.2 21.1 -11.2
14 Mistral-Small-2503 Open 19.4 19.0 -0.4
- Random Baseline Baseline 15.8 15.8 0.0

Dataset on Hugging Face

The complete Polaris-Bench evaluation dataset (all 10,800 multimodal problem instances across 53 tasks and 4 coordinate systems) is hosted on the Hugging Face Hub:

https://huggingface.co/datasets/google/polaris-bench

You can load the full dataset directly using the Hugging Face datasets library:

from datasets import load_dataset

# Load the full 10,800 evaluation problem instances
dataset = load_dataset("google/polaris-bench", split="test")

# Inspect an example
sample = dataset[0]
print(f"Task: {sample['task']} | Coordinate System: {sample['question_type']}")
print(f"Question: {sample['question']}")
print(f"Answer: {sample['answer']}")
# sample["image"] is automatically decoded as a PIL Image object

Offline sample tasks: If you want to explore the benchmark locally without downloading the full Hugging Face image dataset, a standalone set of 20 representative paired Cartesian-Polar tasks is included directly in this repository under examples/sample_tasks/.

Repository Structure

polaris-bench/
β”œβ”€β”€ evaluation/              # Core evaluation module and benchmark scoring
β”‚   β”œβ”€β”€ __init__.py          # Package exports (PolarisDataLoader, Evaluator)
β”‚   β”œβ”€β”€ data_loader.py       # Dataset loader (supports Hugging Face, local JSON, and sample tasks)
β”‚   └── evaluate.py          # Benchmark evaluation script (accuracy by coordinate system and category)
β”œβ”€β”€ examples/                # Quick start guide and local offline sample instances
β”‚   β”œβ”€β”€ quick_start.py       # Standalone demo script (loads data, paired tasks, and runs mock evaluation)
β”‚   └── sample_tasks/        # 20 representative paired tasks for local offline inspection
β”‚       β”œβ”€β”€ sample_tasks.json# Paired Cartesian vs. Polar questions and ground truth
β”‚       β”œβ”€β”€ README.md        # Documentation for sample tasks
β”‚       └── images/          # PNG images organized by task
β”œβ”€β”€ docs/                    # Project webpage source (served via GitHub Pages)
β”‚   β”œβ”€β”€ index.html           # Interactive project page with leaderboard and task viewer
β”‚   └── static/              # Paper figures, teasers, and website assets
β”œβ”€β”€ tests/                   # Unit test suite
β”‚   └── test_benchmark.py    # 9 automated tests for data loading, matching, and metrics
β”œβ”€β”€ pyproject.toml           # Python package build configuration
β”œβ”€β”€ requirements.txt         # Minimal dependencies (datasets, Pillow, etc.)
β”œβ”€β”€ LICENSE                  # Apache 2.0 license
└── README.md                # Project documentation and getting started guide

Installation

git clone https://github.com/google-deepmind/polaris-bench.git
cd polaris-bench
pip install -r requirements.txt

Quick Start

You can inspect the benchmark using either the HuggingFace Hub dataset or the local offline sample data:

from evaluation.data_loader import PolarisDataLoader

# 1. Load from local file or HuggingFace Hub ('google/polaris-bench')
loader = PolarisDataLoader("examples/sample_tasks/sample_tasks.json")

# 2. Filter by task and coordinate system
cart_examples = loader.get_dataset(task="sudoku", question_type="cartesian")
print(f"Loaded {len(cart_examples)} Cartesian Sudoku tasks.")

# 3. Retrieve paired instances (Cartesian vs. Polar)
pairs = loader.get_paired_dataset(task="sudoku")
cart_sample, polar_sample = pairs[0]

print("Cartesian Image:", cart_sample["image_path"])
print("Polar Image:    ", polar_sample["image_path"])
print("Ground Truth:   ", cart_sample["answer"])

Offline sample data: A standalone set of 20 representative paired tasks is provided in examples/sample_tasks/ for quick local inspection without downloading the full dataset.

Evaluation

Evaluating a model on Polaris-Bench consists of two simple steps: model inference and benchmark scoring.

Step 1: Generate Predictions from Your Model

Query your multimodal model (e.g. Gemini, GPT, Claude, or local open-weights) with each task's image and question prompt. Collect the outputs into a predictions JSON file:

import json
from evaluation.data_loader import PolarisDataLoader

loader = PolarisDataLoader("examples/sample_tasks/sample_tasks.json")

# Run inference with your model and save predictions
predictions = {}
for example in loader:
    key = f"{example['task']}::{example['index']}::{example['question_type']}"
    # Call your model inference function here:
    # predictions[key] = your_model.generate(image=example['image'], prompt=example['question'])
    predictions[key] = "A"

with open("predictions.json", "w") as f:
    json.dump(predictions, f, indent=2)

Supported Prediction Formats:

  1. 3-part key (task::index::question_type):
{
  "sudoku::example_001::cartesian": "D",
  "sudoku::example_001::polar": "B",
  "maze::example_042::polar": "A"
}
  1. Nested by coordinate system:
{
  "cartesian": {
    "sudoku::example_001": "D"
  },
  "polar": {
    "sudoku::example_001": "B"
  }
}
  1. List of prediction records:
[
  {
    "task": "sudoku",
    "index": "example_001",
    "question_type": "polar",
    "prediction": "B"
  }
]

Step 2: Score Predictions Against Ground Truth

Run the evaluation script to compute accuracy metrics and the Cartesian-to-Polar drop:

python -m evaluation.evaluate \
  --predictions predictions.json \
  --ground-truth polaris_bench.json

To evaluate only a specific coordinate system, pass --question-type:

python -m evaluation.evaluate \
  --predictions predictions.json \
  --question-type polar

The script displays a formatted terminal summary table and exports a detailed JSON report (evaluation_report.json), illustrated below on GPT-5.2 evaluation results:

==================================================================================
                         POLARIS-BENCH EVALUATION REPORT
==================================================================================
Total Evaluated: 10600 | Overall Accuracy: 58.3%
Cartesian Acc:   77.4% | Polar Acc: 39.2% | Drop (Cartesian - Polar): 38.2 pt
----------------------------------------------------------------------------------
Category                             | Cartesian (%) | Polar (%)  | Drop (pt) 
----------------------------------------------------------------------------------
Algorithmic Logic & Simulation       | 70.7          | 46.7       | 24.0      
Combinatorics & Probability          | 84.8          | 30.7       | 54.1      
Navigation & Routing                 | 82.2          | 42.8       | 39.4      
Spatial Transformation & Geometry    | 76.5          | 41.6       | 34.9      
Visual Pattern Matching              | 70.0          | 34.5       | 35.5      
==================================================================================

Dataset Structure

Each record contains 6 fields (+ image in the HuggingFace Parquet version):

Field Type Description
task string Task name (e.g., sudoku, maze, shape_fitting)
index string Sample index (e.g., example_001)
question_type string cartesian, polar, hexagonal, or octagonal
image PIL.Image Evaluation image (RGBA, ~3000Γ—3000 to 4000Γ—3600)
image_path string Relative path (e.g., images/sudoku/sudoku_example_001_cartesian.png)
question string Evaluation question text
answer string Ground-truth answer (option labels, digits, coordinates, strings, or lists)

Task Taxonomy

53 tasks organized into 5 cognitive categories:

Category Tasks Count
Visual Pattern Matching pattern_completion, shape_fitting, layer_completion, fragment_matching, template_matching, jigsaw_matching, shape_completion, odd_piece_out, anomaly_detection, letter_collection, pattern_prediction 11
Spatial Transformation & Geometry rotation_matching, mirror_reflection, grid_rotation, rotation_center, pivot_rotation, impossible_shape, grid_folding, four_color, area_counting, pipe_lengths, uncut_cells, area_balancing, curve_length 13
Navigation & Routing maze, shortest_path, longest_path, bounded_path_finding, wrapping_path_finding, wrapping_navigation, egocentric_navigation, absolute_navigation, wall_follower, rule_based_navigation, monotonic_path, turn_counting, word_search, largest_number_path 14
Combinatorics & Probability path_counting, bounded_diagonal_paths, bounded_knight_paths, checkpoint_paths, lattice_paths, wrapping_diagonal_paths, knight_paths, edge_counting, random_walk 9
Algorithmic Logic & Simulation n_queens, sudoku, minimum_flips, maximum_collection, collision_detection, bouncing_point 6

Citation

@misc{hu2026cartesianshortcutreevaluatevision,
  title         = {The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space},
  author        = {Xia Hu and Zhenrui Yue and Brian Potetz and Howard Zhou and Leonidas Guibas and Chun-Ta Lu and Zhicheng Wang},
  year          = {2026},
  eprint        = {2605.09883},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2605.09883}
}

Licensing & Disclaimer

Copyright 2026 Google LLC

All software is licensed under the Apache License, Version 2.0 (Apache 2.0); you may not use this file except in compliance with the Apache 2.0 license. You may obtain a copy of the Apache 2.0 license at: https://www.apache.org/licenses/LICENSE-2.0

All other materials are licensed under the Creative Commons Attribution 4.0 International License (CC-BY). You may obtain a copy of the CC-BY license at: https://creativecommons.org/licenses/by/4.0/legalcode

Some data was created with inspiration from:

Unless required by applicable law or agreed to in writing, all software and materials distributed here under the Apache 2.0 or CC-BY licenses are distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the licenses for the specific language governing permissions and limitations under those licenses.

This is not an official Google product.

About

No description, website, or topics provided.

Resources

Contributing

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages