MIT’s VISTA Harness Just Scored a Perfect 100 on ARC-AGI-3 — By Giving Models Eyes Instead of Smarter Brains
October 6, 2026 — The biggest benchmark result of the month didn’t come from a bigger model. It came from better perception.
A paper uploaded to arXiv on October 1 by MIT researchers Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu, and Kaiming He introduces VISTA, a visual harness that lifted Claude Opus 5.0’s Relative Human Action Efficiency score on ARC-AGI-3 from 40.68 to a perfect 100.00. The model cleared all 25 public games and 183 levels using 57.4% fewer actions than first-time human players.
How VISTA works
The team’s argument is that multimodal models were never bad at reasoning — they were blind. The official ARC-AGI-3 interface feeds models a 64×64 grid of numbers: 4,096 digits where a human sees characters, mechanisms, and paths. VISTA instead lets the model perceive the environment through raw screenshots and keeps a lossless visual memory — every frame is archived in original form, and the model can retrieve any of them mid-reasoning, zoom into corners, or read exact pixel values. Flipping back through the archive doesn’t count as a step; only real actions in the game do.
It needs no trained components — just a four-sentence prompt shared across every game — and the code is public.
The numbers that matter
The headline is the perfect score, but the more surprising findings are about cost and bottlenecks. Swapping the official numeric grid for plain 512×512 screenshots took GPT-5.6 Sol from 13.33 to 47.32 with nothing else changed — perception was the bottleneck, not intelligence. Images also used less compute: 30.7 million tokens per game versus 71.9 million for the text version. And more isn’t always better: extending context from 200K to 780K dropped the score from 99 to 93.9.
Why it matters
If these results generalize, the industry has been misdiagnosing agent failures. When an agent flails at a computer-use task, the instinct is to reach for a bigger model — but VISTA suggests the bottleneck may be what the model can see and remember, not how much it can reason. That’s a much cheaper problem to fix, and it reframes the whole “which model is smartest” race into a harness-design question. One honest caveat from the authors: their models postdate the public games, so real generalization will be tested on the private set. Perfect on public games is a milestone, not a finish line.
FAQ
What is ARC-AGI-3?
An interactive benchmark from the ARC Prize Foundation (founded by François Chollet) that measures how efficiently AI agents learn to play unfamiliar video-game-like scenarios with no instructions.
What did VISTA achieve?
A perfect 100.00 Relative Human Action Efficiency score with Claude Opus 5.0 — all 25 public games and 183 levels, using 57.4% fewer actions than first-time human players.
Does this mean models are already as smart as humans at these tasks?
Not exactly — the authors note the models postdate the public games, so scores on the private game set will be the real test of generalization.
Is VISTA open source?
Yes — the code is public, and the harness needs no trained components.
Sources: arXiv paper (Han, Hu, Qiu, Wu, He, MIT); AI Daily Digest (DEV Community).

