A new benchmark that judges AI models on how well they can spot mistakes in assembled furniture from a photo is drawing attention as a fresh way to measure visual reasoning skills. OpenAI's latest model, GPT-6 Astra, posted a far higher score on the test than earlier models, underscoring how quickly AI's spatial and visual reasoning has advanced in less than a year.

A new test built around furniture assembly errors

The benchmark works by showing an AI model photos taken during furniture assembly and asking it to identify where the steps went wrong. Three furniture pieces of varying complexity were used, each built from kits with multiple assembly steps. Models were given the assembly manual, the ability to zoom into photos, and access to a code interpreter, then asked to compare each photo against the instructions, pinpoint every mistake, and describe what went wrong. Scoring runs through three stages: finding the errors, confirming the correct steps, and validating the model's written answers.

The task looks simple, but it demands that a model accurately read fine details such as part orientation, placement, and missing screws straight from a photo, making it a genuinely difficult combination of visual and spatial reasoning. A human could spot the same mistakes at a glance by comparing the photo with the manual, but a model has to recognize the relevant parts and their orientation within the image and then put that into words, which requires several distinct capabilities working together rather than simple image classification.

GPT-6 Astra hits 80%, a sharp jump in just ten months

GPT-6 Astra scored 80% on the test, answering each case in roughly three minutes, faster than rival models. Anthropic's Claude Fable 5.1 came in at 70% and Claude Opus 5 at 61%, showing that several companies have reached a high level around the same time.

What stands out is the pace of improvement. Ten months earlier, the top score on the same test belonged to Anthropic's Claude Opus 4.5, which managed only 28%. That means the leading score has nearly tripled in under a year, a clear sign of how fast AI models' image understanding is progressing. At the same time, open-weight models released by major Chinese companies reportedly trail the frontier by about seven months, suggesting the gap in this kind of visual reasoning capability remains significant.

Remaining hurdles and where the skill could be applied

Since GPT-6 Astra still takes about three minutes per photo, it is not yet fast enough to guide someone through assembly in real time. Even so, this kind of visual and spatial reasoning is seen as potentially useful for tasks with a similar structure, such as diagnosing car problems or pinpointing what has gone wrong inside a broken appliance. GPT-6 Astra is also said to perform well on visual tasks tied to robotics, pointing to a broader spatial-understanding ability that goes beyond simple image recognition.

GPT-6 Astra is OpenAI's newest flagship model, rolled out in September 2026, and it has already picked up a series of feature expansions, from voice support to built-in financial data. This benchmark result suggests that alongside those practical upgrades, the model's core ability to read fine detail out of realistic photos has also been steadily improving.

Summary

On a benchmark built to catch furniture assembly mistakes from photos, OpenAI's GPT-6 Astra scored 80%, well above rival models. The same test's top score stood at just 28% ten months ago, highlighting how quickly AI models' visual and spatial reasoning is advancing. Real-time use is still limited by speed, but the underlying skill is expected to find uses well beyond furniture assembly, including car repair and appliance diagnostics.