On September 11, OpenAI published a customer story featuring Cognition, the company behind the autonomous software engineer Devin[1]. With GPT-6 Astra built in, Devin now tests the code it writes and returns a recording of the software running alongside a report showing what was checked and what was left untested. The goal is to cut down how much code humans have to read by hand.

Letting the Author Handle the Testing

Cognition builds Devin, used by companies ranging from large banks to tech-native startups. As its engineering teams write more code, reviewing that work has itself become the bottleneck[1]. The faster code gets generated, the harder it is for review to keep up.

That is why GPT-6 Astra's behavior caught the company's attention. Walden Yan, co-founder of Cognition, points to Astra's ability to test its work and prove that it functions as expected as the model's big improvement[1]. The value here is not raw coding speed, but whether the model can supply evidence that what it wrote actually runs.

Recording the Run and Declaring What Was Not Covered

One concrete example is the testing of Otter Run, an iPhone game. Devin uses Astra to test the game and returns a recording of it running in a simulator, together with a report listing the checks that passed and the areas left untested[1].

The two artifacts play different roles. The recording shows how the application actually behaved, while the report documents the scope the testing covered. By reading them side by side, engineers can grasp both the real behavior and the remaining gaps without walking through the code line by line[1].

The benefits are not limited to development. When a customer sends a screenshot of a bug, the team can pass it straight to Devin through Astra and get back a screenshot of the fixed result. Yan says this has made responses to customers much quicker[1].

Rolled Out Across Devin, the CLI, and the Desktop App

Cognition is applying Astra across its product lineup. Beyond the core cloud agent, the company says it is using the model to improve its CLI and desktop products as well[1]. According to Cognition's own announcement, Astra is already available in Devin Desktop and Devin CLI, and it runs as part of the model mixture in Devin Cloud[2].

The company has also published benchmark numbers. On FrontierCode 1.1, its proprietary benchmark that grades models on real-world engineering tasks by quality and mergeability, GPT-6 Astra scored 64.5. That puts it above Claude Fable 5.1 at 63.6 and within 0.4 points of the leader, Claude Fable 5 at 64.9, while costing 64 percent less[2].

Cognition adds that when Astra powers Devin's testing capabilities specifically, it reaches state-of-the-art results on the company's internal testing benchmark. The tests it produces are more comprehensive, the reports are clearer, and the video evidence is easier to work with[2].

Whether Review Itself Changes Shape

What Cognition is aiming at is a rebuild of the review process. If Devin can test its own work and submit the results as evidence, engineers can evaluate changes while reading far less code manually[1]. Yan expects that over time the team will have to look at less code by hand and ship more in the end[1].

That said, the whole idea rests on the quality of the evidence. If test coverage is thin, recordings and reports turn into material that merely feels like verification. The fact that untested areas are declared explicitly reads as a design choice made with that trap in mind. How far the output can be trusted in practice is something teams will have to measure for themselves.

Summary

Cognition has built GPT-6 Astra into Devin's testing capabilities, so the agent now returns a run recording paired with a report on what it covered. On FrontierCode 1.1 the model comes within 0.4 points of Claude Fable 5 at 64 percent lower cost, and it is said to reach state-of-the-art results on the company's internal testing benchmark. Against a situation where writing code keeps getting faster while verification piles up, this is a case of handing the evidence-gathering side over to the model.

[1] Source: https://openai.com/index/cognition-devin-testing-with-astra

[2] Source: https://devin.ai/blog/gpt-6-astra