In August 2026, an internal engineering team at Anthropic ran a two-week sprint focused on speeding up claude.ai and the Claude desktop app, and announced that several key interactions became about 3.1 times faster on average (geometric mean). What stands out is that Claude itself, the AI model, handled most of the work, from finding bottlenecks to writing the fixes.
Over 3,000 changes in two weeks, with gains across every major metric
The team picked four user journeys covering 95% of overall usage and tracked 13 metrics throughout the sprint. As a result, a fresh load of claude.ai dropped from roughly 3,085 milliseconds to 550 milliseconds, Claude Code startup went from 837 milliseconds to 347 milliseconds, and loading Claude Cowork fell from 2,566 milliseconds to 728 milliseconds. More than 3,000 changes were merged in those two weeks, and the team says none of them caused a customer-facing incident or a rollback.
One Slack channel, with Claude handling everything from measurement to fixes
The whole project ran out of a single Slack channel, with more than 150 threads active at once during peak periods. The work relied on an internal research model, roughly comparable in capability to Opus 5.5, accessed through the Claude Tag beta. Human engineers set the goals, made the technical calls, and gave final approval, while Claude mostly handled identifying bottlenecks, building benchmarks, writing the fixes, opening pull requests, and watching how each release behaved afterward.
Because wall-clock timing varies too much from run to run to be a reliable optimization target, the team instead optimized against more reproducible numbers such as CPU instruction counts. In one example, cutting instruction counts by 48% in the code that assembles the message list reduced actual processing time by 78%, an improvement felt as more than four times faster. The team also added an automated check that rejects any change that pushes instruction counts back up, so the gains would not quietly slip away.
The search for smaller issues turned up some surprises. One bug froze the screen for a moment right after syntax highlighting finished on a code block. The cause turned out to be em dashes and curly quotes, characters that fall outside the basic Latin-1 range. Whenever such characters appeared, the browser's JavaScript engine switched to storing the entire reply as two-byte characters, which pushed the syntax-highlighting regular expressions onto a much slower path. A fix of about 20 lines brought the time for the first syntax highlight to render down from around one second to about 0.35 seconds.
Other fixes were just as unglamorous but added up: cutting sidebar re-render work by 90%, revising a CSS selector to shave off tens of milliseconds, and eliminating hundreds of thousands of hidden page reloads that were happening every day.
About 200 feature flags kept the rollout safe
Riskier changes were placed behind short-lived feature flags and rolled out in stages, first to employees, then to a small slice of users, and finally to everyone. Around 200 flags were added over the two weeks, and more than half had already been removed by the time the sprint ended. Every change went through automated review plus sign-off from at least one human engineer before reaching production. Shared layout components were also tested across 14 different viewport sizes to confirm that nothing shifted by more than a single pixel.
Summary
The sprint stands out as a case of an AI company putting its own AI model to work on a famously unglamorous problem, its own product's speed, and getting measurable results in a short span of time. With humans setting direction and giving final approval while the AI handled measurement, implementation, and monitoring, the project offers an early look at how much responsibility an AI agent can realistically take on in day-to-day software development.
