On June 28 (local time), xAI began offering its new AI assistant "Grok 4.5" as a private beta to its subsidiaries SpaceX and Tesla. At its core sits a new 1.5-trillion-parameter foundation model (a large-scale model that serves as a base for many uses) called "V9." Elon Musk said that early evaluations show performance "close to, and perhaps exceeding," Anthropic's top model Opus, but the model has not been released to outside parties and no third party has verified the claim. The roadmap disclosed alongside it — shipping a brand-new foundation model built from scratch every month — is also drawing attention.
A New 1.5-Trillion-Parameter Foundation, "V9"
V9, xAI's ninth-generation foundation model, completed pre-training (the process of building a model's base on massive data) on May 26. At 1.5 trillion parameters, it is three times the size of "v8-small" (500 billion), which currently runs in production on X and in Tesla vehicles.
Size alone does not determine performance, however. In V9, the workflow data from the coding tool Cursor was reportedly added not from the start of pre-training but in a later, supplemental stage. An xAI engineer acknowledged that this is "not quite as good as having it in initial training," and the company says its next model, at 2 trillion parameters, is designed to incorporate Cursor data from the pre-training stage. In other words, the freshly beta-tested Grok 4.5 may carry a structural weakness that its successor is specifically built to address.
V9 was optimized for NVIDIA's Blackwell-generation GPUs and trained on xAI's large-scale compute cluster "Colossus" in Memphis, Tennessee. Reinforcement learning (a method of improving performance through trial and error) reportedly continues to refine the model even during the private beta.
The "Close to Opus" Claim Rests on Internal Tests Only
What deserves caution in Musk's claim is that all of the supporting evaluations were conducted internally. The teams running the tests belong to SpaceX and Tesla — both part of the same SpaceX corporation that now owns xAI — so there is no independent third party involved.
Moreover, xAI has not submitted the current Grok generation to any major third-party benchmark. The AI industry has repeatedly documented gaps between developers' self-reported figures and externally measured ones; in one case, a developer claimed 50 percent on a difficult exam while independent researchers measured just 29.4 percent.
The comparison target is also a moving one. On the "Artificial Analysis Intelligence Index," a composite of nine evaluations run by an independent body, Claude-family models such as Claude Fable 5 and Claude Opus 4.8 currently sit at the top, and Grok does not appear. On "SWE-bench Verified," which measures coding ability, Claude Code running on Opus 4.7 scored 87.6 percent, while xAI's older model reached 70.8 percent. V9 is surely a step up from the older model, but as long as it is neither public nor benchmarked, "close to Opus" remains an unverifiable claim.
Why SpaceX and Tesla Come First
Behind the decision to route Grok 4.5 to internal teams before the public lies a change in corporate structure. SpaceX acquired xAI in February 2026, combining the rocket company and the AI lab into a single entity valued at a combined 1.25 trillion USD (about 201 trillion yen). ※1 USD = 161 JPY (as of July 3, 2026)
This merger created an unusual loop for model development. Most AI companies test models on public benchmarks and synthetic tasks, but SpaceX and Tesla generate the demands of real work — satellite trajectory calculations, vehicle manufacturing workflows, and software written by thousands of engineers. The aim is to obtain signals that standardized exam scores cannot capture.
Even so, this is not a product ordinary users can access. The public API that developers use still runs on Grok 4.3, built on the "v8-small" foundation — a model Musk himself once described as having "many fundamental flaws." The gap between the model sitting in SpaceX's engineering bay and the one available through the API amounts to roughly a trillion parameters.
A Plan to Ship a New Foundation Model Every Month
Perhaps the more consequential part of the announcement is not the beta itself but the plan embedded within it. xAI says it will ship new foundation models — not upgraded versions, but ones trained from scratch — every month through the end of 2026.
Monthly software releases are nothing unusual, but monthly foundation-model releases are another matter. Training a foundation model costs hundreds of millions of dollars and requires the sequential stages of pre-training, supervised fine-tuning, and reinforcement learning. If xAI truly runs this every month, it amounts to an unprecedented claim about Colossus's compute capacity.
For developers, this raises a concern separate from capability. If a new model appears every 30 days just after they have evaluated one and built it into a product, the cost and effort of switching become hard to ignore. A rapid release cadence may in fact make choosing a model harder for developers, not easier.
How the Cursor Acquisition Fits In
xAI's moves are also tied to its acquisition of the coding tool Cursor. That deal, struck with developer Anysphere, is a 60-billion-USD (about 9.7-trillion-yen) all-stock transaction signed in June 2026 and expected to close during the third quarter. If completed, it would give xAI access to Cursor's workflow data and its base of more than one million developers.
The near-term constraint, though, is timing. Cursor has maintained a "model-agnostic" stance, able to handle several models at once — Anthropic's Claude, OpenAI's GPT, and its own Composer. Until the merger closes and V9-based models reach the public API, those options are expected to remain unchanged.
Summary
Grok 4.5 is an ambitious model built on V9, xAI's largest foundation model yet, but much of its true shape remains hidden. It is confined to a private beta at SpaceX and Tesla, no public release date has been given, and it has faced no third-party benchmark. On top of that, the public API that ordinary users can actually touch still runs the older model whose flaws the developer itself admits. Whether the "close to Opus" performance claim is real will have to wait for public release and independent evaluation. If the plan to ship a new model from scratch every month comes to pass, it would reset the industry's sense of tempo — but for now, a gap of a trillion parameters remains between what is claimed and what is actually usable.
