What drives up the bill for an AI coding agent is not how smart the model is, but how it walks around a codebase. Sonar, known for its static analysis tools, has published measurements taken inside its own repository. A single change of roughly 800 lines accumulated 156 million context tokens. Everything the agent gathered through grep and file reads stayed in the conversation and was billed again on every later turn, and the numbers put a price on that structure.

An Agent's Detours Are Billed on Every Turn

A large language model receives the entire conversation so far as input each time it plans its next step. Prompt caching cuts the unit price of those resent tokens to roughly 10 percent of the input price, but it never brings them to zero. The real cost of a token is therefore not its size, but its size multiplied by the number of turns it survives.

An agent that opens a 600-line file on turn 40 has not paid for 600 lines once. It has paid for 600 lines across the remaining 470 turns. The more files it reads, and the longer the session runs, the more an early over-read compounds like interest.

156 Million Tokens for an 800-Line Change

Sonar builds its own semantic analysis engine with the help of AI agents, and it keeps the full traces of that work. The example it published is a single pull request containing a backend change of about 800 lines.

That pull request consumed 156 million context tokens, of which cache reads accounted for 152.8 million. The context window peaked at 459,000 tokens, and the cost came to about 41 USD (about 6,400 yen). The final diff itself was something a person could read in five minutes.

What the agent was doing amounts to the "go to definition" feature that any IDE provides for free. To find the target of a call, it ran a command like this:

grep -rn "resolve_return_type"

Back came a list mixing the Python, TypeScript, Java, Rust, C# and shared-core implementations all at once. A regular expression has no way to decide which one a given call binds to. So the agent opened a file, read a generous slice, went looking for the helper that produced the return type, and opened another file. Every one of those reads stayed in the conversation.

One Over-Read Turns Into 2.7 Million Tokens

One case breaks down clearly. The agent only wanted to understand a 67-line function, but because it could not tell where the function started and ended, it read the entire 618-line file. It loaded 6,472 tokens and used the equivalent of roughly 700. About 5,770 tokens were wasted on the spot.

That read entered the conversation on turn 42 and stayed for the remaining 470 turns. Multiplied out, 5,770 by 470 comes to roughly 2.7 million tokens. At a cache-read price of about 0.2 USD (about 31 yen) per million tokens, that single read cost about 0.54 USD (about 85 yen). The same kind of over-read happened about 10 times in that one pull request, alongside dozens of tree-wide greps, several of which came back empty and forced a second, wider search.

Averaged across 18 comparable pull requests in the same repository, each one consumed about 234 million context tokens and about 65 USD (about 10,000 yen), with a median of about 52 USD (about 8,000 yen). Each involved roughly 700 model round-trips, and context windows peaked anywhere between 450,000 and 975,000 tokens. Touching the 1 million token ceiling forces compaction, and earlier context is lost. The width of that spread tracks how much of the codebase the agent had to traverse, not the size of the final diff.

※1 USD = 157 JPY (as of September 3, 2026)

Grep Finds Strings, Not Meaning

The second side effect is more troublesome than the money. A regular expression only finds the spelling you thought of. Calls that arrive through an interface, references that go via an alias, and calls that come from another language all slip through when the search term differs.

This particular pull request was about porting call-site type resolution, already implemented in C#, over to Python. But the C# function is named resolve_type_node, while the Python side uses resolve_return_type. Searching for one will never surface the other. Missed call sites come back as failed builds, which mean another continuous integration run and more rework, and each round pays the context tax again.

Replacing Text Search with a Graph Query

Sonar Vortex sets out to replace that navigation entirely. It sits inside the agent's loop and answers questions from a graph of the codebase rather than from the filesystem.

The graph is produced by SemSitter, the semantic navigation engine Sonar developed in house. It holds functions, methods, classes, fields and parameters as nodes, connects them with typed edges such as calls, references, returns, has-param, is-type, contains and extends, and keeps this Unified Dependency Graph (UDG) up to date on every change.

The question the agent asks changes as well. Instead of "which files mention resolve_return_type," it asks for the definition this call binds to, the type that owns it, its return type and its callers. What comes back is one method body plus the edges that answer the rest. No surrounding file is carried along, and there is nothing to widen. The over-read above would have cost roughly 330,000 tokens as a graph lookup, about one ninth of the original.

Two further edge types make the difference larger. The first is documented_by, linking code to prose, so touching a function surfaces the single relevant paragraph of a design document instead of an entire wiki. That inverts the familiar problem of an agent being led astray by stale documentation. The second is semantically_related, which links equivalent implementations across languages even when the names differ. With that edge in place, the C# and Python mismatch above becomes a single hop.

Dropping Into an Existing Coding Stack

Sonar Vortex is built around MCP (Model Context Protocol) integration, so it can be added to tools teams already use, including Claude Code, Cursor, GitHub Copilot, Windsurf, Google Gemini CLI and OpenAI Codex CLI. Verification reuses existing SonarQube quality profiles, so there is no new rule set to write. Injecting context and constraints before code is written, and verifying changes as they are made, run as one continuous loop.

Summary

What stalls AI coding agents in large codebases is not the quality of their reasoning but the way they move. Sonar's measurements show 156 million context tokens for a single 800-line pull request, about 65 USD (about 10,000 yen) per pull request averaged across 18 of them, and context windows repeatedly approaching the 1 million token ceiling. On top of that, text search misses equivalent implementations in other languages, and that shows up later as rework. Replacing navigation with graph queries is an optimization that reduces what the agent carries rather than adding to it. Because the gap widens with the size of the codebase, teams running agents day to day may find it worth reviewing the line items on their bill.

※ The thumbnail image is an AI-generated illustration.