A team of 9 researchers from Harvard University, Stanford University, Together AI, and Caltech has released HAWKEYE, a framework that gives AI coding agents GPU-specific knowledge so they can generate and optimize compute kernels on their own. Tested across 4 GPU architectures from NVIDIA and AMD, it reports a geometric mean speedup of up to 18.9x on emerging attention operations.
Every new chip means rewriting the software by hand
New AI accelerators arrive almost every year, but the software that extracts their performance has to be rebuilt each time. Instruction sets change, memory hierarchies change, and cache behavior changes, so kernel authors end up retuning their work for every generation. The pool of engineers who can do this is small, and the lag between a chip shipping and its software catching up shows up directly as lost effective performance.
Attempts to have large language models write this code are not new, but low-level GPU programming has been a weak spot. A model can produce syntactically valid code while still missing the practical instincts that matter here, such as which memory access will become the bottleneck or which instruction happens to be fast only on a particular generation.
HAWKEYE aims at that gap. Instead of feeding an agent thick technical manuals or entire production kernels, it hands over a structured account of the optimization patterns that come up again and again.
Teaching with about 10 unit tests instead of a manual
The team built a taxonomy covering standard optimization strategies such as vectorized memory loads, asynchronous data movement, and warp-level operations. The agent uses that taxonomy to decide which technique to reach for.
Adapting the system to a new architecture requires humans to supply only about 10 expert-written unit tests plus reference kernels. From there the agent works in an environment with real GPU access, editing code, running it, verifying correctness, and profiling, then layering multiple optimizations according to the bottlenecks it finds.
From the human side, the cost of teaching shrinks from "read everything" to "show the key points as tests." That difference compounds as the number of chips you want to support grows.
How to read the 18.9x figure
The reported numbers mean different things depending on what is being measured.
For newer attention operations such as Linear Attention, where existing libraries offer thin coverage, the result is a geometric mean of 18.9x over torch.compile. Read the other way, the gap is that wide precisely because the existing implementations were not well optimized. For ordinary BF16 workloads the figure settles at 1.13x over torch.compile, and 1.28x for low-precision formats. Not every GPU workload gets 18.9x faster, and that caveat is worth keeping in mind.
The comparison against hand-written expert kernels is more grounded. Against a Triton implementation of Causal Linear Attention the result was 0.92x, and for Forgetting Attention it was 1.05x. That is roughly even with human experts, winning some and losing some. The value of automation here looks less like beating people and more like spreading kernels of comparable quality into corners that experts never have time to reach.
Whether more attempts keep paying off
Another interesting result concerns what happens as the agent is given more attempts. The evaluation used Gemini 3.1 Pro and GPT-5.4.
According to Arya Tschand, one of the co-first authors, HAWKEYE kept improving as attempts accumulated, whereas approaches that simply hand over hardware documentation and complex implementation examples plateaued early. More information does not make an agent smarter; what changes the ceiling is whether it was given a structure to search against. It is a concrete example of how the design of what you feed an agent determines its performance.
Summary
HAWKEYE classifies GPU-specific optimization patterns and hands them to AI coding agents so they can optimize compute kernels automatically. The human effort needed to support a new architecture drops to roughly 10 unit tests and reference implementations, and the team reports a geometric mean of 18.9x over torch.compile on new attention operations along with 1.13x to 1.28x on ordinary workloads. Results land roughly even with hand-written expert kernels, and expert knowledge is still required for the first tests and reference implementations on an unfamiliar architecture. Even so, as a way to shorten the gap between a chip generation and the software that catches up to it, this is worth watching.
