METR, the nonprofit known for evaluating AI capabilities, has disclosed two security incidents from 2026. In the first, an API key leaked from a researcher's personal setup was abused for three weeks, burning credits with a commercial value of roughly 600,000 USD (about 93 million yen). In the second, external attackers systematically probed the organization's public infrastructure. METR says it found no evidence that sensitive information was accessed in either case.
※1 USD = 155 JPY (as of September 4, 2026)
It Started With a Small App the Researcher Had an AI Write
The first incident dates back to March 2026. A researcher without access to sensitive systems was running agents on an Amazon EC2 instance under a personal account. The instance was deliberately made reachable from the internet, sitting behind Google authentication. It also held an API key for METR's public-models account.
The app itself had been vibe-coded, and its authentication contained a fail-open flaw: when the check failed, it let the request through instead of blocking it. Because the failure was silent, nobody noticed. The environment sat wide open for several days.
METR believes the attacker located it by mining certificate transparency logs for recently registered domains, then filtering for sites whose names hinted at language models or agents. The goal was exactly what was sitting there: a model provider API key. Personal test environments have become a shopping list for credential hunters.
The intrusion itself was very much of this era. The attacker prompted the running agent directly and had it disclose the API key it was holding, added an SSH key to keep access, and then spent three weeks consuming large volumes of tokens with the stolen credentials.
Why It Took Three Weeks to Notice
An organization built around evaluation missed this for weeks, and METR points to three reasons.
First, heavy token usage was routine. Running large-scale evaluations against pre-release models produces a steady stream of rate limit and API errors, many of them spurious, and the team had grown used to them.
Second, the internal usage dashboard at the time did not surface rate-limited requests across all users, so the signal never appeared where anyone was looking.
Third, and most decisively, the credits had been granted to METR free of charge by the model developer. There was no invoice to act as a natural ceiling, and at the time there was no way to set a spending limit on a key like that. The 600,000 USD figure is the commercial value of the consumed credits, not money METR actually lost. Still, the fact remains that nobody was watching the number precisely because nobody was paying it.
May Brought a Full Campaign From Outside
The second incident came in May of the same year. METR was tipped off that it was being targeted by attackers who appeared financially motivated and were likely after access to frontier models.
Their method was exhaustive reconnaissance of public infrastructure: credential stuffing against authentication providers, attempts to obtain OAuth token grants, scanning of newly deployed services, and phishing aimed at staff. Notably, much of the vulnerability discovery work was automated with agents.
At the same time, METR had left a door ajar. Its public transcript viewer had inadvertently exposed a read-only SQL query mechanism. By default the queries were scoped to public data, but a bug could be exploited to reach unpublished evaluation data. Worse, that database had accidentally been loaded with sensitive model data that was never supposed to be there.
An independent security researcher found the flaw and disclosed it responsibly. METR took the API offline and paid a bounty. The attackers had probed the endpoint in passing, but there is no evidence they discovered the bug or reached any non-public data.
Separating Public Services From Internal Systems, and Hiring for Security
In response, METR reworked its posture. It clarified and expanded the policies that apply to all employees and contractors regarding placing METR credentials or data on non-METR infrastructure and devices. It also formalized a security review process for any researcher deploying an application publicly. Monitoring coverage was widened while noisy alerts were pruned, and spend alerts were added to keys wherever possible.
Structurally, public-facing applications now run in a dedicated production environment that is architecturally separated from internal infrastructure, so a misconfiguration on the public side cannot expose internal data.
On staffing, METR hired a security lead and plans to expand further, including a full-time security engineer. Legacy infrastructure that had been widening the attack surface was shut down. Logging around database queries and API usage was increased, and monitoring for unusual API key activity was put in place. Credential lifetimes were shortened and permission scopes narrowed. METR notes that the measures described were accurate as of July 30, 2026.
Summary
What stands out in METR's disclosure is not the size of the loss but the conditions that let it go unseen. In an organization where heavy API usage is normal, abnormal token consumption blends into the background. Free credits come without the simple alarm bell of a bill. And the initial entry point was a small personal app written by an AI. A key handed to an agent is a key an agent will hand over if asked. Rethinking where keys live and how long they stay valid looks like the slow but reliable fix.
