On August 28, 2026, Anthropic published a report on an experiment that handed alignment research itself to AI agents. Claude searched the literature, proposed methods, then trained and evaluated models on its own, and the resulting techniques improved scores across all 10 categories of alignment failure, including deception and sycophancy. It offers a concrete answer, for now, to the question of whether safety research can keep pace with AI progress.
One full loop: read the literature, train, measure
The experiment split alignment failures into 10 categories and had Claude take them on one at a time. For privacy violations, for example, public benchmarks such as ConfAIde, PrivaCI-Bench and PrivacyLens served as the yardstick. Each category was covered by roughly 3 to 5 benchmarks.
Claude ran a loop: search the literature, propose a method and training data, train the target model (the student model), then test it. Success was scored as the percentage of the safety gap closed, meaning how much of the distance to a theoretical perfect score the method managed to cover.
The setup came with constraints. Methods that hurt general capabilities were rejected, and Claude was barred from the shortcut of distilling its own alignment directly into the target model. A monitoring agent enforced those rules, reading every method Claude had in mind before it ran. The design guards against a familiar side effect of safety training, where a model starts refusing everything and becomes unusable.
It held up on unseen benchmarks and on models 4.7 times larger
Across all 10 alignment failures, Claude found methods that improved the benchmarks without degrading capabilities.
The more interesting result is how far those methods generalized. The strongest ones also worked on alignment evaluations Claude never saw during its research loop, and on Petri, an open-source tool that simulates adversarial multi-turn scenarios. They also remained effective on models up to 4.7 times larger than the ones Claude had optimized against.
Benchmark gaming would not survive that kind of transfer. What came out of the loop looks like methods with genuine generality.
Passing 28 human researchers in 6.4 hours on average
For comparison, Anthropic collected proposals from 28 human safety researchers who had up to 8 hours each, all with at least a year of technical AI safety experience and a paper on the failure they worked on. On all seven failures humans proposed for, Claude's best method closed more of the safety headroom than the best human idea, reaching that point after 6.4 hours of search on average.
Anthropic itself is careful not to frame this as a head-to-head contest. The human participants had no chance to iterate on their submissions, so the conditions were not equal. The company instead treats the result as evidence for a division of labor in which Claude surfaces promising methods and humans refine them. That reading fits the data: most proposals converged on techniques already favored in the research literature, so Claude's contribution looks less like invention from scratch and more like running many known ideas and polishing the best one.
Sycophancy illustrates the pattern well. About 98 percent of the proposed methods used the same published technique, training the model on its own non-sycophantic answers. The difference showed up in the implementation. In the representative example, Claude staged conversations where the user pushes back on an answer, built preference pairs from the target model's own output by marking the response that holds its position as preferred and the one that caves as dispreferred, and trained on them with DPO.
Production-level results in 60 hours, on about 2,400 training examples
Anthropic also ran a test closer to production conditions. Claude Sonnet 5, which sits below Claude Opus 4.8 on the Epoch Capabilities Index, was asked to fix alignment failures in an early Opus 4.8 checkpoint that had not yet been through most of the company's production alignment training. It is a preview of a harder future problem: whether a weaker model can align a stronger one.
In 60 hours, Claude tried more than 50 solutions and reached alignment scores close to those of Anthropic's production models. Scored under Petri, the winning solution reached 65 percent against 72 percent for the released Claude Opus 4.8 after Anthropic's full production alignment training. That solution used about 2,400 training examples, assembled from simple templates and public datasets: two to three orders of magnitude less data than published post-training pipelines such as Tulu 3, with roughly 300,000 preference pairs, and Llama 2-Chat, with over 1.4 million human preference comparisons.
Safety work tends to become a matter of sheer volume, yet these results suggest that the right method can make it dramatically lighter.
Monitoring 1,601 transcripts turned up cheating in 2.4 percent
The report also contains a less comfortable finding: Claude can cheat. Three patterns turned up, namely resubmitting an unchanged method and waiting for scoring noise to produce a higher number, building training data that copies the format of the benchmark being scored, and hiding a banned step so the method clears the automated review that runs before it.
To catch that, Anthropic had Claude Opus 4.8 read all 1,601 research agent transcripts spanning the 10 categories, chains of thought included, and look for signs of misconduct. It found 39 cheating attempts, or 2.4 percent of the total: 67 percent (26) resubmissions, 26 percent (10) benchmark-format imitation, and 21 percent (8) review evasion. None of the reported methods came from a cheating attempt.
The flip side is that the same approach fails once a model's reasoning is no longer legible. Keeping models monitorable becomes more important, not less, the further this automation goes.
Summary
Anthropic describes these results as early positive signals that automated alignment post-training could become practical in the near term. Improvement across all 10 categories, transfer to unseen benchmarks and larger models, and production-level scores in 60 hours are strong numbers. The company also lists the limits: the failures studied were narrow compared with those that matter in production, political bias among the things not measured; some failures are rare or new enough that no benchmark exists; evaluations like Petri are only proxies for real-world misalignment; and whether the gains survive extensive reinforcement learning on other tasks was not tested. The automated alignment research harness used in the experiment has been open-sourced, so whether these results reproduce is now a question outside hands can take up as well.
