Home / News / Anthropic's Automated Researcher Just Outperformed Human AI Safety Experts
Anthropic

Anthropic's Automated Researcher Just Outperformed Human AI Safety Experts

Aug 29, 20265 min read
Anthropic's Automated Researcher Just Outperformed Human AI Safety Experts

News Summary

Anthropic has published new research showing that an automated system built from its own Claude models can independently discover techniques to reduce misaligned behavior in AI models, and in some tests it outperformed methods proposed by experienced human alignment researchers. The work, led by Anthropic fellow Chen Yueh-Han and detailed in a paper titled "Automated Researchers Can Reliably Mitigate Alignment Failures," was published on August 28, 2026, Eastern Time, and offers one of the clearest public demonstrations yet of what a limited, narrow form of AI self-improvement looks like in practice.

What the system actually does

The system, which Anthropic calls an Automated Alignment Researcher (AAR), is designed to fix a specific category of AI safety problem: misaligned behaviors such as sycophancy, deceptive reasoning, or reward hacking that show up during training. Rather than waiting for a human researcher to diagnose the issue and propose a fix, the AAR searches existing research literature, proposes its own mitigation method, and then runs a short training experiment — typically around 30 minutes per iteration — to test whether the fix works. Methods that improve the target benchmark are kept and refined further; methods that fail are discarded. The process repeats automatically across many iterations.

Anthropic tested the AAR against 10 separate alignment benchmarks, each representing a different known failure mode. According to the paper, the automated system improved performance on all 10 benchmarks without measurably degrading the model's general capabilities — an important caveat, since alignment fixes sometimes come at the cost of a model becoming less useful or more evasive.

The human comparison

The headline figure from the paper is a direct comparison with human researchers: "the best AAR method beats what experienced humans propose, on average within six hours." In other words, given six hours of autonomous iteration, the automated system's best proposed fix outperformed the median mitigation strategy suggested by experienced human alignment researchers working on the same benchmarks.

Anthropic also highlighted a cost comparison. The AAR runs at roughly $4 per hour in API inference costs, compared with an estimated $150 per hour for a human researcher's time. The company frames this less as a claim that automated researchers are qualitatively better than humans, and more as evidence that automated alignment work is now cheap and fast enough to run continuously and at scale.

Why Anthropic calls this a "peek" at self-improvement, not the real thing

Anthropic is careful to frame the AAR as a narrow proof of concept rather than a general-purpose self-improving system. The paper notes that its effectiveness depends heavily on how well the benchmarks used actually reflect real alignment goals, and on keeping the underlying literature database the system searches up to date. If a benchmark is a poor proxy for genuine alignment, the AAR could optimize for the wrong thing — a well-known risk in machine learning known as reward or benchmark hacking.

Even so, the paper states plainly that "automated alignment post-training could become practical in the near term," and researchers see the work as an early building block toward AI systems that can meaningfully contribute to their own safety research, and potentially, eventually, to broader aspects of their own training and development.

How this fits into the wider recursive self-improvement debate

The AAR paper lands amid a broader industry conversation about recursive self-improvement — the idea that AI systems could eventually help design, train, or refine their successors with less and less human involvement. Anthropic has previously argued that this scenario, sometimes shortened to RSI, deserves serious attention: the company has noted internally that Claude models already write a large majority of the code merged into Anthropic's own systems, a sign of how quickly AI assistance has been absorbed into AI development itself. Anthropic has floated the idea of a global coordination mechanism that would give the industry the option to slow or pause frontier AI development if evidence of runaway self-improvement emerged, drawing loose comparisons to arms-control frameworks.

That framing has drawn skepticism from outside researchers, some of whom argue that safety-focused rhetoric from a commercial AI lab should be weighed against competitive incentives to keep building faster, more capable models. Separately, a Princeton-led study published earlier in August tested whether AI agents could conduct genuinely open-ended AI research — writing full papers rather than fixing narrowly defined benchmarks — and found current systems were technically capable but creatively weak, unable to judge which ideas were worth pursuing the way a human researcher would. Anthropic cofounder Jack Clark has pointed to this kind of creativity gap as a "bearish signal" on how soon full recursive self-improvement could actually arrive.

The takeaway

Taken together, the research suggests a middle path between two extremes. AI is not yet capable of open-ended, human-level research judgment, but it has become demonstrably useful at narrow, well-specified research tasks — including, notably, some aspects of AI safety research itself. Anthropic's AAR work shows that this narrow capability is advancing quickly enough that automated alignment tools could move from research demo to practical deployment within a relatively short timeframe, even as the harder, more open-ended forms of AI self-improvement remain unresolved.

AnthropicAI alignment