August 30, 2026, (Inside AI) — Anthropic has published early evidence that AI systems can reliably train other AI models to be safer, a capability long considered a milestone on the path to artificial general intelligence. The study, titled 'Automated Researchers Can Reliably Mitigate Alignment Failures,' was released on Friday, August 28.
The research, led by Chen Yueh-Han, a researcher in Anthropic's fellows programme, details how automated alignment researchers, or AARs, built with Claude Opus 4.8, improved a model's performance across 10 alignment benchmarks without degrading overall performance.
Each AAR searched literature, proposed a training method, and trained the target model for about 30 minutes on an Nvidia H200 GPU. Over several iterations, the system kept effective methods and discarded ineffective ones, allowing faster operation at greater scale.
The best AAR method beat what experienced human AI researchers proposed, on average within six hours. The automated systems also cost roughly $4 per hour in API inference, compared to $150 per hour paid to human researchers.
"Overall, these results provide early evidence that automated alignment post-training could become practical in the near term," the paper read.
Recursive Self-Improvement Inches Closer
The findings mark a step toward recursive self-improvement, where AI systems enhance their own capabilities. This concept is often cited as a key indicator of AGI, defined by OpenAI as "highly autonomous systems that outperform humans at most economically valuable work."
The timing is notable. OpenAI CEO Sam Altman has said AGI could happen by the end of this year, according to a report by Time. Mark Chen, OpenAI's chief research officer, reportedly estimated the company is "80 per cent of the way" there.
However, the Anthropic paper acknowledges limitations. AARs only work insofar as benchmarks reflect actual alignment goals. Significant work remains in establishing and maintaining those benchmarks. Expanding the AI research literature that automated researchers draw from will also require human counterparts.
Benchmarks Versus Real-World Alignment
The study's reliance on benchmarks raises a critical question: do improvements on specific alignment tests translate to safer real-world behavior? Benchmarks are proxies, and their validity depends on how well they capture the nuanced goals of alignment.
Anthropic's own paper concedes this gap. Human researchers are not obsolete, because the automated systems cannot yet define or update the benchmarks themselves. That task remains firmly in human hands.
Still, the cost and speed advantages are striking. AARs operated at a fraction of the cost of human researchers and completed tasks in hours, not days. If scaled, such systems could accelerate safety research across the industry.
The study also arrives amid broader skepticism. Yann LeCun, often called the 'Godfather of AI,' has called AGI overrated and said scaling AI won't work. His view contrasts sharply with the optimism from OpenAI and the incremental progress shown by Anthropic.
The paper's release on a Friday afternoon suggests Anthropic is not overstating the results. The findings are early, limited to a specific set of benchmarks, and far from a declaration of AGI. But they provide a concrete data point in an otherwise speculative debate.