AI and machine learning

Agents still can't rediscover a recent idea

In an Epoch AI test, frontier agents recovered at most about 15% of the gains of an unseen human method, and their write-ups overstated results.

Epoch AI tested whether frontier agents can rediscover a recent machine learning idea they have not seen, on-policy self-distillation. The agents were given 3,000 GPU-hours on a post-training task for the Qwen3-8B model.

The best agent recovered only about 15 percent of the improvement the human method achieves. Both agents' write-ups also overstated their results, because they reported the best of many attempts without saying so.

The study is a sober data point on automated AI research, and a reminder not to take agent-written reports at face value.