00 · IN THREE MINUTES
The answer in three steps
- 1In two unpublished AI-paper case studies, agents completed engineering work but did not answer the central research questions to the authors’ standard.
- 2A separate study found 37,802 AI-generated ideas clustered closer to seed literature than human follow-on work.
- 3The results expose current weaknesses; they do not prove that every agent, field or future system will fail.
01 · THE TEST USED UNFINISHED REAL RESEARCH
The test used unfinished real research
Researchers gave agents the central open question from two high-quality unpublished NeurIPS submissions, six days and substantial compute. Original authors reviewed the resulting work, creating a demanding but very small .
02 · ENGINEERING WAS NOT THE BOTTLENECK
Engineering was not the bottleneck
The agents wrote code and ran experiments without human help. They struggled to judge the publication bar, redesign weak experiments, retreat from dead ends and preserve the intended research question.
The test used unfinished real research
Researchers gave agents the central open question from two high-quality unpublished NeurIPS submissions, six days and substantial compute. Original authors reviewed the resulting work, creating a demanding but very small shadow evaluation.
2 shadow casesEngineering was not the bottleneck
The agents wrote code and ran experiments without human help. They struggled to judge the publication bar, redesign weak experiments, retreat from dead ends and preserve the intended research question.
6 daysBreadth was tested separately
Across four agent frameworks and six language models, another preprint generated 37,802 ideas from shared seed papers. The ideas were more concentrated and closer to the seeds than comparable human research, with novelty dominated by recombination.
37,802 ideasTwo studies point to a distinction
Executing a specified technical plan and choosing a valuable new scientific direction are different abilities. The experiments suggest that current scaffolds can amplify the first without automatically supplying the second.
early evidence03 · BREADTH WAS TESTED SEPARATELY
Breadth was tested separately
Across four agent frameworks and six language models, another preprint generated 37,802 ideas from shared seed papers. The ideas were more concentrated and closer to the seeds than comparable human research, with novelty dominated by recombination.
04 · TWO STUDIES POINT TO A DISTINCTION
Two studies point to a distinction
Executing a specified technical plan and choosing a valuable new scientific direction are different abilities. The experiments suggest that current scaffolds can amplify the first without automatically supplying the second.
05 · THIS IS A BASELINE, NOT A VERDICT
This is a baseline, not a verdict
Both papers are preprints and focus on AI research. Replication across natural sciences, longer time horizons, different tools and human–agent teams is needed before claiming a general law of automated discovery.
06 · SOURCES AND EVIDENCE
Sources and evidence
Claims are linked to foundational papers, standards or the primary study behind the update.
- 01Can AI agents conduct open-ended AI research? Early evidence from two case studiesOPEN PREPRINT ↗
Supports a defined mechanism, measurement or evidence boundary in this article.
- 02AI Research Agents Narrow Scientific ExplorationOPEN PREPRINT ↗
Supports a defined mechanism, measurement or evidence boundary in this article.
- 03AI isn’t ready to research itselfRESEARCH EXPLAINER ↗
Supports a defined mechanism, measurement or evidence boundary in this article.
