A different evaluation
A July 29 preprint introduces shadow evaluations for open-ended AI research. An agent receives the central question of a strong unpublished paper while the answer paper stays hidden; the original authors then grade the work.
The researchers ran the method on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute.
What worked and what failed
The agents completed the experimental engineering without human help, but made no substantial progress on either central question. The original authors unambiguously rejected both outputs.
Recurring failures included poor judgment about publishable quality, uncreative responses to design flaws, ineffective backtracking, weak resource awareness, and instruction drift. A robustness check with another model and scaffold reproduced similar problems.
What it means
The result is early evidence that current agents can accelerate implementation and experiments without yet replacing the judgment required to frame questions, interpret negative results, and redirect a research program.
Research teams should treat agents as experiment and analysis assistants, keeping hypotheses, stop rules, resource budgets, and major pivots under human review. Failed hypotheses and execution logs should remain part of the record.
Limitations
There were only two cases, and the original authors graded their own questions. This is a preprint, not a peer-reviewed estimate across disciplines or agent architectures. The released reviews, repositories, and logs nonetheless support follow-up replication.