Data & Decision ScienceAug 20, 2026

A "shadow evaluation" test finds agents still cannot do open-ended AI research

MIT Technology Review reported on 18 August 2026 on a study testing whether AI agents can do the open-ended part of research — choosing hypotheses, deciding what evidence would settle a question, knowing when to start over. The researchers propose shadow evaluation: put the agent to a research question drawn from a high-quality unpublished paper, so the answer cannot have been memorised. Running Claude Opus 4.8 against questions from two papers submitted to NeurIPS 2026, they found agents are not yet capable of conducting open-ended AI research, which puts recursive self-improvement further out than the loudest forecasts assume.

What it means The gap is in framing the question, not executing the work — which is the argument for keeping a human on hypothesis selection and letting agents run the parts where the question is already settled.

Where it came from MIT Technology Review

Back to the Stream