1/4

Shadow evaluations expose the research gap in AI agents

A practical brief on shadow evaluations: test whether an agent can choose, recover, and defend a research direction after it has learned to run the tools.

Shadow evaluations ask a harder product question than "can the agent use tools?" Give the agent the central research question from an unpublished paper, hide the paper and findings, then have the original authors grade the output like a conference submission. In two NeurIPS 2026 case studies, the agents completed the engineering work but the final papers scored 2/6 and 1/6; both were rejected by the authors. 1
The first prototype is an evaluation gate: run one internal question with a fixed model, tools, time, compute budget, and stop rules; then let subject-matter experts score the artifact and replay the log for shortcuts, dead ends, budget use, and instruction drift. The paper releases expert reviews, survey responses, agent repositories, and run logs, but its evidence is still preliminary: two cases, non-blinded grading, and scaffold effects remain open. 1

Related content

Comments

Sign in to comment.