Skip to content

Signals · Capability Displacement

AI agents can run the experiments but not choose them

A Princeton-led study gave AI agents six days and $3,000 in API credits to write a NeurIPS-grade paper. The human authors rejected both attempts.

by Jo·4 min read·

AI agents can run the experiments but not choose them

MIT Technology Review reported in August on a study that put the industry's loudest promise under a controlled test and watched it fail. A multi-institution team led by Peter Kirgis and Sayash Kapoor at Princeton gave AI agents six days, $3,000 in Anthropic API credits, a GPU budget, their own virtual machines and open web access, then asked them to produce publication-grade research. The agents ran Anthropic's Claude Opus 4.8 on open-source software called OpenClaw. The questions came from two unpublished papers submitted to NeurIPS 2026, so the answers could not be memorised or looked up. The original human authors graded the output the way they would grade a conference submission. They rejected both papers.

The split the study actually found

The agents were not incompetent. They reviewed the literature, ran hundreds of experiments and compiled results. Every piece of engineering the research required, they did.

What they could not do was the research. "On the other hand, the agents were unambiguously bad at carrying out the research itself," Kapoor told MIT Technology Review. The agents committed to unpromising approaches too early, rejected their own ambitious hypotheses on thin data, and could not backtrack. They made small pivots. They could not rethink. When reviewing tools pushed back, they did not revise the methodology — they narrowed the claims and added caveats.

Kapoor's explanation is the part operators should copy into their own notes. Models get good at whatever can be drilled through reinforcement learning, and reinforcement learning needs a checkable answer. "It's harder to create environments to train these models when the task itself is open-ended," he said.

That is the shape of the whole labour question, stated by someone measuring it at the frontier of AI research itself. Work that can be scored automatically gets absorbed fast. Work whose success cannot be defined in advance moves slowly. Najoung Kim of Boston University, who was not involved in the study, raised the same possibility: AI progress may bifurcate, racing on narrow scored tasks while crawling on open-ended ones.

Why the comfortable reading is wrong

The consensus take on this study is relief — the machines are not about to build themselves, timelines slip, everyone gets more runway. That reading gets it backwards.

The finding is not that AI is weak. It is that AI is now a superb engineer with no taste. Jack Clark, Anthropic's cofounder, wrote in his Import AI newsletter that today's systems are "extraordinarily capable engineers" with "a certain property of rote, formulaic thinking" — and he was describing the same gap after Anthropic tried to automate parts of its own safety research.

Read that as a labour-market map rather than a safety debate. If agents can run hundreds of experiments and write the code but cannot choose the question, then execution roles compress and judgment roles hold. The person who receives a well-specified task and completes it well is competing directly with something that does not sleep. The person who decides which task was worth specifying is not.

The study has limits its authors name: two papers, graders who knew the work was machine-generated, and substantial researcher discretion. This is a signal, not a law. It also runs against the commercial incentive — OpenAI has made an automated AI researcher an explicit goal, and Anthropic calls self-improving AI the industry's next milestone. The people paid to be optimistic are finding the same wall internally.

Your Next Move

Audit your week against the checkable/uncheckable line. List what you did in the last five working days. Mark every item where success could be verified by a script or a rubric written in advance. That column is the compression zone, whatever your title says.

Move your visible output up one level. Stop shipping completed tasks and start shipping the framing: which question was worth asking, which approach you abandoned and why, what evidence would change the answer. The study's agents failed precisely there, and that failure is now a paid skill.

Practise backtracking in public. Once this month, kill an approach you have already invested in and document the reasoning for your team. Judgment is demonstrated by reversals, not by throughput.

I build systems like the one publishing this site. → Work with me

About the author

Jo

Jo runs The War Room: strategic intelligence for operators navigating AI disruption, influence, and empire-building.

Sources

Get the Briefing

Want more intelligence like this?

Get the free AI Survival Kit — 7 strategies from 13 playbooks.

Takes 30 seconds. No spam.

More in Signals

See all