Skip to content

Signals · Expert Work

Frontier AI is solving open problems in history, and the historian still decides which problems count

Historian Benjamin Breen used GPT-6 and Opus 5.5 to find a new Newton source. He still chose the archive and judged the proof.

by ·3 min read·

New here? Start with the free AI Survival Kit →

Frontier AI is solving open problems in history, and the historian still decides which problems count

0:00 / 6:00

This voice is generated by AI.

Part of the guide Will AI take your job? How to find out for your role, in an afternoon

Frontier AI is solving open problems in history, and the historian still decides which problems count

Historian Benjamin Breen wrote on his Substack, Res Obscura, on Thursday that frontier models can now help solve open historical problems. Until now their main use was transcribing documents. He wrote it in the week GPT-6 Sol and Opus 5.5 were released. Using OpenAI's GPT-6 Astra, he identified the French alchemical passage that Isaac Newton had freely translated into Latin. Anthropic's Opus 5.5 then turned up strong evidence that Newton drew on a manuscript in the papers of Samuel Hartlib. Breen's conclusion is that historians working in groups alongside current frontier models would produce numerous advances in historical knowledge. He adds that this "was not the case as recently as last year."

What the headline reading misses

The easy reading is that the model did the historian's job. The detail says otherwise. Breen sets out when these models do well. Experts have already named the problems. The data is digitised and accessible. The work suits multilingual reasoning, maths or bespoke code. Most importantly, a solution can be clearly proven or disproven. Breen puts these as questions an expert asks before starting, and he chose the Hartlib papers himself because they passed.

The Hartlib run shows how the work splits. Opus 5.5 downloaded over 5,000 primary source files from Hartlib's archive. It then created sub-agents to search Google Books and other archive sites for Hartlib's unidentified sources across languages. It noticed that Newton and Hartlib each used a coded form for the same ingredient, Hungarian vitriol. Newton's was a true anagram, and Hartlib's was closer to backwards writing. Breen's first reaction was that it "felt like a stretch." The model then shared the marginal annotation that had explained the cipher to 17th-century readers, and matching quantities across the two texts completed the case. He rates the top findings as "the sort of thing I could imagine spending a week of research on."

The misses are just as useful. When Breen set both models on WW1 and WW2 encrypted messages, they came up empty. Breen's reading is that "the low hanging fruit here seems to have been plucked." On John Dee's coded book, Liber Loagaeth, Astra concluded that the text is almost entirely nonsense syllables. Breen's own verdict is that this was not a meaningful breakthrough.

The judgement layer moved up

Two things changed for the expert.

First, the model now proposes the projects. GPT-6 suggested on its own that it trace Darwin's informants through his writings. Breen says the idea matched his professional sense of a worthwhile research project. In the past, he writes, models struck him as more useful for making data visualisations.

Second, the reasoning is harder to check. Frode Weierud, a historical cryptology researcher, writes that his team is still analysing the Astra logs from the break of a July 1941 Enigma message. The archive file references the model cited are correct. They are not available on the Crypto Cellar Research website, and it is not clear where the model found them. Breen calls these models "maniacally determined" on problems they consider tractable.

That is the operator's position in any expert field. AI does the routine work, such as downloading thousands of files, reading across languages and testing anagrams. You do the thinking. You choose the problem and decide what counts as proof. You reject what is a stretch, and you trace the evidence when the model cannot explain how it got there.

Your Next Move

  • Write your field's list of tractable problems. Use Breen's test: problems experts have already named, with accessible data and an answer that can be proven or disproven. That list is where frontier models pay off first, and you are qualified to write it.
  • Make every AI-assisted finding traceable. Require the model to cite its sources. Check any source you cannot trace yourself before the finding goes further, as Weierud's team is doing with the Astra logs.
  • Sort a week of your own work. Separate the tasks a model can now run from the calls that stay with you. The afternoon method for finding out which parts of your role AI is absorbing walks through the sort step by step.

I build systems like the one publishing this site. → Work with me

About the author

Jo

Jo runs The War Room: one signal a day on how AI is changing work, and what to do about it.

Sources

Get the Briefing

Want more intelligence like this?

The weekly briefing, free — and the AI Survival Kit with it.

Takes 30 seconds. No spam.

More in Signals

See all