NewsAI research
AI checked Fermat's Last Theorem in 11 days. The labs want to slow down.
In the same week, one AI lab showed its models finishing expert work in days, and another's chief scientist argued that nobody has yet made that safe to scale at full speed.
What Claude did
On 4 September, Anthropic published research describing how Claude produced a complete, computer-checked proof of Fermat's Last Theorem in Lean, a language whose checker accepts a proof only if every step is valid. The run took 11 days of wall-clock time, wrote about 13 million lines of Lean and proved some 30,300 theorems along the way.
It did not start from nothing. The work built on Mathlib, the community's shared library, on the Imperial College London project led by Kevin Buzzard, and on Prove2Me, a tool designed at Columbia University. Buzzard reviewed the result and confirmed that it proves the theorem from the axioms of mathematics alone. Anthropic is open that the proof is far longer than it needs to be.
What OpenAI's chief scientist wrote
Two days later, OpenAI published an essay by its chief scientist, Jakub Pachocki, arguing that no lab has solved alignment and monitoring well enough to keep scaling at maximum speed. He called for safety thresholds enforced by outside auditors or public bodies, and said he expects voluntary slowdowns across the industry until such standards exist.
Why the two belong together
The first story is about capability: AI agents working for days on a hard problem, with a strict checker at the end. The second is about control: whether the people building these systems can still see what is happening inside them. Both are true at once.
For a business, the useful lesson sits between them. AI is most trustworthy where its output can be checked (a proof checker, a booking calendar, a stock count) and least trustworthy where nobody verifies the result. Put it to work in the first kind of place, and keep a person in the second.








