The project resulted in 13 million lines of code and the proof of nearly 30,000 intermediate theorems.
This milestone is significant because it demonstrates that AI can handle "autoformalization," the process of converting complex human reasoning into rigorous, machine-checked data.
Traditionally, verifying a groundbreaking mathematical proof is a grueling manual process that can take experts years to complete.
By automating this verification, AI infrastructure could drastically reduce the time needed to peer-review new discoveries and help identify errors in existing mathematical literature, ensuring a more reliable foundation for future research.
The success of the project relied on a multi-agent system where dozens of Claude instances collaborated through a platform called Prove2Me, which managed the complex logical steps and prevented the AI from losing track of the overall goal.
While the effort required roughly six billion output tokens from a high-end research model, Anthropic suggests that similar collaborative formalization is becoming possible even with consumer-level AI subscriptions.
This technology is expected to become a standard tool for mathematicians, allowing them to provide a "computer-checked" seal of approval alongside traditional written research.