Skip to main content
Back to News
news/AI Research

OpenAI Releases 722 Math Proofs From One Agent Prompt

OpenAI published 722 manuscripts from an unreleased frontier model, nearly all generated by a single prompt to a single agent, reigniting the math debate.

Stefan Trbojevic

Stefan Trbojevic

7 October 20262 min read
LinkedIn
Abstract network of glowing nodes radiating from a bright central hub in deep navy and cyan

The takeaway

Single-agent prompting, not multi-agent orchestration, produced OpenAI's latest math haul, and formal verification is becoming the release format of record.

Why it matters for builders

A single agent with a large thinking budget produced hundreds of math results, so the premium moves from multi-agent orchestration to context, verification and reproducibility. Lean proofs and machine-checkable artifacts are now the release format frontier claims arrive in.

OpenAI Releases 722 Math Proofs From One Agent Prompt

OpenAI has published 722 mathematical manuscripts organized into 372 result families, all produced by an internal frontier model the company has still not released. According to a spokesperson, nearly every result came from a single prompt handed to a single AI agent, a sharp contrast with the large agent swarms behind the lab's earlier mathematics breakthroughs.

What happened

The collection now sits in a public GitHub repository, openai/math, and includes papers, supporting proof artifacts, Lean formalizations for many results, and ten abridged summaries of the model's reasoning.

OpenAI says it evaluated roughly 4,000 problems and that the average published result consumed compute equivalent to about three hours of ChatGPT Pro thinking, according to The Verge. The repository warns explicitly that the collection mixes results at different stages of verification and that some unformalized results "could have issues."

The release was expected. In September, OpenAI said the same internal model had resolved more than 100 long-standing open problems, following its Navier-Stokes announcement. This batch replaces hints with documents, complete with revision and citation protocols.

A broad field of glowing data nodes narrowing into a single bright processing node

Why it matters for builders

The headline number is 372 families, but the operational detail is more interesting: a single agent, driven by a single prompt, produced almost all of them. That undercuts the assumption that frontier research work requires sprawling multi-agent orchestration. If one agent with enough thinking budget can work through thousands of open problems, the engineering premium shifts from orchestration to context management, verification, and reproducibility.

Verification is where the release is thinnest. Lean formalizations cover many but not all manuscripts, and OpenAI itself flags that unverified results may contain errors. AGMAI, the independent Advisory Group on Mathematics and AI that OpenAI set up in September, published recommendations urging labs to release results promptly, disclose model names, prompts and compute costs, and stop treating math releases as marketing vehicles. The group is now the de facto referee for how these drops get judged.

Not everyone is satisfied. WIRED reports that mathematicians describe a "perception of mobster behavior" from leading AI companies, and that OpenAI had assured attendees at an August gathering it would not publish everything at once.

What comes next

The mathematical community will need months to assess the papers. For builders, the practical signal is the release format rather than the result count: proof artifacts and machine-checkable certificates are becoming the default way frontier research claims are shipped, and any agent stack expected to work on hard problems will need to treat formal verification as a first-class step rather than an afterthought.

Share𝕏

The Automation Brief

Read 5 AI stories instead of 50.

The essential moves in AI agents, models, automation and infrastructure — filtered for builders and operators, with the part that actually matters.

No noise. Unsubscribe anytime.

Editorial notes

Reported by

Stefan Trbojevic

Edited by

n8n Lab Editorial

Published

7 October 2026

Updated

7 October 2026

AI disclosure: AI assisted with research and drafting. Factual claims are reviewed by an editor.

n8n Lab is an independent service provider. We are not affiliated with, endorsed by, or sponsored by n8n GmbH. “n8n” is a trademark of n8n GmbH and is used here only to describe the platform-specific implementation and automation services we provide.