You are currently viewing Leading AI Models Struggle With Original Math Problems, Study Finds

Leading AI Models Struggle With Original Math Problems, Study Finds

Artificial intelligence is becoming a common tool in modern mathematics, helping researchers scan literature, verify proofs, and check manuscripts. But when it comes to solving original math problems, even the most advanced AI systems are still falling short.

That’s the conclusion of a new study published on the arXiv preprint server, where a group of mathematicians set out to test how well leading AI models perform on genuine, unpublished research questions — problems that have never appeared online or in textbooks.

Unlike earlier evaluations that relied on contest-style questions or well-known exercises, this study focused entirely on original math problems drawn directly from the researchers’ own work, eliminating the possibility that AI models could rely on memorized solutions from training data.

How the AI Was Tested

Each mathematician involved in the study submitted one original problem and solved it independently to confirm it was tractable. To prevent information leakage, the researchers encrypted the solutions so they would not appear in public sources accessible to AI systems.

The final test included ten original math problems spanning diverse areas such as stochastic analysis, spectral graph theory, symplectic geometry, and algebraic topology. Leading AI systems, including GPT-5.1 Pro and Gemini 3 Pro, were given just one attempt per question, with no follow-up prompts, clarifications, or hints.

The experiment, called First Proof, focused on a specific stage of mathematical research. As the authors explained, the goal was to evaluate AI performance at “the final and most well-specified stage of math research,” where the framework is understood but creative reasoning is still required.

Why AI Still Falls Short

The results suggest that fears of AI replacing mathematicians are premature. While AI models excel at summarizing existing knowledge and handling structured, contest-like tasks, they struggled to solve original math problems in a single attempt.

According to the researchers, the models lacked the intuition and creative reasoning required to navigate unfamiliar mathematical territory. The study concludes that current AI systems remain far better at pattern recognition than at producing genuinely new mathematical insights.

What Comes Next

The research team plans to release the encrypted solutions on February 13 and is already preparing a second set of original math problems. Their long-term goal is to turn First Proof into a permanent benchmark for evaluating AI’s ability to tackle real-world mathematical research.

As the authors noted, “We hope to use this understanding to design a more formal benchmark,” one that continues to test the boundaries of what AI can — and cannot — do in advanced mathematics.

Goodle Preferred Source

Don’t miss out on our latest news—follow us for the latest AI newsbreakthroughs, and insights that matter.

Leave a Reply