25 Fields Medalists Warn AI Firms Are ‘Destroying’ Mathematics by Misusing Problems as Benchmarks
Twenty-five Fields Medal winners have signed a joint statement accusing leading AI companies, including OpenAI, of abusing unsolved mathematical problems as performance benchmarks — a practice they say risks corrupting both the integrity of mathematical research and the public’s understanding of machine intelligence.
The Core Dispute
At issue is the growing tendency of AI labs to showcase model capabilities by pointing to results on famous open problems — from the Riemann Hypothesis to longstanding combinatorial conjectures. The signatories argue that when AI systems produce partial or heuristic outputs on such problems and these are framed as breakthroughs, it creates a distorted feedback loop: benchmarks get optimized against, funding flows toward headline-grabbing demos, and genuine mathematical rigor is sidelined.
A Fields Medalist who responded exclusively to the controversy described the situation as a category error. Mathematics, this person argued, is not a leaderboard. Problems are open because they encode deep structural questions, not because no one has thrown enough compute at them. Treating them as benchmark targets conflates the map with the territory — and misleads policymakers, investors, and the public about what AI can actually do.
Why This Matters for the AI-Crypto Nexus
The dispute lands squarely in the middle of the crypto-AI convergence. Decentralized compute networks, AI agent protocols, and model-tokenization platforms increasingly market themselves on “benchmark performance” — often citing math and reasoning scores as proof of utility. If the academic community formally rejects the validity of these benchmarks, the marketing foundation for a substantial slice of the AI-crypto sector weakens.
- Decentralized compute networks that advertise training or inference on frontier models may face harder scrutiny over how they validate capability claims.
- AI agent tokens whose value propositions rest on “reasoning benchmarks” could see credibility erosion.
- Data marketplaces selling mathematical or scientific datasets may need to rethink how they certify quality.
Forward-Looking Perspective
The statement is unlikely to stop AI labs from using math problems as evaluation targets — benchmarks are too useful commercially. But it may accelerate a split: rigorous, peer-reviewed evaluation standards on one side, and marketing-driven benchmark culture on the other. For crypto-AI projects, the smart move is to align with the former. Verifiable computation, transparent evaluation methodologies, and on-chain attestation of model outputs could become genuine differentiators — not because regulators demand it, but because the academic backlash makes unverifiable claims increasingly costly.
The deeper question the Fields Medalists raise is philosophical: if AI’s value is measured by problems it cannot fully solve, then the industry’s obsession with benchmark saturation may be measuring the wrong thing entirely.




