A new benchmark created by 64 mathematicians reveals that advanced AI systems confidently attempt to solve math problems with no solution, exposing a critical gap in reasoning capability. Google's Gemini 3 Pro achieved 30 percent accuracy on research-level problems but no model exceeded 50 percent when identifying deliberately unsolvable tasks.
Researchers unveiled SOOHAK, a math benchmark containing 439 handwritten problems designed to evaluate AI reasoning at scale. The dataset includes 99 intentionally unsolvable tasks, serving as a crucial test for whether AI systems recognize the limits of problems rather than simply generate plausible-sounding answers.
The results highlight a significant weakness: while increased computational power improves performance on solvable problems, it does not enhance models' ability to acknowledge when a problem cannot be solved. This disconnect suggests that scaling alone cannot address fundamental reasoning gaps in current AI systems.
Google's Gemini 3 Pro led the benchmark across research-level problems, but the broader finding remains troubling. No model achieved even 50 percent accuracy on the unsolvability detection task—a baseline that would be expected from systems claiming advanced mathematical reasoning.
The benchmark addresses a gap between flashy individual successes and the sustained, broad research skills required for authentic mathematical problem-solving. AI systems that confidently attempt unsolvable problems pose real risks in applied domains where recognizing problem constraints is essential.
SOOHAK represents an effort to establish more rigorous evaluation standards for mathematical AI. Rather than measuring only success on solvable problems, the benchmark forces systems to demonstrate judgment about problem feasibility—a capability that current models struggle to develop, regardless of their overall performance levels.
The findings suggest that future AI development must address not just computational scale, but fundamental differences in how models approach reasoning tasks and recognize epistemic boundaries.
StemDeck is a new open-source AI stem separator that runs locally on your machine without cloud dependencies. The free tool splits audio into individual instrument tracks.
A firsthand look at China's AI development reveals both countries pursuing remarkably similar technological paths, despite geopolitical tensions. The competition appears less ideological and more focused on matching capabilities.
Anthropic will permanently increase Claude Code's weekly limits by 25% starting September 14, but users will see a 17% net reduction once a current 50% temporary boost expires.
A developer who scraped artwork for AI training is now collaborating with Cara, a creator platform designed to prevent unauthorized AI data collection, as the service faces ongoing attacks from trolls attempting to breach and publish its data.