GOOGLE'S GEMINI-SQL2 DOMINATES TEXT-TO-SQL BENCHMARKS
■ AI-SUMMARIZED FROM 3 SOURCES ▸ TIMELINE
Google Research's Gemini-SQL2 achieves 80.04% accuracy on the BIRD benchmark, significantly outperforming competitors from OpenAI and Anthropic. The system converts natural language queries into executable SQL code.
■ MORE FROM THE AI DESK
Denmark has implemented a requirement for students to orally defend their written work as a countermeasure against AI-generated assignments. The policy aims to verify authentic student comprehension and authorship.
Anthropic is making Auto Mode the default setting in Claude Code for Pro, Max, and Team plans starting August 14. The company argues the automated safety classifier is more effective at catching dangerous commands than human reviewers.
A study of over 2,500 readers found they cannot distinguish AI-generated short stories from human-written ones. Participants rated the machine-written texts higher—until they learned the truth.
DeepMind has released an open source weather prediction model that produces accurate hurricane forecasts using lower-resolution data, surprising meteorologists with its efficiency gains.