OPENAI DISPUTES GPT-5.6 SOL BENCHMARK RESULTS
■ AI-SUMMARIZED FROM 5 SOURCES ▸ TIMELINE
OpenAI claims its GPT-5.6 Sol model outperforms Anthropic's Opus 5 on the ARC-AGI-3 benchmark when using OpenAI's latest API, contradicting official test results that showed the model scoring significantly lower.
■ SOURCES
► Bloomberg Tech► Bleeping Computer► Bloomberg Tech► The Decoder► Rest of World■ SUMMARY WRITTEN BY AI FROM THE LINKS ABOVE
■ MORE FROM THE AI DESK
Denmark has implemented a requirement for students to orally defend their written work as a countermeasure against AI-generated assignments. The policy aims to verify authentic student comprehension and authorship.
Anthropic is making Auto Mode the default setting in Claude Code for Pro, Max, and Team plans starting August 14. The company argues the automated safety classifier is more effective at catching dangerous commands than human reviewers.
A study of over 2,500 readers found they cannot distinguish AI-generated short stories from human-written ones. Participants rated the machine-written texts higher—until they learned the truth.
DeepMind has released an open source weather prediction model that produces accurate hurricane forecasts using lower-resolution data, surprising meteorologists with its efficiency gains.