OPENAI WITHDRAWS ENDORSEMENT OF FLAWED AI CODING TEST
■ AI-SUMMARIZED FROM 2 SOURCES ▸ TIMELINE
OpenAI discovered that approximately 30 percent of tasks in SWE-Bench Pro, a widely used benchmark for measuring AI programming capabilities, are broken. The company has withdrawn its earlier endorsement of the test.
■ MORE FROM THE AI DESK
The Electronic Frontier Foundation has urged courts to resist pressure to fundamentally alter copyright law in response to artificial intelligence developments. The organization argues that existing legal frameworks are sufficient to address AI-related concerns.
Anthropic has unveiled Fable 5.1, a new upgrade to its Mythos-class Claude model that reduces operational costs by 25% for standard workloads and up to 45% for complex agentic tasks.
The Fermi Explorer Mission, a nonprofit organization, plans to launch a spacecraft to Alpha Centauri—the Sun's nearest star system—by the end of 2029. The mission will use a trajectory discovered by an AI system developed by PSI.
World Labs has released Atlas, a world model designed to understand and reason about spatial environments. The system represents a step toward AI systems capable of comprehending 3D space and physical interactions.