OPENAI DISPUTES GPT-5.6 SOL BENCHMARK RESULTS
OpenAI claims its GPT-5.6 Sol model outperforms Anthropic's Opus 5 on the ARC-AGI-3 benchmark when using OpenAI's latest API, contradicting official test results that showed the model scoring significantly lower.
Moonshot's Kimi K3 has reignited debate over China's role in the AI race, but Beringea Chief Investment Officer Karen Mc…
OpenAI counters Anthropic's ARC-AGI-3 record: GPT-5.6 Sol scores 38.3 percent, but only with its own API features instea…
Governments can now deploy top-tier AI locally, bypassing costly U.S. cloud rentals.
OpenAI: OpenAI says using its Responses API harness with GPT-5.6 Sol tripled its ARC-AGI-3 score with fewer output token…
Article URL: https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/ Comments URL: https://news.…
Newley Purnell / Bloomberg: DeepSeek rolls out the official V4 Flash API in public beta, touting enhanced agent capabili…
Three Claude models attacked real companies during cybersecurity tests after a misconfiguration gave them internet acces…
Artificial Analysis: DeepSeek V4 Flash scores 50 on the Artificial Analysis Intelligence Index, matching Gemini 3.6 Flas…
OpenAI: OpenAI disputes Apple's account of pre-lawsuit contacts, rejects claims that ex-staff used confidential informat…
A unit of OpenAI has reached a settlement with the US Justice Department to resolve allegations that the California-base…
OpenAI and Anthropic have confirmed that their AI models were involved in separate, newly disclosed third-party cybersec…