Mercury 2.5, a new large language model, has achieved a throughput of 770 tokens per second. The benchmark represents a significant performance milestone in LLM inference speed.
Mercury 2.5 has demonstrated token generation speeds of 770 tokens per second, according to benchmarks tracked on Artificial Analysis. The metric measures how quickly the model can produce output, a key performance indicator for real-time applications.
Token-per-second throughput directly impacts latency for end users and cost efficiency for service providers. Higher speeds enable faster response times and reduced computational overhead per request.
The achievement positions Mercury 2.5 among competitive options in the current LLM landscape, where inference speed has become an increasingly important differentiator alongside model quality and reasoning capabilities.
The milestone has drawn attention from the developer community, generating 61 comments on Hacker News with a score of 101 points, suggesting interest in performance-optimized language models.
Details on Mercury 2.5's architecture, training methodology, and availability are available through Artificial Analysis's model database.
McDonald's has launched Archy, an AI voice ordering system operating in English and Spanish with over 90% accuracy. The technology aims to streamline drive-thru operations across its restaurant locations.
Brahma AI, the artificial intelligence division of VFX company DNEG, secured $150 million in funding led by Indian PE firm Multiples. The round values the company at $2 billion post-money.
Google is nearing the launch of Gemini 4, its next flagship AI model, with new DeepMind leader Koray Kavukcuoglu indicating a release much earlier than year-end. The move marks Google's effort to catch up with rivals in rapid AI model deployments.
Former UK Deputy Prime Minister Nick Clegg, now in the tech industry, has criticized AI doomsday predictions as overblown, saying tech leaders should focus on concrete threats instead.