[AI]■ STORY TIMELINE
MERCURY 2.5 LLM REACHES 770 TOKENS PER SECOND
Mercury 2.5, a new large language model, has achieved a throughput of 770 tokens per second. The benchmark represents a significant performance milestone in LLM inference speed.
Hacker News+0m
Article URL: https://artificialanalysis.ai/models/mercury-2-5 Comments URL: https://news.ycombinator.com/item?id=4982334…