Researchers have demonstrated Maple-Preview, a 20 billion parameter mixture-of-experts model, executing at 120 tokens per second on an iPhone. The achievement signals substantial progress in running large language models on mobile devices.
Maple-Preview represents a significant advancement in on-device AI inference. The ternary quantized model achieves competitive performance while maintaining practical speed on consumer smartphones, eliminating the need for cloud processing for certain tasks.
The 120 tokens-per-second throughput enables real-time interaction with a model that would traditionally require dedicated server infrastructure. This performance level makes the system viable for privacy-critical applications where keeping data local matters.
The mixture-of-experts architecture allows selective activation of model parameters, reducing computational overhead compared to dense models of equivalent size. Ternary quantization further compresses the model by limiting weights to three discrete values.
The project, shared on Hacker News, generated 120 points and 34 comments, indicating strong community interest in mobile AI capabilities. The development reflects growing momentum toward practical edge deployment of large language models.
CopilotKit released Channels SDK, a new tool that enables AI agents to deploy across multiple messaging platforms. The open-source SDK connects agents to Slack, Microsoft Teams, and other channels without rebuilding core logic.
The core team managing Nixpkgs, the package repository for the NixOS Linux distribution, has officially disbanded. The announcement was made on the NixOS Discourse forum.
Oracle has implemented a policy prohibiting AI-generated code contributions to OpenJDK, the open-source Java platform. The ban applies to code created by large language models and similar AI systems.
A website operator discovered that nearly all of their traffic came from automated bots rather than human visitors, raising questions about how internet metrics are measured and reported.