Nvidia is integrating d-Matrix's inference-focused XPUs with its GPUs via NVLink Fusion, marking the chipmaker's embrace of specialized AI hardware as demand shifts from training to inference workloads.
d-Matrix, an AI chip startup, will collaborate with Nvidia to combine its XPU processors with Nvidia's GPU architecture. The partnership leverages NVLink Fusion interconnect technology to enable seamless communication between the two systems.
Inference—the process of running trained AI models for predictions—has emerged as a critical compute bottleneck. Unlike training, which demands raw processing power, inference requires optimized systems for speed, efficiency, and cost-effectiveness at scale.
Nvidia's move signals recognition that no single chip architecture dominates inference workloads. By integrating specialized inference processors alongside its GPUs, the company addresses diverse customer needs across data centers.
d-Matrix CEO Sid Sheth highlighted inference as the next major battleground in AI compute, where customized silicon gains competitive advantage over generalized solutions.
This partnership reflects broader industry trends: major AI players increasingly develop custom chips for specific workloads, while established chipmakers partner with startups to remain competitive in rapidly fragmenting markets.
Google has released a native Gemini application for Windows PCs, bringing feature parity with the existing macOS version. The app includes a keyboard shortcut for quick access.
Nvidia co-founder and CEO Jensen Huang identified cybersecurity as the next major market for artificial intelligence, predicting the technology will fundamentally reshape how computer systems are defended.
Cognition has released SWE-2, a new AI model designed to compete with Anthropic's Claude 5.1 and OpenAI's GPT-Astra. The model targets software engineering tasks and represents Cognition's latest push in the generative AI space.