TechBriefe
Ai

CHIPSMORE: New Chiplet Design Boosts LLM Inference with Compute-in-Interconnect Technology

Technical Paper Link 19.09.2026

CHIPSMORE reverses this by embedding arithmetic logic within the interconnect fabric itself

Researchers at the National University of Singapore have introduced CHIPSMORE, a novel hardware architecture designed to accelerate large language model inference by integrating computation directly into interconnect and memory chiplets. The approach targets multi-request and multi-mode workloads, aiming to reduce latency and energy consumption during AI processing. The design was detailed in a recent technical paper presented at a computer architecture conference. The CHIPSMORE system uses heterogeneous chiplets that combine logic, memory, and interconnect layers to perform compute-in-memory and compute-in-interconnect operations. This allows data to be processed closer to where it is stored, minimizing data movement—a major bottleneck in traditional AI accelerators. By supporting multiple inference modes and handling concurrent requests, the architecture improves throughput for real-time LLM applications such as chatbots and code generation tools. How Compute-in-Interconnect Reduces Data Movement Traditional systems shuttle data between separate compute and memory units, consuming significant time and power.

CHIPSMORE reverses this by embedding arithmetic logic within the interconnect fabric itself, enabling operations as data flows between chiplets. This reduces the need for frequent off-chip memory accesses, which are slow and energy-intensive. Experimental results show the design cuts inference latency by up to 40% compared to baseline systems under multi-request scenarios. Can This Approach Scale to Larger Models? Scalability remains a key focus for the research team, who note that CHIPSMORE’s modular chiplet design allows incremental expansion by adding more compute or memory units. The architecture supports various model sizes through configurable partitioning, though challenges remain in balancing thermal density and manufacturing yield. Researchers suggest future work will explore 3D stacking and advanced packaging to further enhance integration. Frequently Asked Questions What makes CHIPSMORE different from other AI accelerators?

Unlike conventional designs that separate compute and memory, CHIPSMORE performs computation within the interconnect and memory layers, reducing data movement and improving efficiency for multi-request LLM inference. Is CHIPSMORE intended for data centers or edge devices? The architecture is primarily targeted at data center environments where high-throughput, low-latency LLM serving is critical, though its energy efficiency could benefit edge applications in later iterations. Does the paper include real-world performance metrics? Yes, the researchers provide simulation-based evaluations showing up to 40% latency reduction and improved throughput under multi-request workloads compared to traditional accelerator baselines.

Share:

More stories: