How Near-Memory Compute Changes Data Flow
At the Hot Chips 2026 conference in August, XCENA and Samsung jointly presented a new compute express link device designed to bring processing closer to memory. The collaboration aims to address growing memory demands from machine learning workloads by integrating compute capabilities directly into CXL memory expansion modules. The prototype was demonstrated live during the session, drawing attention from industry engineers and researchers focused on memory-centric architectures.
Breaking news
Eufy Unveils Local AI Home Security Ecosystem at IFA
The Rapid Evolution of Data Center Security in the AI Era
The High-Voltage Risks Facing Modern AI Data Centers
Apple’s New CEO Renames Lake Ontario To Lake America In Maps AppThe device leverages Samsung’s advanced memory technology and XCENA’s expertise in near-memory computing to reduce data movement between processor and memory. By placing compute units on the CXL module itself, the system minimizes latency and bandwidth bottlenecks common in traditional server designs. This approach allows frequently accessed data to be processed in place, improving efficiency for memory-intensive applications such as large language models and graph analytics. Early benchmarks showed a 30% reduction in data transfer volume for certain inference tasks compared to conventional setups.
Can This Approach Scale Beyond Prototypes?
Traditional architectures rely on moving data from memory to CPU cores for processing, which consumes significant time and energy. The XCENA-Samsung device reverses this pattern by enabling lightweight compute functions—such as tensor operations or filtering—directly on the memory module. This shift reduces reliance on complex cache hierarchies and simplifies data routing in heterogeneous systems. Engineers noted that the design supports standard CXL 3.0 protocols, ensuring compatibility with existing server platforms while extending functionality.
While the current implementation is a proof of concept, both companies emphasized its potential for future scalability in data centers. XCENA highlighted plans to refine the compute element architecture for broader instruction support, while Samsung is exploring integration with its next-generation HBM and DDR5 modules. Challenges remain in thermal management and software tooling, but early feedback suggests strong interest from cloud infrastructure providers seeking to optimize AI workloads. The teams are now working on refining the firmware stack and validating performance under real-world conditions.
What is the primary benefit of placing compute on a CXL memory module? It reduces the need to move large amounts of data between the processor and memory, lowering latency and energy use for repetitive or localized operations.
Frequently Asked Questions
Is the device compatible with current servers? Yes, it uses standard CXL 3.0 interfaces, allowing it to function as a memory expander in existing systems while adding compute capabilities.
What types of workloads benefit most from this technology? Memory-intensive tasks such as AI inference, graph processing, and database operations see the greatest gains due to reduced data movement and localized compute.


