Skip to main content

SMART KV Cache Technology - Brochure

Page 1

Accelerating AI Inference with KV Cache and CXL® Memory Designing a Cost-effective KV Cache using Penguin Solutions’ SMART CXL Memory.

Breaking the AI memory wall: Scalable memory architecture for next-generation AI infrastructure As enterprises deploy large language models (LLMs), copilots, and generative AI applications at scale, the focus of AI infrastructure is shifting from model training to high-performance inference.

Key Benefits

Eliminate GPU memory bottlenecks

While GPUs provide enormous compute capability, inference performance is increasingly limited by memory capacity and bandwidth. Architects rely on Key-Value (KV) cache, a memory optimization that stores precomputed keys (K) and values (V) and avoids recomputation of new tokens. Increasingly KV caches are being designed with CXL memory which is an optimal solution to cost effectively increase memory capacity and bandwidth using PCIe slots. A major driver of this demand is Key-Value (KV) cache, a memory structure used by transformer models to accelerate token generation during inference. As prompts grow longer, conversations extend across multiple interactions, and thousands of users simultaneously interact with models, memory limitations create a bottleneck, known as the AI memory wall, which prevents GPUs from operating at full efficiency. In many AI inference environments, GPUs spend a significant portion of time waiting for memory rather than performing computation. As models grow and workloads scale, memory architecture becomes a critical factor in AI performance and infrastructure efficiency. By combining Compute Express Link® (CXL®) memory expansion technology, Penguin Solutions’ SMART CXL Memory Add-in-Cards (AICs), and Penguin Solutions MemoryAI™ KV Cache Server, organizations can build scalable AI infrastructure capable of supporting modern inference workloads. © 2026 Penguin Solutions KV Cache Technology Brochure - 5.26

Enable longer context windows for AI models

Support greater inference concurrency

Improve GPU utilization and efficiency

Scale memory capacity independently from compute

Featured Solutions

SMART 1TB CXL Memory Add-in-Cards (AICs)

Penguin Solutions MemoryAI™ KV Cache Server


Turn static files into dynamic content formats.

Create a flipbook
SMART KV Cache Technology - Brochure by Penguin Solutions - Issuu