The Physical Limits of Current Memory Scaling
Building on our previously published insights into the hidden bottleneck behind AI scaling, the urgency surrounding AI memory architectures has only intensified. Modern compute infrastructure is increasingly defined not by raw arithmetic capacity, but by memory bandwidth and data movement constraints. As AI workloads transition from basic interactive queries to continuous, multi-step agentic planning, memory bandwidth and data movement latency have become primary design bottlenecks.
This three-part blog series examines the physical limits of traditional memory scaling, evaluates current industry alternatives, and presents monolithic vertical integration as the path forward for next-generation AI accelerators.
Clustering or stacking additional HBM dies provides essential bandwidth gains, but it increases thermal density, packaging complexity, and manufacturing costs. Overcoming these physical trade-offs requires exploring alternative memory architectures, including direct Back-End-of-Line (BEOL) integration above compute logic.
The Inference Explosion & The Transit Tax
The industry has moved past the theoretical phase of the « inference explosion. » Capital and engineering resources have decidedly shifted from model training to real-time inference deployment. This shift is highlighted by massive recent investments in purpose-built inference hardware, such as Cerebras securing $1 billion in 2025, d-Matrix securing $275 million for hyperscale chiplet deployments in 2025, and UK-based startup Olyx raising over $200 million in 2026.
What is the common point between them? They introduced new architectures with massive SRAM usage or In-Memory computing to reduce data traffic, showing that AI inference is increasingly becoming a memory-bandwidth-bound problem, where efficient data movement is the primary performance driver rather than raw FLOPs.
This shift in focus is further shown by recent industry perspectives, such as those from Micron and Qualcomm, which identify the shift toward Agentic AI as a catalyst transforming the traditional memory wall into a critical, system-level crisis that legacy hierarchies struggle to contain.
Fetching data from off-chip DRAM consumes up to three orders of magnitude more energy than the computation itself
In layman’s terms, today’s AI workloads are data movement problems with computation embedded inside. When customers pay API bills for AI inference, they assume they are paying for computation (FLOPs). In reality, arithmetic is cheap; they are paying a massive « transit tax. » Fetching data from off-chip DRAM consumes up to three orders of magnitude more energy than the computation itself.
To illustrate the scale of this cost: every single word an AI generates requires pulling the model’s entire memory across the chip. When hundreds of users query the AI at once, the system doesn’t run out of mathematical horsepower, it runs out of highway lane space to move that data. Data traffic becomes the ultimate speed limit of the system.

Model prefill sits in the high-intensity compute-bound plateau, whereas real-time token decode sits squarely on the steep memory-bandwidth-bound slope. Increasing raw compute cores provides zero speedup for autoregressive decode without a proportional increase in memory bandwidth.
The Context Tax & KV Caching
This dynamic is aggressively amplified by the Context Tax. As context windows for agentic AI expand into the millions of tokens, the memory required to hold the Key-Value (KV) cache grows exponentially. Keeping these massive context windows « warm » in scarce, off-chip memory drains system efficiency and forces cloud providers to charge high residency fees. Inference economics now universally comes down to power per token and latency per token.
As illustrated in Figure 2, auto-regressive token generation recalculates attention across all previous tokens at each step when un-cached. By caching the Key and Value states from earlier iterations, the model only computes attention for the newest token, significantly shrinking matrix dimensions and drastically speeding up the decode phase.

(Prefill vs. Decode with KV Caching)
As depicted in Figure 3, the prefill phase achieves high compute utilization (>80%) by processing input tokens in parallel across compute engines. Conversely, during the decode phase, sequential token generation drops compute utilization below 20% due to the memory streaming bottleneck pipeline, forcing arithmetic registers to wait on weight transit across memory links.

Illustrating why agentic loops, generating hundreds of sequential tokens, starve high-power arithmetic cores while waiting on memory transit.
To put simply, modern AI infrastructure faces a profound memory bottleneck where real-time inference efficiency is dictated by data movement rather than arithmetic capacity. The heavy transit tax of moving model weights across long interconnects, coupled with the exponential context tax of managing multi-million-token KV caches, creates a severe operational wall for scaling next-generation agentic AI.
What’s Next
In Part 2, we will explore how the industry is currently attempting to bypass these physical constraints through custom chiplets, wafer-scale engines, or near/ in-memory compute architectures, and analyze the fundamental trade-offs each approach entails.
In Part 3, we will then introduce Vertical Compute’s VIM™ architecture, a breakthrough solution that bypasses these structural trade-offs, through true 3D Back-End-of-Line (BEOL) monolithic integration.
Follow us on Linkedin for latest news and content updates