# Vera Rubin’s DRAM Cut in Half: Is the Memory Cycle Really Over?

### What disappeared, demand or the boundary between HBM, SOCAMM, and storage?

On June 4, SemiAnalysis reported that NVIDIA is lowering the standard configuration of SOCAMM, the memory module attached to the CPU side of Vera Rubin NVL72.

The standard build, which had been expected to center on 192GB modules, shifts to 96GB modules, and CPU-side memory per rack drops from roughly 54TB to 28TB, cut in half. The next day, memory and storage stocks fell across the board, Micron and SanDisk among them.

The questions followed immediately.

*If CPU orchestration is supposed to matter more, why cut CPU-side DRAM in half first?*

*HBM4 capacity is not rising much over the previous generation either, so isn’t this a sign that memory demand is weakening?*

The last three pieces each took on one slice of this question.

The starting point of the first piece was that the bottleneck in AI inference is moving outside the GPU, toward the CPU and system memory.

**Vera’s DRAM cut is the event where those three slices meet at a single point.**

This piece starts with how far that 28TB number should be trusted. From there it follows whether the half of demand that appears to have vanished has actually vanished, whether HBM, SOCAMM, and storage were ever things that belonged together for the same reason, and what the reduced DRAM is really pointing at.

**What the market priced in and what the SemiAnalysis report actually said were not the same thing.**

## 1. The First Misreading the DRAM Cut Created

NVIDIA has announced nothing.

The [official reference spec](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/) still lists up to 54TB, and what SemiAnalysis reported is a supply chain account that the actual shipping standard build is dropping toward 96GB. Eight SOCAMM modules attach to each Vera CPU, and a rack holds 36 Vera CPUs, so 192GB modules line up with the official spec’s 54TB, while 96GB modules come to roughly 28TB.

The background the report gave is supply and cost.

With the supply of LPDDR5X chips going into SOCAMM modules not keeping up enough to fill the full standard build, the move is to lower the standard configuration to 96GB and get racks installed first rather than delaying rack shipments over memory, leaving the high-capacity build as a custom option.

By the report’s figures, this adjustment lowers the cost of a single rack from about $7.6M to $6.8M, and TCO per GPU comes down with it. So 54TB is the number that says how far it can go, and 28TB is the report of how it actually gets installed within that constraint.

**Of the AI inference state, which part must stay inside HBM4, and which part can move outside it?**

## 2. The Fastest Lane Did Not Weaken

What shrank is the CPU-side memory, not HBM4.

GB300 NVL72 put forward 20.7TB of HBM3e and 576TB/s of GPU memory bandwidth.

Vera Rubin NVL72 puts forward 20.7TB of HBM4 and 1,580TB/s of GPU memory bandwidth.

Capacity did not rise much, but bandwidth gets far stronger. HBM4’s change is in the speed of the corridor rather than the size of the warehouse.

**If HBM4 is that important, why does the memory outside HBM become more important?**

## 3. A Strong Hot Lane Creates a Buyer Problem

That HBM4 is fast and strong is not purely good news for the platform owner. It is an expensive lane with tight supply, so you cannot put everything on it. That much more, you have to choose harder what to keep on HBM4 and what to push outside it.

Deciding this placement is the orchestration the CPU handles. But saying the CPU matters more does not mean CPU-side DRAM has to get bigger. Its importance lies in the power to control what goes where, rather than in capacity.

Think it through again from the platform owner’s position with this view, and Vera’s DRAM cut is not a strange event. It is the platform owner concentrating the expensive HBM4 lane on hot state and testing how far the less hot state can be sent outside.

So reading 96GB straight as a memory demand collapse is premature, and so is reading it as a trivial launch optimization.

Which one it is can only be known by watching how many systems actually get installed in what configuration.
