Cerebras announced the CS-4 this morning: three WSE-3 Turbo wafer-scale engines in one rack, 750 PFLOPS of AI compute, 129.6 petabytes per second of memory bandwidth, 2-microsecond wafer-to-wafer latency, support for models "over 50 trillion parameters," and first shipments this quarter. The company claims up to 30× faster per-user inference than GPU systems and up to 10× the throughput per watt of the CS-3.
Everything in that paragraph is the vendor's own number, from the vendor's own launch. Treat the benchmark ratios accordingly — 30× and 10× are marketing comparatives with no stated baseline model, batch size or competing configuration, and they do no work below. The architectural facts are different in kind: a wafer-scale engine gets its bandwidth from SRAM etched on the wafer, not from stacked DRAM sitting beside the die. A CS-4 rack has no high-bandwidth memory in it. That is a design choice Cerebras has made for four product generations, and it is checkable from the spec sheet rather than from a benchmark.
That makes this the piece of hardware a different argument has been waiting for. The claim circulating this week — relayed second-hand from a broadcast remark by Cathie Wood, and we have not seen the primary clip — is that inference chips are beginning to engineer out the need for high-bandwidth memory, which is why memory stocks are worth avoiding. Every version of that argument we have seen is a forecast about silicon that does not exist yet. The CS-4 is silicon that does, and it is shipping this quarter.
So take the argument seriously and run it. Two numbers decide it, and they are the chart and table above: how much memory a machine with no HBM actually needs, and how much of it Cerebras would have to sell before Micron noticed. Both are computable. Neither favours the claim.
132 gigabytes, against 50 trillion parameters
The first chart above is the whole finding.
A WSE-3 carries 44 GB of on-wafer SRAM — Cerebras' published figure, and the launch describes WSE-3 Turbo as the same silicon clocked roughly twice as fast rather than a new part, so we take the 44 GB as unchanged. Three engines in a CS-4 rack gives 132 GB of memory on the wafers.
Now the model. Cerebras advertises support for models "over 50 trillion parameters." Parameters have to be stored as bytes, and the launch does not say at what precision, so here is all three:
| Precision | Bytes per parameter | Weights of a 50T model | Against 132 GB on wafer |
|---|---|---|---|
| 4-bit | 0.5 | 25,000 GB | 189× |
| FP8 | 1 | 50,000 GB | 379× |
| FP16 | 2 | 100,000 GB | 758× |
At the most generous precision anyone runs in production, the wafers hold 0.53% of the model. At FP8, 0.26%.
This is not a gotcha about the product. Cerebras has never claimed the weights sit on the wafer at that scale — its architecture streams them from an external weight store, and that is the point of the design. It is a gotcha about the phrase. "Engineered out the need for high-bandwidth memory" is true and "engineered out the need for memory" is not, and the second is what the investment argument requires. Something is holding 25 to 100 terabytes of weights beside every rack, and at that capacity and that price the something is commodity DRAM.
Which is a Micron product line. Just not the one everybody was watching.
What 129.6 PB/s is, and what it is not
The headline bandwidth figure is 43.2 PB/s per wafer, three wafers, 129.6 PB/s for the rack. It is roughly three orders of magnitude above what a single HBM-fed accelerator delivers, and it is not comparable to an HBM figure without saying why.
Divide it by the memory it addresses: 129.6 PB/s across 132 GB means the entire on-wafer memory can be read about a million times a second. That is what SRAM adjacency buys, and it is genuinely a different regime from stacked DRAM. It is also a statement about 132 GB. An HBM bandwidth number describes a much larger and much slower pool, which is why the two figures cannot be ranked against each other — they are bandwidth over different amounts of stuff.
The honest version of the architectural claim is narrow and still interesting: for the fraction of a model that fits on wafer, wafer-scale removes the memory wall entirely, and it does so without HBM. Everything outside that fraction is a memory-hierarchy problem exactly like everyone else's.
The threshold: what Cerebras would have to sell
Now size it. This is where most analysis quietly invents a number — the dollar content of HBM in a GPU rack, which no vendor discloses. We are not going to. We will use a ceiling that nothing can exceed instead.
Assume, impossibly, that every dollar of Cerebras revenue removes a dollar of Micron cloud-memory revenue. That cannot be true — memory is a fraction of any system's price, so a dollar of accelerator revenue can displace at most a fraction of a dollar of memory revenue — and it is useful precisely because it cannot be beaten. Whatever the real displacement rate is, it is smaller.
Micron's cloud-memory business unit did $13,769M in the quarter ended 2026-05-28, the basis quarter of our Micron model. Cerebras did $180.1M in the June quarter, as we covered in its first full quarter as a public company.
| $M per quarter | Multiple of Cerebras today | |
|---|---|---|
| Cerebras, 2026 Q2 | 180 | 1× |
| A 10% dent in Micron cloud memory | 1,377 | 7.6× |
| A 25% dent | 3,442 | 19.1× |
| Micron cloud memory, FQ3-26 | 13,769 | 76.5× |
Under the assumption most favourable to the bear case that arithmetic permits, Cerebras' entire company revenue currently displaces 1.3% of one Micron business unit for one quarter. Its whole trailing-twelve-month revenue of $680.7M is 2.2% of that unit's trailing year.
And the ceiling is doing a lot of hiding. Micron's cloud-memory line is not only HBM — it carries high-capacity server DRAM, 256GB RDIMMs, LP5X SOCAMM2 modules and data-centre SSDs. An architecture that removes HBM and keeps a 25-terabyte weight store leaves most of that list intact and adds to some of it.
What it does to our Micron model
Our Micron model prices exactly this exposure, so the claim can be tested against it rather than argued about.
The claim lands on the cloud memory vertical — modelled on capacity rather than growth, because Micron's constraint is supply: 2,400 PB in the basis quarter, 95% utilisation, capacity contracted ahead of production. One number there is ours and not Micron's, and it matters here: the petabyte figure is not disclosed. It is $13,769M divided by an assumed $6.04 per gigabyte. Micron publishes no bit volumes, so the level is a calibration of ours and the argument has to be made on price, not on bits.
That vertical carries $131.14B of the model's present value, against a base-case fair value of $706.33 per share on 1.145B diluted shares. So the sensitivity is direct arithmetic:
| Haircut to cloud-memory present value | Effect per share | Share of base fair value |
|---|---|---|
| 10% | −$11.45 | 1.6% |
| 25% | −$28.63 | 4.1% |
| 50% | −$57.27 | 8.1% |
Set that against the ordinary assumptions in the same model. The base case already has revenue peaking in fiscal 2028 and declining after, and the bear case takes the exit multiple from 3.5× revenue to 2.0× and the discount rate from 11% to 14% — moves that reprice the whole company far harder than halving the present value of its largest vertical does.
So the model does not change, and the size of the no is the finding. An architectural shift that removed a quarter of Micron's cloud-memory economics — which would require Cerebras to be nineteen times its current size, under a displacement assumption that overstates its effect — costs about 4% of fair value. The multiple and the cycle decide this stock. The architecture, on today's numbers, does not.
There is a real timing caveat and it cuts in Micron's favour. The largest customer agreements are take-or-pay with floor prices, with $22B of cash deposits and related commitments disclosed in the 10-Q for the same quarter. Demand destruction inside the contract term hits the uncontracted book first. An attach-rate shock has to outlive those contracts before it reaches the P&L, and nothing announced today is within years of that.
The company making the argument
We published a bruising piece on Cerebras two days ago: revenue down 7% sequentially, gross margin from 44.6% to 14.2%, free cash flow of −$477M on $417M of capital expenditure in a single quarter. Both readings are true at once and it is worth saying how they fit.
The CS-4 is evidence that a real, shipping alternative architecture exists — and $417M of quarterly capital expenditure against $180.1M of revenue is what it costs Cerebras to put that architecture in front of customers. A company spending more than twice its revenue on datacentre buildout, at 14.2% gross margin, is not currently displacing anything at scale. That is not a contradiction of the architecture story; it is the price of it, and it is why the threshold table above reads the way it does.
The thing to watch is not whether wafer-scale works. It does. It is whether Cerebras can sell it at a gross margin that funds the next generation.
What to watch
- Whether the CS-4 rack has an external DRAM tier, and how large. Cerebras' prior systems paired the wafer with a separate weight store. If the CS-4 datasheet quantifies it, the "engineered out memory" framing collapses into "moved the memory," and the terabytes involved are a demand number rather than a threat.
- The precision the 50-trillion-parameter claim assumes. The difference between 4-bit and FP16 is 75,000 GB of weight storage per model. Nobody has said which.
- Cerebras' CS-4 order book or revenue guidance. The threshold table has no left-hand side until there is a forward number. $1.38B a quarter is the level at which this stops being an architecture story.
- HBM gigabytes per accelerator across the next GPU generation. If attach per unit is flat or rising while unit volumes rise, the "engineering out" claim is a forecast about silicon that does not exist, not an observation about silicon that does.
- Micron's next disclosure of contracted versus uncontracted cloud-memory revenue. That split, not the architecture, decides whether any attach-rate shock reaches the income statement before the take-or-pay agreements run off.
CS-4 specifications — three WSE-3 Turbo engines, 750 PFLOPS, 129.6 PB/s of memory bandwidth, 2µs wafer-to-wafer latency, support for models over 50 trillion parameters, first shipments this quarter, and the 30×-versus-GPU and 10×-throughput-per-watt comparatives — are Cerebras' own claims from its 19 August 2026 launch, relayed to us through market-recap posts and the company's launch page; none of it is in a filing and the benchmark ratios have no published baseline. The 44 GB of on-wafer SRAM is Cerebras' published WSE-3 figure, carried forward to WSE-3 Turbo because the launch describes it as the same silicon at a higher clock — if the Turbo part changed the SRAM, the 132 GB changes with it. Weight-storage figures are 50 trillion parameters times 0.5, 1 and 2 bytes; the precision is our enumeration, not a company statement. Cerebras revenue of $180.1M for the June 2026 quarter, $680.7M trailing twelve months, 14.2% gross margin, $416.9M of capital expenditure and −$476.7M of free cash flow are from our coverage of the print, which sources them to the 12 August release. Micron cloud-memory revenue of $13,769M for the quarter ended 2026-05-28, the $131.14B segment present value, the $706.33 base-case fair value, 1.145B diluted shares, the 3.5× exit multiple and 11% discount rate, and the bear case's 2.0× and 14%, are from our Micron forward model as revised 2026-08-18; the model's own figures are assumptions of ours, not disclosures, and the 2,400 PB implied volume in particular is $13,769M divided by an assumed $6.04 per gigabyte because Micron publishes no bit volumes. The take-or-pay floor prices and the $22B of customer deposits and commitments are quoted in that model from Micron's 10-Q for the same quarter. The Cathie Wood remark is a third-party paraphrase of a broadcast comment; we have not seen the primary source and nothing in this piece rests on it. No price or market capitalisation is quoted here.