Cisco added Supermicro rack-scale compute to its Secure AI Factory with NVIDIA on August 25. Beginning in October, Cisco's authorised channel will be able to sell Supermicro's liquid- and air-cooled dense GPU systems alongside Cisco networking, security, observability and deployment services as one pre-validated architecture.
That is a real expansion of what Cisco can package. It is not a new GPU, a new rack or an order. The announcement disclosed no customer purchase, no rack count, no contract value, no revenue-sharing terms and no margin for either company. That does not mean the channel relationship has no value.
The useful way to read it is to separate the layers. NVIDIA defines the accelerator and its scale-up fabric. Supermicro builds and cools the compute racks. Cisco supplies the Ethernet fabrics around those racks, the operating and security plane, validation and the route to market. The table above is the announcement without the phrase "full stack" doing all the work.
The compute belongs to Supermicro; the architecture begins with NVIDIA
The flagship option is Vera Rubin NVL72, a third-generation MGX rack containing 72 Rubin GPUs and 36 Vera CPUs. ConnectX-9 SuperNICs carry scale-out traffic and BlueField-4 DPUs offload infrastructure services. Inside the rack, NVLink 6 connects all 72 GPUs into one scale-up domain.
NVIDIA's preliminary specification gives that rack:
- 3,600 PFLOPS, or 3.6 exaflops, of dense NVFP4 inference peak;
- 20.7 TB of HBM4 at 1,580 TB/s of aggregate memory bandwidth;
- 54 TB of CPU LPDDR5X, taking total fast CPU-plus-GPU memory to 74.7 TB;
- 260 TB/s of aggregate NVLink bandwidth across the rack; and
- 28.8 TB/s of scale-out network bandwidth.
Those categories matter. The often-repeated 75 TB is not 75 TB of HBM4; it is approximately 20.7 TB of GPU memory plus 54 TB of CPU memory. The 1.58 PB/s figure is HBM4 bandwidth, not capacity and not NVLink. And 3.6 exaflops is a preliminary dense low-precision peak, not a measured result for every model.
NVIDIA says Vera Rubin NVL72 can deliver 10 times more tokens per megawatt and one-tenth the cost per million tokens than GB200 NVL72. Its own footnotes make those workload-specific comparisons: the token claims use Kimi-K2-Thinking with defined input and output sequence lengths, and the results remain subject to change. "10x throughput per watt" is a useful shorthand only if the benchmark travels with it.
The smaller option is Supermicro's 2U HGX Rubin NVL8. Each system holds eight Rubin GPUs; as many as nine systems fit in a rack, reaching the same 72-GPU count without turning the whole rack into one NVL72 scale-up domain. NVIDIA specifies 400 PFLOPS of NVFP4 inference, 2.3 TB of HBM4 at 176 TB/s, and 28.8 TB/s of NVLink switch bandwidth per eight-GPU system. It can pair Rubin with a Vera CPU or with next-generation AMD or Intel x86 hosts. That is the modular choice for customers who value CPU flexibility and a conventional server boundary over the NVL72's rack-wide fabric.
Scale-up is NVIDIA. Scale-out is where Cisco gets paid
The networking names sound interchangeable until the boundary is drawn.
NVLink 6 is the scale-up fabric. It is NVIDIA technology inside NVL72, connecting the 72 GPUs at 3.6 TB/s per GPU and 260 TB/s across the rack. Cisco does not replace it.
Ethernet is the scale-out fabric. It joins racks to one another, to storage and to front-end services. Cisco's N9100 switches use NVIDIA Spectrum-X silicon for the back-end GPU network; its N9300 systems use Cisco Silicon One for the front end. Cisco Nexus One unifies both, with NX-OS or SONiC as the network operating system. Cisco describes itself as the first NVIDIA technology partner with an NCP-compliant reference architecture spanning its own Silicon One systems and partner-developed Spectrum-X systems.
Around that network, Cisco adds the portion Supermicro does not specialise in: Cloud Control and Nexus One for operations; Intersight integration for compute management, due in the fourth quarter of 2026; Hybrid Mesh Firewall policy enforcement on BlueField DPUs; AI Defense; and Splunk observability across compute, NICs, optics, networks, jobs and agents.
The new Cisco Validated Infrastructure Services, or CVIS, is the operational promise underneath "pre-validated." It is aligned with NVIDIA Infrastructure Services and is meant to take a deployment from discovery and design through automated provisioning, reference-architecture checks and an evidence report. Cisco says the automation can reduce validation from months to weeks. That is a target for the service, not a result disclosed for a Supermicro deployment today.
Support is coordinated rather than magically consolidated. Cisco will handle L0 and L1 triage and route component issues to the appropriate partner; Supermicro remains responsible for its products. A customer gets one front door, not one manufacturer.
Cooling is not an accessory above 200 kW
Cisco's technical FAQ says modern NVL72 racks can exceed 200 kW. At that density, cooling and networking are no longer facility choices made after the server order. They are part of the system design.
Supermicro's DLC-2 stack puts cold plates on the processors and connects them through vertical manifolds to in-rack or in-row coolant distribution units. Its June blueprint specifies in-row CDUs of up to 1.8 MW, with rear-door heat exchangers and liquid-to-air sidecars for sites without facility water. Cisco adds liquid-cooled N9000 networking so the thermal design continues from compute rack to fabric.
The scale also explains why the announcement targets neoclouds and sovereign clouds as much as conventional enterprises. Supermicro's smallest published DCBBS blueprint is a 1,152-GPU unit: 16 NVL72 compute racks, six networking racks and six storage racks, inside a site design starting at 5 MW. It carries 331 TB of HBM4. The arithmetic reconciles: 1,152 GPUs × 288 GB is 331.8 TB, rounded by Supermicro to 331 TB.
Divide 5 MW by 1,152 GPUs and the starting facility envelope is about 4.34 kW per GPU. That is not GPU board power — the denominator also supports networking, storage, cooling and power conversion — but it is the right planning number. A rack specification without the site around it is not an AI factory.
What changed on August 25, and what did not
Supermicro announced the racks in January, detailed the systems in March and published the 1,152-GPU DCBBS blueprint in June. Cisco launched Secure AI Factory with NVIDIA in 2025 and expanded its Spectrum-X, Silicon One, security and edge options in March 2026. August 25 did not combine new silicon with new cooling. It combined two existing portfolios into a Cisco-validatable, Cisco-manageable and Cisco-channel-orderable configuration.
That matters because Cisco's channel already has a denominator. Cisco reported $9.3 billion of FY2026 AI infrastructure orders from hyperscalers, including $4 billion in its July quarter, and about $4 billion of FY2026 AI infrastructure revenue. It expects $7.5 billion in FY2027. The Supermicro architecture is aimed partly at a different pool — enterprises, neoclouds and sovereign buyers — but Cisco disclosed no portion of those order or revenue figures attributable to it.
Supermicro's denominator is larger and more direct. It guided FY2027 revenue to $65–72 billion, up 66–84% from FY2026, and said it entered the year with more than $60 billion of new orders. The Cisco agreement adds a sales and validation route into that plan. It does not add another dollar to the disclosed order book yet.
This is why describing the deal as Cisco "moving into compute" needs care. Cisco is adding Supermicro compute to what it sells and supports; it is not designing or manufacturing the GPU servers. Its own FAQ says UCS remains the foundation of enterprise compute, while Supermicro extends the portfolio into the densest rack-scale systems. Cisco broadens the catalogue. Supermicro supplies the rack.
The Supermicro model already contains this story
Our Supermicro model puts the Cisco-targeted business inside OEM Appliance & Large Data Center, the vertical that includes liquid-cooled SuperClusters and DCBBS cooling, power, networking and deployment services. It already assumes an NVL-class rack equivalent priced at $2.75 million, rising 1% a quarter near term as GPU content and DCBBS attachment increase. That rack price is ours; Supermicro discloses neither rack volumes nor average selling price.
The model's base case is $60.79 a share. Move only the near-term quarterly rack-price drift and the sensitivity is small:
| Rack-price drift | Fair value | Change from base |
|---|---|---|
| 0% a quarter | $59.26 | −$1.53 |
| 1% a quarter — base | $60.79 | — |
| 2% a quarter | $62.40 | +$1.61 |
The partnership could help the attachment rate: a validated Cisco package may make it easier to sell more infrastructure around each Supermicro rack. It could also place Cisco networking and services where Supermicro might otherwise have captured more of the DCBBS bill. The commercial split is undisclosed, so choosing either direction today would invent a number.
The model does not change. Even moving the price-drift assumption a full percentage point changes fair value by only 2.5–2.6%, while the announcement supplies no evidence to move it at all. The investable event is the first Cisco-channel customer order, not the availability date.
What to watch
- The first named order and its boundary. Rack count, GPU count and contract value will show whether Cisco is reselling compute, leading a complete design or merely supplying the network around a Supermicro-led build.
- Who recognises which revenue. Cisco has not said whether Supermicro hardware passes through its accounts, through authorised partners or only appears in a validated bill of materials. That decides whether the announcement expands Cisco revenue or mostly its networking pull-through.
- A CVIS evidence report. Job completion time, token throughput and cluster performance from a delivered system would turn "pre-validated" from a process claim into a measured result.
- October availability. The systems are expected to become orderable then. Orderable is not shipped, accepted or revenue-recognised; those dates will determine which fiscal quarter gets any benefit.
- Supermicro's customer mix and gross margin. Its base economics remain a thin-margin integrator's. If Cisco-led enterprise and sovereign deployments attach more cooling, power and service content, the evidence should appear in mix and margin before it appears in a partnership slide.
Cisco's August 25 release and FAQ are the sources for the October order window, product and vendor split, NCP architecture, CVIS, support model, management roadmap and the statement that NVL72 racks can exceed 200 kW. Cisco's $9.3B of FY2026 hyperscaler AI orders, ~$4B of FY2026 AI revenue and $7.5B FY2027 expectation are from its July-quarter release. Supermicro's rack configurations, DLC-2 cooling and 1,152-GPU blueprint are company disclosures from January, March and June 2026; its $65–72B guide and >$60B of new orders are from its June-quarter release. NVIDIA's NVL72 and NVL8 specifications and performance comparisons are preliminary, workload-specific where stated and subject to change. The 16-rack reconciliation, 4.34 kW facility envelope per GPU and HBM total are our arithmetic. The $2.75M rack price, price drift and fair values are assumptions and outputs from our Supermicro model, not company disclosures. No order value or revenue split was disclosed.