Sunday, September 13, 2026

RACK-SCALE AI INFRASTRUCTURE - NVIDIA Vera Rubin v/s AMD Helios

The War for AI Compute Has Left the Chip Behind

Trying to put 72 modern GPUs into a single rack is like trying to park a space shuttle at your house. You cannot just wheel it into your garage — you would have to re-engineer the entire building around it. NVIDIA and AMD took that exact approach, and in doing so have shifted the battleground for AI infrastructure from the chip to the rack.

For years, the industry measured competition by GPU performance numbers. Today, the battleground has expanded to rack-scale AI systems that integrate compute, memory, networking, and software into a single platform. AMD's Helios and NVIDIA's Vera Rubin NVL72 are the clearest expression of that shift — and they represent fundamentally opposite bets on how to build an AI factory.

"The AI race is no longer about building the fastest chip. It is about delivering the most complete AI infrastructure platform — from silicon to cloud."

Two Racks. Two Philosophies.

Vera Rubin NVL72
  • 72 Rubin GPUs + 36 Vera CPUs
  • 🧠3.6 EFLOPS sparse NVFP4 inference
  • 💾~21 TB HBM3e memory
  • 🔗NVLink 6 fabric — proprietary
  • 190–230 kW power draw
  • 🏗️800V DC architecture required
  • 🔒Vertically integrated closed stack
  • 🛠️CUDA + full software ecosystem
Helios AI Rack
  • 72 Instinct MI455X GPUs + 6th Gen EPYC Venice CPUs
  • 🧠">~2.4 EFLOPS (FP8 estimated)
  • 💾31 TB HBM4 memory — 50% more
  • 🔗UALink + open Ethernet scale-out
  • ~140 kW power draw
  • 🏗️Meta Open Rack Wide (ORW) spec
  • 🔓Open Compute Project (OCP) standard
  • 🛠️ROCm + open software stack
SPECIFICATION
VERA RUBIN NVL72
HELIOS
GPU Count
72 Rubin GPUs
72 MI455X GPUs
CPU
36 × Vera (Grace successor)
6th Gen EPYC Venice
AI Performance
3.6 EFLOPS (NVFP4)
~2.4 EFLOPS (FP8 est.)
Memory (per rack)
~21 TB HBM3e
31 TB HBM4 (+50%)
Memory per GPU
288 GB (est.)
432 GB HBM4
Interconnect
NVLink 6 (proprietary, ultra-low latency)
UALink + open Ethernet
Power Draw
190–230 kW
~140 kW
Cooling
Full liquid cooling required
Liquid manifold + quick-disconnect
Power Architecture
800V DC (facility retrofit req.)
Standard — no 800V mandate
Rack Standard
NVIDIA proprietary
Meta ORW / OCP open spec
Software Maturity
CUDA — industry standard
ROCm — rapidly maturing
Est. System Cost
$3M+ (NVL72 ref.)
Competitively positioned

Speed vs Memory: Two Different Bets on the AI Bottleneck

STACK ARCHITECTURE — NVIDIA VERA RUBIN NVL72 VS AMD HELIOS
CUDA / CUDA-X Software
Proprietary · Deeply optimised · Broadest library support
NVLink 6 Fabric
Proprietary copper spine · Ultra-low latency · All-to-all GPU bandwidth
72 × Rubin GPU + 36 × Vera CPU
Co-designed silicon · Shared memory hierarchy
~21 TB HBM3e
288 GB per GPU · Optimised for throughput
NVIDIA — Vertically Integrated
VS
ROCm / Open Software Stack
Open source · OCP aligned · Vendor-neutral tooling
UALink + Open Ethernet
Open standard · Mix-and-match networking · Scale-out flexibility
72 × MI455X GPU + EPYC Venice CPU
Meta ORW double-wide · Hot-swap serviceability
31 TB HBM4
432 GB per GPU · 50% more than Vera Rubin
AMD — Open Coalition

NVIDIA: Built for Speed, Unified Control

Vera Rubin is engineered for pure throughput and low-latency reasoning. By tying 72 Rubin GPUs together with 36 Vera CPUs over NVLink 6 — NVIDIA's proprietary copper spine fabric — the entire rack behaves as one massive, hyper-optimised compute engine. The hardware, interconnect, and CUDA software stack are co-designed from the ground up. That tight integration is precisely why NVIDIA can claim 3.6 EFLOPS of sparse NVFP4 inference. The tradeoff: you play entirely by NVIDIA's rules.

AMD: Built for Memory, Built for Openness

AMD made a different wager. Instead of competing solely on raw compute numbers, Helios bets that memory capacity is the real AI bottleneck — especially for large language models with massive context windows. At 31 TB of HBM4 across the rack and 432 GB per GPU, Helios offers 50% more memory than Vera Rubin. AMD also chose open standards throughout: Open Rack Wide physical spec, UALink interconnect, and standard Ethernet for scale-out — giving data centres the freedom to mix and match without vendor lock-in.


Memory Is the New Differentiator

For large language model inference — especially models requiring long context windows, multi-modal inputs, or large batch sizes — memory capacity has become just as important as raw floating-point throughput. A model that does not fit in GPU memory cannot run efficiently, regardless of how fast the compute units are.

~21
TB HBM3e per rack
NVIDIA Vera Rubin NVL72
288 GB per GPU
31
TB HBM4 per rack
AMD Helios
432 GB per GPU · +50% vs Vera Rubin

AMD's next-generation MI450 GPU is reportedly designed with 432 GB of HBM4, significantly ahead of the 288 GB expected on NVIDIA's Vera Rubin. For organisations running frontier models, agentic AI workloads, or memory-intensive inference pipelines, that extra capacity is not a marketing statistic — it directly determines what you can serve without model sharding or offloading.


What It Actually Takes to Run These Systems

Neither of these racks can be wheeled into a legacy raised-floor data centre. The facility requirements are substantial — and represent a real capital cost that belongs in any architecture decision.

NVIDIA Vera Rubin NVL72190–230 kW
NVL72
AMD Helios~140 kW
Helios
NVIDIA Rubin Ultra (2027, projected)~600 kW
Rubin Ultra

NVIDIA's NVL72 draws 190–230 kW and drives an 800V DC architecture that forces electrical microgrid retrofits in most existing facilities. AMD's Helios, built on Meta's Open Rack Wide spec, draws roughly 140 kW and avoids the 800V mandate — a meaningful operational advantage for organisations deploying into existing infrastructure. The cooling story is unambiguous for both: air cooling is no longer viable at this density. Vera Rubin requires full liquid cooling; Helios uses a cooling manifold with quick-disconnect trays. What is striking is the roadmap: NVIDIA's Rubin Ultra, projected for 2027, is expected to draw approximately 600 kW per rack — a number that will force an entirely new generation of data centre design.


Walled Garden vs Open Coalition

The single biggest architectural difference between these two platforms is not the GPU count, the memory, or the power draw. It is the ecosystem philosophy — and that choice has long-term implications that extend well beyond the hardware purchase.

🔒 NVIDIA — The Integrated Stack
  • CUDA remains the dominant AI software standard with the deepest library ecosystem
  • NVLink provides unmatched intra-rack GPU-to-GPU bandwidth
  • Grace CPU + Rubin GPU co-designed for unified memory access
  • Validated reference architecture reduces deployment risk
  • Omniverse, TensorRT, NeMo — a complete AI software platform
  • Tradeoff: full vendor lock-in at every layer of the stack
🔓 AMD — The Open Coalition
  • Open Compute Project (OCP) and UALink — no proprietary lock-in
  • Standard Ethernet scale-out — mix and match networking vendors
  • ROCm open-source software stack, increasingly PyTorch compatible
  • Meta, Microsoft, Oracle and OpenAI already deploying Helios
  • Pensando networking for programmable data-plane control
  • Tradeoff: software maturity still catching CUDA in some areas

Microsoft's decision to deploy Helios across Azure — for frontier AI inference, Azure AI services, and enterprise workloads — is the most significant validation of AMD's open-stack bet to date. It joins Meta, OpenAI, and Oracle in adopting AMD's next-generation AI platform, signalling that the hyperscaler tier is no longer treating AMD as a backup option.


AI Is Reshaping the Semiconductor Supply Chain

The competition between these two racks is not just a story about two companies. It is a signal about how the entire AI supply chain is reorganising. Success in rack-scale AI now requires the ability to coordinate silicon design, advanced packaging, memory production, networking, and manufacturing at a scale that few organisations in the world can manage.

TSMC 2nm
EPYC Venice + MI450
SK Hynix / Samsung
HBM4 Production
AMD Instinct MI455X
432 GB HBM4
Pensando
Network Silicon
Helios Rack
Meta ORW
Azure / Cloud
Enterprise AI

Samsung and SK Hynix are both expanding HBM4 production capacity specifically to supply the next wave of AI accelerators. TSMC's advanced 2nm process will power AMD's EPYC Venice CPUs and MI450 accelerators. The ability to secure these inputs — and assemble them reliably at rack scale — is becoming as strategically important as the chip designs themselves.


Which One Is Right for You?

✅ CHOOSE NVIDIA VERA RUBIN IF…
  • You need the absolute fastest, tightly integrated, low-latency inference performance
  • 🛠️Your teams are deeply invested in CUDA and the NVIDIA software ecosystem
  • 🏗️You can retrofit your facility for 800V DC power infrastructure
  • 📦You want a fully validated, single-vendor reference architecture with minimal deployment risk
  • 🔬Your workloads are compute-bound rather than memory-bound
✅ CHOOSE AMD HELIOS IF…
  • 💾Memory capacity is your primary bottleneck — large models, long context, big batch sizes
  • 🔓Vendor diversification and avoiding lock-in are strategic priorities
  • Lower power draw and avoiding 800V facility upgrades matters operationally
  • 🔧Your ops team values physical serviceability and the double-wide hot-swap design
  • 💰Competitive pricing pressure against NVIDIA is a factor in your procurement

For most buyers today, NVIDIA Vera Rubin is the safer near-term choice — it offers the broadest software maturity, the most tuned reference architecture, and the clearest deployment path. Helios is the more compelling choice if you are thinking about where the constraints will be in 18–36 months: memory capacity, energy efficiency, openness, and the leverage that comes from not depending on a single vendor.


The Rack Has Become the Computer

What NVIDIA and AMD have built is not just a denser version of yesterday's server. It is a new unit of compute — one where the rack itself is the system, where the interconnect is as important as the processor, and where facility design, memory architecture, and software ecosystem are all part of the same purchasing decision.

The question for AI infrastructure teams in 2025 is not which GPU has the best benchmark. It is which rack-scale philosophy — closed and fast, or open and memory-rich — better matches the direction your workloads and your organisation are heading.



Vera Rubin is the speed champion.
Helios is the memory and openness play.

#NVIDIAVeraRubin#AMDHelios#AIInfrastructure#HBM4#RackScaleAI#GPUComputing#DataCenter#EnterpriseAI#EPYC#CUDA#OpenComputing

Quantum Doesn't Replace Classical Computing. It extends and Completes It

 The Biggest Myth in Quantum Computing

Quantum computing isn't here to replace classical computing. In fact, it can't — at least not on its own. Every practical quantum computer in operation today depends on a powerful classical computer working alongside it.

This is the part that rarely makes the headlines. The narrative around quantum has been dominated by superlatives: exponential speedup, unbreakable encryption, problems solved in seconds that would take classical machines millennia. While these claims have technical basis in specific contexts, they obscure a more important and immediately actionable truth: quantum and classical computing are partners, not competitors.

Understanding how that partnership works — and why it is necessary — is the foundation for making any serious enterprise quantum strategy.

"The future of computing isn't a battle between classical and quantum. It's a partnership where each system does what it does best."

Inside the Hybrid Loop

Hybrid quantum-classical computing is not a marketing term. It describes a precise architectural pattern in which two fundamentally different types of processors exchange information in a continuous loop, each doing the work the other cannot do efficiently.

Here is how a single hybrid computation actually runs:

THE HYBRID QUANTUM–CLASSICAL EXECUTION LOOP
💻
Classical
Prepares the problem & generates circuit
⚛️
Quantum
Executes quantum operations on qubits
📊
Classical
Collects & analyses measurement results
🔄
Classical
Decides next parameters or terminates
↺ Loop repeats until convergence

The classical computer is the conductor. It formats the problem, compiles the quantum circuit, sends instructions to the quantum hardware, reads the probabilistic measurement outputs, and feeds updated parameters back into the next quantum run. The quantum processor does one thing — but it does it in a way no classical system can efficiently replicate: it manipulates qubits in superposition and entanglement to explore solution spaces that would require exponential classical resources to traverse.

💻
Classical Computer's Role
  • Problem formulation & pre-processing
  • Quantum circuit generation & compilation
  • Hardware control & error management
  • Measurement result collection
  • Post-processing & parameter optimisation
  • Decision logic for next iteration
⚛️
Quantum Processor's Role
  • Superposition — explore many states simultaneously
  • Entanglement — correlate qubits across the register
  • Interference — amplify correct answers
  • Execute circuits classical hardware cannot efficiently simulate
  • Return probabilistic measurement samples
Classical
Precision · Control · Reliability
+
Quantum
Superposition · Entanglement · Interference
=
Hybrid Model
Neither could achieve alone
CLASSICAL PROVIDES THE FOUNDATION. QUANTUM EXTENDS THE FRONTIER.

Where Quantum Computing Stands Today

The current era of quantum hardware is described by researchers as the Noisy Intermediate-Scale Quantum (NISQ) era. Today's quantum processors carry between 100 and 1,000+ qubits, but those qubits are imperfect — they accumulate errors over time (decoherence), and gate operations introduce noise. This is precisely why the classical computer's role is so critical: it handles error mitigation, circuit optimisation, and the statistical aggregation of many noisy quantum runs into reliable outputs.

1,000+
Qubits on IBM's largest quantum processors (2024)
~100μs
Typical qubit coherence time — classical control must operate within this window
2033+
Projected timeline for fault-tolerant quantum advantage in commercial problems

IBM's quantum roadmap targets fault-tolerant quantum computing through logical qubit architectures — where multiple physical qubits are combined to form a single reliable logical qubit. Until that milestone is reached, every quantum workload will be a hybrid workload.


Real-World Applications of Hybrid Computing

The use cases gaining the most traction are those where the problem structure maps naturally to quantum mechanics — where the difficulty is not raw compute speed, but the ability to explore enormous combinatorial or quantum-mechanical solution spaces.

🧪
Molecular Simulation
Drug discovery and materials science — simulating electron interactions that are intractable for classical systems.
📦
Combinatorial Optimisation
Logistics, supply chain, and scheduling problems where the solution space grows exponentially with inputs.
💰
Financial Modelling
Portfolio optimisation, risk analysis, and Monte Carlo acceleration using quantum amplitude estimation.
🤖
Quantum ML
Variational quantum circuits as trainable model layers — hybrid quantum-classical neural networks.
🔐
Cryptography
Post-quantum cryptographic standards (NIST PQC) — preparing classical infrastructure for the quantum threat.
Energy Systems
Grid optimisation, battery chemistry simulation, and quantum-assisted climate modelling.

The key algorithms powering most of these use cases — Variational Quantum Eigensolver (VQE), Quantum Approximate Optimisation Algorithm (QAOA), and Quantum Phase Estimation — are all natively hybrid. They require a classical optimiser running in the outer loop, updating quantum circuit parameters after each quantum execution.


From NISQ to Fault-Tolerant: The Journey

1
NOW — NISQ ERA
Noisy, error-prone hardware — hybrid is essential
Classical computers compensate for quantum noise through error mitigation, circuit optimisation, and statistical averaging. Most quantum advantage demonstrations are narrow, domain-specific, and require heavy classical scaffolding.
2
MID-TERM — ERROR-CORRECTED ERA
Logical qubits emerge — hybrid becomes more capable
Multiple physical qubits are combined into logical qubits with built-in error correction. The classical overhead required per quantum operation decreases. Useful quantum advantage starts to emerge for specific problem classes in chemistry and optimisation.
3
LONG-TERM — FAULT-TOLERANT ERA
Fault-tolerant quantum — hybrid remains the model
Even with fault-tolerant quantum processors, the classical computer does not disappear. Problem formulation, result interpretation, system control, and integration with enterprise data and applications will remain classical responsibilities. The partnership endures.

What Enterprises Should Be Doing Now

The most common mistake enterprises make with quantum is treating it as a future problem — something to revisit in five years when the hardware is "ready." That framing misses two immediate imperatives.

🏢 Enterprise Quantum Readiness Checklist
  • Inventory your cryptographic exposure. Quantum computers will eventually break current RSA and ECC encryption. NIST's post-quantum cryptography standards (finalised in 2024) define the path to quantum-safe infrastructure — migration is a multi-year program that should start now.
  • Identify hybrid-ready problem classes. Audit your optimisation, simulation, and machine learning workloads for problems that are structurally suited to hybrid quantum-classical approaches.
  • Build quantum literacy in technical teams. IBM Quantum and open-source tools like Qiskit lower the barrier significantly. Developers can write and run real quantum circuits on IBM's cloud-accessible quantum hardware today.
  • Engage with cloud quantum platforms. IBM Quantum Network, Amazon Braket, and Azure Quantum all provide cloud access to quantum hardware. Experimentation is low-cost relative to the strategic value of understanding the technology.
  • Integrate quantum into AI and HPC roadmaps. Hybrid quantum-classical fits naturally alongside GPU computing and AI workloads — it is not a separate initiative, it is an extension of existing high-performance computing strategy.

One Shift in Thinking Changes Everything

The organisations that will extract real value from quantum computing in the next decade are not the ones waiting for a standalone quantum breakthrough. They are the ones building hybrid competency now — understanding how classical and quantum systems divide labour, where the boundaries of each lie, and how that boundary shifts as hardware matures.

Quantum computing is a powerful extension to the classical computing stack. It is not a replacement. It is not a revolution that arrives overnight. It is a partnership — and like all productive partnerships, it rewards those who invest in understanding both sides.

Classical provides the foundation.
Quantum extends the frontier.
Neither system replaces the other. Together, they create a computing model that neither could achieve alone. The question for enterprise technologists is not whether to engage with quantum — it is how to build the hybrid capability that positions them for what comes next.

How Does a Variational Quantum Circuit "See" an Image?

A Variational Quantum Circuit (VQC) does not look at an image the way a Convolutional Neural Network does. There is no sliding kernel, no pooling layer, no spatial hierarchy built from local patterns. Instead, image information is encoded into qubits, transformed through trainable quantum operations, and measured to produce features that a classical system can act on.

Understanding this pipeline is key to grasping both the promise and the present constraints of quantum image processing — and why it represents one of the most active research frontiers in hybrid quantum-classical AI.

🖼️
Image
✂️
Preprocess
🔡
Encode
⚛️
Quantum Transform
🔗
Entangle
📐
Measure
🧠
Classical Decision
🔁
Learn
1
Image → Smaller Information Units
CLASSICAL PRE-PROCESSING
A full-resolution image cannot be fed directly into today's quantum hardware — current processors have a limited number of usable, low-noise qubits. The first step is therefore a classical one: resize, normalise, and if necessary divide the image into patches. Each patch will map to a manageable number of qubits. This constraint is not a fundamental limitation of quantum computing — it is a property of the current NISQ hardware era.
256×256 image
Resize / Normalise
Patch extraction
n feature values
2
Classical → Quantum Encoding
AMPLITUDE / ANGLE ENCODING
Pixel values or extracted features are mapped onto quantum states — a process called encoding. The most common approach is angle encoding: a feature value θ is used as the rotation angle of a quantum gate, placing each qubit into a superposition that carries the information. Alternatively, amplitude encoding can represent an exponentially large feature vector in a small number of qubits — though preparing such states efficiently is itself an open research problem.
|0⟩
Rᵧ(θ₁)
|0⟩
Rᵧ(θ₂)
|0⟩
Rᵧ(θₙ)
3
Trainable Quantum Transformations
PARAMETERISED GATES
The encoded qubits pass through parameterised rotation gates — Rₓ(α), Rᵧ(β), R_z(γ) — where the angles α, β, γ are learnable parameters, updated by a classical optimiser during training. These are the quantum analogue of weights in a neural network layer. The parameters are initialised randomly and refined through the hybrid feedback loop.
Rₓ(α)
Rᵧ(β)
R_z(γ)
CX
Rₓ(α')
← trainable parameters
4
Entanglement Creates Feature Interaction
CNOT / CZ GATES
Entangling gates (CNOT, CZ) link qubits together so that the state of one qubit influences another. This is what separates a VQC from a simple collection of independent rotations — entanglement allows the circuit to construct representations that capture relationships between encoded features, not just individual feature values. The structure of the entanglement pattern (which qubits are connected, and how) is an architectural design choice analogous to the connectivity pattern in a neural network.
q₀
q₁
q₂
q₃
← entangled — jointly encode feature relationships
5
Repeated Layers Build Richer Representations
CIRCUIT DEPTH
A single layer of rotations and entanglement produces a limited transformation. Multiple alternating layers — rotation block, entanglement block, rotation block, entanglement block — progressively transform the initial encoded state into a richer, task-dependent quantum representation. Increasing circuit depth increases expressive power, but also increases sensitivity to hardware noise. Balancing depth against decoherence is a central engineering challenge of VQC design on real quantum hardware.
Encode
Layer 1
Layer 2
Layer L
|ψ(θ)⟩
6
Measurement Bridges Quantum and Classical Worlds
QUANTUM → CLASSICAL
Quantum states cannot be directly passed to a classical neural network — they must be measured. When a qubit is measured, it collapses to a 0 or 1. Repeating the circuit many times (called "shots") and computing expectation values — the average measurement outcome — produces a set of classical real numbers. These become the feature vector that the downstream classical layer processes. The measurement step is irreversible and introduces statistical noise, which is another reason why many circuit runs are needed per training step.
|ψ(θ)⟩
Measure ×N shots
⟨Z₀⟩, ⟨Z₁⟩ … ⟨Zₙ⟩
Classical vector
7
Classical Processing Produces the Result
DOWNSTREAM TASKS
The measured expectation values — now a compact classical feature vector — are passed to a classical layer for the final decision. Depending on the application, this produces a classification label, a segmentation map, a denoised output, or an enhanced version of the original image. The quantum circuit's role is not to produce the answer directly — it is to generate a feature representation that is difficult for a classical network to construct on its own.
QVC features
Dense / Softmax
Classification
/
Segmentation
/
Enhancement
8
The Circuit Learns Through Feedback
CLASSICAL OPTIMISER
The output is compared against the ground truth using a loss function. A classical optimiser — typically Adam, SPSA, or a gradient-free method — computes how the circuit parameters should change to reduce the loss. The updated parameters are loaded back into the quantum circuit, and the entire pipeline runs again. This outer classical optimisation loop is what makes the circuit variational — it is the learning engine of the system. Gradients can be computed using the parameter-shift rule, a technique native to quantum circuits that enables exact gradient estimation without backpropagation.
Output
Loss L(θ)
∇L → update θ
↺ New circuit run

🧠 THE CORE INSIGHT
A VQC is not a "quantum CNN." Its role is better understood as a trainable quantum feature transformation inside a hybrid pipeline.
The quantum circuit does not replace the classical neural network — it contributes a fundamentally different kind of feature space. Whether that quantum feature space is more expressive, more data-efficient, or better suited to certain visual tasks than a purely classical representation is the question driving current research. The honest answer: in some carefully designed scenarios, early evidence suggests yes. In general, under realistic noise constraints, the jury is still out.
COMPLETE HYBRID VQC PIPELINE — CLASSICAL & QUANTUM ZONES
CLASSICAL DOMAIN
Image loading & pre-processing
Patch extraction / feature reduction
Circuit parameter initialisation
Measurement result collection
Dense / output layer
Loss computation & optimisation
QUANTUM DOMAIN
Angle / amplitude encoding
Parameterised rotation gates
Entangling gates (CNOT / CZ)
Repeated rotation + entanglement layers
Qubit measurement (expectation values)
🖼️ Image
Pre-process
Encode
⚛️ VQC
Measure
Classify
Loss
↺ Classical optimiser updates θ → repeat
💡 The Real Research Question
"Can carefully designed quantum feature spaces provide useful representations for visual-learning tasks under realistic qubit, noise, and measurement constraints?"
The question is no longer simply "Can quantum computers process images?" — current hybrid pipelines demonstrate that they can. The open challenge is whether the quantum feature representations they produce offer a genuine advantage over classical alternatives in terms of expressibility, sample efficiency, or generalisation. That question makes Quantum Image Processing + VQCs one of the most exciting research directions in hybrid quantum–classical AI today.