At first glance, the network inside a modern AI factory can look very similar to a traditional data center network. You’ll see BGP, ECMP, high-radix switches, Clos fabrics, and, in many Ethernet environments, familiar technologies like EVPN-VXLAN.
That may sound a lot like the data center networks we’ve been building for years, and in many ways, it is. The difference is not that AI suddenly replaces traditional data center networking with something completely new. Instead, AI changes what the network is being asked to do.
In other words, technical differences with design and configuration aside, the real difference to understand is the actual goal of the network.
What a Modern Traditional Data Center Looks Like
First, we need to define “traditional.”
A modern data center in 2026 doesn’t necessarily mean the old access-aggregation-core architectures many network engineers grew up with (including me). Large enterprise and cloud data centers have already moved toward IP-based Clos fabrics, typically implemented as leaf-spine architectures.
Servers and appliances connect to leaf switches, every leaf connects to every spine, and layer 3 ECMP provides multiple equal-cost paths through the fabric. Adding leaves increases endpoint capacity, and adding spines increases available fabric bandwidth.
On top of that IP fabric, EVPN-VXLAN has become a common architecture for network virtualization. The underlay provides IP connectivity between VTEPs. VXLAN provides encapsulation and network segmentation, and MP-BGP EVPN distributes MAC and IP reachability information through the control plane. Instead of stretching physical VLANs throughout the data center, operators can create logical Layer 2 and Layer 3 networks over a routed IP fabric.
This architecture works extremely well for conventional workloads such as virtualization clusters, Kubernetes platforms, databases, web applications, and typical enterprise services. In a traditional data center, the network is designed primarily to provide connectivity between applications, services, users, and storage.
Also, modern data centers already carry plenty of east-west traffic. Microservices communicate with other microservices, applications query distributed databases and storage systems, and virtual machines and containers communicate across racks. The old characterization of the data center as primarily north-south traffic hasn’t accurately described many environments for years.
What changes with AI is not simply the direction of the traffic. It’s the nature of that traffic and its relationship to the workload. In a large GPU cluster, the network is part of the distributed computing system itself.
AI Takes East-West Traffic to Another Level
Large-scale AI training takes thousands (or many tens of thousands) of GPUs and attempts to make them behave like one enormous computing system. Rather than running mostly independent workloads, GPUs repeatedly exchange data during distributed training through collective communication operations such as AllReduce, AllGather, ReduceScatter, and All-to-All. This produces an unusual network workload.
Thousands of GPUs may communicate at the same time, moving enormous amounts of data between servers and racks as part of the same computational operation. These aren’t simply independent application flows that happen to be traveling east-west like in a traditional data center. In an AI data center network, the communication is coordinated and often tightly synchronized as part of the same distributed computation.
Imagine 1,024 GPUs participating in an AllReduce operation. Each GPU performs computation locally and then exchanges data with other GPUs so the results can be combined and distributed before the next stage of computation proceeds.
If one GPU or network path becomes a straggler during a collective operation, progress across the larger operation can be constrained by that slower participant. The result can be delayed synchronization and longer job completion time.
The network has effectively become part of the synchronization mechanism of the distributed computing system, and this creates a direct relationship between network performance and compute performance:
Bandwidth matters because GPUs need to exchange enormous volumes of data. Latency matters because communication frequently occurs within the critical path of the workload. Packet loss matters because recovery can interrupt efficient data movement. Congestion matters because synchronized traffic can create hotspots and bursts. And consistency matters because the performance of the slowest participant can affect the larger collective operation.
The issue, then, isn’t simply that AI creates more east-west traffic. It creates high-bandwidth, synchronized, performance-sensitive east-west traffic that is directly tied to the progress of computation.
That’s a very different economic equation when the endpoints are thousands of very expensive accelerators.
The AI Network Is Part of the Computer
This is why large AI factories and neoclouds often have a network specifically dedicated to GPU-to-GPU communication. That backend GPU-to-GPU compute fabric is generally separate from the frontend network used for more conventional services like workload access, API communication, internet connectivity, and other application traffic. Also, storage may have its own network, not to mention multiple service tenants, out-of-band networks, and for neoclouds, multiple third-party customer tenants.
AI architectures can therefore contain several distinct fabrics, with more specialization than we typically see in conventional data centers.
That difference is one of the first major architectural differences between conventional data centers and AI factories.
Scale-Up and Scale-Out
Scale-Up
There’s another networking layer that doesn’t really have an equivalent in conventional data center network architectures. Inside modern GPU systems, accelerators communicate through extremely high-bandwidth scale-up interconnects such as NVIDIA NVLink. With rack-scale systems such as NVL72 architectures, this interconnect can extend across many GPUs within the rack.
Scale-up provides the extremely high-bandwidth, low-latency connectivity GPUs need to communicate with minimal overhead.
Unlike conventional networks, these interconnects eliminate much of the traditional network stack and use mechanisms such as credit-based flow control and link-level error recovery.
As link speeds increase, technologies such as PAM4 signaling, FEC, and signal processing make maintaining predictable, deterministic latency very challenging.
Scale-Out
Scale-out networking connects GPU systems across the larger AI cluster using technologies such as Ethernet with RoCE or InfiniBand. Unlike tightly coupled scale-up interconnects, scale-out designs prioritize scalability and flexibility while still requiring very high bandwidth and consistently low latency.
Because distributed AI workloads frequently exchange and synchronize data, congestion, packet loss, and latency can leave GPUs sitting idle. Technologies like RDMA, ECN, PFC, and adaptive routing help keep communication efficient and predictable across large clusters containing thousands of GPUs.
The important difference is that scale-up and scale-out solve different parts of the same problem. Scale-up connects GPUs as tightly as possible within a system or rack, while scale-out connects those systems across the larger cluster.
Together, they create a hierarchy of GPU connectivity that doesn’t really exist in a conventional data center. An AI factory may therefore have an extremely fast scale-up fabric within GPU systems, a separate scale-out fabric connecting those systems, and additional networks for storage, frontend connectivity, and management.
Rail-Optimized Networks
The physical topology can change as well. In a traditional leaf-spine network, a server may have two NICs connected to a redundant pair of leaf switches. The goal is resilient server connectivity. GPU systems can have many high-speed network interfaces because individual GPUs or groups of GPUs need enormous amounts of bandwidth.
That leads to rail-optimized designs. Imagine every GPU server contains eight GPU-facing network interfaces. Instead of connecting all eight interfaces arbitrarily into the fabric, NIC 0 from every server can connect into one set of leaf switches, NIC 1 into another, and so forth.
Each of these groups forms a rail. The topology can therefore be aligned with how GPUs and collective communication libraries distribute traffic. Physical network design becomes closely connected to the architecture of the compute system itself. That is very different from designing a conventional server access network.
Oversubscription Becomes Too Expensive
Traditional data center networks often use oversubscription because not every endpoint needs maximum bandwidth simultaneously. The exact ratio can vary, but the point is that traditional data center networks can tolerate significant oversubscription because not every endpoint is expected to send data at line rate simultaneously. On the other hand, large AI training fabrics are often designed for 1:1 or near-1:1 bandwidth, minimizing or eliminating oversubscription within the GPU scale-out fabric.
In an AI data center network, a bottleneck doesn’t just slow a few flows. It can leave large numbers of very expensive GPUs waiting idle, which explains the prevalence of non-blocking or near-non-blocking fabrics.
This implies that the economics of traditional data center networks and AI cluster networks are different. In a traditional network, additional switching capacity is an infrastructure expense, but in an AI factory, insufficient network capacity can cost dramatically more in lost revenue as GPUs sit idle.
The NIC Becomes Part of the Architecture
The server NIC is also much more important. Traditional servers commonly use Ethernet NICs primarily to connect the operating system to the network. AI systems often use sophisticated adapters such as NVIDIA ConnectX NICs, BlueField DPUs, and SuperNICs that participate directly in RDMA, congestion management, telemetry, security, isolation, and infrastructure offload.
This creates a more tightly coupled system, so the endpoint and network fabric can’t be treated as completely separate engineering domains. That also means troubleshooting an AI networking problem may require correlating switch telemetry, NIC counters, GPU behavior, collective communication performance, optical health, and workload placement.
Building the Network Changes Too
A conventional data center network can largely be designed as its own infrastructure domain. Design the fabric, deploy the underlay and overlay, configure services and policies, validate it, and connect the servers.
Modern automation improves this process with Infrastructure-as-code, device and config templates, CI/CD pipelines, intent-based systems, and continuous validation, all of which are already well-established practices in conventional data center networking.
An AI factory adds another layer of coordination. GPU topology, NIC placement, switch radix, rail design, optics, cabling, storage architecture, collective communication patterns, power availability, rack density, and network oversubscription all influence one another.
For an AI factory, you can’t design the network independently and then just connect the GPU servers later. The network has to be co-designed with the compute architecture in mind right from the start.
Large AI GPU cluster deployments are also often built around repeatable units of infrastructure. A unit or pod may define a particular number and type of GPU systems along with the switches, ports, optics, cabling, storage connectivity, and power required to support them. Expansion is a matter of deploying additional validated units rather than designing each new portion of the cluster independently.
For neoclouds, the end customer is yet another factor. A neocloud doesn’t just build one large GPU cluster. Just like in a public cloud, it has to repeatedly allocate pieces of that infrastructure to different customers, potentially across Ethernet, InfiniBand, storage, host networking, and newer scale-up domains. Provisioning therefore becomes an ongoing operational workflow rather than a one-time network build project.
Day 2 Operations is Part of GPU Economics
Day 2 operations may be the biggest operational difference. Traditional network operations focus heavily on availability and the network being up. With AI infrastructure, that’s not enough. An AI network can be technically “up” while performing poorly enough to affect the economics of the cluster.
A dirty optical transceiver, degraded link, congestion hotspot, bad traffic distribution, incorrectly configured RoCE parameters, or poorly performing rail can reduce collective communication performance without producing a traditional outage.
This makes performance observability much more important. Operators need visibility into switches and interfaces, but they also need visibility potentially into SuperNICs, DPUs, GPUs, collective communication behavior, queue utilization, congestion events, optical performance, and the relationship between network behavior and workload.
Traditional polling intervals may also miss transient events. A congestion event lasting milliseconds or microseconds would probably disappear long before the next SNMP polling cycle even though it affected an AI workload. Traditional polling can easily miss these transient conditions, which increases the importance of streaming telemetry, hardware counters, queue and congestion telemetry, and other high-resolution observability mechanisms.
Therefore, operations begin to look less like traditional device management and more like operating a distributed computing fabric. So instead of asking if the network is up and working, operators ask if the network is allowing GPU clusters to perform as designed.
The Same Foundations but a Different Engineering Problem
AI factories don’t invalidate decades of data center networking knowledge. BGP still matters as does Clos designs, ECMP, cabling discipline, routing domains, automation, management networks, EVPN-VXLAN, and so on.
The real difference is that the optimization target has changed. A traditional data center network is primarily designed to provide scalable, resilient connectivity between computing systems, whereas an AI factory network is designed to make thousands of GPUs behave like one computing system.
That changes how much bandwidth we provision, how we build the topology, how endpoints participate in the network, how we handle congestion, how we isolate tenants, how we monitor performance, and how the infrastructure is deployed and operated.
For network engineers, that’s probably the most useful way to understand the transition. We’re not throwing away traditional data center networking. Instead, we’re taking many of the same technologies and engineering principles and applying them to a workload where the network itself has become part of the computer.
Thanks,
Phil





Leave a comment