Three DGX Sparks in a ring, and the two places NVIDIA's guide fights your LAN
A practical field guide to adapting NVIDIA's DGX Spark ring setup to a real office network
Three MSI EdgeXpert boxes (a DGX Spark with a different badge), three QSFP cables, and a playbook whose header says "1 HR". It took two days. Both days went to the same root cause: the guide assumes your network does not look like everybody's network.
This is the guide as I actually ran it, from Step 1 to the number at the end, with the two places where it broke on an ordinary office LAN and what fixed them. NVIDIA's original is Connect Three DGX Spark in a Ring Topology; the patched script lives on my branch.
The setup
Hardware
- 3ร MSI EdgeXpert, NVIDIA GB10, hostnames
edgexpert-94af(node 1),edgexpert-a71b(node 2),edgexpert-a724(node 3) - One ConnectX-7 per box with two QSFP ports, 200 GbE each. Each physical port shows up as two logical interfaces in Linux (
enp1s0f0np0+enP2p1s0f0np0for Port0,enp1s0f1np1+enP2p1s0f1np1for Port1), so a ring needs four IPs per node - The 10 GbE RJ45 port (
enP7s7) is not cabled. Remember that; it matters in Step 6 - A MacBook Pro on macOS 26.6, with three tmux sessions, each an SSH shell into one node
Software
- Ubuntu 24.04.4 LTS, aarch64
- Open MPI 4.1.6 from apt (
libopenmpi-dev) - NCCL: for a three-node ring the setup script builds a ring-specific fork,
github.com/zyang-dev/ncclbranchdgxspark-3node-ring, plus NVIDIA'snccl-testswithMPI=1 - Tailscale on the Mac and on all three Sparks
Network
The office and the machine room have separate network uplinks and physically separate LANs. The Mac is on the office LAN at 192.168.0.92/24, with a router at 192.168.0.1. The Sparks are on the machine-room LAN, where node 1's Wi-Fi address is 192.168.0.81/24 and its router is also 192.168.0.1. Those matching private ranges don't make them the same LAN; Tailscale is the path I use between them.
Keep the Sparks' subnet and gateway in mind, because NVIDIA's ring addressing plan is about to hand them out again on the same machine.
Tailscale mesh (100.64.0.0/10): how the Mac reaches everything
โ
LAN 0: office โ LAN 1: machine room, Wi-Fi 192.168.0.0/24
โโโโโโโโโโโโโโโ โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ MacBook Pro โโโโโโโโโโ โ โ
โ tmux 0/1/2 โ โ node 1 edgexpert-94af โ
โโโโโโโโโโโโโโโ โ wifi 192.168.0.81 ts 100.87.56.74 โ
โ Port0 โ โ Port1 โ
โ โ โ โ
โ link 0 (200 GbE) link 1 (200 GbE) โ
โ 192.168.200/201.x 192.168.202/203.x โ
โ โ โ โ
โ Port1 โ โ Port0 โ
โ node 2 edgexpert-a71b node 3 edgexpert-a724โ
โ ts 100.124.39.67 ts 100.106.52.66 โ
โ Port0 โโโโโ link 2 โโโโโโ Port1 โ
โ 192.168.204/205.x โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
Two networks matter here and they must not be confused:
- Management: the Tailscale addresses (
100.87.56.74,100.124.39.67,100.106.52.66). Every node can reach every other node on it. This is what goes into the script's config file and whatmpirun -Huses. - Fabric: the three QSFP links. Each link is its own point-to-point subnet. Node 3 cannot reach node 1's link-0 address, and it never should; that is what a ring is.
Step 1. Same username
whoami on all three boxes. Ours is aiplux everywhere, so nothing to do. If they differ, the guide's useradd/usermod -aG sudo recipe is fine.
Step 2. Cabling
Port0 is the QSFP port next to the RJ45, Port1 is the far one. The ring is:
| Cable | From | To |
|---|---|---|
| link 0 | node 1 Port0 | node 2 Port1 |
| link 2 | node 2 Port0 | node 3 Port1 |
| link 1 | node 3 Port0 | node 1 Port1 |
ibdev2netdev should list all four roce* devices as (Up) on every node. The setup script checks the same thing itself and prints:
Checking UP CX7 interfaces...
Found UP CX7 interfaces ['enp1s0f0np0', 'enP2p1s0f0np0', 'enp1s0f1np1', 'enP2p1s0f1np1'] on 100.87.56.74. Checking other nodes...
Checking CX7 interface link speed...
It also refuses to continue unless ethtool reports 200000 (200 Gb/s) on each interface. If a cable is seated badly this is where you find out.
Step 3. Network interface configuration, and trap number one
The guide offers two options: run NVIDIA's spark_cluster_setup script (recommended), or write netplan files by hand. Both hand out the same addresses, and this is the plan for node 1, copied from the guide:
network:
version: 2
ethernets:
enp1s0f0np0:
dhcp4: false
addresses:
- 192.168.0.1/24
enP2p1s0f0np0:
dhcp4: false
addresses:
- 192.168.1.1/24
enp1s0f1np1:
dhcp4: false
addresses:
- 192.168.2.1/24
enP2p1s0f1np1:
dhcp4: false
addresses:
- 192.168.3.1/24
Node 2 gets 192.168.4.1, 192.168.5.1, 192.168.0.2, 192.168.1.2; node 3 gets 192.168.2.2, 192.168.3.2, 192.168.4.2, 192.168.5.2. Six /24s, 192.168.0.0 through 192.168.5.0. The script's ip_for_3node_ring_link() computed exactly the same thing: 192.168.{link_index * 2 + local_index_in_pair}.{node_id}/24.
Look at node 1's first line again: 192.168.0.1/24. That is our Wi-Fi router. Here is what the kernel does with it:
- Before the change, node 1 has
192.168.0.81/24onwlP9s9anddefault via 192.168.0.1 dev wlP9s9 metric 600. Every host on the LAN, and the router itself, is reached through Wi-Fi. netplan applyputs192.168.0.1/24onenp1s0f0np0. The kernel now has two routes for192.168.0.0/24: Wi-Fi at metric 600 and the QSFP cable at metric 102. Lower wins. Anything addressed to a host on your LAN, an SSH session from a laptop on that LAN included, is now sent down the cable to node 2.192.168.0.1is also one of node 1's own addresses now. Packets for your router are delivered to node 1 itself, and node 1 answers ARP for the router's address on the cable.- Node 2 gets
192.168.0.2/24on its end of the same cable: the same two routes, the same misrouting, plus an address collision with whatever your DHCP server handed.2to.
On September 8, the symptom was straightforward: the network connection dropped. The fabric configuration overlapped the subnet used by the existing Wi-Fi/AP network. On node 1, it also assigned the gateway's address to a local interface. The conflict was between interfaces on the Spark, not between the office and machine-room LANs.
The problem starts in step 2. netplan apply can return cleanly while the new routes break access to the existing network.
The fix is a third octet nobody else uses. In node_scripts/detect_and_configure_cluster_networking.py:
# Keep the CX7 fabric away from the common 192.168.0.0/24 LAN subnet.
# Each logical interface gets its own /24; link 0 uses .200/.201 and
# link 1 uses .202/.203.
FABRIC_SUBNET_BASE_OCTET = 200
def ip_for_3node_ring_link(link_index: int, node_id: int, local_index_in_pair: int) -> str:
subnet_octet = FABRIC_SUBNET_BASE_OCTET + link_index * 2 + local_index_in_pair
return f"192.168.{subnet_octet}.{node_id}/24"
Which gives this address plan:
| Link | Interfaces | node 1 | node 2 | node 3 |
|---|---|---|---|---|
| link 0 (1โ2) | Port0 on node 1, Port1 on node 2 | 192.168.200.1, 192.168.201.1 | 192.168.200.2, 192.168.201.2 | |
| link 1 (1โ3) | Port1 on node 1, Port0 on node 3 | 192.168.202.1, 192.168.203.1 | 192.168.202.2, 192.168.203.2 | |
| link 2 (2โ3) | Port0 on node 2, Port1 on node 3 | 192.168.204.1, 192.168.205.1 | 192.168.204.2, 192.168.205.2 |
If you write netplan by hand, do the same: keep the guide's files, change the third octet to something your LAN will never use. 192.168.200 to 192.168.205 was ours.
With that patched, the script run is the one the guide describes. Config file first, with the management addresses (Tailscale, in our case), not anything on the fabric:
{
"nodes_info": [
{"ip_address": "100.87.56.74", "port": 22, "user": "aiplux", "password": "..."},
{"ip_address": "100.124.39.67", "port": 22, "user": "aiplux", "password": "..."},
{"ip_address": "100.106.52.66", "port": 22, "user": "aiplux", "password": "..."}
]
}
Then, from node 1:
cd dgx-spark-playbooks/nvidia/multi-sparks-through-switch/assets/spark_cluster_setup
bash spark_cluster_setup.sh -c config/spark_config_ring.json --run-setup
The wrapper creates a venv, installs paramiko and scp, and runs spark_cluster_setup.py, which SSHes into every node with the password from the JSON. The topology detection is a nice piece of work: each node broadcasts custom Ethernet frames (EtherType 0x88B5) out of both CX7 ports for 20 seconds, learns which MAC answers on which port, and reports to node 1 on port 9999. Three machines, each seeing a different neighbor on each port, means ring3, and the netplan files are generated and applied. Then it waits 10 seconds (longer on retries), verifies every interface has exactly one address, and pings every neighbor across every link.
Steps 4 and 5. SSH and hostname checks
The script does these for you. It generates ~/.ssh/id_ed25519_shared, copies it to every node, appends the public key to each authorized_keys, and adds a Host * / IdentityFile ~/.ssh/id_ed25519_shared block to each node's SSH config:
Generating shared SSH key for all nodes...
Setting up shared SSH access across all nodes...
Configuring shared SSH key on node 100.87.56.74...
Successfully configured 100.87.56.74 with shared key
Configuring shared SSH key on node 100.124.39.67...
Successfully configured 100.124.39.67 with shared key
Configuring shared SSH key on node 100.106.52.66...
Successfully configured 100.106.52.66 with shared key
Shared SSH keys configured successfully.
Spark cluster setup completed successfully.
Nothing broke here. Moving on.
Step 6. NCCL test, and trap number two
The same script continues into the NCCL test: apt install libopenmpi-dev on every node, clone and build the NCCL ring fork with NVCC_GENCODE="-gencode=arch=compute_121,code=sm_121", clone and build nccl-tests, then run a 16 GiB all_gather_perf across the three nodes. The builds take a few minutes and finish fine. Then this:
Successfully setup NCCL dependencies on all nodes...
Running NCCL test...
NCCL test command: export CUDA_HOME="/usr/local/cuda" && export MPI_HOME="/usr/lib/aarch64-linux-gnu/openmpi" && export NCCL_HOME="$HOME/nccl_spark_cluster/build/" && export LD_LIBRARY_PATH="$NCCL_HOME/lib:$CUDA_HOME/lib64/:$MPI_HOME/lib:$LD_LIBRARY_PATH" && mpirun -np 3 -H 100.87.56.74:1,100.124.39.67:1,100.106.52.66:1 --mca plm_rsh_agent "ssh -o UserKnownHostsFile=/dev/null -o StrictHostKeyChecking=no" -x LD_LIBRARY_PATH=$LD_LIBRARY_PATH -x UCX_NET_DEVICES=enp1s0f0np0,enP2p1s0f0np0,enp1s0f1np1,enP2p1s0f1np1 -x NCCL_SOCKET_IFNAME=enp1s0f0np0,enP2p1s0f0np0,enp1s0f1np1,enP2p1s0f1np1 -x OMPI_MCA_btl_tcp_if_include=enp1s0f0np0,enP2p1s0f0np0,enp1s0f1np1,enP2p1s0f1np1 -x NCCL_IB_HCA=rocep1s0f0,rocep1s0f1,roceP2p1s0f0,roceP2p1s0f1 -x NCCL_IB_SUBNET_AWARE_ROUTING=1 -x NCCL_IB_MERGE_NICS=0 -x NCCL_NET_PLUGIN=none $HOME/nccl-tests_spark_cluster/build/all_gather_perf -b 16G -e 16G -f 2
^CTraceback (most recent call last):
File ".../spark_cluster_setup.py", line 911, in <module>
main()
File ".../spark_cluster_setup.py", line 903, in main
if not run_nccl_test(config.get("nodes_info", []), ring_topology, up_interfaces):
File ".../spark_cluster_setup.py", line 282, in run_nccl_test
exit_code, output, error = paramiko_run_command_with_output(ssh, mpirun_cmd)
File ".../spark_cluster_setup.py", line 94, in paramiko_run_command_with_output
time.sleep(0.1)
KeyboardInterrupt
It sits there until you press Ctrl-C. No output, no error, no timeout.
A word on that command line first. NVIDIA's stock script does not use the CX7 names there. It pins UCX_NET_DEVICES, NCCL_SOCKET_IFNAME and OMPI_MCA_btl_tcp_if_include to enP7s7, the 10 GbE RJ45 port, on the assumption that every Spark is plugged into the same management switch through it. Ours is not cabled; we manage the boxes over Tailscale. So an earlier patch on the branch swapped in the CX7 interfaces, first one, then all four. That is what you see above, and it is the wrong direction. To see why, here is what mpirun does before a single byte of NCCL runs:
mpirunon node 1 (the head node process) opens a listening TCP port and builds a URI with every IPv4 address it owns, in interface order.- It SSHes to the other two nodes and starts an
orteddaemon on each, handing it that URI. - Each
ortedwalks the address list and callsconnect()on them one at a time until one answers. Only then are the ranks launched. - The ranks run
MPI_Init, exchange the NCCL unique id over Open MPI's TCP BTL, NCCL bootstraps its own sockets onNCCL_SOCKET_IFNAME, and only after all that does the 16 GiB all-gather go over RoCE onNCCL_IB_HCA.
The problem lies in step 3. With --mca oob_base_verbose 5 you can watch it happen. This is the list node 1 advertised:
[edgexpert-a724:312632] [[39497,0],2] oob:tcp: working peer [[39497,0],0] address tcp://127.0.0.1,192.168.200.1,192.168.202.1,192.168.201.1,192.168.203.1,192.168.0.81,100.87.56.74,172.17.0.1:48683
Loopback, then four fabric addresses, then Wi-Fi, then Tailscale, then docker0. Node 2 shares link 0 with node 1, so its second attempt lands:
[edgexpert-94af:308572] [[39497,0],0] connection_handler: working connection (27, 0) 192.168.200.2:60452
Node 3 does not share link 0. It has no route to 192.168.200.0/24, so the SYN goes out its default route, to the Wi-Fi router, and dies quietly. Open MPI waits for the connect timeout, retries, waits again. Run by hand, it prints this within a minute and keeps waiting; inside the setup script you see nothing at all, because the script only shows a command's output after the command exits:
------------------------------------------------------------
A process or daemon was unable to complete a TCP connection
to another process:
Local host: edgexpert-a724
Remote host: edgexpert-94af
This is usually caused by a firewall on the remote host. Please
check that any firewall (e.g., iptables) has been disabled and
try again.
------------------------------------------------------------
It is not a firewall. Do not go looking for one.
Two more things you will see and can ignore. Lines saying Authorization required, but no authorization protocol specified are SSH's X11 forwarding complaining; they have nothing to do with the hang. And a bare mpirun -np 3 -H ... hostname hangs the same way, which is the fastest way to prove the problem is Open MPI's bootstrap and not NCCL.
Dead ends, in the order I tried them
--mca oob_tcp_if_include 100.87.56.74/32,100.124.39.67/32,100.106.52.66/32, one /32 per node. Silently ignored on Open MPI 4.1.6. The advertised list did not change at all.- Pinning only the daemon channel:
--mca oob_tcp_if_include tailscale0while leaving the four CX7 names inOMPI_MCA_btl_tcp_if_include.hostnamenow works, but the NCCL test dies in step 4, because the rank-to-rank TCP layer plays the same game with the same subnets:
WARNING: Open MPI failed to TCP connect to a peer MPI process. This
should not happen.
Your Open MPI job may now hang or fail.
Local host: edgexpert-a724
PID: 313704
Message: connect() to 192.168.204.1:1026 failed
Error: Operation now in progress (115)
The fix
Everything that is "bootstrap" rides the management network. Only NCCL_IB_HCA stays on the fabric. Both tailscale0 and the CIDR 100.64.0.0/10 worked for oob_tcp_if_include; NCCL_SOCKET_IFNAME wants an interface name anyway, so the script uses the name everywhere.
export OMPI_MCA_plm_rsh_agent="ssh -o UserKnownHostsFile=/dev/null -o StrictHostKeyChecking=no"
M=tailscale0
mpirun -np 3 -H 100.87.56.74:1,100.124.39.67:1,100.106.52.66:1 \
--mca oob_tcp_if_include $M \
-x LD_LIBRARY_PATH \
-x UCX_NET_DEVICES=$M -x NCCL_SOCKET_IFNAME=$M -x OMPI_MCA_btl_tcp_if_include=$M \
-x NCCL_IB_HCA=rocep1s0f0,rocep1s0f1,roceP2p1s0f0,roceP2p1s0f1 \
-x NCCL_IB_SUBNET_AWARE_ROUTING=1 -x NCCL_IB_MERGE_NICS=0 -x NCCL_NET_PLUGIN=none \
$HOME/nccl-tests_spark_cluster/build/all_gather_perf -b 16G -e 16G -f 2
In the script, run_nccl_test() no longer hardcodes anything. It asks node 1 which interface owns the management IP from the config file and uses that:
cmd = f"ip -o -4 addr show | awk '$4 ~ /^{node0['ip_address']}\\// {{print $2}}'"
exit_code, output, error = paramiko_run_command_with_output(ssh, cmd)
mgmt_iface = output.split()[0] if output.split() else ""
...
f"--mca oob_tcp_if_include {mgmt_iface} "
f"-x UCX_NET_DEVICES={mgmt_iface} "
f"-x NCCL_SOCKET_IFNAME={mgmt_iface} "
f"-x OMPI_MCA_btl_tcp_if_include={mgmt_iface} "
On a stock Spark managed over the RJ45 port, that resolves to enP7s7 and behaves exactly like NVIDIA's version. On ours it resolves to tailscale0.
The result
A rerun of just the test:
bash spark_cluster_setup.sh -c config/spark_config_ring.json --run-nccl-test
Using management interface tailscale0 for MPI/NCCL bootstrap
...
Avg bus bandwidth from NCCL test: 22.732 GB/s
NCCL test BW is as expected
NCCL test completed.
And the raw nccl-tests table from a manual run with the same flags:
# Rank 2 Group 0 Pid 314342 on edgexpert-a724 device 0 [000f:01:00] NVIDIA GB10
#
# out-of-place in-place
# size count type redop root time algbw busbw #wrong time algbw busbw #wrong
# (B) (elements) (us) (GB/s) (GB/s) (us) (GB/s) (GB/s)
17179869168 1431655764 float none -1 510560 33.65 22.43 0 502258 34.21 22.80 0
# Out of bounds values : 0 OK
# Avg bus bandwidth : 22.6181
A 16 GiB all-gather across three nodes in about 510 ms. The script's pass mark for a ring is 10 GB/s of bus bandwidth (80 Gb/s); the switch topology is held to 21.875 GB/s (175 Gb/s), and the ring cleared that bar too. Note that the script builds NVIDIA's ring-specific NCCL fork for this layout; the switch topology gets stock NCCL 2.28.
After the fixes, the 200 GbE fabric worked normally again. The successful script run above measured 22.732 GB/s of NCCL bus bandwidth, about 182 Gb/s. The 200 Gb/s link rate and the NCCL bus-bandwidth measurement describe different things; this isn't a measured before-and-after speedup from the broken setup.
Fun fact: the whole of day two was Claude Code driving my tmux session over send-keys and capture-pane, because the Mac has no SSH key for the Sparks and two of the three sessions were busy pulling a 360 GB model. It only ever typed into session 0.
What to keep
- Pick fabric subnets your LAN cannot collide with before you run anything.
192.168.0.0/24and192.168.1.0/24are what most routers ship with, and NVIDIA's plan uses both. We use192.168.200.0/24through192.168.205.0/24. - The MPI bootstrap needs a network every node can reach. In a ring, the fabric is not that network; the management network is. Point
oob_tcp_if_include,btl_tcp_if_include,NCCL_SOCKET_IFNAMEandUCX_NET_DEVICESthere, and leave onlyNCCL_IB_HCAon the CX7 devices. - When
mpirunhangs, run it withhostnameinstead of your program and add--mca oob_base_verbose 5. Read the address list it advertises. The firewall message at the end lies. - Per-host
/32entries inoob_tcp_if_includedo nothing on Open MPI 4.1.6. Use an interface name, or a CIDR wide enough to match (100.64.0.0/10for Tailscale). - Loopback and
docker0are in every advertised list. Restricting the interface is the fix; deleting Docker is not.
That's it. Nearly every home router ships answering to 192.168.0.1 or 192.168.1.1, which is why any guide that hands those addresses to a network card will eventually meet someone's gateway.