Three DGX Sparks in a ring, and the two places NVIDIA's guide fights your LAN

A practical field guide to adapting NVIDIA's DGX Spark ring setup to a real office network

<!-- 2026-09-09 ยท tags: dgx-spark, connectx-7, nccl, open-mpi, netplan, tailscale, ring-topology -->

Three MSI EdgeXpert boxes (a DGX Spark with a different badge), three QSFP cables, and a playbook whose header says "1 HR". It took two days. Both days went to the same root cause: the guide assumes your network does not look like everybody's network.

This is the guide as I actually ran it, from Step 1 to the number at the end, with the two places where it broke on an ordinary office LAN and what fixed them. NVIDIA's original is Connect Three DGX Spark in a Ring Topology; the patched script lives on my branch.

The setup

Hardware

  • 3ร— MSI EdgeXpert, NVIDIA GB10, hostnames edgexpert-94af (node 1), edgexpert-a71b (node 2), edgexpert-a724 (node 3)
  • One ConnectX-7 per box with two QSFP ports, 200 GbE each. Each physical port shows up as two logical interfaces in Linux (enp1s0f0np0 + enP2p1s0f0np0 for Port0, enp1s0f1np1 + enP2p1s0f1np1 for Port1), so a ring needs four IPs per node
  • The 10 GbE RJ45 port (enP7s7) is not cabled. Remember that; it matters in Step 6
  • A MacBook Pro on macOS 26.6, with three tmux sessions, each an SSH shell into one node

Software

  • Ubuntu 24.04.4 LTS, aarch64
  • Open MPI 4.1.6 from apt (libopenmpi-dev)
  • NCCL: for a three-node ring the setup script builds a ring-specific fork, github.com/zyang-dev/nccl branch dgxspark-3node-ring, plus NVIDIA's nccl-tests with MPI=1
  • Tailscale on the Mac and on all three Sparks

Network

The office and the machine room have separate network uplinks and physically separate LANs. The Mac is on the office LAN at 192.168.0.92/24, with a router at 192.168.0.1. The Sparks are on the machine-room LAN, where node 1's Wi-Fi address is 192.168.0.81/24 and its router is also 192.168.0.1. Those matching private ranges don't make them the same LAN; Tailscale is the path I use between them.

Keep the Sparks' subnet and gateway in mind, because NVIDIA's ring addressing plan is about to hand them out again on the same machine.

              Tailscale mesh (100.64.0.0/10): how the Mac reaches everything
                        โ”‚
   LAN 0: office        โ”‚        LAN 1: machine room, Wi-Fi 192.168.0.0/24
 โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”        โ”‚      โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
 โ”‚ MacBook Pro โ”‚โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜      โ”‚                                                  โ”‚
 โ”‚ tmux 0/1/2  โ”‚               โ”‚           node 1  edgexpert-94af                 โ”‚
 โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜               โ”‚           wifi 192.168.0.81  ts 100.87.56.74     โ”‚
                               โ”‚          Port0 โ—                 โ— Port1         โ”‚
                               โ”‚                โ”‚                 โ”‚               โ”‚
                               โ”‚     link 0 (200 GbE)      link 1 (200 GbE)       โ”‚
                               โ”‚     192.168.200/201.x     192.168.202/203.x      โ”‚
                               โ”‚                โ”‚                 โ”‚               โ”‚
                               โ”‚          Port1 โ—                 โ— Port0         โ”‚
                               โ”‚   node 2  edgexpert-a71b   node 3  edgexpert-a724โ”‚
                               โ”‚   ts 100.124.39.67         ts 100.106.52.66      โ”‚
                               โ”‚          Port0 โ—โ”€โ”€โ”€โ”€ link 2 โ”€โ”€โ”€โ”€โ”€โ— Port1         โ”‚
                               โ”‚               192.168.204/205.x                  โ”‚
                               โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Two networks matter here and they must not be confused:

  • Management: the Tailscale addresses (100.87.56.74, 100.124.39.67, 100.106.52.66). Every node can reach every other node on it. This is what goes into the script's config file and what mpirun -H uses.
  • Fabric: the three QSFP links. Each link is its own point-to-point subnet. Node 3 cannot reach node 1's link-0 address, and it never should; that is what a ring is.

Step 1. Same username

whoami on all three boxes. Ours is aiplux everywhere, so nothing to do. If they differ, the guide's useradd/usermod -aG sudo recipe is fine.

Step 2. Cabling

Port0 is the QSFP port next to the RJ45, Port1 is the far one. The ring is:

CableFromTo
link 0node 1 Port0node 2 Port1
link 2node 2 Port0node 3 Port1
link 1node 3 Port0node 1 Port1

ibdev2netdev should list all four roce* devices as (Up) on every node. The setup script checks the same thing itself and prints:

Checking UP CX7 interfaces...
Found UP CX7 interfaces ['enp1s0f0np0', 'enP2p1s0f0np0', 'enp1s0f1np1', 'enP2p1s0f1np1'] on 100.87.56.74. Checking other nodes...
Checking CX7 interface link speed...

It also refuses to continue unless ethtool reports 200000 (200 Gb/s) on each interface. If a cable is seated badly this is where you find out.

Step 3. Network interface configuration, and trap number one

The guide offers two options: run NVIDIA's spark_cluster_setup script (recommended), or write netplan files by hand. Both hand out the same addresses, and this is the plan for node 1, copied from the guide:

network:
  version: 2
  ethernets:
    enp1s0f0np0:
      dhcp4: false
      addresses:
        - 192.168.0.1/24
    enP2p1s0f0np0:
      dhcp4: false
      addresses:
        - 192.168.1.1/24
    enp1s0f1np1:
      dhcp4: false
      addresses:
        - 192.168.2.1/24
    enP2p1s0f1np1:
      dhcp4: false
      addresses:
        - 192.168.3.1/24

Node 2 gets 192.168.4.1, 192.168.5.1, 192.168.0.2, 192.168.1.2; node 3 gets 192.168.2.2, 192.168.3.2, 192.168.4.2, 192.168.5.2. Six /24s, 192.168.0.0 through 192.168.5.0. The script's ip_for_3node_ring_link() computed exactly the same thing: 192.168.{link_index * 2 + local_index_in_pair}.{node_id}/24.

Look at node 1's first line again: 192.168.0.1/24. That is our Wi-Fi router. Here is what the kernel does with it:

  1. Before the change, node 1 has 192.168.0.81/24 on wlP9s9 and default via 192.168.0.1 dev wlP9s9 metric 600. Every host on the LAN, and the router itself, is reached through Wi-Fi.
  2. netplan apply puts 192.168.0.1/24 on enp1s0f0np0. The kernel now has two routes for 192.168.0.0/24: Wi-Fi at metric 600 and the QSFP cable at metric 102. Lower wins. Anything addressed to a host on your LAN, an SSH session from a laptop on that LAN included, is now sent down the cable to node 2.
  3. 192.168.0.1 is also one of node 1's own addresses now. Packets for your router are delivered to node 1 itself, and node 1 answers ARP for the router's address on the cable.
  4. Node 2 gets 192.168.0.2/24 on its end of the same cable: the same two routes, the same misrouting, plus an address collision with whatever your DHCP server handed .2 to.

On September 8, the symptom was straightforward: the network connection dropped. The fabric configuration overlapped the subnet used by the existing Wi-Fi/AP network. On node 1, it also assigned the gateway's address to a local interface. The conflict was between interfaces on the Spark, not between the office and machine-room LANs.

The problem starts in step 2. netplan apply can return cleanly while the new routes break access to the existing network.

The fix is a third octet nobody else uses. In node_scripts/detect_and_configure_cluster_networking.py:

# Keep the CX7 fabric away from the common 192.168.0.0/24 LAN subnet.
# Each logical interface gets its own /24; link 0 uses .200/.201 and
# link 1 uses .202/.203.
FABRIC_SUBNET_BASE_OCTET = 200

def ip_for_3node_ring_link(link_index: int, node_id: int, local_index_in_pair: int) -> str:
    subnet_octet = FABRIC_SUBNET_BASE_OCTET + link_index * 2 + local_index_in_pair
    return f"192.168.{subnet_octet}.{node_id}/24"

Which gives this address plan:

LinkInterfacesnode 1node 2node 3
link 0 (1โ†”2)Port0 on node 1, Port1 on node 2192.168.200.1, 192.168.201.1192.168.200.2, 192.168.201.2
link 1 (1โ†”3)Port1 on node 1, Port0 on node 3192.168.202.1, 192.168.203.1192.168.202.2, 192.168.203.2
link 2 (2โ†”3)Port0 on node 2, Port1 on node 3192.168.204.1, 192.168.205.1192.168.204.2, 192.168.205.2

If you write netplan by hand, do the same: keep the guide's files, change the third octet to something your LAN will never use. 192.168.200 to 192.168.205 was ours.

With that patched, the script run is the one the guide describes. Config file first, with the management addresses (Tailscale, in our case), not anything on the fabric:

{
    "nodes_info": [
        {"ip_address": "100.87.56.74",  "port": 22, "user": "aiplux", "password": "..."},
        {"ip_address": "100.124.39.67", "port": 22, "user": "aiplux", "password": "..."},
        {"ip_address": "100.106.52.66", "port": 22, "user": "aiplux", "password": "..."}
    ]
}

Then, from node 1:

cd dgx-spark-playbooks/nvidia/multi-sparks-through-switch/assets/spark_cluster_setup
bash spark_cluster_setup.sh -c config/spark_config_ring.json --run-setup

The wrapper creates a venv, installs paramiko and scp, and runs spark_cluster_setup.py, which SSHes into every node with the password from the JSON. The topology detection is a nice piece of work: each node broadcasts custom Ethernet frames (EtherType 0x88B5) out of both CX7 ports for 20 seconds, learns which MAC answers on which port, and reports to node 1 on port 9999. Three machines, each seeing a different neighbor on each port, means ring3, and the netplan files are generated and applied. Then it waits 10 seconds (longer on retries), verifies every interface has exactly one address, and pings every neighbor across every link.

Steps 4 and 5. SSH and hostname checks

The script does these for you. It generates ~/.ssh/id_ed25519_shared, copies it to every node, appends the public key to each authorized_keys, and adds a Host * / IdentityFile ~/.ssh/id_ed25519_shared block to each node's SSH config:

Generating shared SSH key for all nodes...
Setting up shared SSH access across all nodes...
Configuring shared SSH key on node 100.87.56.74...
Successfully configured 100.87.56.74 with shared key
Configuring shared SSH key on node 100.124.39.67...
Successfully configured 100.124.39.67 with shared key
Configuring shared SSH key on node 100.106.52.66...
Successfully configured 100.106.52.66 with shared key
Shared SSH keys configured successfully.
Spark cluster setup completed successfully.

Nothing broke here. Moving on.

Step 6. NCCL test, and trap number two

The same script continues into the NCCL test: apt install libopenmpi-dev on every node, clone and build the NCCL ring fork with NVCC_GENCODE="-gencode=arch=compute_121,code=sm_121", clone and build nccl-tests, then run a 16 GiB all_gather_perf across the three nodes. The builds take a few minutes and finish fine. Then this:

Successfully setup NCCL dependencies on all nodes...
Running NCCL test...
NCCL test command: export CUDA_HOME="/usr/local/cuda" && export MPI_HOME="/usr/lib/aarch64-linux-gnu/openmpi" && export NCCL_HOME="$HOME/nccl_spark_cluster/build/" && export LD_LIBRARY_PATH="$NCCL_HOME/lib:$CUDA_HOME/lib64/:$MPI_HOME/lib:$LD_LIBRARY_PATH"  && mpirun -np 3 -H 100.87.56.74:1,100.124.39.67:1,100.106.52.66:1 --mca plm_rsh_agent "ssh -o UserKnownHostsFile=/dev/null -o StrictHostKeyChecking=no" -x LD_LIBRARY_PATH=$LD_LIBRARY_PATH -x UCX_NET_DEVICES=enp1s0f0np0,enP2p1s0f0np0,enp1s0f1np1,enP2p1s0f1np1 -x NCCL_SOCKET_IFNAME=enp1s0f0np0,enP2p1s0f0np0,enp1s0f1np1,enP2p1s0f1np1 -x OMPI_MCA_btl_tcp_if_include=enp1s0f0np0,enP2p1s0f0np0,enp1s0f1np1,enP2p1s0f1np1 -x NCCL_IB_HCA=rocep1s0f0,rocep1s0f1,roceP2p1s0f0,roceP2p1s0f1 -x NCCL_IB_SUBNET_AWARE_ROUTING=1 -x NCCL_IB_MERGE_NICS=0 -x NCCL_NET_PLUGIN=none $HOME/nccl-tests_spark_cluster/build/all_gather_perf -b 16G -e 16G -f 2
^CTraceback (most recent call last):
  File ".../spark_cluster_setup.py", line 911, in <module>
    main()
  File ".../spark_cluster_setup.py", line 903, in main
    if not run_nccl_test(config.get("nodes_info", []), ring_topology, up_interfaces):
  File ".../spark_cluster_setup.py", line 282, in run_nccl_test
    exit_code, output, error = paramiko_run_command_with_output(ssh, mpirun_cmd)
  File ".../spark_cluster_setup.py", line 94, in paramiko_run_command_with_output
    time.sleep(0.1)
KeyboardInterrupt

It sits there until you press Ctrl-C. No output, no error, no timeout.

A word on that command line first. NVIDIA's stock script does not use the CX7 names there. It pins UCX_NET_DEVICES, NCCL_SOCKET_IFNAME and OMPI_MCA_btl_tcp_if_include to enP7s7, the 10 GbE RJ45 port, on the assumption that every Spark is plugged into the same management switch through it. Ours is not cabled; we manage the boxes over Tailscale. So an earlier patch on the branch swapped in the CX7 interfaces, first one, then all four. That is what you see above, and it is the wrong direction. To see why, here is what mpirun does before a single byte of NCCL runs:

  1. mpirun on node 1 (the head node process) opens a listening TCP port and builds a URI with every IPv4 address it owns, in interface order.
  2. It SSHes to the other two nodes and starts an orted daemon on each, handing it that URI.
  3. Each orted walks the address list and calls connect() on them one at a time until one answers. Only then are the ranks launched.
  4. The ranks run MPI_Init, exchange the NCCL unique id over Open MPI's TCP BTL, NCCL bootstraps its own sockets on NCCL_SOCKET_IFNAME, and only after all that does the 16 GiB all-gather go over RoCE on NCCL_IB_HCA.

The problem lies in step 3. With --mca oob_base_verbose 5 you can watch it happen. This is the list node 1 advertised:

[edgexpert-a724:312632] [[39497,0],2] oob:tcp: working peer [[39497,0],0] address tcp://127.0.0.1,192.168.200.1,192.168.202.1,192.168.201.1,192.168.203.1,192.168.0.81,100.87.56.74,172.17.0.1:48683

Loopback, then four fabric addresses, then Wi-Fi, then Tailscale, then docker0. Node 2 shares link 0 with node 1, so its second attempt lands:

[edgexpert-94af:308572] [[39497,0],0] connection_handler: working connection (27, 0) 192.168.200.2:60452

Node 3 does not share link 0. It has no route to 192.168.200.0/24, so the SYN goes out its default route, to the Wi-Fi router, and dies quietly. Open MPI waits for the connect timeout, retries, waits again. Run by hand, it prints this within a minute and keeps waiting; inside the setup script you see nothing at all, because the script only shows a command's output after the command exits:

------------------------------------------------------------
A process or daemon was unable to complete a TCP connection
to another process:
  Local host:    edgexpert-a724
  Remote host:   edgexpert-94af
This is usually caused by a firewall on the remote host. Please
check that any firewall (e.g., iptables) has been disabled and
try again.
------------------------------------------------------------

It is not a firewall. Do not go looking for one.

Two more things you will see and can ignore. Lines saying Authorization required, but no authorization protocol specified are SSH's X11 forwarding complaining; they have nothing to do with the hang. And a bare mpirun -np 3 -H ... hostname hangs the same way, which is the fastest way to prove the problem is Open MPI's bootstrap and not NCCL.

Dead ends, in the order I tried them

  • --mca oob_tcp_if_include 100.87.56.74/32,100.124.39.67/32,100.106.52.66/32, one /32 per node. Silently ignored on Open MPI 4.1.6. The advertised list did not change at all.
  • Pinning only the daemon channel: --mca oob_tcp_if_include tailscale0 while leaving the four CX7 names in OMPI_MCA_btl_tcp_if_include. hostname now works, but the NCCL test dies in step 4, because the rank-to-rank TCP layer plays the same game with the same subnets:
WARNING: Open MPI failed to TCP connect to a peer MPI process.  This
should not happen.
Your Open MPI job may now hang or fail.
  Local host: edgexpert-a724
  PID:        313704
  Message:    connect() to 192.168.204.1:1026 failed
  Error:      Operation now in progress (115)

The fix

Everything that is "bootstrap" rides the management network. Only NCCL_IB_HCA stays on the fabric. Both tailscale0 and the CIDR 100.64.0.0/10 worked for oob_tcp_if_include; NCCL_SOCKET_IFNAME wants an interface name anyway, so the script uses the name everywhere.

export OMPI_MCA_plm_rsh_agent="ssh -o UserKnownHostsFile=/dev/null -o StrictHostKeyChecking=no"
M=tailscale0
mpirun -np 3 -H 100.87.56.74:1,100.124.39.67:1,100.106.52.66:1 \
  --mca oob_tcp_if_include $M \
  -x LD_LIBRARY_PATH \
  -x UCX_NET_DEVICES=$M -x NCCL_SOCKET_IFNAME=$M -x OMPI_MCA_btl_tcp_if_include=$M \
  -x NCCL_IB_HCA=rocep1s0f0,rocep1s0f1,roceP2p1s0f0,roceP2p1s0f1 \
  -x NCCL_IB_SUBNET_AWARE_ROUTING=1 -x NCCL_IB_MERGE_NICS=0 -x NCCL_NET_PLUGIN=none \
  $HOME/nccl-tests_spark_cluster/build/all_gather_perf -b 16G -e 16G -f 2

In the script, run_nccl_test() no longer hardcodes anything. It asks node 1 which interface owns the management IP from the config file and uses that:

cmd = f"ip -o -4 addr show | awk '$4 ~ /^{node0['ip_address']}\\// {{print $2}}'"
exit_code, output, error = paramiko_run_command_with_output(ssh, cmd)
mgmt_iface = output.split()[0] if output.split() else ""
...
f"--mca oob_tcp_if_include {mgmt_iface} "
f"-x UCX_NET_DEVICES={mgmt_iface} "
f"-x NCCL_SOCKET_IFNAME={mgmt_iface} "
f"-x OMPI_MCA_btl_tcp_if_include={mgmt_iface} "

On a stock Spark managed over the RJ45 port, that resolves to enP7s7 and behaves exactly like NVIDIA's version. On ours it resolves to tailscale0.

The result

A rerun of just the test:

bash spark_cluster_setup.sh -c config/spark_config_ring.json --run-nccl-test
Using management interface tailscale0 for MPI/NCCL bootstrap
...
Avg bus bandwidth from NCCL test: 22.732 GB/s
NCCL test BW is as expected
NCCL test completed.

And the raw nccl-tests table from a manual run with the same flags:

#  Rank  2 Group  0 Pid 314342 on edgexpert-a724 device  0 [000f:01:00] NVIDIA GB10
#
#                                                              out-of-place                       in-place
#       size         count      type   redop    root     time   algbw   busbw  #wrong     time   algbw   busbw  #wrong
#        (B)    (elements)                               (us)  (GB/s)  (GB/s)             (us)  (GB/s)  (GB/s)
 17179869168    1431655764     float    none      -1   510560   33.65   22.43       0   502258   34.21   22.80       0
# Out of bounds values : 0 OK
# Avg bus bandwidth    : 22.6181

A 16 GiB all-gather across three nodes in about 510 ms. The script's pass mark for a ring is 10 GB/s of bus bandwidth (80 Gb/s); the switch topology is held to 21.875 GB/s (175 Gb/s), and the ring cleared that bar too. Note that the script builds NVIDIA's ring-specific NCCL fork for this layout; the switch topology gets stock NCCL 2.28.

After the fixes, the 200 GbE fabric worked normally again. The successful script run above measured 22.732 GB/s of NCCL bus bandwidth, about 182 Gb/s. The 200 Gb/s link rate and the NCCL bus-bandwidth measurement describe different things; this isn't a measured before-and-after speedup from the broken setup.

Fun fact: the whole of day two was Claude Code driving my tmux session over send-keys and capture-pane, because the Mac has no SSH key for the Sparks and two of the three sessions were busy pulling a 360 GB model. It only ever typed into session 0.

What to keep

  • Pick fabric subnets your LAN cannot collide with before you run anything. 192.168.0.0/24 and 192.168.1.0/24 are what most routers ship with, and NVIDIA's plan uses both. We use 192.168.200.0/24 through 192.168.205.0/24.
  • The MPI bootstrap needs a network every node can reach. In a ring, the fabric is not that network; the management network is. Point oob_tcp_if_include, btl_tcp_if_include, NCCL_SOCKET_IFNAME and UCX_NET_DEVICES there, and leave only NCCL_IB_HCA on the CX7 devices.
  • When mpirun hangs, run it with hostname instead of your program and add --mca oob_base_verbose 5. Read the address list it advertises. The firewall message at the end lies.
  • Per-host /32 entries in oob_tcp_if_include do nothing on Open MPI 4.1.6. Use an interface name, or a CIDR wide enough to match (100.64.0.0/10 for Tailscale).
  • Loopback and docker0 are in every advertised list. Restricting the interface is the fix; deleting Docker is not.

That's it. Nearly every home router ships answering to 192.168.0.1 or 192.168.1.1, which is why any guide that hands those addresses to a network card will eventually meet someone's gateway.