Exercise solutions: Multi-GPU Fine-Tuning: DDP, ZeRO, FSDP and the Wire Between the Cards
These are the worked solutions for the exercises in Part 14, Multi-GPU Fine-Tuning: DDP, ZeRO, FSDP and the Wire Between the Cards. Read the exercise first; coming here before you have tried it defeats the point.
Exercise 1
Over NVLink at 900 GB/s: 3.08 divided by 900 is about 0.00342 seconds, 3.42 milliseconds.
Over split PCIe at 64 GB/s: 3.08 divided by 64 is about 0.0481 seconds, 48.1 milliseconds.
The ratio is 48.1 divided by 3.42, about 14.1, the same as simply dividing the two bandwidths, 900 divided by 64, about 14.1. This follows directly from the definition of transfer time, bytes divided by bandwidth: for any fixed number of bytes, the ratio between two transfer times equals the ratio between the two bandwidths, since the byte count cancels out of the comparison. This is the arithmetic behind the article’s claim that PCIe is roughly an order of magnitude slower than datacentre interconnect. It holds for a 1.5B model’s gradient exactly as it would for a 70B model’s, because it is a statement about the wire, not about the model.
Exercise 2
With accumulation steps at 1, an all-reduce happens on every micro-batch, moving about 14 GB each time. With accumulation steps at 8, gradients are accumulated locally across 8 micro-batches before a single all-reduce synchronises the accumulated result, once per 8 micro-batches instead of once per 1. That is an 8-fold drop in the number of all-reduce events per epoch, and since each one still moves the same roughly 14 GB, it is also an 8-fold drop in total gradient bytes crossing the wire per epoch.
The forward and backward passes still process exactly the same number of training examples per epoch either way, so the actual compute is unchanged. What changes is purely how often the two cards have to stop and talk to each other. On a fast interconnect this barely matters, because communication was already mostly hidden behind computation. On a slow one, every synchronisation point is time spent waiting rather than computing, so cutting the number of synchronisation points by 8 directly cuts the wasted waiting time by roughly the same factor. That is exactly why the article calls gradient accumulation the single most effective knob on a slow interconnect.
Exercise 3
SHARD_GRAD_OP corresponds to ZeRO stage two: gradients and optimizer state are still sharded across the cards, but parameters are not. Each card now holds a full, un-sharded copy of every weight at rest, rather than only a shard that gets gathered just in time.
Memory usage goes up relative to full shard, because the weights themselves, 2 bytes per parameter in bf16, are now replicated on every card instead of split between them.
Communication goes down, because the constant per-layer all-gather of parameter shards during the forward and backward pass, which full sharding needs to reconstruct each layer’s weights, disappears entirely. What remains is gradient and optimizer synchronisation, a coarser, less frequent pattern than full sharding’s continuous gathering.
Prefer this setting over full sharding specifically when the model already nearly fits with parameters fully replicated, since you get most of stage three’s memory relief from the optimizer and gradient tenants without paying stage three’s steep communication tax. That is exactly the article’s own advice to use the lowest sharding stage that makes a run fit rather than reaching for full sharding by default.
Exercise 4
The first check is to set NCCL_DEBUG=INFO and rerun, so the log states outright which transport NCCL actually chose for this job, rather than guessing.
The most likely actual cause is NCCL_P2P_DISABLE, or a stricter-than-necessary NCCL_P2P_LEVEL, inherited from the datacentre container. On a two-card desktop board, direct peer-to-peer transfer over the shared PCIe lanes is often the fastest path available, and either of these settings can silently force traffic onto a slower path, through host memory rather than directly between the cards, with no error message anywhere in the run. That would explain a job that works but underperforms an equivalent machine, exactly the symptom described.
NCCL_IB_DISABLE is the red herring here. It turns off InfiniBand transport, which a desktop board never had in the first place, so whether it is set to disable it or not changes nothing on this hardware. A colleague auditing the inherited environment file who flags this variable as the suspect is looking in the wrong place. The fix is to check and, if necessary, clear the P2P-related variables instead, then confirm with the debug log that peer-to-peer transport is actually being used.
Exercise 5
This is necessarily an approximation, since real multi-GPU traffic on a shared root complex does not divide perfectly evenly the way this estimate assumes. As a first-order estimate, though: two cards splitting 16 lanes get 8 lanes each; four cards splitting the same 16 lanes get roughly 4 lanes each, about half the per-card bandwidth of the two-card case, somewhere in the neighbourhood of 32 GB/s per card-pair link rather than 64 GB/s. Every strategy that talks inside every layer, tensor parallelism especially, and full FSDP sharding to a lesser extent, gets punished harder as this number falls, while gradient accumulation, whose cost is fixed per synchronisation rather than per layer, becomes relatively more valuable as bandwidth drops.
For a job that already fits in a quarter of one card’s memory: run plain data parallelism across all four, replicate the model on each, split the batch four ways, and lean on gradient accumulation to keep the number of all-reduces low. This sidesteps the will-not-fit sharding machinery entirely, and the periodic all-reduce is the only traffic on the wire, which amortises well.
For a job that genuinely needs all four cards’ combined memory: full sharding is unavoidable, but it should be paired with the largest gradient accumulation step count the workflow can tolerate and the lowest sharding stage that actually makes the model fit. That is the article’s general rule, just applied under a worse bandwidth budget than the two-card case it was originally stated for.