Running a 320B MoE on two DGX Sparks: what actually went wrong
Serving GLM-5.3-Flash (320B-A18B, NVFP4) at 262K context across a 2-node GB10 cluster — and the eleven things that had to be fixed first.
The setup
Two NVIDIA DGX Sparks, linked by 2×200 GbE ConnectX-7. Each node: one GB10 (SM121, compute capability 12.1) and 128 GB of unified memory. That last word does most of the work in this post.
The target: zai-org/GLM-5.3-Flash — 320B total parameters, 18B active per token, a natively multimodal MoE with NoPE MLA and Manifold-Constrained Hyper-Connections — served as weight-only NVFP4 via LibertAIDAI/GLM-5.3-Flash-NVFP4, at tensor-parallel 2 over Ray.
The arithmetic that shapes everything:
checkpoint 181 GiB
resident at TP=2 88.63 GiB per node
budget at util 0.85 103.4 GiB per node
headroom ~12.9 GiB per node for KV + activations + everything else
On a discrete GPU, "headroom" means VRAM. On GB10 it means host RAM too, because the memory is unified. Every GPU allocation is a host allocation. That single fact explains almost every failure below.
Why this model needs its own container
Stock vLLM cannot run GLM-5.3-Flash on GB10 at all. The model uses NoPE MLA — qk_rope_head_dim=0 — and the stock sparse-attention kernel assumes DeepSeek's pe_dim=64. It doesn't degrade; it doesn't run.
So the serving image is a digest-pinned community build carrying day-0 SM121 fixes (SM90 NoPE sparse-MLA extended to SM121 via FA2, FlashInfer pinned to 0.6.18 because 0.6.17 produced NaN at batch 64–256 rows, NCCL 2.30.7, an FA2 fp8-KV tile cap that is what makes fp8 KV usable here), with a thin ray[default] layer on top because the base ships no Ray.
Pin by digest, never tag — a third-party tag can move under you. And gate the layer:
bash glm/verify-image.sh # fails if the ray layer changed or removed ANY base pin
That gate is not ceremony. Those pins are load-bearing for correctness: if pip install ray[default] had dragged FlashInfer off 0.6.18, the symptom would be wrong numbers at certain batch sizes, not a build error. It caught a real case on first use.
Method: climb a ladder, don't jump
With ~12.9 GiB of headroom and a documented history on this cluster of a memory-starved run taking sshd down while ICMP still replied — recovery required a power cycle — the approachwas a context ladder: 32K → 131K → 262K, with hard abort criteria (MemAvailable below ~4 GB on either node, sshd latency degrading, UVM or OOM in dmesg), and every rung logged with its measurements.
If a rung trips the criteria, the previous rung is the working configuration. Record it and stop.
Pitfall 1: your OOM killer is probably inert
earlyoom was installed and systemctl is-active said active. It was doing nothing.
earlyoom's default thresholds are an AND across memory and swap:
SIGTERM when mem <= 10% and swap <= 10%
SIGKILL when mem <= 5% and swap <= 5%
This cluster also sets vm.swappiness=0 as a UVM-livelock defence on unified memory. So swap stays ~100% free, the swap condition is never true, and earlyoom never fires. The two hardening measures silently disarm each other, and the status output gives you false comfort.
sudo sed -i 's/^EARLYOOM_ARGS=.*/EARLYOOM_ARGS="-r 3600 -m 4,2 -s 100,100"/' /etc/default/earlyoom
sudo systemctl restart earlyoom
-s 100,100 makes the swap side always true for both signals, reducing each AND to its memory condition. -m 4,2 puts SIGTERM at ~4.9 GiB and SIGKILL at ~2.5 GiB of 124608 MiB.
Give both percentages explicitly. With a bare -m 4 -s 100, earlyoom halves both for SIGKILL and you get mem <= 2.00% and swap <= 50.00% — which swappiness=0 makes unreachable. SIGTERM would work; escalation would be disarmed for precisely the case it exists for, a process wedged in UVM livelock that ignores SIGTERM.
Also: reading journalctl -u earlyoom immediately after a restart can return the previous run's thresholds. Filter by --since and confirm by timestamp, or you'll conclude your change failed.
Pitfall 2: every memory failure presents as something else
This is the single most useful operational lesson from the whole exercise. Out-of-memory on this cluster never says "out of memory". Observed presentations:
- A Ray
SYSTEM_ERROR connection error code 2. End of file.- A worker dying with no message at all
Compilation terminated.with no compiler diagnostic
That last one is worth internalising: nvcc always prints a reason for a real compile error. A bare "Compilation terminated" means the compiler was killed by a signal.
journalctl -u earlyoom is the first thing you check, not the last. It names the process and the threshold every time:
sending SIGTERM to process 1848676 uid 0 "cicc": badness 1358, VmRSS 5284 MiB
Pitfall 3: the driver/kernel lockstep trap
These machines have no DKMS. Kernel modules come only from prebuilt linux-modules-nvidia-<branch>-<kernel> packages. The kernel and the driver are separate source packages on independent, per-machine phased-update schedules. One apt upgrade can pull a new kernel while holding the driver back — on one node but not the other.
Reboot into that gap and the running kernel has no nvidia.ko. Symptom: nvidia-container-cli: initialization error: nvml error: driver not loaded.
The subtle version bit harder. A health check reported metapkg lockstep OK throughout an outage, because it compared the two metapackages to each other and never to the running kernel. A pinpoint linux-image-7.0.0-1019-nvidia had been installed ahead of the metapackage pair and booted. Both metapackages agreed with each other, at a different version than the kernel actually running.
The fix was a kernel covered check comparing the running kernel's ABI against the kernel the modules metapackage targets — direction-aware, so "upgraded but not yet rebooted" doesn't cry wolf.
Pitfall 4: flags that look right and aren't
Three in a row, each costing a ~9-minute model load to discover:
--distributed-executor-backend ray is not optional. This build defaults to mp and does not infer Ray from a live cluster. Without it: "World size (2) is larger than the number of available GPUs (1)".
--kv-cache-memory is not a registered flag. It only ever worked through argparse prefix-abbreviation. The real name is --kv-cache-memory-bytes. The abbreviation breaks silently the moment another --kv-cache-memory* option appears.
--skip-mm-profiling does not give you a text-only run. It skips the engine's profiling pass. The API server then warms the vision processor anyway, which took 51 s + 24 s andgot rank 0 OOM-killed after Application startup complete. The actual gate is:
mm_limits = {k: v for k, v in allowed_mm_limits.items() if v > 0}
So the limits must be set to zero, not omitted:
--limit-mm-per-prompt {"image":0,"video":0}
Confirmed working when the log says running in text-only mode and there is no multimodal warmup phase at all.
Pitfall 5: a repo ID is not always a model path
Glm5NextProcessor.from_pretrained does a raw open(os.path.join(model_path, "processor_config.json")) instead of resolving through the Hub. It only accepts a local directory. Passing the repo ID gives you a FileNotFoundError well into startup.
Fix: resolve via snapshot_download first and pass the snapshot path, keeping --served-model-name so clients never notice.
Pitfall 6: there is no shared filesystem
Ray TP=2 means each rank loads its own shard from its own node's disk. The HuggingFace cache is bind-mounted per node. A checkpoint present only on the head fails ~30 s into engine init with a traceback that names the path but not the reason:
ray::RayWorkerProc.initialize_worker() (ip=10.0.0.2)
RuntimeError: Cannot find any model weights with '/root/.cache/.../snapshots/...'
…while rank 0 happily logs Checkpoint size: 181.30 GiB. So you need ~181 GiB on both Sparks, ~362 GiB total.
One wrinkle: the cache directories are created by the container as root, so a host-side hf download or an rsync fails on permissions. The fetch script downloads from inside a throwaway container so ownership matches the serving path and no sudo is needed.
Better still, preflight it. The launcher probes every Ray node via NodeAffinitySchedulingStrategy and reports magnetar dir=True safetensors=0 MISSING before serving, rather than failing inside the engine nine minutes later.
Pitfall 7: the JIT compiler is a memory bomb
MAX_JOBS defaults to the CPU count (20 here) and drives ninja's parallel workers in FlashInfer's JIT. Each worker spawns a cudafe++ at ~1.17 GiB RSS — roughly 23 GiB of compiler memory.
That lands on top of 88.63 GiB of resident weights, because the JIT storm runs during determine_available_memory() — after the model has loaded. Both nodes hit earlyoom's threshold simultaneously, and the workers died with "connection error code 2", which reads like a crash rather than an OOM.
MAX_JOBS=2, forwarded to both nodes' containers, fixed it.
The CUTLASS saga: three layers of the same wall
The marlin MoE backend logs a warning:
Your GPU does not have native support for FP4 computation. Weight-only FP4 compression will be used leveraging the Marlin kernel.
That is correct and expected for a weight-only NVFP4-A16 checkpoint — but it does mean the FP4 tensor cores sit idle. flashinfer_cutlass is the path that uses them. Getting there took three fixes.
Layer 1 — the header. FlashInfer's JIT includes <nvrtc.h>, which this image ships only inside the pip wheel at dist-packages/nvidia/cu13/include/, and FlashInfer's nvcc command line doesn't add that directory:
fatal error: nvrtc.h: No such file or directory
Symlink every wheel header the CUDA include dir is missing. Done.
Layer 2 — the memory. With the header fixed, the module now tries to build: 97 nvcc translation units of heavy CUTLASS templates, triggered on the first MoE forward — i.e. at KV-cache profiling, with the weights already resident and ~10 GiB of headroom. A single cicc on the worst file measured 5284 MiB RSS, 4.5× the cudafe++ figure MAX_JOBS=2 was sized against. earlyoom SIGTERMed it at object 20 of 97, ~40 minutes in.
The fix inverts the constraint: build it when no model is loaded, where there is ~110 GiB free instead of ~10 GiB. Then MAX_JOBS can go up to 10 rather than being throttled to 2. Whole thing takes ~20–25 minutes, once.
This also exposed an embarrassing assumption of mine. I had claimed the JIT result was "cached into the bind-mounted flashinfer cache". It wasn't — only ~/.cache/huggingface was mounted. The JIT output lived in the container's writable layer and died with the container, so every restart paid the build again. Mounting ~/.cache/flashinfer was the actual prerequisite for any of this to be worth doing.
Layer 3 — the library. All 97 objects then compiled, and the build died at the final link:
/usr/bin/ld: cannot find -lnvrtc: No such file or directory
The image ships libnvrtc.so.13 — the runtime SONAME — but not the unversioned libnvrtc.so that ld -lnvrtc resolves. That's the devel symlink normally supplied by the CUDA toolkit package.
The lesson generalises: fixing the headers got the units to compile and moved the failure one layer down, where it costs 20 minutes of successful compilation to discover. Worth checking both halves of a toolchain gap at once. (libcudart.so was present and libcuda.so comes from the stubs/ directory already on the link line — nvrtc was the only gap.)
Optimization: two experiments, one negative
CUTLASS MoE: works, buys nothing
| CUTLASS | marlin | |
|---|---|---|
| decode, single stream | 14.4 tok/s | ~14.7 |
| prefill @ 60,028 tok | ~1,268 tok/s | — |
| decode, 8 concurrent | 57.0 tok/s aggregate | — |
| memory under load | identical | identical |
No speedup. (Honest caveat: the marlin figure came from an ad-hoc single request rather than the same harness, so the 2% gap is inside the methodological difference. The supported claim is the negative one.)
Why, and this is the useful part: single-stream decode here is memory-bandwidth bound, not compute bound. ~18B active parameters at ~0.5 byte each is 9–10 GB of weights read per token against LPDDR5X. FP4 tensor cores cannot accelerate a GEMM that is waiting on memory.
The near-linear 3.96× scaling to 8 concurrent streams is the same fact from the other side: extra streams reuse one weight read, so throughput scales almost linearly until compute finally matters.
Which also retires the worry the marlin warning creates. It never cost anything.
Speculative decoding: 2.8×
If decode is bandwidth-bound, the lever is not faster math — it's fewer weight reads per token. That is exactly what speculative decoding does: draft several tokens cheaply, verify them in one pass of the big model.
DFlash2 (incoai/GLM-5.3-Flash-DFlash2, num_speculative_tokens=7):
| DFlash2 | baseline | |
|---|---|---|
| decode, warm median (5 runs) | 40.6 tok/s | 14.4 |
| peak run | 48.9 tok/s | — |
2.8×, with the peak exceeding the published 46.9.
And it wins despite acceptance being well under spec:
accepted / drafted : 965 / 1638 = 58.9% (published 74.1%)
mean accepted/draft : 4.12 of 7 → ~5.12 tokens per verify step
per-position: pos0 88.9% pos1 76.9% pos2 66.7% pos3 56.8%
pos4 50.0% pos5 40.6% pos6 32.5%
Fifteen points below the published acceptance and still 2.8×. That table also shows where the remaining headroom is: positions 5 and 6 land only 40.6% and 32.5% of the time, so the last two draft slots mostly burn compute. num_speculative_tokens=5 is the obvious next experiment.
The cost is KV. The drafter is 2.34 GB of weights — more than the entire host margin on the tighter node — so the launcher automatically trades the KV budget 6 GiB → 3 GiB when speculation is on. That takes KV from 925,447 tokens (3.53× concurrency at 262K) to 310,292 (1.18×). Fine for single-user research; you hold roughly one full-length request instead of three.
The memory asymmetry nobody predicted
Measured at 262K, text-only, model resident:
| pulsar | magnetar | |
|---|---|---|
| idle | 6.6 GB | 11.1 GB |
| during 60K-token prefill | 6.1 GB | 10.6 GB |
| earlyoom SIGTERM at | ~4.9 GB | ~4.9 GB |
Two nominally identical machines, and one consistently runs ~4.5 GB tighter — every single time, across every configuration tested. Working margin on the tighter node is 1.2–1.7 GB.
If you have a heterogeneous-in-practice pair like this, the tighter node is your real budget. Averaging them will get you killed.
One more, because it cost a restart
The first request after startup runs at ~7 tok/s. The next ones run at ~40.
That is mhc_fused_tilelang and the xqa decode path compiling on first inference. It looks exactly like a regression and it is warmup. Never benchmark the first request — and put that in the README, because it's the kind of thing that lives in one person's head otherwise.
And one open defect
No reasoning parser in this build surfaces the chain-of-thought. deepseek_r1 and glm47 behave identically: content is always correct and never contaminated, reasoning_content is always None — across chat_template_kwargs variants and streaming.
The state machine is working as designed (the adapter sets initial_state=REASONING, which is why content stays clean), but REASONING events never reach reasoning_content in either aggregation path. A defect in this day-0 build, not a misconfiguration. It costs nothing for ordinary use; the chain-of-thought is simply discarded. If you need it, call /v1/com pletions with the rendered prompt.
The trap: judging a parser by whether content looks clean. It always does. Judge it by whether reasoning_content holds the text that precedes </think> in the raw generation.
Outcome
| Context | 262,144, with 60,028-token recall verified |
| Decode | 40.6 tok/s warm median, 48.9 peak |
| Prefill | ~1,268 tok/s at 60K |
| Concurrency | 57.0 tok/s aggregate at 8 streams |
| KV | 310,292 tokens, 1.18× at 262K |
| Margin | 6.5 GB on the tighter node, no earlyoom activity |
Starting it is three commands:
source glm/cluster-env.sh && make head PROFILE=glm # node 1
source glm/cluster-env.sh && make worker PROFILE=glm # node 2
make serve PROFILE=glm # node 1
PROFILE is mandatory and has no default. That's deliberate: the profiles carry different container images and different forwarded environment, and a wrong default does not error; it hangs in an NCCL collective when rank 1 picks a different backend from rank 0. A loud failure beats a silent wrong-image bring-up.
What I'd tell my past self
- Check the OOM killer actually fires. Installed ≠ armed. Two hardening measures can cancel each other and both report healthy.
- On unified memory, the compiler is a first-class memory consumer. Not just the model.
- Invert the constraint instead of squeezing into it. Building the kernels when nothing else is resident turned a 40-minute failure into a 20-minute success — and let parallelism go up rather than down.
- A bare "Compilation terminated" means a signal, not a syntax error.
- Measure before optimising, and publish the negative result. CUTLASS was the intuitive win and delivered nothing; the bandwidth analysis it produced is what predicted speculative decoding would work. The failed experiment paid for the successful one.
- Assume nothing is shared. Weights, images, JIT caches — all per node, all silently fatal when missing.
- Write down which node is the tight one. Averages lie.