> ## Content Index
> Fetch the complete content index at: https://corti.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Running a 320B MoE on two DGX Sparks: what actually went wrong
- URL: https://corti.com/running-a-320b-moe-on-two-dgx-sparks-what-actually-went-wrong/
- Published: 2026-09-29T13:15:16.000Z
- Updated: 2026-09-29T13:15:16.000Z
- Author: Sascha Corti

*Serving GLM-5.3-Flash (320B-A18B, NVFP4) at 262K context across a 2-node GB10 cluster — and the eleven things that had to be fixed first.*

## The setup

Two NVIDIA DGX Sparks, linked by 2×200 GbE ConnectX-7\. Each node: one GB10 (SM121, compute capability 12.1) and **128 GB of unified memory**. That last word does most of the work in this post.

The target: `zai-org/GLM-5.3-Flash` — 320B total parameters, 18B active per token, a natively multimodal MoE with NoPE MLA and Manifold-Constrained Hyper-Connections — served as weight-only NVFP4 via `LibertAIDAI/GLM-5.3-Flash-NVFP4`, at tensor-parallel 2 over Ray.

The arithmetic that shapes everything:

```
checkpoint            181 GiB
resident at TP=2      88.63 GiB per node
budget at util 0.85   103.4 GiB per node
headroom              ~12.9 GiB per node for KV + activations + everything else

```

On a discrete GPU, "headroom" means VRAM. On GB10 it means **host RAM too**, because the memory is unified. Every GPU allocation is a host allocation. That single fact explains almost every failure below.

## Why this model needs its own container

Stock vLLM cannot run GLM-5.3-Flash on GB10 *at all*. The model uses NoPE MLA — `qk_rope_head_dim=0` — and the stock sparse-attention kernel assumes DeepSeek's `pe_dim=64`. It doesn't degrade; it doesn't run.

So the serving image is a digest-pinned community build carrying day-0 SM121 fixes (SM90 NoPE sparse-MLA extended to SM121 via FA2, FlashInfer pinned to 0.6.18 because 0.6.17 produced NaN at batch 64–256 rows, NCCL 2.30.7, an FA2 fp8-KV tile cap that is what makes fp8 KV usable here), with a thin `ray[default]` layer on top because the base ships no Ray.

Pin by **digest, never tag** — a third-party tag can move under you. And gate the layer:

```bash
bash glm/verify-image.sh   # fails if the ray layer changed or removed ANY base pin

```

That gate is not ceremony. Those pins are load-bearing for *correctness*: if `pip install ray[default]` had dragged FlashInfer off 0.6.18, the symptom would be wrong numbers at certain batch sizes, not a build error. It caught a real case on first use.

## Method: climb a ladder, don't jump

With \~12.9 GiB of headroom and a documented history on this cluster of a memory-starved run taking sshd down while ICMP still replied — recovery required a power cycle — the approachwas a **context ladder**: 32K → 131K → 262K, with hard abort criteria (`MemAvailable` below \~4 GB on either node, sshd latency degrading, UVM or OOM in `dmesg`), and every rung logged with its measurements.

If a rung trips the criteria, the previous rung is the working configuration. Record it and stop.

---

## Pitfall 1: your OOM killer is probably inert

`earlyoom` was installed and `systemctl is-active` said `active`. It was doing nothing.

earlyoom's default thresholds are an **AND across memory and swap**:

```
SIGTERM when mem <= 10% and swap <= 10%
SIGKILL when mem <=  5% and swap <=  5%

```

This cluster also sets `vm.swappiness=0` as a UVM-livelock defence on unified memory. So swap stays \~100% free, the swap condition is never true, and **earlyoom never fires**. The two hardening measures silently disarm each other, and the status output gives you false comfort.

```bash
sudo sed -i 's/^EARLYOOM_ARGS=.*/EARLYOOM_ARGS="-r 3600 -m 4,2 -s 100,100"/' /etc/default/earlyoom
sudo systemctl restart earlyoom

```

`-s 100,100` makes the swap side always true for both signals, reducing each AND to its memory condition. `-m 4,2` puts SIGTERM at \~4.9 GiB and SIGKILL at \~2.5 GiB of 124608 MiB.

Give **both** percentages explicitly. With a bare `-m 4 -s 100`, earlyoom halves both for SIGKILL and you get `mem <= 2.00% and swap <= 50.00%` — which `swappiness=0` makes unreachable. SIGTERM would work; escalation would be disarmed for precisely the case it exists for, a process wedged in UVM livelock that ignores SIGTERM.

Also: reading `journalctl -u earlyoom` immediately after a restart can return the *previous* run's thresholds. Filter by `--since` and confirm by timestamp, or you'll conclude your change failed.

## Pitfall 2: every memory failure presents as something else

This is the single most useful operational lesson from the whole exercise. Out-of-memory on this cluster never says "out of memory". Observed presentations:

- A Ray `SYSTEM_ERROR`
- `connection error code 2. End of file.`
- A worker dying with no message at all
- `Compilation terminated.` with **no compiler diagnostic**

That last one is worth internalising: nvcc always prints a reason for a real compile error. A bare "Compilation terminated" means the compiler was killed by a signal.

**`journalctl -u earlyoom` is the first thing you check, not the last.** It names the process and the threshold every time:

```
sending SIGTERM to process 1848676 uid 0 "cicc": badness 1358, VmRSS 5284 MiB

```

## Pitfall 3: the driver/kernel lockstep trap

These machines have no DKMS. Kernel modules come only from prebuilt `linux-modules-nvidia-<branch>-<kernel>` packages. The kernel and the driver are **separate source packages on independent, per-machine phased-update schedules**. One `apt upgrade` can pull a new kernel while holding the driver back — on one node but not the other.

Reboot into that gap and the running kernel has no `nvidia.ko`. Symptom: `nvidia-container-cli: initialization error: nvml error: driver not loaded`.

The subtle version bit harder. A health check reported `metapkg lockstep OK` throughout an outage, because it compared the two metapackages *to each other* and never to the **running kernel**. A pinpoint `linux-image-7.0.0-1019-nvidia` had been installed ahead of the metapackage pair and booted. Both metapackages agreed with each other, at a different version than the kernel actually running.

The fix was a `kernel covered` check comparing the running kernel's ABI against the kernel the modules metapackage targets — direction-aware, so "upgraded but not yet rebooted" doesn't cry wolf.

## Pitfall 4: flags that look right and aren't

Three in a row, each costing a \~9-minute model load to discover:

**`--distributed-executor-backend ray` is not optional.** This build defaults to `mp` and does not infer Ray from a live cluster. Without it: *"World size (2) is larger than the number of available GPUs (1)"*.

**`--kv-cache-memory` is not a registered flag.** It only ever worked through argparse prefix-abbreviation. The real name is `--kv-cache-memory-bytes`. The abbreviation breaks silently the moment another `--kv-cache-memory*` option appears.

**`--skip-mm-profiling` does not give you a text-only run.** It skips the *engine's* profiling pass. The API server then warms the vision processor anyway, which took 51 s + 24 s andgot rank 0 OOM-killed *after* `Application startup complete`. The actual gate is:

```python
mm_limits = {k: v for k, v in allowed_mm_limits.items() if v > 0}

```

So the limits must be **set to zero**, not omitted:

```
--limit-mm-per-prompt {"image":0,"video":0}

```

Confirmed working when the log says `running in text-only mode` and there is no multimodal warmup phase at all.

## Pitfall 5: a repo ID is not always a model path

`Glm5NextProcessor.from_pretrained` does a raw `open(os.path.join(model_path, "processor_config.json"))` instead of resolving through the Hub. It only accepts a **local directory**. Passing the repo ID gives you a `FileNotFoundError` well into startup.

Fix: resolve via `snapshot_download` first and pass the snapshot path, keeping `--served-model-name` so clients never notice.

## Pitfall 6: there is no shared filesystem

Ray TP=2 means each rank loads *its own shard from its own node's disk*. The HuggingFace cache is bind-mounted per node. A checkpoint present only on the head fails \~30 s into engine init with a traceback that names the path but not the reason:

```
ray::RayWorkerProc.initialize_worker() (ip=10.0.0.2)
RuntimeError: Cannot find any model weights with '/root/.cache/.../snapshots/...'

```

…while rank 0 happily logs `Checkpoint size: 181.30 GiB`. So you need \~181 GiB on **both** Sparks, \~362 GiB total.

One wrinkle: the cache directories are created by the container as root, so a host-side `hf download` or an `rsync` fails on permissions. The fetch script downloads *from inside a throwaway container* so ownership matches the serving path and no sudo is needed.

Better still, preflight it. The launcher probes every Ray node via `NodeAffinitySchedulingStrategy` and reports `magnetar dir=True safetensors=0 MISSING` **before** serving, rather than failing inside the engine nine minutes later.

## Pitfall 7: the JIT compiler is a memory bomb

`MAX_JOBS` defaults to the **CPU count** (20 here) and drives ninja's parallel workers in FlashInfer's JIT. Each worker spawns a `cudafe++` at \~1.17 GiB RSS — roughly 23 GiB of compiler memory.

That lands on top of 88.63 GiB of resident weights, because the JIT storm runs during `determine_available_memory()` — *after* the model has loaded. Both nodes hit earlyoom's threshold simultaneously, and the workers died with "connection error code 2", which reads like a crash rather than an OOM.

`MAX_JOBS=2`, forwarded to both nodes' containers, fixed it.

---

## The CUTLASS saga: three layers of the same wall

The marlin MoE backend logs a warning:

> Your GPU does not have native support for FP4 computation. Weight-only FP4 compression will be used leveraging the Marlin kernel.

That is correct and expected for a weight-only NVFP4-A16 checkpoint — but it does mean the FP4 tensor cores sit idle. `flashinfer_cutlass` is the path that uses them. Getting there took three fixes.

**Layer 1 — the header.** FlashInfer's JIT includes `<nvrtc.h>`, which this image ships *only* inside the pip wheel at `dist-packages/nvidia/cu13/include/`, and FlashInfer's nvcc command line doesn't add that directory:

```
fatal error: nvrtc.h: No such file or directory

```

Symlink every wheel header the CUDA include dir is missing. Done.

**Layer 2 — the memory.** With the header fixed, the module now *tries* to build: **97 nvcc translation units** of heavy CUTLASS templates, triggered on the first MoE forward — i.e. at KV-cache profiling, with the weights already resident and \~10 GiB of headroom. A single `cicc` on the worst file measured **5284 MiB RSS**, 4.5× the `cudafe++` figure `MAX_JOBS=2` was sized against. earlyoom SIGTERMed it at object 20 of 97, \~40 minutes in.

The fix inverts the constraint: build it **when no model is loaded**, where there is \~110 GiB free instead of \~10 GiB. Then `MAX_JOBS` can go *up* to 10 rather than being throttled to 2\. Whole thing takes \~20–25 minutes, once.

This also exposed an embarrassing assumption of mine. I had claimed the JIT result was "cached into the bind-mounted flashinfer cache". It wasn't — only `~/.cache/huggingface` was mounted. The JIT output lived in the container's writable layer and died with the container, so **every restart paid the build again**. Mounting `~/.cache/flashinfer` was the actual prerequisite for any of this to be worth doing.

**Layer 3 — the library.** All 97 objects then compiled, and the build died at the final link:

```
/usr/bin/ld: cannot find -lnvrtc: No such file or directory

```

The image ships `libnvrtc.so.13` — the runtime SONAME — but not the unversioned `libnvrtc.so` that `ld -lnvrtc` resolves. That's the devel symlink normally supplied by the CUDA toolkit package.

The lesson generalises: **fixing the headers got the units to compile and moved the failure one layer down, where it costs 20 minutes of successful compilation to discover.** Worth checking both halves of a toolchain gap at once. (`libcudart.so` was present and `libcuda.so` comes from the `stubs/` directory already on the link line — nvrtc was the only gap.)

---

## Optimization: two experiments, one negative

### CUTLASS MoE: works, buys nothing

|                       | CUTLASS              | marlin    |
| --------------------- | -------------------- | --------- |
| decode, single stream | **14.4 tok/s**       | \~14.7    |
| prefill @ 60,028 tok  | \~1,268 tok/s        | —         |
| decode, 8 concurrent  | 57.0 tok/s aggregate | —         |
| memory under load     | identical            | identical |

No speedup. (Honest caveat: the marlin figure came from an ad-hoc single request rather than the same harness, so the 2% gap is inside the methodological difference. The supported claim is the negative one.)

**Why**, and this is the useful part: single-stream decode here is **memory-bandwidth bound, not compute bound**. \~18B active parameters at \~0.5 byte each is 9–10 GB of weights read *per token* against LPDDR5X. FP4 tensor cores cannot accelerate a GEMM that is waiting on memory.

The near-linear **3.96× scaling to 8 concurrent streams** is the same fact from the other side: extra streams reuse one weight read, so throughput scales almost linearly until compute finally matters.

Which also retires the worry the marlin warning creates. It never cost anything.

### Speculative decoding: 2.8×

If decode is bandwidth-bound, the lever is not faster math — it's **fewer weight reads per token**. That is exactly what speculative decoding does: draft several tokens cheaply, verify them in one pass of the big model.

DFlash2 (`incoai/GLM-5.3-Flash-DFlash2`, `num_speculative_tokens=7`):

|                              | DFlash2        | baseline |
| ---------------------------- | -------------- | -------- |
| decode, warm median (5 runs) | **40.6 tok/s** | 14.4     |
| peak run                     | **48.9 tok/s** | —        |

**2.8×**, with the peak exceeding the published 46.9.

And it wins *despite* acceptance being well under spec:

```
accepted / drafted  : 965 / 1638 = 58.9%     (published 74.1%)
mean accepted/draft : 4.12 of 7 → ~5.12 tokens per verify step
per-position: pos0 88.9%  pos1 76.9%  pos2 66.7%  pos3 56.8%
              pos4 50.0%  pos5 40.6%  pos6 32.5%

```

Fifteen points below the published acceptance and still 2.8×. That table also shows where the remaining headroom is: positions 5 and 6 land only 40.6% and 32.5% of the time, so the last two draft slots mostly burn compute. `num_speculative_tokens=5` is the obvious next experiment.

**The cost is KV.** The drafter is 2.34 GB of weights — *more than the entire host margin on the tighter node* — so the launcher automatically trades the KV budget 6 GiB → 3 GiB when speculation is on. That takes KV from 925,447 tokens (3.53× concurrency at 262K) to 310,292 (1.18×). Fine for single-user research; you hold roughly one full-length request instead of three.

---

## The memory asymmetry nobody predicted

Measured at 262K, text-only, model resident:

|                          | pulsar     | magnetar |
| ------------------------ | ---------- | -------- |
| idle                     | **6.6 GB** | 11.1 GB  |
| during 60K-token prefill | **6.1 GB** | 10.6 GB  |
| earlyoom SIGTERM at      | \~4.9 GB   | \~4.9 GB |

Two nominally identical machines, and one consistently runs **\~4.5 GB tighter** — every single time, across every configuration tested. Working margin on the tighter node is 1.2–1.7 GB.

If you have a heterogeneous-in-practice pair like this, **the tighter node is your real budget.** Averaging them will get you killed.

## One more, because it cost a restart

The first request after startup runs at **\~7 tok/s**. The next ones run at \~40.

That is `mhc_fused_tilelang` and the xqa decode path compiling on first inference. It looks exactly like a regression and it is warmup. **Never benchmark the first request** — and put that in the README, because it's the kind of thing that lives in one person's head otherwise.

## And one open defect

No reasoning parser in this build surfaces the chain-of-thought. `deepseek_r1` and `glm47` behave identically: `content` is always correct and never contaminated, `reasoning_content` is always `None` — across `chat_template_kwargs` variants and streaming.

The state machine is working as designed (the adapter sets `initial_state=REASONING`, which is *why* content stays clean), but REASONING events never reach `reasoning_content` in either aggregation path. A defect in this day-0 build, not a misconfiguration. It costs nothing for ordinary use; the chain-of-thought is simply discarded. If you need it, call `/v1/com pletions` with the rendered prompt.

The trap: judging a parser by whether `content` looks clean. It always does. Judge it by whether `reasoning_content` holds the text that precedes `</think>` in the raw generation.

---

## Outcome

|             |                                                  |
| ----------- | ------------------------------------------------ |
| Context     | **262,144**, with 60,028-token recall verified   |
| Decode      | **40.6 tok/s** warm median, 48.9 peak            |
| Prefill     | \~1,268 tok/s at 60K                             |
| Concurrency | 57.0 tok/s aggregate at 8 streams                |
| KV          | 310,292 tokens, 1.18× at 262K                    |
| Margin      | 6.5 GB on the tighter node, no earlyoom activity |

Starting it is three commands:

```bash
source glm/cluster-env.sh && make head PROFILE=glm     # node 1
source glm/cluster-env.sh && make worker PROFILE=glm   # node 2
make serve PROFILE=glm                                 # node 1

```

`PROFILE` is mandatory and has no default. That's deliberate: the profiles carry different container images *and* different forwarded environment, and a wrong default does not error; it hangs in an NCCL collective when rank 1 picks a different backend from rank 0\. A loud failure beats a silent wrong-image bring-up.

## What I'd tell my past self

1. **Check the OOM killer actually fires.** Installed ≠ armed. Two hardening measures can cancel each other and both report healthy.
2. **On unified memory, the compiler is a first-class memory consumer.** Not just the model.
3. **Invert the constraint instead of squeezing into it.** Building the kernels when nothing else is resident turned a 40-minute failure into a 20-minute success — and let parallelism go *up* rather than down.
4. **A bare "Compilation terminated" means a signal, not a syntax error.**
5. **Measure before optimising, and publish the negative result.** CUTLASS was the intuitive win and delivered nothing; the bandwidth analysis it produced is what predicted speculative decoding *would* work. The failed experiment paid for the successful one.
6. **Assume nothing is shared.** Weights, images, JIT caches — all per node, all silently fatal when missing.
7. **Write down which node is the tight one.** Averages lie.