2AM Homelab

Two GPUs on Unraid: one for Plex, one for a local LLM

Aug 10, 2026 · unraid, gpu, ollama, local-llm, plex

Running Plex transcodes and a local LLM on the same Unraid box works well, right up until both want the same card. This is how I split two GPUs by role and kept them that way — including the pinning mistake that silently breaks after a reboot, the driver bug that wedged both cards, and the model-sizing arithmetic that turned a “too slow to use” setup into one that answers in seconds.

The hardware here is deliberately unglamorous: a Tesla V100 16 GB doing inference and an RTX 2080 8 GB doing Plex transcodes. Both are cheap used cards in 2026. The lessons apply to any two-GPU split.

Why split by role at all

The obvious approach — let both containers see both GPUs — is the wrong one. A Plex transcode starting mid-inference will happily grab VRAM on the card your model is resident in, and the failure mode is ugly: either the transcode fails, or your model gets evicted and the next request takes a minute to reload 15 GB of weights.

Pinning each container to one specific card avoids all of it. The cards never contend, and each workload gets predictable performance. The cost is that neither can burst onto the other’s card, which for a homelab is a trade worth making every time.

1. Pin by UUID, never by index

This is the single most important thing in this post.

Every guide shows you --gpus device=0 or NVIDIA_VISIBLE_DEVICES=0. Device indexes are not stable. They can change when a card is reseated, when the driver enumerates in a different order, or when one card fails to initialize and the remaining one shifts up. The day that happens, your Plex container silently starts transcoding on the card holding your model — or your inference container grabs the transcode GPU and nothing works.

Get the UUIDs:

nvidia-smi -L
GPU 0: Tesla V100-PCIE-16GB (UUID: GPU-aaaa1111-bbbb-2222-cccc-333344445555)
GPU 1: NVIDIA GeForce RTX 2080 (UUID: GPU-dddd6666-eeee-7777-ffff-888899990000)

Then pin each container to its UUID rather than its index. In an Unraid container template, that’s the NVIDIA_VISIBLE_DEVICES variable:

NVIDIA_VISIBLE_DEVICES = GPU-aaaa1111-bbbb-2222-cccc-333344445555

UUIDs survive reboots, reseats and driver reloads. I can vouch for this personally: after physically reseating the V100, the card came back with the same UUID, so the inference container’s pin needed no edit at all. Had it been pinned to an index, it would have come back pointed at the wrong card.

⚠️ Put the pin in the template, not just the running container. A GUI “Apply” rebuilds the container from its template and silently drops anything you added by hand — the same trap that eats bind mounts on hand-built containers.

Related Unraid gotcha: the same “don’t trust the label across a reboot” rule applies to disks. After that same reboot, device letters shifted — a data disk moved from sdg to sdf and watching the old letter made a running operation look stalled. Resolve devices from /proc/mdstat or disks.ini, never from a letter you remember.

2. Expect Xid errors, and set up persistence mode

Two separate cards on this machine wedged the same way, months apart. The symptom is a container that won’t start, with a CDI error like:

failed to get device handle from UUID: Not Found

The card is still physically present — lspci sees it — but the driver is stuck in a loop:

NVRM: RmInitAdapter failed! (0x24:0x72:1603)

and nvidia-smi only enumerates the other GPU. Walking back through the syslog showed the real sequence: an Xid 62 (internal micro-controller halt), then Xid 154 (GPU reset required), then days later an Xid 32 during a transcode, and finally the RmInitAdapter loop.

Things worth knowing:

The mitigation that actually helped was persistence mode. Running nvidia-persistenced at boot keeps both GPUs continuously initialized, so the on-demand re-initialization that was failing simply never runs:

# from your Unraid go file
nvidia-persistenced

Verify with nvidia-smi -q | grep -i persistence — you want Enabled on both cards.

⚠️ Check driver availability before planning a downgrade. I couldn’t move off the suspect driver: for the running kernel, the only Volta-compatible build available was the exact version I suspected. Newer driver branches drop Volta support entirely — the 580 branch is the last one supporting Maxwell, Pascal and Volta. If you’re buying an older datacenter card, check which driver branches still support it and whether those branches build for your kernel.

A watchdog is worth ten minutes of setup

Since the wedge recurs, detect it rather than discover it. A User Script on a */10 * * * * schedule doing two checks covers it:

  1. Full wedge — the expected GPU no longer appears in nvidia-smi -L → alert.
  2. Early warning — the count of NVRM: Xid or RmInitAdapter failed lines in dmesg has increased since the last run → warn, so you can reboot at a convenient hour instead of discovering it during a movie.

Keep the state file in /tmp so a recovery reboot resets it cleanly and you don’t get spurious “recovered” notifications.

⚠️ Unraid cron gotcha: editing schedule.json for a User Script does not regenerate the cron fragment. Write the fragment and run update_cron, or your schedule silently never fires.

3. Choosing a model that fits: MoE beats dense

With the card pinned, the interesting question is what actually runs well on 16 GB. I benchmarked three options with the same setup (GPU plus host-RAM offload):

ModelSplit (GPU/CPU)Generation
9B dense, q8full GPU60 tok/s
27B dense81/199.1 tok/s
35B MoE (a3b)61/3928.8 tok/s

Read that middle row carefully. The 35B model is three times faster than the 27B model, despite being larger. Once any part of a dense model spills into host RAM, every single token pays the RAM-bandwidth tax. A Mixture-of-Experts model only activates a fraction of its parameters per token, so the spill costs far less.

For hybrid GPU+RAM inference, pick MoE over dense. Model size matters much less than architecture once you’re over your VRAM budget.

4. Quantization: measure before you trade away context

The 35B ran at 25.6 tok/s in a 4-bit quant, and the card was drawing 37 W of a 250 W budget at 28% utilization. That’s not a slow GPU — that’s a starving one, idling while it waits for weights to stream over PCIe.

Two things I got wrong on the way to fixing it, both worth stealing:

Misreading the split. ollama ps reported 39%/61% CPU/GPU, which looks like layers running on the CPU. They weren’t — the load log showed all 41 layers offloaded to the GPU. What sat in host RAM was ~8.7 GB of MoE expert weights being streamed on demand. The number in ollama ps doesn’t mean what it appears to mean.

Assuming context length was the lever. I proposed cutting the context window to free VRAM. It couldn’t have worked: this model’s KV cache is only 180 MiB at 32K context, so going down to 8K would have freed ~135 MiB against a ~6 GB shortfall. For contrast, a 14B model I measured had a 1,728 MiB KV cache at the same context — because it uses far more KV heads. So context reduction is a real lever on some models and almost useless on others. Measure your model’s actual KV cache size before sacrificing context for it.

The real fix was a smaller quantization that fit entirely in VRAM. Same prompt (~7,200 tokens), same card:

IQ3_XXS (14.68 GB)Q4_K_M (23 GB)
Split11%/89% GPU39%/61%
Prompt eval451 tok/s97 tok/s
Generation47.7 tok/s25.6 tok/s
Wall clock16 s74 s

4.6× the prompt throughput, from nothing but making the weights fit.

The sizing arithmetic

Usable VRAM is less than the sticker number. On a 16 GB card:

15.9 GB usable
 - 0.18 GB KV cache (measure yours!)
 - 0.25 GB compute buffer
 ≈ 15.4 GB for weights

So a 14.68 GB quant fits with ~800 MB of headroom. The next size up (15.28 GB) is uncomfortably tight, and 15.94 GB doesn’t fit at all. There is a real cliff here, and being 300 MB over it costs you two thirds of your throughput.

Did quality suffer?

This is where most quantization posts hand-wave, so: I ran 20 prompts with objectively checkable answers at temperature 0, across the categories quantization degrades first — arithmetic, logic and negation, counting, long-context recall, strict format compliance, and honesty under uncertainty.

The result was an exact tie: 17/20 for both, and identical per-category. The three failures were the same three on both models, which makes them model-level limits rather than quantization damage.

I’ll add the honest caveat: an earlier three-question probe suggested the smaller quant was better, which would have been a fun claim to make. It didn’t reproduce at temperature 0. That was noise, not a finding, and I’m reporting it because the temptation to keep the exciting version of a result is exactly how bad benchmarks get published.

5. Storage: keep models off the user share

A subtle but large win. The model store was originally on a normal Unraid user share (/mnt/user/appdata/...). Moving it to bind directly to the underlying pool (/mnt/cache/appdata/...) cut model load time from 92 s to 31 s — the FUSE layer that makes user shares work is expensive for multi-gigabyte sequential reads.

Two more storage notes:

The same reasoning applies to any large, latency-sensitive workload on Unraid. Game servers are the other obvious case: giving them their own cache-only share keeps them off the parity-protected array and out of nightly appdata backups.

6. Assorted gotchas

Troubleshooting

SymptomCause
Container won’t start, failed to get device handle from UUIDPinned GPU has wedged, or the UUID changed (card replaced)
nvidia-smi shows only one of two cardsDriver stuck in RmInitAdapter loop — needs a full power cycle, not a module reload
Wrong container using the wrong GPU after a rebootPinned by device index instead of UUID
Inference much slower than benchmarks suggestModel spilling to host RAM; check the split and fit a smaller quant
GPU at low wattage and low utilizationStarving on PCIe transfer, not compute-bound — the model doesn’t fit
Model load takes 90+ secondsModel store on a FUSE user share; bind to the pool directly
Empty response right after switching modelsLoad time exceeding your client timeout; retry once resident

What I’d do differently

  1. Pin by UUID from the very first container. Retro-fitting this after a card shuffles is how you find out it mattered.
  2. Check driver-branch support before buying an older datacenter card. A cheap V100 is a great deal until the only compatible driver has a bug and the newer branches have dropped your architecture entirely.
  3. Fit the model in VRAM before tuning anything else. Every other optimization is noise next to crossing that threshold.
  4. Benchmark with checkable answers. “It seems fine” is not a quality assessment, and it took twenty scored prompts to be confident the faster setup wasn’t a downgrade.