Split out of the review on #442, which moved the Strix Halo preflight onto GTT when the VRAM pool is a BIOS carveout. That change is scoped to which pool is measured; this issue is about what the floor means once measured.
The gap
Every self-hosted lane gates on a fixed absolute floor — 8 GiB on the Strix lanes, 16 GiB on the discrete ones — applied to whatever pool is reported. The step exists to catch a serve left holding the GPU, but against a large pool a fixed floor only trips near exhaustion:
| lane |
pool |
floor |
trips when held exceeds |
| Strix Halo (GTT aperture) |
~62 GiB |
8 GiB |
~54 GiB |
| Strix Halo (old large carveout) |
~64 GiB |
8 GiB |
~56 GiB |
| MI300X |
192 GiB |
16 GiB |
176 GiB |
A leftover llama-server holding 10–20 GiB clears all three. This is not a regression introduced by #442 — the same floor sat against a similarly large VRAM pool on the hosts where the lane previously passed — but it means the check is much coarser than "catches a leftover serve" suggests.
Options
- Proportional floor — require some fraction of the aperture free (e.g. 50%), so sensitivity scales with the host.
- Absolute-held check — fail when used exceeds a threshold, which is closer to what "a serve is still holding it" actually means.
- Process-based — the reclaim step already knows the process patterns; assert none survive rather than inferring from memory.
- Leave it — accept that the floor only catches near-exhaustion and say so in the step comment.
(3) is arguably the most direct expression of the intent, but the pools are also consumed by non-E2E tenants on shared hosts, so it may under-catch.
Why not in #442
It changes behaviour on every lane including the discrete ones, where the current floor has been stable for a long time. #442 keeps the existing floor semantics and only corrects which pool they apply to; its comment now states the effective threshold rather than claiming the check is more sensitive than it is.
Related: #294 (these preflight blocks are duplicated across six jobs in two files, so any change here lands six times until that is addressed).
Split out of the review on #442, which moved the Strix Halo preflight onto GTT when the VRAM pool is a BIOS carveout. That change is scoped to which pool is measured; this issue is about what the floor means once measured.
The gap
Every self-hosted lane gates on a fixed absolute floor — 8 GiB on the Strix lanes, 16 GiB on the discrete ones — applied to whatever pool is reported. The step exists to catch a serve left holding the GPU, but against a large pool a fixed floor only trips near exhaustion:
A leftover
llama-serverholding 10–20 GiB clears all three. This is not a regression introduced by #442 — the same floor sat against a similarly large VRAM pool on the hosts where the lane previously passed — but it means the check is much coarser than "catches a leftover serve" suggests.Options
(3) is arguably the most direct expression of the intent, but the pools are also consumed by non-E2E tenants on shared hosts, so it may under-catch.
Why not in #442
It changes behaviour on every lane including the discrete ones, where the current floor has been stable for a long time. #442 keeps the existing floor semantics and only corrects which pool they apply to; its comment now states the effective threshold rather than claiming the check is more sensitive than it is.
Related: #294 (these preflight blocks are duplicated across six jobs in two files, so any change here lands six times until that is addressed).