Repository navigation
Conversation
Contributor
Dagger benchmarks:
|
| Backend | Regressions | Improvements | Within noise | Inconclusive time | Suites |
|---|---|---|---|---|---|
| Threads | 0 | 0 | 8 | 2 | 4/4 |
| Distributed | 4 | 4 | 4 | 0 | 4/4 |
| Distributed+Threads | 0 | 0 | 1 | 4 | 4/4 |
| MPI | 0 | 0 | 0 | 0 | 4/4 |
| MPI+Threads | 3 | 5 | 18 | 2 | 4/4 |
| Total | 7 | 9 | 31 | 8 | 20/20 |
Regressions across all backends
| Backend | Benchmark | Metric | Change |
|---|---|---|---|
| MPI+Threads | linalg/dagger/N=1024 (block 512)/cholesky |
time | +186.2% |
| MPI+Threads | linalg/dagger/N=1024 (block 512)/solve (A\b via lu) |
time | +167.7% |
| MPI+Threads | linalg/dagger/N=1024 (block 512)/lu |
time | +123.9% |
| Distributed | linalg/dagger/N=1024 (block 128)/syrk (A'*A) |
allocs | +74.4% |
| Distributed | linalg/dagger/N=1024 (block 128)/syrk (A'*A) |
memory | +51.2% |
| Distributed | linalg/dagger/N=1024 (block 128)/matmul (A*A) |
allocs | +37.6% |
| Distributed | linalg/dagger/N=1024 (block 128)/matmul (A*A) |
memory | +36.2% |
Improvements across all backends
| Backend | Benchmark | Metric | Change |
|---|---|---|---|
| MPI+Threads | stencil/dagger/N=256 (block 128)/assign (const) |
time | -74.5% |
| Distributed | sparse/dagger/N=256 (block 16)/cg solve (laplacian) |
memory | -61.5% |
| Distributed | sparse/dagger/N=256 (block 16)/cg solve (laplacian) |
allocs | -58.7% |
| MPI+Threads | stencil/dagger/N=1024 (block 512)/multi-expr |
time | -48.6% |
| Distributed | sparse/dagger/N=1024 (block 64)/cg solve (laplacian) |
memory | -48.4% |
| MPI+Threads | array/dagger/N=256 (block 128)/transpose (permutedims) |
time | -47.1% |
| Distributed | sparse/dagger/N=1024 (block 64)/cg solve (laplacian) |
allocs | -44.5% |
| MPI+Threads | sparse/dagger/N=1024 (block 64)/spgemm (S*S) |
time | -37.5% |
| MPI+Threads | sparse/dagger/N=256 (block 16)/spmv (S*x) |
time | -35.2% |
Threads (1 process × 4 threads) — 0 regressions, 0 improvements
| Suite | Regressions | Improvements | Within noise | Inconclusive time |
|---|---|---|---|---|
| array | 0 | 0 | 3 | 0 |
| linalg | 0 | 0 | 1 | 0 |
| sparse | 0 | 0 | 0 | 0 |
| stencil | 0 | 0 | 4 | 2 |
Within noise
| Benchmark | Metric | Change |
|---|---|---|
array/dagger/N=1024 (block 512)/broadcast (X .+ 1) |
time | +49.9% |
linalg/dagger/N=256 (block 128)/syrk (A'*A) |
time | +39.2% |
array/dagger/N=1024 (block 128)/norm |
time | +34.9% |
stencil/dagger/N=1024 (block 512)/assign (const) |
time | +25.3% |
stencil/dagger/N=1024 (block 512)/neighbors (Reflect) |
time | -25.5% |
stencil/dagger/N=1024 (block 128)/assign (const) |
time | -29.6% |
stencil/dagger/N=256 (block 128)/assign (const) |
time | -33.1% |
array/dagger/N=1024 (block 512)/reduce (sum) |
time | -63.0% |
Inconclusive timing changes
| Benchmark | Metric | Change |
|---|---|---|
stencil/dagger/N=256 (block 128)/update (+) |
time | +65.0% |
stencil/dagger/N=256 (block 128)/multi-expr |
time | +35.3% |
Distributed (4 processes × 1 thread) — 4 regressions, 4 improvements
| Suite | Regressions | Improvements | Within noise | Inconclusive time |
|---|---|---|---|---|
| array | 0 | 0 | 1 | 0 |
| linalg | 4 | 0 | 3 | 0 |
| sparse | 0 | 4 | 0 | 0 |
| stencil | 0 | 0 | 0 | 0 |
Regressions
| Benchmark | Metric | Change |
|---|---|---|
linalg/dagger/N=1024 (block 128)/syrk (A'*A) |
allocs | +74.4% |
linalg/dagger/N=1024 (block 128)/syrk (A'*A) |
memory | +51.2% |
linalg/dagger/N=1024 (block 128)/matmul (A*A) |
allocs | +37.6% |
linalg/dagger/N=1024 (block 128)/matmul (A*A) |
memory | +36.2% |
Improvements
| Benchmark | Metric | Change |
|---|---|---|
sparse/dagger/N=256 (block 16)/cg solve (laplacian) |
memory | -61.5% |
sparse/dagger/N=256 (block 16)/cg solve (laplacian) |
allocs | -58.7% |
sparse/dagger/N=1024 (block 64)/cg solve (laplacian) |
memory | -48.4% |
sparse/dagger/N=1024 (block 64)/cg solve (laplacian) |
allocs | -44.5% |
Within noise
| Benchmark | Metric | Change |
|---|---|---|
linalg/dagger/N=1024 (block 128)/syrk (A'*A) |
time | +135.7% |
linalg/dagger/N=1024 (block 128)/qr |
time | +40.6% |
array/dagger/N=1024 (block 512)/broadcast (X .+ 1) |
time | -40.7% |
linalg/dagger/N=1024 (block 128)/solve (A\b via lu) |
time | -61.2% |
Distributed+Threads (2 processes × 2 threads) — 0 regressions, 0 improvements
| Suite | Regressions | Improvements | Within noise | Inconclusive time |
|---|---|---|---|---|
| array | 0 | 0 | 0 | 1 |
| linalg | 0 | 0 | 0 | 2 |
| sparse | 0 | 0 | 0 | 0 |
| stencil | 0 | 0 | 1 | 1 |
Within noise
| Benchmark | Metric | Change |
|---|---|---|
stencil/dagger/N=1024 (block 128)/multi-expr |
time | +39.2% |
Inconclusive timing changes
| Benchmark | Metric | Change |
|---|---|---|
array/dagger/N=1024 (block 512)/alloc (rand) |
time | +142.3% |
stencil/dagger/N=1024 (block 512)/neighbors (Reflect) |
time | +47.0% |
linalg/dagger/N=256 (block 256)/matmul (A*A) |
time | +38.7% |
linalg/dagger/N=256 (block 256)/lu |
time | +37.5% |
MPI (4 ranks × 1 thread) — 0 regressions, 0 improvements
| Suite | Regressions | Improvements | Within noise | Inconclusive time |
|---|---|---|---|---|
| array | 0 | 0 | 0 | 0 |
| linalg | 0 | 0 | 0 | 0 |
| sparse | 0 | 0 | 0 | 0 |
| stencil | 0 | 0 | 0 | 0 |
MPI+Threads (2 ranks × 2 threads) — 3 regressions, 5 improvements
| Suite | Regressions | Improvements | Within noise | Inconclusive time |
|---|---|---|---|---|
| array | 0 | 1 | 7 | 2 |
| linalg | 3 | 0 | 3 | 0 |
| sparse | 0 | 2 | 0 | 0 |
| stencil | 0 | 2 | 8 | 0 |
Regressions
| Benchmark | Metric | Change |
|---|---|---|
linalg/dagger/N=1024 (block 512)/cholesky |
time | +186.2% |
linalg/dagger/N=1024 (block 512)/solve (A\b via lu) |
time | +167.7% |
linalg/dagger/N=1024 (block 512)/lu |
time | +123.9% |
Improvements
| Benchmark | Metric | Change |
|---|---|---|
stencil/dagger/N=256 (block 128)/assign (const) |
time | -74.5% |
stencil/dagger/N=1024 (block 512)/multi-expr |
time | -48.6% |
array/dagger/N=256 (block 128)/transpose (permutedims) |
time | -47.1% |
sparse/dagger/N=1024 (block 64)/spgemm (S*S) |
time | -37.5% |
sparse/dagger/N=256 (block 16)/spmv (S*x) |
time | -35.2% |
Within noise
| Benchmark | Metric | Change |
|---|---|---|
linalg/dagger/N=1024 (block 512)/matvec (A*x) |
time | +389.3% |
stencil/dagger/N=1024 (block 512)/assign (const) |
time | +124.7% |
array/dagger/N=256 (block 128)/map (sin.(X)) |
time | +119.6% |
array/dagger/N=1024 (block 512)/broadcast (X .+ 1) |
time | +103.9% |
linalg/dagger/N=256 (block 256)/syrk (A'*A) |
time | +65.6% |
stencil/dagger/N=256 (block 128)/alloc (neighbors Wrap) |
time | +56.0% |
array/dagger/N=256 (block 128)/reduce (sum) |
time | +54.3% |
stencil/dagger/N=256 (block 256)/multi-expr |
time | +48.1% |
array/dagger/N=1024 (block 128)/broadcast (X .+ 1) |
time | +44.2% |
stencil/dagger/N=256 (block 128)/neighbors (Pad) |
time | +44.1% |
linalg/dagger/N=1024 (block 512)/matmul (A*A) |
time | +41.4% |
stencil/dagger/N=1024 (block 128)/assign (const) |
time | -36.5% |
stencil/dagger/N=1024 (block 512)/update (+) |
time | -36.8% |
array/dagger/N=1024 (block 512)/map (sin.(X)) |
time | -38.3% |
stencil/dagger/N=256 (block 128)/update (+) |
time | -40.4% |
array/dagger/N=1024 (block 512)/norm |
time | -42.4% |
array/dagger/N=1024 (block 512)/reduce (sum) |
time | -48.2% |
stencil/dagger/N=256 (block 256)/update (+) |
time | -55.7% |
Inconclusive timing changes
| Benchmark | Metric | Change |
|---|---|---|
array/dagger/N=256 (block 128)/add (X + X) |
time | +106.4% |
array/dagger/N=256 (block 256)/add (X + X) |
time | +38.1% |
Full results and plots (download the benchmark-results-* artifacts).
jpsamaroo
force-pushed
the
jps/fix-ipc-ownership
branch
from
October 1, 2026 00:51
05c62fe to
3f45600
Compare
Keep device arrays on their owners during datadeps slot creation and source fetches for in-place copies. Passing Chunks through the backend transport avoids deserializing an array while retaining a remote processor/context. Replace CUDA's legacy borrowed IPC mapping with the shared staging/import hooks. Hold export storage until the receiver has copied into its own array, release it in finally, and close handles through the CUDA-version-aware driver module. Small transfers and disabled IPC use host staging. Add shared CI coverage for two workers using device 1 on CUDA, ROCm, oneAPI, Metal, and OpenCL. Explicit placements exercise both transfer directions, source independence, datadeps updates, broadcast, and tiled multiplication on the backends supporting it. Record the ownership and IPC lifecycle rules. Validation: - Cyclops, Julia 1.12.6 / CUDA 6.4: unchanged EthanBug/repro.jl completes and its result matches A*B; all 39 shared CUDA regression checks pass. - The old code fails the new broadcast with CUipcMemHandle undefined and tiled multiplication with the original remote-context assertion. - OpenCL/pocl: all 38 shared checks pass; allocation suite: 24 checks pass. - 100 local GPU moves, 10 warmup batches and minimum of 5 GC-counter runs: before 100 allocations / 1,600 bytes; after 0 / 0. The failing cross-worker baseline cannot supply a successful-operation allocation comparison.
`aliasing(::CuArray)` used `pointer(x)`, which calls `take_ownership!` for the calling task's active device. On a worker owning two GPUs without P2P access, the Phase-1 aliasing RPC runs on device 0 while the array lives on device 1, so it threw "cannot take the GPU address of inaccessible device memory". It also re-stamped stream ownership as a side effect. Aliasing only needs the address, so read it from the underlying memory directly. `pin_buffer!` called `CUDA.pin`, which dedupes per context, while host registration is process-wide: moving one host array to two GPUs of one process threw HOST_MEMORY_ALREADY_REGISTERED. Skip ranges that are already page-locked, under a lock so two contexts can't race the check. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The same-worker device-to-device remainder path launches one kernel that reads `from` and writes `to`, which needs both on one device. Two GPUs of one worker need not be peer-accessible (e.g. 2x L40S), so a stencil halo exchange between them failed in `pointer`. Take the direct path only when source and destination share a memory space; other pairs use the existing host-staged gather/scatter. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Distributed `collect` concatenated fetched tiles directly, so tiles on different GPUs of one process met in one device-side `cat` kernel and failed without peer access. The result is a host `Array` anyway, so copy each GPU tile to host first. Add a regression testset for two GPUs owned by one process covering datadeps aliasing on the non-active device (and double host pinning), a cross-device stencil halo exchange, and collect of mixed-device tiles. It errors without the preceding fixes and passes with them on 2x L40S (Julia 1.13.1, CUDA 6.4); single-GPU agents skip it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
GPU `execute!` returns with its kernels still queued on Dagger's per-device stream. Distributed `collect` then fetched each tile and copied it to host from the caller's task, i.e. on a different, task-local stream. That is only correct when the array library synchronizes the previous owner on a cross-stream access: CUDA, AMDGPU and OpenCL do, but oneAPI does not, and it also copies on the calling task's device rather than the tile's. Passing CUDA tests did not show that the wait existed. Synchronize the tile's processor with `gpu_synchronize` and copy under its `with_context`, on the chunk's owner. The wait then happens where the stream lives, and a remote tile is sent as a host array instead of being serialized as a device array. Plain values in `chunks` (e.g. from `sort`) pass through. Add a test that collects right after a deliberately slow kernel, with tiles owned by the driver and by remote workers. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
One-shot HtoD sources were page-locked with CUDA.pin, whose unregistration rides a GC finalizer. A registration could therefore outlive its transfer and collide with whatever the allocator placed at that address next, giving ERROR_HOST_MEMORY_ALREADY_REGISTERED, or ERROR_INVALID_VALUE once a stale finalizer unregistered a newer owner's range mid-copy. Multi-chunk DArrays pin many large tiles per build, so they hit it first. Register through a Dagger-owned refcounted table (keyed by base pointer, since registration is process-wide) and unregister in a `finally` after the copy has synchronized, while the caller still holds the buffer. Overlap is detected from the registration error itself (endpoint probes miss an interior registration) and falls back to registering a private copy. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Same defect as the CUDA fix, in a harsher form. HtoD sources were registered
by `pin_buffer!`, which unregisters from a GC finalizer. AMDGPU's unregister
takes a lock, and a finalizer may not switch tasks, so under contention the
unregister was dropped ("task switch not allowed from inside gc finalizer")
and the range stayed registered after its array died, leaking pinned host
memory and exposing whatever the allocator put there next.
Track registrations in a Dagger-owned table and unregister in a `finally`
after the copy has synchronized. HIP registration is neither refcounted nor
exclusive (duplicates are idempotent and one unregister drops the pin), so we
never unregister a range someone else registered: a buffer whose first or last
byte is already pinned gets a private copy to register instead.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… finalizer `pin_buffer!` page-locked staging buffers with `CUDA.pin`, whose GC finalizer unregisters the range under a `ReentrantLock`. When that lock is contended the finalizer throws "task switch not allowed from inside gc finalizer" and the unregistration is lost, while GC frees the memory anyway. The freed range stays registered with the driver: a later allocation at that address fails to register or is DMA'd through stale pages. A 4-GPU GEMM benchmark crashed this way (segfault a few iterations in, during Datadeps' host-staged cross-device remainder copies). ROCExt's `pin_buffer!` had the same flaw: its own finalizer called `AMDGPU.Mem.unpin`, which takes AMDGPU's registration lock. 2-GPU Datadeps tests on RX 6900 XTs segfaulted inside `hipMemcpyWithStream` after the same finalizer error. Register buffers ourselves and unregister from a finalizer that only `trylock`s; on contention it re-arms itself, which keeps the buffer alive until a later collection can unregister it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`aliasing(::SubArray)` located a view through `pointer` on its parent and on the view. For a CUDA array that takes CUDA.jl stream ownership for the *active* device, which throws when the array lives on another GPU of the process without peer access (a view of a host array rebuilt on one GPU while another GPU's task is current). Add a `data_address` hook, defaulting to `pointer`, and compute the view's address from its parent's; CUDAExt reads the raw address instead. AMDGPU.jl's `pointer` likewise takes stream ownership (synchronizing the stream the array was last used on), so ROCExt reads the raw address too, in `aliasing(::ROCArray)` as well. MetalExt's `data_address` reports the same address as its `aliasing(::MtlArray)`, so a view's spans line up with its parent's. oneAPI.jl's and OpenCL.jl's `pointer` have no side effects and keep the default. Also fix `aliasing(::CuArray)`, which added `x.offset` (an element count) to the pointer as bytes. (AMDGPU.jl's `offset` counts elements in older versions and bytes in newer ones; ROCExt asks AMDGPU's `derive` which.) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…m host Two host<->device paths broke on data Datadeps produces when it rebuilds a view of a host array on a GPU: - The out-of-place host-to-device `move` of a `Chunk` labelled with a host processor assumed the payload was host memory, but here it is the parent device array already moved to the GPU; page-locking that as host memory throws (CUDA: from `CUDA.pin`; ROCm: `hipHostRegister` fails with hipErrorHostMemoryAlreadyRegistered). Hand device arrays to the device-array `move`. - In-place `move!` between a host array and a strided view of a device array used `copyto!`, which indexes the view element by element from the host (scalar indexing). Stage through a dense device buffer and gather/scatter with a broadcast kernel. Both fixes apply to every GPU extension (CUDA, ROCm, oneAPI, Metal, OpenCL), whose host<->device paths share this shape. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…r copies Two GPUs of one process need not be peers (PCIe-attached GPUs behind a host bridge often are not). CUDA.jl's `copyto!` between them then stages through a fresh *pageable* host `Vector` with a synchronous DtoH, 1.4 GB/s on L40S; Datadeps' remainder copies took a host round trip of their own, and a KA kernel touching both buffers is illegal outright. `cuMemcpyPeerAsync` works without peer access -- the driver pipelines it through its own pinned buffers -- at 21.7 GB/s on the same machine. Route every same-process device-to-device copy through it: in-place `move!`, the out-of-place `move`, and Datadeps remainder copies, through a new core hook `device_remainder_copy!`. Many short spans (halos) are packed on the source and unpacked on the destination, so a copy is not thousands of host-staged transfers. The copies also run on a per-device copy stream, so they overlap the destination's kernels instead of queueing behind them. They are ordered per allocation rather than per device: the copy waits only for the source's last write and the destination's last write and read, not for kernels that merely read the data being sent (a GEMM panel's owner is busy reading it while other GPUs ask for it). Each GPU task records an event for the allocations it writes and reads, using a new `writes` task option in which Datadeps lists the written positional arguments; allocations with unknown history fall back to waiting on the whole device stream. The copying task waits on the host for its copy, so every existing synchronization point stays valid. Its own record then skips the destination it filled (an event on the device stream would make later copies out of it wait for unrelated kernels): the copy stamps the allocation with its task, so the skip can never apply to a later allocation at the same address -- an unnamed "skip the next write" flag, left on short-lived packing buffers, made a later write go unrecorded and a copy read stale data. On 4x L40S, peer copies alone took a 4-GPU Float32 GEMM (N=56736) from 27 s per call to 4.3 s; overlapping them with compute took it from ~3.9 s to ~3.1-3.4 s (under GreedyScheduler, with the next commit's traversal). ROCExt gets the same machinery, with HIP events and `hipMemcpyPeerAsync`. HIP does not stage through pageable memory (between two RX 6900 XTs without peer access, AMDGPU's `copyto!` and `hipMemcpyPeerAsync` both reach ~5.5 GB/s), but the old path synchronized the source's whole device stream and queued the copy behind the destination's kernels, and remainder copies still went through the host. A 2-GPU Float32 GEMM (N=12288, GreedyScheduler) went from ~5.8-6.1 s to ~4.6-4.8 s per call. The cross-GPU tests run for ROCm too. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Host-to-device `move`s registered the caller's host array with `hipHostRegister` for the length of one copy and unregistered it right after (`pin_host!`/`unpin_host!`). Under multi-threaded Datadeps that corrupted data: a 2-GPU Float32 GEMM (1024x1024, 256x256 tiles) read whole tile columns back as NaN in 6 of 6 stress runs, on one GPU as well as two, and under RoundRobin and Greedy alike. With the uploads unregistered, 8 of 8 runs were clean; registering a private copy instead was clean too, but uploaded at 0.8-7 GB/s. Registration buys nothing here: HIP stages pageable uploads through its own pinned buffers, at ~11 GB/s for 1-64 MiB on an RX 6900 XT, the same as from registered memory. Upload directly. DtoH staging buffers (`pin_buffer!`) stay registered. It is faster too: registering and unregistering a tile per upload cost more than the upload. A 2-GPU Float32 GEMM (N=12288, GreedyScheduler, host-resident operands) went from ~4.6-4.8 s to ~1.7-1.9 s per call. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rkers Use the active test environment on every worker to prevent stale LLVM and GPUCompiler modules after backend installation. Honor Pkg.test test_args so worker counts and test selections also apply to package-driven runs. CUDA CI suites pass: 281 GPU checks and 106 stencil checks.
The one-thread GPU run stopped in AMDGPU hostcall_host_wait with no exception. Two threads allow compilation and scheduling to continue on the other thread while the persistent host-call task polls. Apply this to the driver and Distributed workers.
MPI refs report this Distributed worker on every rank, so the old worker-id guard inspected remote placeholders during GPU copy tasks. Use the acceleration ownership predicate for CUDA and ROCm. Validation: Julia 1.11 two-rank CUDA MPI suite passed (221/227 checks). After ten warmup iterations and GC, the minimum of five measurements for 10,000 local inspections was 0 allocations / 0 bytes before and after. The ownership predicate infers Bool; storage unwrap remains intentionally dynamic.
A standalone HIP program reproduces missing asynchronous kernel/copy writes on this gfx1036 host. Document the diagnostic serialization flags so this runtime issue does not become a production stencil synchronization cost.
GPU workers can abort outside Julia exception handling. Flush the selected suites and completed test summaries before entering the next suite so redirected CI logs preserve the last completed boundary.
On Xe2, the default v2 adapter aborts while SYCL wraps oneAPI.jl native queues for GEMM. The same oneAPI-only GEMM passes with the documented legacy adapter. Select that adapter consistently for Intel tests and benchmarks and record the diagnosis. Validation: Julia 1.11 oneAPI GPU body passed all 281 assertions (233 in the single-process body and 48 in the two-worker body), including sparse solvers. The split avoids this 14GiB host exhausting GPU-backed system memory; Intel stencils remain the existing skip.
Level Zero IPC metadata embeds a process-local file descriptor. Serializing its number makes imports fail on another rank; DAGGER_IPC cannot provide the missing SCM_RIGHTS transport. Disable eligibility, preserve the export/import primitives, and assert the fallback in the shared MPI suite. Validation: Julia 1.11 two-rank oneAPI suite passed 171 and 174 assertions with the legacy Level Zero adapter. The eligibility predicate infers Bool and measured 0 allocations / 0 bytes before and after over 10k calls (10 warmups, GC, minimum of five runs). Transfers now use existing host staging; the previously broken IPC transport cannot provide a valid throughput baseline.
The two-rank oneAPI suite now passes 171 and 174 assertions after disabling invalid IPC and selecting the legacy Level Zero adapter. Restore its dedicated test job with that adapter. Leave MPI benchmark jobs disabled pending their own validation.
The two-rank AMDGPU 2.8 suite passes but its handle-cache task finalizer calls a missing binding. Record that a zero exit status does not establish successful resource teardown, and locate the defect in the dependency before changing Dagger lifecycle code.
AMDGPU 2.1 on Julia 1.11 uses the embedded LLVM backend unless the external backend package is already loaded. That backend computes incorrect mixed-boundary halo indices on gfx1036. Install and load the modern backend before kernel compilation in ordinary/MPI tests and both benchmark revision environments, preserving the generic stencil implementation. Validation: AMDGPU 2.1 passed 281 GPU and 51 stencil assertions with external codegen; embedded codegen passed 49 stencil assertions and raised two bounds errors. AMDGPU 2.8 already passes the same bodies. Benchmark bootstrap selected external codegen on driver/two workers and passed a mixed-boundary stencil (four assertions). After 10 warmup batches and GC, minima of five runs over 100 kernel launches were unchanged: 5,504 allocations / 198,464 bytes before and after.
jpsamaroo
force-pushed
the
jps/fix-ipc-ownership
branch
from
October 9, 2026 00:24
e3a9775 to
21f5974
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes some old CUDA IPC code, and various other GPU-related issues.
Written by Claude