Skip to content

Fix GPU IPC and collect - #758

Open
jpsamaroo wants to merge 25 commits into
masterfrom
jps/fix-ipc-ownership
Open

jpsamaroo wants to merge 25 commits into
masterfrom
jps/fix-ipc-ownership

Conversation

@jpsamaroo

Copy link
Copy Markdown
Member

Fixes some old CUDA IPC code, and various other GPU-related issues.

Written by Claude

Base automatically changed from jps/mpi-full-threads to master September 30, 2026 23:15
@jpsamaroo jpsamaroo changed the title Fix GPU IPC Fix GPU IPC and collect Sep 30, 2026
@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Dagger benchmarks: dirty vs master

Counts are benchmark metrics (time, allocations, bytes). Timing changes require non-overlapping median ± IQR bands; both revisions need at least five timed samples, and timing regressions need an independent confirmation run.

Backend Regressions Improvements Within noise Inconclusive time Suites
Threads 0 0 8 2 4/4
Distributed 4 4 4 0 4/4
Distributed+Threads 0 0 1 4 4/4
MPI 0 0 0 0 4/4
MPI+Threads 3 5 18 2 4/4
Total 7 9 31 8 20/20

Regressions across all backends

Backend Benchmark Metric Change
MPI+Threads linalg/dagger/N=1024 (block 512)/cholesky time +186.2%
MPI+Threads linalg/dagger/N=1024 (block 512)/solve (A\b via lu) time +167.7%
MPI+Threads linalg/dagger/N=1024 (block 512)/lu time +123.9%
Distributed linalg/dagger/N=1024 (block 128)/syrk (A'*A) allocs +74.4%
Distributed linalg/dagger/N=1024 (block 128)/syrk (A'*A) memory +51.2%
Distributed linalg/dagger/N=1024 (block 128)/matmul (A*A) allocs +37.6%
Distributed linalg/dagger/N=1024 (block 128)/matmul (A*A) memory +36.2%

Improvements across all backends

Backend Benchmark Metric Change
MPI+Threads stencil/dagger/N=256 (block 128)/assign (const) time -74.5%
Distributed sparse/dagger/N=256 (block 16)/cg solve (laplacian) memory -61.5%
Distributed sparse/dagger/N=256 (block 16)/cg solve (laplacian) allocs -58.7%
MPI+Threads stencil/dagger/N=1024 (block 512)/multi-expr time -48.6%
Distributed sparse/dagger/N=1024 (block 64)/cg solve (laplacian) memory -48.4%
MPI+Threads array/dagger/N=256 (block 128)/transpose (permutedims) time -47.1%
Distributed sparse/dagger/N=1024 (block 64)/cg solve (laplacian) allocs -44.5%
MPI+Threads sparse/dagger/N=1024 (block 64)/spgemm (S*S) time -37.5%
MPI+Threads sparse/dagger/N=256 (block 16)/spmv (S*x) time -35.2%
Threads (1 process × 4 threads) — 0 regressions, 0 improvements
Suite Regressions Improvements Within noise Inconclusive time
array 0 0 3 0
linalg 0 0 1 0
sparse 0 0 0 0
stencil 0 0 4 2

Within noise

Benchmark Metric Change
array/dagger/N=1024 (block 512)/broadcast (X .+ 1) time +49.9%
linalg/dagger/N=256 (block 128)/syrk (A'*A) time +39.2%
array/dagger/N=1024 (block 128)/norm time +34.9%
stencil/dagger/N=1024 (block 512)/assign (const) time +25.3%
stencil/dagger/N=1024 (block 512)/neighbors (Reflect) time -25.5%
stencil/dagger/N=1024 (block 128)/assign (const) time -29.6%
stencil/dagger/N=256 (block 128)/assign (const) time -33.1%
array/dagger/N=1024 (block 512)/reduce (sum) time -63.0%

Inconclusive timing changes

Benchmark Metric Change
stencil/dagger/N=256 (block 128)/update (+) time +65.0%
stencil/dagger/N=256 (block 128)/multi-expr time +35.3%
Distributed (4 processes × 1 thread) — 4 regressions, 4 improvements
Suite Regressions Improvements Within noise Inconclusive time
array 0 0 1 0
linalg 4 0 3 0
sparse 0 4 0 0
stencil 0 0 0 0

Regressions

Benchmark Metric Change
linalg/dagger/N=1024 (block 128)/syrk (A'*A) allocs +74.4%
linalg/dagger/N=1024 (block 128)/syrk (A'*A) memory +51.2%
linalg/dagger/N=1024 (block 128)/matmul (A*A) allocs +37.6%
linalg/dagger/N=1024 (block 128)/matmul (A*A) memory +36.2%

Improvements

Benchmark Metric Change
sparse/dagger/N=256 (block 16)/cg solve (laplacian) memory -61.5%
sparse/dagger/N=256 (block 16)/cg solve (laplacian) allocs -58.7%
sparse/dagger/N=1024 (block 64)/cg solve (laplacian) memory -48.4%
sparse/dagger/N=1024 (block 64)/cg solve (laplacian) allocs -44.5%

Within noise

Benchmark Metric Change
linalg/dagger/N=1024 (block 128)/syrk (A'*A) time +135.7%
linalg/dagger/N=1024 (block 128)/qr time +40.6%
array/dagger/N=1024 (block 512)/broadcast (X .+ 1) time -40.7%
linalg/dagger/N=1024 (block 128)/solve (A\b via lu) time -61.2%
Distributed+Threads (2 processes × 2 threads) — 0 regressions, 0 improvements
Suite Regressions Improvements Within noise Inconclusive time
array 0 0 0 1
linalg 0 0 0 2
sparse 0 0 0 0
stencil 0 0 1 1

Within noise

Benchmark Metric Change
stencil/dagger/N=1024 (block 128)/multi-expr time +39.2%

Inconclusive timing changes

Benchmark Metric Change
array/dagger/N=1024 (block 512)/alloc (rand) time +142.3%
stencil/dagger/N=1024 (block 512)/neighbors (Reflect) time +47.0%
linalg/dagger/N=256 (block 256)/matmul (A*A) time +38.7%
linalg/dagger/N=256 (block 256)/lu time +37.5%
MPI (4 ranks × 1 thread) — 0 regressions, 0 improvements
Suite Regressions Improvements Within noise Inconclusive time
array 0 0 0 0
linalg 0 0 0 0
sparse 0 0 0 0
stencil 0 0 0 0
MPI+Threads (2 ranks × 2 threads) — 3 regressions, 5 improvements
Suite Regressions Improvements Within noise Inconclusive time
array 0 1 7 2
linalg 3 0 3 0
sparse 0 2 0 0
stencil 0 2 8 0

Regressions

Benchmark Metric Change
linalg/dagger/N=1024 (block 512)/cholesky time +186.2%
linalg/dagger/N=1024 (block 512)/solve (A\b via lu) time +167.7%
linalg/dagger/N=1024 (block 512)/lu time +123.9%

Improvements

Benchmark Metric Change
stencil/dagger/N=256 (block 128)/assign (const) time -74.5%
stencil/dagger/N=1024 (block 512)/multi-expr time -48.6%
array/dagger/N=256 (block 128)/transpose (permutedims) time -47.1%
sparse/dagger/N=1024 (block 64)/spgemm (S*S) time -37.5%
sparse/dagger/N=256 (block 16)/spmv (S*x) time -35.2%

Within noise

Benchmark Metric Change
linalg/dagger/N=1024 (block 512)/matvec (A*x) time +389.3%
stencil/dagger/N=1024 (block 512)/assign (const) time +124.7%
array/dagger/N=256 (block 128)/map (sin.(X)) time +119.6%
array/dagger/N=1024 (block 512)/broadcast (X .+ 1) time +103.9%
linalg/dagger/N=256 (block 256)/syrk (A'*A) time +65.6%
stencil/dagger/N=256 (block 128)/alloc (neighbors Wrap) time +56.0%
array/dagger/N=256 (block 128)/reduce (sum) time +54.3%
stencil/dagger/N=256 (block 256)/multi-expr time +48.1%
array/dagger/N=1024 (block 128)/broadcast (X .+ 1) time +44.2%
stencil/dagger/N=256 (block 128)/neighbors (Pad) time +44.1%
linalg/dagger/N=1024 (block 512)/matmul (A*A) time +41.4%
stencil/dagger/N=1024 (block 128)/assign (const) time -36.5%
stencil/dagger/N=1024 (block 512)/update (+) time -36.8%
array/dagger/N=1024 (block 512)/map (sin.(X)) time -38.3%
stencil/dagger/N=256 (block 128)/update (+) time -40.4%
array/dagger/N=1024 (block 512)/norm time -42.4%
array/dagger/N=1024 (block 512)/reduce (sum) time -48.2%
stencil/dagger/N=256 (block 256)/update (+) time -55.7%

Inconclusive timing changes

Benchmark Metric Change
array/dagger/N=256 (block 128)/add (X + X) time +106.4%
array/dagger/N=256 (block 256)/add (X + X) time +38.1%

Full results and plots (download the benchmark-results-* artifacts).

@jpsamaroo
jpsamaroo force-pushed the jps/fix-ipc-ownership branch from 05c62fe to 3f45600 Compare October 1, 2026 00:51
jpsamaroo and others added 23 commits October 8, 2026 14:35
Keep device arrays on their owners during datadeps slot creation and source
fetches for in-place copies. Passing Chunks through the backend transport
avoids deserializing an array while retaining a remote processor/context.

Replace CUDA's legacy borrowed IPC mapping with the shared staging/import
hooks. Hold export storage until the receiver has copied into its own array,
release it in finally, and close handles through the CUDA-version-aware
driver module. Small transfers and disabled IPC use host staging.

Add shared CI coverage for two workers using device 1 on CUDA, ROCm, oneAPI,
Metal, and OpenCL. Explicit placements exercise both transfer directions,
source independence, datadeps updates, broadcast, and tiled multiplication
on the backends supporting it. Record the ownership and IPC lifecycle rules.

Validation:
- Cyclops, Julia 1.12.6 / CUDA 6.4: unchanged EthanBug/repro.jl completes and
  its result matches A*B; all 39 shared CUDA regression checks pass.
- The old code fails the new broadcast with CUipcMemHandle undefined and
  tiled multiplication with the original remote-context assertion.
- OpenCL/pocl: all 38 shared checks pass; allocation suite: 24 checks pass.
- 100 local GPU moves, 10 warmup batches and minimum of 5 GC-counter runs:
  before 100 allocations / 1,600 bytes; after 0 / 0. The failing cross-worker
  baseline cannot supply a successful-operation allocation comparison.
`aliasing(::CuArray)` used `pointer(x)`, which calls `take_ownership!` for
the calling task's active device. On a worker owning two GPUs without P2P
access, the Phase-1 aliasing RPC runs on device 0 while the array lives on
device 1, so it threw "cannot take the GPU address of inaccessible device
memory". It also re-stamped stream ownership as a side effect. Aliasing only
needs the address, so read it from the underlying memory directly.

`pin_buffer!` called `CUDA.pin`, which dedupes per context, while host
registration is process-wide: moving one host array to two GPUs of one
process threw HOST_MEMORY_ALREADY_REGISTERED. Skip ranges that are already
page-locked, under a lock so two contexts can't race the check.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The same-worker device-to-device remainder path launches one kernel that
reads `from` and writes `to`, which needs both on one device. Two GPUs of
one worker need not be peer-accessible (e.g. 2x L40S), so a stencil halo
exchange between them failed in `pointer`. Take the direct path only when
source and destination share a memory space; other pairs use the existing
host-staged gather/scatter.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Distributed `collect` concatenated fetched tiles directly, so tiles on
different GPUs of one process met in one device-side `cat` kernel and
failed without peer access. The result is a host `Array` anyway, so copy
each GPU tile to host first.

Add a regression testset for two GPUs owned by one process covering
datadeps aliasing on the non-active device (and double host pinning), a
cross-device stencil halo exchange, and collect of mixed-device tiles. It
errors without the preceding fixes and passes with them on 2x L40S
(Julia 1.13.1, CUDA 6.4); single-GPU agents skip it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
GPU `execute!` returns with its kernels still queued on Dagger's per-device
stream. Distributed `collect` then fetched each tile and copied it to host
from the caller's task, i.e. on a different, task-local stream. That is only
correct when the array library synchronizes the previous owner on a
cross-stream access: CUDA, AMDGPU and OpenCL do, but oneAPI does not, and it
also copies on the calling task's device rather than the tile's. Passing
CUDA tests did not show that the wait existed.

Synchronize the tile's processor with `gpu_synchronize` and copy under its
`with_context`, on the chunk's owner. The wait then happens where the stream
lives, and a remote tile is sent as a host array instead of being serialized
as a device array. Plain values in `chunks` (e.g. from `sort`) pass through.

Add a test that collects right after a deliberately slow kernel, with tiles
owned by the driver and by remote workers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
One-shot HtoD sources were page-locked with CUDA.pin, whose unregistration
rides a GC finalizer. A registration could therefore outlive its transfer and
collide with whatever the allocator placed at that address next, giving
ERROR_HOST_MEMORY_ALREADY_REGISTERED, or ERROR_INVALID_VALUE once a stale
finalizer unregistered a newer owner's range mid-copy. Multi-chunk DArrays
pin many large tiles per build, so they hit it first.

Register through a Dagger-owned refcounted table (keyed by base pointer, since
registration is process-wide) and unregister in a `finally` after the copy has
synchronized, while the caller still holds the buffer. Overlap is detected from
the registration error itself (endpoint probes miss an interior registration)
and falls back to registering a private copy.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Same defect as the CUDA fix, in a harsher form. HtoD sources were registered
by `pin_buffer!`, which unregisters from a GC finalizer. AMDGPU's unregister
takes a lock, and a finalizer may not switch tasks, so under contention the
unregister was dropped ("task switch not allowed from inside gc finalizer")
and the range stayed registered after its array died, leaking pinned host
memory and exposing whatever the allocator put there next.

Track registrations in a Dagger-owned table and unregister in a `finally`
after the copy has synchronized. HIP registration is neither refcounted nor
exclusive (duplicates are idempotent and one unregister drops the pin), so we
never unregister a range someone else registered: a buffer whose first or last
byte is already pinned gets a private copy to register instead.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… finalizer

`pin_buffer!` page-locked staging buffers with `CUDA.pin`, whose GC
finalizer unregisters the range under a `ReentrantLock`. When that lock
is contended the finalizer throws "task switch not allowed from inside gc
finalizer" and the unregistration is lost, while GC frees the memory
anyway. The freed range stays registered with the driver: a later
allocation at that address fails to register or is DMA'd through stale
pages. A 4-GPU GEMM benchmark crashed this way (segfault a few iterations
in, during Datadeps' host-staged cross-device remainder copies).

ROCExt's `pin_buffer!` had the same flaw: its own finalizer called
`AMDGPU.Mem.unpin`, which takes AMDGPU's registration lock. 2-GPU
Datadeps tests on RX 6900 XTs segfaulted inside `hipMemcpyWithStream`
after the same finalizer error.

Register buffers ourselves and unregister from a finalizer that only
`trylock`s; on contention it re-arms itself, which keeps the buffer alive
until a later collection can unregister it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`aliasing(::SubArray)` located a view through `pointer` on its parent
and on the view. For a CUDA array that takes CUDA.jl stream ownership for
the *active* device, which throws when the array lives on another GPU of
the process without peer access (a view of a host array rebuilt on one
GPU while another GPU's task is current). Add a `data_address` hook,
defaulting to `pointer`, and compute the view's address from its
parent's; CUDAExt reads the raw address instead.

AMDGPU.jl's `pointer` likewise takes stream ownership (synchronizing the
stream the array was last used on), so ROCExt reads the raw address too,
in `aliasing(::ROCArray)` as well. MetalExt's `data_address` reports the
same address as its `aliasing(::MtlArray)`, so a view's spans line up
with its parent's. oneAPI.jl's and OpenCL.jl's `pointer` have no side
effects and keep the default.

Also fix `aliasing(::CuArray)`, which added `x.offset` (an element count)
to the pointer as bytes. (AMDGPU.jl's `offset` counts elements in older
versions and bytes in newer ones; ROCExt asks AMDGPU's `derive` which.)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…m host

Two host<->device paths broke on data Datadeps produces when it rebuilds
a view of a host array on a GPU:

- The out-of-place host-to-device `move` of a `Chunk` labelled with a
  host processor assumed the payload was host memory, but here it is the
  parent device array already moved to the GPU; page-locking that as
  host memory throws (CUDA: from `CUDA.pin`; ROCm: `hipHostRegister`
  fails with hipErrorHostMemoryAlreadyRegistered). Hand device arrays to
  the device-array `move`.
- In-place `move!` between a host array and a strided view of a device
  array used `copyto!`, which indexes the view element by element from
  the host (scalar indexing). Stage through a dense device buffer and
  gather/scatter with a broadcast kernel.

Both fixes apply to every GPU extension (CUDA, ROCm, oneAPI, Metal,
OpenCL), whose host<->device paths share this shape.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…r copies

Two GPUs of one process need not be peers (PCIe-attached GPUs behind a
host bridge often are not). CUDA.jl's `copyto!` between them then stages
through a fresh *pageable* host `Vector` with a synchronous DtoH, 1.4 GB/s
on L40S; Datadeps' remainder copies took a host round trip of their own,
and a KA kernel touching both buffers is illegal outright.
`cuMemcpyPeerAsync` works without peer access -- the driver pipelines it
through its own pinned buffers -- at 21.7 GB/s on the same machine.

Route every same-process device-to-device copy through it: in-place
`move!`, the out-of-place `move`, and Datadeps remainder copies, through a
new core hook `device_remainder_copy!`. Many short spans (halos) are
packed on the source and unpacked on the destination, so a copy is not
thousands of host-staged transfers.

The copies also run on a per-device copy stream, so they overlap the
destination's kernels instead of queueing behind them. They are ordered
per allocation rather than per device: the copy waits only for the
source's last write and the destination's last write and read, not for
kernels that merely read the data being sent (a GEMM panel's owner is
busy reading it while other GPUs ask for it). Each GPU task records an
event for the allocations it writes and reads, using a new `writes` task
option in which Datadeps lists the written positional arguments;
allocations with unknown history fall back to waiting on the whole device
stream. The copying task waits on the host for its copy, so every
existing synchronization point stays valid. Its own record then skips the
destination it filled (an event on the device stream would make later
copies out of it wait for unrelated kernels): the copy stamps the
allocation with its task, so the skip can never apply to a later
allocation at the same address -- an unnamed "skip the next write" flag,
left on short-lived packing buffers, made a later write go unrecorded and
a copy read stale data.

On 4x L40S, peer copies alone took a 4-GPU Float32 GEMM (N=56736) from
27 s per call to 4.3 s; overlapping them with compute took it from ~3.9 s
to ~3.1-3.4 s (under GreedyScheduler, with the next commit's traversal).

ROCExt gets the same machinery, with HIP events and `hipMemcpyPeerAsync`.
HIP does not stage through pageable memory (between two RX 6900 XTs
without peer access, AMDGPU's `copyto!` and `hipMemcpyPeerAsync` both
reach ~5.5 GB/s), but the old path synchronized the source's whole
device stream and queued the copy behind the destination's kernels,
and remainder copies still went through the host. A 2-GPU Float32 GEMM
(N=12288, GreedyScheduler) went from ~5.8-6.1 s to ~4.6-4.8 s per
call. The cross-GPU tests run for ROCm too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Host-to-device `move`s registered the caller's host array with
`hipHostRegister` for the length of one copy and unregistered it right
after (`pin_host!`/`unpin_host!`). Under multi-threaded Datadeps that
corrupted data: a 2-GPU Float32 GEMM (1024x1024, 256x256 tiles) read
whole tile columns back as NaN in 6 of 6 stress runs, on one GPU as
well as two, and under RoundRobin and Greedy alike. With the uploads
unregistered, 8 of 8 runs were clean; registering a private copy
instead was clean too, but uploaded at 0.8-7 GB/s.

Registration buys nothing here: HIP stages pageable uploads through its
own pinned buffers, at ~11 GB/s for 1-64 MiB on an RX 6900 XT, the same
as from registered memory. Upload directly. DtoH staging buffers
(`pin_buffer!`) stay registered.

It is faster too: registering and unregistering a tile per upload cost
more than the upload. A 2-GPU Float32 GEMM (N=12288, GreedyScheduler,
host-resident operands) went from ~4.6-4.8 s to ~1.7-1.9 s per call.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rkers

Use the active test environment on every worker to prevent stale LLVM and GPUCompiler modules after backend installation. Honor Pkg.test test_args so worker counts and test selections also apply to package-driven runs. CUDA CI suites pass: 281 GPU checks and 106 stencil checks.
The one-thread GPU run stopped in AMDGPU hostcall_host_wait with no exception. Two threads allow compilation and scheduling to continue on the other thread while the persistent host-call task polls. Apply this to the driver and Distributed workers.
MPI refs report this Distributed worker on every rank, so the old worker-id guard inspected remote placeholders during GPU copy tasks. Use the acceleration ownership predicate for CUDA and ROCm.

Validation: Julia 1.11 two-rank CUDA MPI suite passed (221/227 checks). After ten warmup iterations and GC, the minimum of five measurements for 10,000 local inspections was 0 allocations / 0 bytes before and after. The ownership predicate infers Bool; storage unwrap remains intentionally dynamic.
A standalone HIP program reproduces missing asynchronous kernel/copy writes on this gfx1036 host. Document the diagnostic serialization flags so this runtime issue does not become a production stencil synchronization cost.
GPU workers can abort outside Julia exception handling. Flush the selected suites and completed test summaries before entering the next suite so redirected CI logs preserve the last completed boundary.
On Xe2, the default v2 adapter aborts while SYCL wraps oneAPI.jl native queues for GEMM. The same oneAPI-only GEMM passes with the documented legacy adapter. Select that adapter consistently for Intel tests and benchmarks and record the diagnosis.

Validation: Julia 1.11 oneAPI GPU body passed all 281 assertions (233 in the single-process body and 48 in the two-worker body), including sparse solvers. The split avoids this 14GiB host exhausting GPU-backed system memory; Intel stencils remain the existing skip.
Level Zero IPC metadata embeds a process-local file descriptor. Serializing its number makes imports fail on another rank; DAGGER_IPC cannot provide the missing SCM_RIGHTS transport. Disable eligibility, preserve the export/import primitives, and assert the fallback in the shared MPI suite.

Validation: Julia 1.11 two-rank oneAPI suite passed 171 and 174 assertions with the legacy Level Zero adapter. The eligibility predicate infers Bool and measured 0 allocations / 0 bytes before and after over 10k calls (10 warmups, GC, minimum of five runs). Transfers now use existing host staging; the previously broken IPC transport cannot provide a valid throughput baseline.
The two-rank oneAPI suite now passes 171 and 174 assertions after disabling invalid IPC and selecting the legacy Level Zero adapter. Restore its dedicated test job with that adapter. Leave MPI benchmark jobs disabled pending their own validation.
The two-rank AMDGPU 2.8 suite passes but its handle-cache task finalizer calls a missing binding. Record that a zero exit status does not establish successful resource teardown, and locate the defect in the dependency before changing Dagger lifecycle code.
AMDGPU 2.1 on Julia 1.11 uses the embedded LLVM backend unless the external backend package is already loaded. That backend computes incorrect mixed-boundary halo indices on gfx1036. Install and load the modern backend before kernel compilation in ordinary/MPI tests and both benchmark revision environments, preserving the generic stencil implementation.

Validation: AMDGPU 2.1 passed 281 GPU and 51 stencil assertions with external codegen; embedded codegen passed 49 stencil assertions and raised two bounds errors. AMDGPU 2.8 already passes the same bodies. Benchmark bootstrap selected external codegen on driver/two workers and passed a mixed-boundary stencil (four assertions). After 10 warmup batches and GC, minima of five runs over 100 kernel launches were unchanged: 5,504 allocations / 198,464 bytes before and after.
@jpsamaroo
jpsamaroo force-pushed the jps/fix-ipc-ownership branch from e3a9775 to 21f5974 Compare October 9, 2026 00:24

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant