Skip to content

64-bit object id redesign: design documents (#3994) - #4015

Draft
lvkale wants to merge 18 commits into
reviewed-with-reconversefrom
objid-redesign
Draft

lvkale wants to merge 18 commits into
reviewed-with-reconversefrom
objid-redesign

Conversation

@lvkale

@lvkale lvkale commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Status and merge gate (2026-10-09). This PR now carries the PR 1a implementation (#4017, squash-merged into objid-redesign) besides the documents, and will carry 1a's fixes (#4028) and 1b (#4021) once they merge there. It leaves draft and merges as one unit, never 1a alone (1a without 1b leaves hashed-kind arrays with the initial tranche only and an abort on restart with a different process count), when:

  1. Object id PR 1a fixes: demand creation by id, buffered index requests, synchronous LB location update #4028 is approved by Aditya and merged into objid-redesign;
  2. Object id PR 1b: tranche allocator on PE 0, refills, deferred insertion, restart with any process count #4021 is rebased onto that, passes its tests including the restart ones, and merges into objid-redesign;
  3. ChaNGa (a ~262k-element bounded array, which today prints the PR 0 "needs 17 bits ... hold 16" note) is rebuilt and run against objid-redesign by Kale or Ritvik, with the note gone and the run unchanged.
  4. A full sweep, since the id layout is a deep change: every test under tests/charm++ and tests/converse, every example under examples/charm++ with a test target, the benchmarks we run (pingpong, jacobi), the ten PPL mini-apps at https://charmplusplus.org/miniApps/ (LeanMD, AMR, Barnes-Hut, DenseLU, HPCCG, Kripke, TS, FFT, RA, EP Stream), and the applications paratreet2 and ChaNGa, each in single-process multi-PE and multi-process configurations, with a balancer where supported and checkpoint/restart where supported, on the Mac and on Anvil. CI's subset is not enough for this merge.

PR 2 (manual) follows as a small PR to the reviewed line; PR 3 and PR 4 are their own PRs there, not passengers here. A checkpoint written before PR 1 cannot be restored after it (id layout and checkpoint format change; pupLayout's marker makes that a clear abort): the release notes must say so.


Design documents for the 64-bit object id redesign discussed in #3994, on the objid-redesign branch where the implementation will be developed. Opened as a draft; the merge gate is below (design doc section 9.1). The documents are in the repository so the work does not depend on anyone's laptop.

doc/objid64-design.md (the design):

  • Id = tag 3 | collection 12 (build-time CMK_OBJID_COLLECTION_BITS; AMPI may need 21 or 24) | payload 49. CMK_OBJID_HOME_BITS goes away: no PE number is stored in the id.
  • Bounded arrays pack their index into the whole 49-bit payload (today: 16 bits).
  • Unbounded arrays: payload = home key of the index (width H fixed per run from the process count and an expand factor, persisted in checkpoints) + a unique number (49-H bits) handed out in tranches: deterministic initial tranche per process, refills from an allocator on PE 0. Local insertion stays synchronous.
  • One home per element, computable from the id or from the index alike, keyed to a process. Delivery by id never needs the index off the source PE; the linear reverse scan goes away. Aditya's handleUnknownByID (e417584) is folded in.
  • Ids are stable across checkpoint/restart including shrink/expand; directories are rebuilt, not restored.
  • A per-process directory and location cache follows as PR 3 (section 7 analyses which tables stay per PE and each race in the shared table).
  • PR sequence (section 9), test matrix (10), open questions for Aditya (11).

doc/objid64-analysis.md (why): the id -> home -> index -> home -> location chain as the code runs it today, the defects that follow from having two homes (D1-D7), and each scheme considered with why it was kept or dropped, including the rejected 128-bit option.

Aditya has confirmed the design subsumes the location-manager work on his rate-aware-gpu-lb branch; ideas and code from there will be cherry-picked as the implementation proceeds. Kale and Eric Bohm for design review.

PR 0 from the design (overflow abort in getNewObjectID, the no-compressor note, and a migrating test for a non-compressible index) changes no behaviour and can land on the reviewed line ahead of everything else.

🤖 Generated with Claude Code

@lvkale

lvkale commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

Rebasing objid-redesign (and the stacked objid-pr1a-layout / objid-pr1b-allocator) onto reviewed-with-reconverse d17228d now, to pick up the checkpoint-restart fix (#4019): on Anvil the 1a+1b branch passed everything except restarts with >1 process and >1 PE per process, which hang in exactly that pre-fix path. Force-pushes follow; no content changes beyond the rebase.

lvkale and others added 11 commits October 1, 2026 01:17
…rray's index is not packed into ids

An array with no index compressor mints element ids from a per-PE counter in
the element field (16 bits by default). Nothing checked the counter: past
65,535 elements on one PE it carried into the home field, so ids silently
collided with another PE's and every home lookup for them went to the wrong
PE. getNewObjectID now aborts at the limit, naming the location manager and
the remedies (insert from more PEs, setBounds, fewer collection bits). The
message is kept under 255 characters because reconverse's CmiAbort formats
into a 256-byte buffer.

Whether an array gets a compressor is easy to get wrong by accident: the
sized CkArrayOptions constructors set bounds, but setNumInitial/setEnd do
not, and bounds that need more than CMK_OBJID_ELEMENT_BITS are refused
silently. The CkLocMgr constructor now prints one note on PE 0 in either
case, with the bit count it needed or the setBounds remedy. A public
FixedArrayIndexCompressor::bitsNeeded(bounds) reports the width; make()
uses it too.

Diagnostics only; no change to ids or delivery. Part of the object id work
tracked in #3994 (PR 0 of doc/objid64-design.md on the objid-redesign
branch).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… index is not packed into ids

Every migrating test in the tree uses a small bounded array, whose index is
compressed into the element id, so the other id scheme (per-PE counter plus
an index<->id map in CkLocMgr) was never exercised together with migration.
This test creates a 2D array with no bounds and no initial size, inserts
every element dynamically from PE 0, and runs a message ring in which each
element migrates to the next PE every five steps; it ends by quiescence.
Default 4x3 elements, count 40; arguments nX nY count.

Registered in the regular test list (anytime_migration is in FTDIRS and
only runs under syncfttest).

Part of the object id work tracked in #3994.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Two documents for the object id redesign discussed on charm #3994:

- doc/objid64-design.md: the detailed design. 64-bit id with a 12-bit
  collection field (build-time) and a 49-bit payload; bounded arrays pack
  their index into the whole payload; unbounded arrays carry a hash key of
  the index (width fixed per run from the process count and an expand
  factor) plus a unique number handed out in tranches, so the home of an
  element is computable from the id or the index alike and no PE number is
  stored in the id. Ids are stable across restart and shrink/expand;
  directories are rebuilt. Delivery by id never needs the index off the
  source PE. A per-process directory and cache follow as a later PR.
  Includes the PR sequence, test matrix, and open questions for Aditya.

- doc/objid64-analysis.md: the analysis behind it: the id -> home -> index
  -> home -> location chain as the code runs it today, the seven defects
  that follow from having two homes, and every scheme considered with why
  it was kept or dropped, including the 128-bit option that was rejected.

Merging the implementation is deferred until the reviewed line is stable
with users and the test suite is broader; the diagnostics and the
hashed-kind migration test (PR 0 in the design) can land first.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ranch

The object id redesign (#3994, design in doc/objid64-design.md) is developed
as a stack of pull requests whose base is objid-redesign, merged to the
reviewed line once at the end. The reconverse workflows trigger only on the
reviewed line, so those pull requests and pushes to the branch would get no
CI. Add the branch to the push and pull_request triggers of both workflows.
Drop this commit, or leave it (harmless), at the final merge.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ge within a run)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ag bits

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…kCallback variant as a later item

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…hared-table ordering, LB global update epoch, restart)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Kale's point (2026-09-30): a message for an element that lives on another PE
of the same process must never go to the index home for its location. Stated
as five guarantees (resident elements resolve locally; departed elements leave
a forwarding entry; one request per process; the home answers from the shard
on any rank; same-PE delivery unchanged) with what each needs from the shared
idx -> id and id -> location shards.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… re-send, lock-free minting, allocator clamp, restart without waiting)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…esign 2 and 4.3 note the sender-repair rule applies to both kinds

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
lvkale and others added 4 commits October 1, 2026 01:28
…ed objid-pr* branches

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…es, sweep on a cap, what eviction costs)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…s with per-process tranches, delivery by id (#4017)

* Object id: no home PE in the id; packed indices get 49 bits; hashed ids with a per-process tranche; delivery by id

PR 1a of the object id redesign (#3994; design in doc/objid64-design.md on
this branch, sections 1-6). The layout, the home function and the delivery
path change here; the tranche allocator with refills and the per-process
directory follow in later changes.

Layout (objid.h). The id is 3 type tag bits, CMK_OBJID_COLLECTION_BITS
(default 12, build-time) and a 49-bit payload. CMK_OBJID_HOME_BITS is gone:
no PE number is stored in the id. An array with bounds that fit packs its
index into the whole payload (16 bits before). An array without a usable
compressor gets a "hashed" payload: a hash key of the index in the top bits
and a number unique within the array in the low bits. The key width is fixed
per run from the process count and the +objid_expand factor (default 8:
the presumed ceiling on later expansion), stored in ck::objid::Layout, one
copy per process, and written into every checkpoint with the readonlies so
a restart keeps the layout the ids were minted under.

One home per element (cklocation.h). CkLocMgr::homePe(idx) and homePe(id)
now agree by construction: a packed-index array uses its map, as before,
decoding the index from the id when needed; a hashed array hashes the index
to a process (key % CkNumNodes(), rank 0 of it) from the id or from the
index alike. The old CkLocMgr::homePe(id), which read home bits, and the
creator-PE home of counter-minted ids are gone, and with them the two homes
that disagreed for every dynamically inserted array. CkLocCache gets its
manager back-pointer so it can compute the home of an id.

Minting (cklocation.C). A hashed id is minted synchronously on the inserting
PE from its process's tranche of unique numbers: the lower half of the
unique space split evenly among the processes present at launch, with one
atomic cursor per process per array. Exhausting the initial tranche aborts
with a message; refills come with the allocator in a later change. The
per-PE idCounter and its silent overflow are gone. The index hash is a
splitmix64 mix of the index words, not CkArrayIndex::hash().

Delivery (ckarray.C). A receive-side cache miss forwards to homePe(id)
through handleUnknownByID (Aditya Bhosale's e417584, folded in): the index
is never recovered off the source PE, so lookupIdx loses its linear scan of
idx2id and aborts if asked about an element this PE has never seen. A sender
that forwards via the home now also asks for the location once
(requestLocationOnce), so later sends do not repeat the detour. The
disabled "after two hops, send home" rule is enabled: sound now that the
home of an id is the home of its index.

Load balancer global update. Under CMK_GLOBAL_LOCATION_UPDATE the emigrating
PE broadcasts the new location with the element's true epoch, in place of
UpdateLocation fabricating one from each receiver's stale cache (which let
an older reply overwrite a newer location).

Restart. Ids never change: packed ids are functions of the index, hashed
ids are restored verbatim. Directories are rebuilt, not restored: restore()
now informs the home like resume() does, and the __FAULT__ branch of
CkLocCache::pup that shipped the home's remote entries is dropped. The
process's tranche cursor travels with rank 0's branch. Until the allocator
lands, restarting with a different number of processes aborts for hashed
arrays, with a message naming the array; packed arrays restart on any
process and PE count, and hashed arrays on any PE count.

New test tests/charm++/unbounded_restart: checkpoint an unbounded array,
restart on the same and on half the PEs, ping every element.

Tested on reconverse-darwin-arm8 (production), single process and with 2-3
processes: pingpong, megatest, hello/4darray, startupTest, anytime_bcastred,
unbounded_migration (also on 6 PEs across 3 processes), the three
load_balancing tests with RotateLB, chkpt (same count, and 4->2 and 2->4 PEs
in one process), unbounded_restart, and a 200,000-element unbounded insert
from one PE that the counter scheme aborted at 65,536. Not tested: a
2-process x 2-PE restart, which hangs on the unmodified branch as well.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* ckarray: the home must not redirect a multi-hop message to itself

recvMsg's "after two hops go home" rule replaced the cached PE with the home PE
whenever the cached PE was not this PE. At the home itself that made the home
re-send the message to itself forever, so an element that had migrated twice
never received it (anytime_bcastred hung at step 3 in the reconverse stress
job, seed 1; every seed on the Mac). Apply the rule only on PEs other than the
home; the home follows its own entry.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* ckarray: a demand-creation request also asks the home for the location

bufferForCreation buffered the message and sent requestDemandCreation to the
home, but for createhome the request named the home itself as the creation
PE, so after creating the element the home told nobody. The requester's
bufferedCreationMsgs entry was never flushed: only messages sent from the
home PE ever reached a demand-created element, on the pre-redesign runtime
too. The packed kind hid it because the source holds the id and forwards the
small message to the home instead of buffering.

Pair the request with a location request by index (new
CkLocMgr::requestLocationAtHome); the home buffers it until the element
exists and its reply flushes the buffered messages. Found by the new
tests/charm++/objid_insert test on the hashed kind.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* tests: objid_home checks the id and home invariants from every PE

Packed 3D array with 64^3 bounds (18 index bits, beyond the old 16-bit
budget) and a hashed 2D array. From every PE, for every element: homePe(idx)
== homePe(id); packed ids are computable on PEs that never saw the index;
hashed ids satisfy key == indexHashKey(idx) and home == rank 0 of key %
CkNumNodes(); all PEs agree on every home; ids are distinct and each hashed
unique part lies in its creating process's initial tranche. Repeated after
two migrations of every element: ids and homes unchanged.

The tranche-size check recomputes the file-local formula of cklocation.C
(initial region halved and split by process count); update it together.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* tests: objid_insert covers insertion, cold sends, stale caches, demand creation, deletion

Runs in both kinds (-b for packed). Every PE inserts elements locally (and
checks ckLocal() is non-null on return) and on the next PE; every PE then
sends by index to every element, including ones it never touched; elements
migrate to random PEs and the sends repeat over one-hop and two-hop stale
caches (the recvMsg "more than one hop, go home" rule), five cycles; a
[createhome] entry sent from every PE to never-inserted indices must create
each element exactly once at its home; half the elements destroy themselves
and the sends repeat over the survivors.

Uses thisProxy[thisIndex].ckDestroy(): a direct ckDestroy() inside a
broadcast entry method segfaults in CkArrayBroadcaster::attemptDelivery on
the base runtime as well (to be filed separately).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* tests: anytime_bcastred -u runs the migration stress on an unbounded array

With -u the array is created without bounds and inserted dynamically, so
its element ids are the hashed kind (home = rank 0 of the process chosen by
the index hash) instead of the packed index. Behaviour without -u is
unchanged. The test targets and the reconverse stress job run both.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* tests: objid_sections multicasts to sections of a hashed-kind array under migration

Two overlapping sections of a 32-element unbounded array are built once;
for 30 rounds main multicasts to both, each member checks it was reached
through a section it belongs to and that its id is unchanged since
insertion, section reductions are checked, and then every element migrates
to a random PE. Section cookies and spanning trees hold element ids, which
the redesign keeps stable for the element's lifetime.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* ckarray: deliver to the bound element just demand-created, not the stale null pointer

deliverInline's bound-sibling branch demand-creates the missing element and
then delivered to the pointer looked up before the creation, which is still
null; CkAssert(elem) in deliverToElement is compiled out in production, so
the first message to a demand-created bound sibling dereferenced null (also on
upstream main and classic netlrts). Re-look the element up after creation and
abort with a message if it is still absent. Found by tests/charm++/objid_bound.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* tests: objid_bound covers bound arrays on a hashed-kind array under manual and LB migration

A is unbounded (hashed ids); B is bound to A and inserted; C is bound to A,
never inserted, and reached only through a [createhome] entry, so its
elements are demand-created on the sibling's PE by deliverInline's
bound-sibling branch (the one remaining user of the reverse lookup). Each
round A[i] messages B[i] and C[i], which check they are co-located with A[i],
and B[i] messages A[i+1]; then A migrates (migrateMe, or AtSync with
+balancer RotateLB / GreedyLB) and B and C must follow: their migration
counts are enforced equal to A's. Runs single-process and under lcrun.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Brings in #4016 (PR 0, squash-merged there; this branch had the original
commits), #4014 CkTreeCacheManager, #4020 trace-summary message bytes, and
the 2026-09-30 triage documents. The three conflicts were PR 0 text that
PR 1a (#4017) had already rewritten; the objid-redesign side is kept in
each, so the 1a files are unchanged by this merge.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@lvkale

lvkale commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

Status 2026-10-09: since the squash merge of #4017 into objid-redesign this afternoon, this PR carries the PR 1a implementation in addition to the design documents. Aditya's must-fix items on #4017 are not yet addressed; a fix PR against objid-redesign follows and needs his approval. This PR stays a draft and must not be merged before that lands and the reviewed line is stable with users, as the description says.

The reported conflicts are PR 0: it was squash-merged on the reviewed line as #4016 while this branch keeps the original commits. Merging reviewed-with-reconverse into objid-redesign clears them.

lvkale and others added 3 commits October 9, 2026 17:43
…= 1a + fixes + 1b as one unit; PR 2-4 as their own PRs)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…niapps, two machines) to the merge gate; ChaNGa run by Kale or Ritvik

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…usplus.org/miniApps

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant