Skip to content

Disclose a retired chain's remaining keys after an executor restart - #378

Merged
ctfbruce merged 36 commits into
mainfrom
b71/tail-disclosure
Oct 7, 2026
Merged

ctfbruce merged 36 commits into
mainfrom
b71/tail-disclosure

Conversation

@ctfbruce

@ctfbruce ctfbruce commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

Executors with a configured TESLA seed now recover the immediately previous chain's pending disclosure tail after a restart. The recovered chain never signs; its due keys accompany the current chain's heartbeat disclosures and are verified against the dispatcher's recorded chain. Executor schema 8 records the disclosure delay required for recovery; upgrade the executor database before starting this build.

Startup removes stale legacy tagger filters only on the selected packet-counting interface and checks the remaining interfaces in the same network namespace. If prior signers cannot be retired, recovery is withheld. Recovered disclosure timing and the 24-hour retention limit follow the monotonic clock.

Recovery requires a matching seed, recorded disclosure delay and ready startup clock. It covers only the immediately previous generation; heartbeat success does not confirm durable dispatcher storage. Older chain records, repeated restarts before delivery and arbitrary backup rollback remain outside this increment. Part of #71; this PR does not close the broader issue.

Validation: focused retired-chain, heartbeat and disclosure race tests; storage policy and both database test suites; regenerated protobuf and SQL bindings from the combined sources. The required kernel lane checks unchanged-interface and changed-interface/fallback restarts.

TheodorAdrienIsaak Mattli added 2 commits October 5, 2026 10:39
Epochs advance on the executor's monotonic clock from the chain origin it
announces, the wall-clock reading at chain start, and a verifier maps a
capture's wall time to an epoch from that origin. Two clock conditions
make that mapping wrong and were not detected: a chain started while the
host clock was not ready announces an origin that may be wrong, and a
wall-clock step or a suspend afterwards moves wall time away from the
monotonic time the epochs follow. The executor kept tagging and reported
attribution as available although honest packets mapped to the wrong
epoch.

The key schedule now reads the host clock readiness at chain start and
measures the drift between wall and monotonic time elapsed since the
origin. While the origin is untrusted (clock_unready, for the chain's
life) or the drift exceeds min(epoch length / 2, 1 s) (clock_drift), no
key signs: CurrentKey, the one decision point of both the pure-Go and the
kernel signing path, returns nothing, so the pure-Go tagger stops at once
and a kernel tagger's slot is removed at its next refresh. The schedule
also never signs with the key of an epoch below one it has already signed
with. Attribution carries the two reasons, the dispatcher accepts them, a
node that tags packets refuses new runs with FailedPrecondition while
either holds, and a node whose tagging mode is none admits them. The
bound is a safety threshold, not a measured uncertainty bound; recovery
is a restart with a ready clock, which starts a new chain.

GitHub issue #72.
Keys live only in the running schedule, so an executor restart before a
chain's final disclosure left the keys of its last epochs unpublished
and the packets they tagged unverifiable. With a configured seed the
previous chain is now re-derived at start from the seed and its recorded
generation, origin, epoch length, chain length and disclosure delay,
which the executor records with each chain from now on, and only when
the re-derived anchor equals the recorded one, the final key is not yet
due and the host clock is ready. The re-derived schedule is disclosure
only: it never returns a signing key, so a recorded generation never
signs again; the heartbeat carries its due key beside the current
chain's until the final key has been disclosed, and the dispatcher
stores such keys under the retired chain's anchor through the path it
already used for anchored disclosures, at most four per heartbeat.

The executor schema records the disclosure delay and its minimum rises
with the migration, since startup writes the column; chains recorded
before it, restarts without a seed and starts on an unready clock still
lose the tail, and the log says why.

GitHub issue #71.
@ctfbruce
ctfbruce changed the base branch from b72/clock-attribution to main October 5, 2026 09:26
@ctfbruce ctfbruce closed this Oct 5, 2026
@ctfbruce ctfbruce reopened this Oct 5, 2026
TheodorAdrienIsaak Mattli and others added 25 commits October 5, 2026 11:48
The installed-upgrade acceptance read the schedule history of a released
executor database with the generated chain query, which now selects the
disclosure delay that a database of the released schema does not have.
Read the six original columns with a plain statement before and after
the upgrade, so both sides of the comparison use the same shape on both
schemas; the new column is covered by the executor's own tests.

GitHub issue #71.
The pure-Go tagger read the schedule's current key once to decide
whether to tag a packet and again, through the per-packet tag
derivation, to derive the tag. With the schedule refusing any epoch
below the highest one it has already signed with, another signing path
reaching the next epoch between the two reads turned a valid packet
into a datagram write error on a healthy clock. The tag is now derived
from the key the first read returned; the epoch fence is unchanged.

The two clock reasons an executor may report now have their own
attribution states in the dispatcher's health metrics and the
executors_attribution_state labels, so every registered executor is
counted in exactly one state again, and the nodes table of the command
names them. The protocol comment, the documentation of a transient
drift that returns within the bound and of the kernel removal timing
are corrected to match the behaviour.

GitHub issue #72.
A legacy tc filter outlives the process that attached it and keeps
signing with its last installed key, so a restart that re-derived the
previous chain could disclose a key a stale filter still used. Each
start on a configured interface now removes the tagger filters an
earlier process left on its egress hook, before this process attaches
any tagger and before the previous chain is re-derived; when a filter
cannot be listed or removed, or a foreign filter sits at the tagger's
priority, no tail is disclosed and the reason is logged. A TCX
attachment ends with its process and needs no sweep.

The re-derived chain advanced on the wall clock, since its recorded
origin carries no monotonic reading. It now advances on the monotonic
clock from the recovery instant, whose wall reading, taken on a ready
clock, places the origin: a wall step after recovery neither advances
nor delays a disclosure, and drift is still reported.

The chain was dropped one epoch after its final disclosure whether or
not a heartbeat had delivered the final key. It is now kept until a
heartbeat that carried k_{L-1} under its anchor succeeded; a later key
covers every earlier one, so only that one must arrive. A chain whose
final key no heartbeat delivered within 24 hours of its final
disclosure is dropped with a warning, and a start within that bound
still re-derives a chain whose final key is already due, so the first
heartbeat carries it.

GitHub issue #71.
Take the single signing decision per packet and the clock reasons in the
health metrics from the corrected base. The schedule, the protocol
definition and its generated bindings are changed by both sides in
different hunks and merge without conflict. The disclosure-only
schedule's test asked the per-run derivation the base removed; it now
asks ComputeTagForPacket, which reads the same key.

GitHub issue #71.
Take the disclosure-lag gauges, the counter-attachment and resource
metrics and the admin account recovery from main. The attribution
state series keeps the clock_unready and clock_drift labels next to
the states main lists, in the metrics table, the exporter's label list
and the two tests that cover them; the other metrics tests from main
stay as they are.

GitHub issue #72.
Take main's disclosure-lag gauges, counter-attachment and resource
metrics and admin account recovery through the corrected base, whose
attribution state series keeps the clock labels.

GitHub issue #71.
Add the OpenSSH allowed_signers trust file for releases@debuglet.netsec.ethz.ch
(Ed25519, SHA256:vETQG+wE6uwMv4MBFfx7dxo8IWpvR8z6/Sw7MT2WwQo) and document how
to confirm and install it. The private key exists only as the release-signing
environment secret.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Publish the production release signer
Production executors all report enforcement_mode fallback with
enforcement_reason not_permitted and tagging.ipv4 none. The role gave the
binary only cap_net_admin,cap_bpf; on kisti-ams (5.15,
unprivileged_bpf_disabled=2) the tagger load then fails in the verifier
with "R2 pointer comparison prohibited", which needs CAP_PERFMON, and
without CAP_NET_RAW the pure-Go fallback tagger and guest ICMP have no raw
sockets either.

- Grant exactly executor_capabilities (cap_net_admin, cap_net_raw,
  cap_perfmon, cap_bpf) with setcap, which replaces rather than adds, in
  the executor role and the serial rollout, and restart the executor when
  they change. The ambient set in the unit follows the same list.
- Do not grant cap_sys_resource: from 5.11 BPF memory is memcg-charged and
  cilium/ebpf leaves RLIMIT_MEMLOCK alone once cap_bpf lets its probe map
  succeed. On an older kernel the unit sets LimitMEMLOCK=infinity instead.
- verify.yml fails on an executor binary without the exact set; dbl doctor
  no longer passes CAP_BPF without CAP_PERFMON.
- Render [executors."<executor_id>"] operator labels on the dispatcher
  from per-host executor_display_* inventory variables, checked by the
  preflight for every executor before the dispatcher is touched.
- Document executor_public_host as the advertised address ip_metadata
  needs, check its shape, and default the IPv4 connectivity reflector to
  the dispatcher's gRPC endpoint when dispatcher_addr is an IPv4 literal,
  which needs no new port. Listener ports stay off.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The provisioner image checks requirements.yml against
DEBUGLET_PROVISIONER_COLLECTIONS_SHA256, so a comment edit is a
provisioner change. Its note on the capabilities module is now stale;
updating it belongs with a provisioner rebuild.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…adata

Grant executors their eBPF capabilities and render vantage metadata
API 1.16 adds is_public, the last dispatcher-observed address_v4 and
address_v6, their prefix and ASN from the existing offline ASN lookup,
and address_observations (via control or reflection, observed_at and
the lookup source) to GET /executors, and records the same fields in
the result's admission snapshot.

Only addresses the dispatcher saw on the executor's authenticated
connections are listed: the control connection's peer, preferred for
its family, and the peer of the existing ReflectAddress call for the
other family. Executors now make that call over each family every ten
minutes, resolving dispatcher.addr and verifying the dispatcher's
certificate for its name; [connectivity] observe_addresses turns it
off, and a family with a configured reflector is not called twice.
Observations end with the control session, like connectivity proof.

[metadata] address_opt_out makes an executor private, as a RIPE Atlas
host can: the public listing withholds its addresses and nothing else.
It is independent of location_opt_out. Executors are public by
default, including older ones, so the privacy note and changelog say
how to stay private before the dispatcher is upgraded. dbl nodes adds
ADDRESS_V4 and ADDRESS_V6 columns.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The compatibility canary pins the exact JSON shape of dbl nodes and
refused the new is_public, address, prefix, ASN and
address_observations fields. Allow them, null for an older
dispatcher, and check the observation records' shape.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
List executors with RIPE Atlas-style addresses (API 1.16)
API 1.17 adds status (connected, disconnected, abandoned after 30
days, never_connected), status_since, first_connected,
last_connected, total_uptime and tags to GET /executors, and a status
query parameter that also lists executors that are not connected. The
default listing is unchanged and degrades to registration times if the
history cannot be read.

Dispatcher schema 24 keeps probe_status and probe_addresses. A
registration records the executor as connected; once a minute the
maintenance loop advances uptime and address runs of registered
executors and closes the others at their last record, which also closes
rows a stopped dispatcher left open. A registration within two minutes
of a connected row continues its streak. The migration fills
probe_status from the recorded TESLA chains, and the build requires
schema 24.

Host tags come from a fixed vocabulary in [metadata] host_tags, also
set with dbl executor join --host-tag; executor and dispatcher both
validate them. System tags are derived from connectivity and the
observed addresses (works, capable), the address history (stable
1d/30d/90d), and the executor's address self-check: whether its local
IPv4 source toward the dispatcher is RFC 1918, reported as a boolean
only, and whether its resolver's A or AAAA answer for the dispatcher
name passed the dispatcher's TLS check. dbl nodes gains --status and
STATUS and TAGS columns, and the client gains Probes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Record executor status history and RIPE Atlas-style tags (API 1.17, schema 24)
Follow RIPE Atlas: an address belongs to the longest prefix announced in
BGP and to that prefix's origin AS. `debuglet-dispatcher
-build-asn-database PATH` downloads RIS's daily whois dumps and RIPE's AS
names, keeps (origin, prefix) pairs at least -ris-min-peers RIS peers see
(default 10), attributes a MOAS prefix to the origin the most peers see
(lowest AS number on a tie), and writes a GeoIP2-ASN-compatible MMDB of
type Debuglet-RIS-ASN whose build epoch is the dumps' generation time. The
record also carries announced_prefix_length so the reported prefix is the
announced one, not a tree node narrowed by more-specifics or widened by
merging. The file is verified through the dispatcher's own loader before
it is renamed into place; short downloads, failed gzip checksums, dumps
without their end marker or too old, and near-empty tables are refused.

The dispatcher checks its configured metadata databases every minute and
switches to a replaced file after verifying it as at startup, keeping the
previous database on failure. Executor control connections are not
dropped; registrations after the switch use the new file and existing
snapshots keep their source.

Ansible installs a daily debuglet-ris-asn timer on the dispatcher host,
builds the first database before the configuration names it, and enables
this for prod and dev.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
vincent10400094 and others added 2 commits October 6, 2026 15:48
Derive executor ASN and prefix from RIPE RIS, reloaded without restart
Chain payments are listed as unsupported in this release; every shipped
configuration keeps them disabled.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@ctfbruce
ctfbruce marked this pull request as ready for review October 7, 2026 07:42
@ctfbruce
ctfbruce merged commit ca27662 into main Oct 7, 2026
16 checks passed
@vincent10400094
vincent10400094 deleted the b71/tail-disclosure branch October 8, 2026 00:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants