Skip to content

ePBS (Gloas) support, part 3: validators' builder config from the Commit-Boost config - #508

Closed
JasonVranek wants to merge 4 commits into
epbs-pr2from
epbs-pr3
Closed

JasonVranek wants to merge 4 commits into
epbs-pr2from
epbs-pr3

Conversation

@JasonVranek

@JasonVranek JasonVranek commented Oct 9, 2026 •

Copy link
Copy Markdown
Collaborator

No description provided.

`commit-boost builder-config` works out each validator key's keymanager
builder config (keymanager-APIs #88) from the Commit-Boost config, routed
as Commit-Boost routes it: a mux key gets its mux's relays, any other key
gets [[relays]]. Mux keys are resolved with Commit-Boost's own startup
code, loaders included, without the relays' header secrets, and with
rpc_url only for a registry loader. Two muxes with one id are refused.
Each relay hostname is one entry, at Commit-Boost's advertised URL, with
the hostname as auth data. Every value is written, so the validator
client's own builder defaults never apply: unset, builder_boost_factor
is 100, min_bid is 0 and the cap is unclamped.

`print` writes the configs as one JSON document, each mux's config once
with its named and loader-fetched keys and the [[relays]] config as the
default, for an operator's own tooling to send through each validator
client's keymanager API, as it sends fee recipients. stdout carries only
the document.

`apply` writes them to the validator clients given as repeated
--vc <keymanager URL>=<token file>, from the config (--config, else
CB_CONFIG) or from a printed document (--from, or - for stdin), so it
also runs where the validator client is and the config is not. The
document must be as print wrote it, for the same --advertised-url, so a
misspelled field cannot drop out of the POST. Each key is written only
to the clients that list it, since some clients accept a write for any
key, and --preserve-entries keeps entries at other URLs.
- Each write names its mux or [[relays]]; each client written to gets a
  line, and a closing tally counts keys written and not written, errors
  and warnings.
- Mux keys no client lists, and keys more than one client lists, are one
  error each that lists the keys. Keys only a loader lists are one
  warning, and both name any client that could not be listed. --partial
  counts unheld mux keys, and a client with no keys, as warnings.
- One client given twice under loopback names is refused. A client that
  leaves 3 writes in a row unanswered, or answers 403 before any write
  succeeds, as Lodestar does under --proposerSettingsFile, gets no
  further writes, and a Lodestar cap refusal stops capped writes to it.
- It exits 2 when it stops before contacting a client, on a bad flag, an
  unreadable, empty or multi-line token file, a config or loader error,
  or a printed document that does not match, and 1 on an error after.
  The loaders' own warnings reach stderr, which a RUST_LOG set for
  another program cannot hide, and a closed pipe cannot stop a run.

Commit-Boost parses the new keys but does not act on them: [pbs]
max_execution_payment_gwei, min_bid_p2p_eth and
builder_boost_factor_p2p; [[mux]] min_bid_eth and builder_boost_factor;
relay max_execution_payment_gwei. builder-config refuses a [pbs] or
[[mux]] key that Commit-Boost does not read, so a typo cannot write a
default.

Commit-Boost parses each [[mux]] on its own when loading its config,
since the flattened mux list reads a mux that fails to parse as no muxes
at all, which left the keys on [[relays]] without an error. A `mux` that
is not an array of tables, such as [mux], is refused too. An ETH amount
written as a TOML float parses from its decimal form, so 0.009 is
exact, and a negative or non-finite amount is a config error.

The ePBS page's setup step 2 gains a builder-config walkthrough: which
subcommand fits each setup, the printed document and the contract for
tooling that sends it, each client's keymanager API and token, how a
proposer settings file overrides the API, Docker recipes and a
Kubernetes sidecar. The Docker, binary, Kubernetes and configuration pages
and the README point ePBS users at it.
Commit-Boost tells apart a bid or preferences request whose auth data
names none of its key's relays. Auth data naming Commit-Boost itself, by
the request's Host header, is a key with no builder config, or Prysm
sending its entry URL for an entry without auth data: it gets 400 at
once instead of a dial. Auth data naming another of Commit-Boost's
relays is a stale config, still dialed. Both are logged, naming up to 5
keys an epoch of 32 slots and counting the rest, with memory bounded at
1024 keys an epoch, and cb_pbs_auth_data_route_total counts each
outcome: relay, dial, stale and no_config. Every series starts at 0, so
an alert on increase() sees the first miss.

The ePBS page gains a section on keys with a missing or stale builder
config, a metrics row and two troubleshooting rows.
PbsConfig::validate asks rpc_url for its chain id with no timeout, so a
stalled RPC hangs Commit-Boost's startup, every config reload and
`commit-boost init`. The check gives up after http_timeout_seconds, the
timeout the mux loaders use, and with rpc_url set that timeout must be
above 0, since the check could not pass at 0. The timeout's error leaves
the URL out, as it often carries an API key.
The PBS service watched its config file, so replacing the file ended the
watch. A Kubernetes ConfigMap update swaps the `..data` symlink the file
points through, and editors and Ansible replace a file by a rename: the
first such update reloaded the config and every later one was missed,
so Commit-Boost kept using the old relays.

The service now also watches the file's directory, and reloads whenever
the file's contents change. The watch on the file stays, since a
single-file bind mount, as `commit-boost init` sets up, shows a write in
place only there; a directory that cannot be watched leaves just the
file watched, with a warning. Opens and reads are not changes, so reading
the file does not wake the watcher. A reload that fails, such as one
whose rpc_url does not answer, leaves the contents unaccepted, and the
next event retries it.

The reload runs on the service's tokio runtime: it ran on the watcher's
own thread, where a config with rpc_url panicked in its chain id check
and took the watcher down with it.

The automatic reload docs say what triggers a reload, that a touch or a
change to a file the config names needs `POST /reload`, and that the
Docker compose single-file mount needs edits made in place.
@JasonVranek JasonVranek closed this Oct 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant