Skip to content

supervisor: one engine for instances and tcm - #1393

Merged
bigbes merged 5 commits into
v3from
bigbes/gh-no-v3-supervisor
Oct 5, 2026
Merged

bigbes merged 5 commits into
v3from
bigbes/gh-no-v3-supervisor

Conversation

@bigbes

@bigbes bigbes commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

tt had two process supervisors that shared nothing: the instance watchdog in cli/running, behind tt start, and lib/watchdog, behind tt tcm start --watchdog. This replaces both with one engine, internal/supervisor, which is also meant to run non-Tarantool processes under tt start later.

The engine runs one child at a time on a single event loop. A Source supplies the spec before every start, so the configuration is re-read on a restart, and decides whether an exited child runs again. The loop owns the pid files, the signal policy (stop signals with escalation to SIGKILL, a reload signal, an ignored set, forwarding of the rest), the restart delay, the periodic integrity check and the cleanup. Every child is started through one replaceable function, where starting integrity-checked bytes will plug in later.

  • A failed integrity check is always a hard stop: the child is killed and nothing restarts it, whatever restart_on_failure says. The old instance watchdog treated that kill as an ordinary exit and restarted an instance with restart_on_failure: true over the files that had failed the check.
  • Pid files are owned through a lock (internal/pidfile), and every tt writer takes part: the instance and TCM watchdogs, tt start --interactive, the daemon, tt tcm start and tt kill. Two tt start of one instance no longer start two watchdogs, and tt kill no longer removes the pid file or the sockets of a watchdog that came up after the one it killed.
  • tt stop sent during the restart pause is no longer lost, and signals sent while an instance is being stopped are handled at once.
  • TCM gets the same watchdog: SIGKILL after 30 seconds, SIGHUP and SIGQUIT stop TCM instead of orphaning it, pid files are removed on exit, a second tt tcm start --watchdog is refused, the periodic integrity check applies, and tt tcm status finds TCM the way tt tcm stop does. The stop commands wait 35 seconds, longer than the watchdog's escalation.
  • lib/watchdog and the lib/ directory are gone.

The engine is covered by tests on real processes and by a model-based fuzzer (FuzzEngine) that replays event scripts against a deterministic clock and fake processes and checks invariants after every step; its seed corpus covers the interleavings the engine must handle: a check failing as the child exits, a stop queued behind a restart, a panic in a callback. The integration tests for running, start, stop, status, kill, restart, logrotate and tcm pass the same set on macOS and Linux as on the base commit. The branch also drops a stray instance snapshot committed to v3 earlier.

@bigbes

bigbes commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator Author

Review by group of changes:

  • Stray snapshot — drops instances.enabled/…/00000000000000000000.snap committed earlier and ignores instances.enabled at the root.
  • The engine — internal/supervisor: event loop, pid files, signal policy, integrity hard stop, panic-safe teardown, tests on real processes and the model-based fuzzer.
  • Pid-file protocol — internal/pidfile, and every tt pid-file writer through it.
  • Instances on the engine — running.Watchdog replaced, socket cleanup for the context tarantool ran with, tt kill cleanup under the lock.
  • TCM on the engine — tt tcm start --watchdog on the engine, lib/watchdog deleted, tt tcm status, stop timeouts.

Whole PR

A snapshot of an instance run in the repository root was committed
together with the removal of tarantoolctl support:
instances.enabled/repro/var/lib/instance001/00000000000000000000.snap.
Nothing reads it. Remove it and ignore instances.enabled at the root,
where running tt in the checkout creates that directory.
@bigbes
bigbes force-pushed the bigbes/gh-no-v3-supervisor branch 2 times, most recently from 70ae689 to db30481 Compare September 30, 2026 16:10
Add internal/supervisor, one engine for the two supervisors tt has:
the instance watchdog in cli/running and the TCM watchdog in
lib/watchdog. Neither is switched over yet.

The engine runs one child at a time on a single event loop. A Source
supplies the Spec before every start, so configuration is re-read on
a restart, and decides whether an exited child runs again. The loop
owns the supervisor and child pid files, the signal policy (stop
signals with escalation to SIGKILL, a reload signal with a hook, an
ignored set, forwarding of everything else), the restart delay, the
periodic check and a cleanup hook. Consumers log through typed
events, and every child is started through one replaceable StartFunc.
Spec.Command runs a Spec once without supervising it.

The guarantees the package doc states:
- a failed check is always a hard stop: the child is SIGKILLed, Run
  ends with OpCheck and the Source is not asked about a restart; a
  check error that races the child's exit, wraps more than the
  cancellation, or comes from a panicking Check counts as a failure;
- a pid file is owned through an exclusive flock held for the
  supervisor's lifetime, so of any number of supervisors starting at
  once exactly one owns it, and a stale file is taken over in place;
- signals that have already arrived are handled before the next start
  and before a restart is decided, so a pending stop always wins;
- a panicking callback or StartFunc still has the child killed and
  waited, the checks stopped, Cleanup run and the pid files removed;
- Detach reaps the process it detaches.

The tests run real child processes: signal forwarding and escalation,
restarts and their delay, signals during startup and the restart
delay, pid-file lifecycle and concurrent takeover, cleanup on every
exit path, process groups, bursts of signals across restarts. A
model-based fuzzer, FuzzEngine, replays event scripts against a fake
clock and fake processes and checks invariants after every step, on
what the engine did as well as on the events it reported; its seeds
cover a check failing as the child exits, a stop queued behind a
restart and a panic in a callback.
The flock protocol for pid files moves from the supervisor to a leaf
package, internal/pidfile, so that every writer and remover of a tt
pid file can follow it; cli/process_utils may not import the
supervisor. Acquire takes a file under a non-blocking exclusive lock
and takes a stale one over in place; Release unlinks and then
unlocks; Keep gives up the lock and leaves the file, for a writer
that exits while the process it wrote down goes on; RemoveFor removes
the file of a process that is gone, under the lock, and only while
the file still names that process.

CreatePIDFile acquires the file and hands the ownership to the caller:
the interactive tt start and tt daemon hold it while they run and
release it when they stop, and the TCM interactive start and the TCM
watchdog keep it. CheckPIDFile still refuses a file naming a live
process with the same message, and no longer removes a stale one. tt
kill removes the pid file of the watchdog it killed only while it
still names that watchdog.
The watchdog tt start detaches for an instance now runs the shared
supervisor engine. Its Source reads the configuration before every
start and builds the Spec of the tarantool process, as the script and
cluster instances did in their Start, and after an exit reads
restart_on_failure again and switches to a new log if its settings
changed. The watchdog owns the pid file of the instance, which names
the watchdog; SIGINT, SIGTERM and SIGQUIT reach tarantool as they were
sent and are followed by SIGKILL 30 seconds later; SIGHUP rotates the
log and reaches tarantool; any other signal is forwarded. tarantool
stays in the watchdog's process group, so tt kill kills both. The log
lines are the watchdog's as before, written from the engine's events.

Behaviour that changes with it: a failed periodic integrity check
stops the instance for good instead of restarting it when
restart_on_failure is set; a tt stop sent during the restart delay is
no longer lost; signals are handled while an instance stops instead
of after it; and a watchdog whose pid file is taken by another cleans
up nothing. Errors found while the Spec is built are logged as
instance creation errors, since that step now covers what the start
used to check.

Sockets are removed for the context tarantool actually ran with, after
every exit and at the end, so a configuration change between the start
and the exit no longer leaves the old sockets behind or removes paths
no process used. tt kill removes the sockets of the watchdog it killed
while it still holds that watchdog's pid-file lock, so a watchdog that
comes up meanwhile keeps its sockets as well as its pid file.

tt start --interactive runs the same Spec without a supervisor, through
Spec.Command. StartWatchdog detaches through supervisor.Detach. The
Watchdog, Provider, Instance and processController types go, with the
script and cluster instance types; their tests move to the Specs and
read the tarantool output without the data race the old tests had, so
cli/running runs under -race.
tt tcm start --watchdog runs TCM under internal/supervisor in the
current process. The watchdog takes watchdog.pid before TCM starts and
tcm.pid right after every start through the pid-file protocol, refuses
to run while either names a running process, and removes both when it
exits. lib/watchdog, the last package under lib/, goes.

TCM runs in a process group of its own and is restarted after every
exit but a failed integrity check. SIGINT, SIGTERM, SIGHUP and SIGQUIT
stop it with SIGTERM, and the group is killed 30 seconds later if it
still runs. The signal policy is a table, so moving a signal from the
stop set to the ignored set is one line. The periodic integrity check
takes its period from --integrity-check-period, one day by default
when --integrity-check is set, as for tt start; a failure stops TCM
for good. The log lines of the old watchdog stay.

tt tcm status finds TCM the way tt tcm stop does, through watchdog.pid
first, so it reports TCM running while the watchdog waits to restart
it, and NOT RUNNING when there is no pid file at all. An interactive
tt tcm start whose tcm.pid is refused stops the TCM it started. tt
stop, tt quit, tt tcm stop and tt daemon stop wait 35 seconds, longer
than the 30 a watchdog gives its child before it kills it, so they no
longer report a failure while the watchdog completes the stop.

Every tt writer of a pid file now takes part in the protocol, which
the internal/pidfile doc says. The tests run the watchdog in a process
of its own over a shell script standing in for TCM.
@bigbes
bigbes merged commit 838feff into v3 Oct 5, 2026
23 checks passed
@bigbes
bigbes deleted the bigbes/gh-no-v3-supervisor branch October 5, 2026 20:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants