Repository navigation
postgres in container with PID 1 aggregating orphaned processes leading to restarts & recovery #1349
Description
Activity
While it is regrettable that extra (crashing) processes within the container can cause PostgreSQL server itself to crash, we will not be installing
dumb-initor similar in thepostgresimages. The reason is already in the linked article, don't run extra processes within the container. It would be best iflivenessProbeandreadinessProbecould be defined via a separate container, butTCPbased should be enough to check that it is "up" and prevent polluting the PID namespace of thepostgrescontainer when it is PID 1. Alternatively increasing their timeout would make the failure less likely.This problem also could not have occurred if the
psqlprocess hadn't been started inside the container. It is good practice to consider a container running a service as a closed unit: you shouldn't start jobs or interactive sessions inside the container. I can understand the appeal of using the container'spsqlexecutable to avoid having to install the PostgreSQL client anywhere else, but it is a shortcut that you shouldn't take.The
--initflag ondocker runis the local solution to this andshareProcessNamespaceshould do the same thing by making the "pause" container be PID 1 for the whole pod (and thus prevent PostgreSQL from inheriting unexpected children).shareProcessNamespaceshould do the same thing by making the "pause" container be PID 1 for the whole pod (and thus prevent PostgreSQL from inheriting unexpected children).The shareProcessNamespace K8S option does works around the primary issue - thanks for sharing that insight.
However:- Today the major providers of open source postgres helm charts do not wield that option - like bitnami
- There is security and possibly other impact of sharing process namespace with sidecar containers
- Any self-managed K8S fabric for this standard postgres docker image will have the same potential issue lurking
The fact that
- postgres depends on child process relationships for health management,
- spawning a postgres provided executable for health management (pg_isready) is common and best practice for health monitoring
- in a default container runtime environment any orphaned process now gets addded as a child
- postgres has no additional 'bookkeeping' on the role of specif child processes - such that any unnatural exists of inherited child processes cause instability of the main process.
this to me suggests a fix at a fundamental container level is better than a patch / workaround in the container orchestration level.
Independent confirmation, plus one thing I think is new: PostgreSQL 18 does not fix this.
We hit this in production on 16 — roughly one crash-recovery cycle per day, and ~4.5/day during a period when a full disk was making the server slow enough to fail healthchecks constantly. It took five weeks to track down, because the victim leaves almost no trace: it isn't a backend, so it logs nothing, it has no activity string (hence no
DETAIL: Failed process was running:), and it receives no signal.Minimal repro, no healthcheck involved:
$ docker run -d --name t -e POSTGRES_PASSWORD=x postgres:16 $ docker exec t sh -c '(sleep 2; exit 2) & exit 0' $ docker logs t LOG: server process (PID 76) exited with exit code 2 LOG: terminating any other active server processes LOG: database system was not properly shut down; automatic recovery in progress
The same orphan exiting 1 is completely silent —
CleanupBackend()whitelists 0 and 1.PG18 changes the label, not the outcome
REL_18_STABLE/postmaster.cnow looks the pid up before deciding, and logs"untracked child process"rather than"server process"— but the crash path is unchanged:if (!EXIT_STATUS_0(exitstatus) && !EXIT_STATUS_1(exitstatus)) HandleChildCrash(pid, exitstatus, _("untracked child process"));
Two practical consequences: upgrading to 18 doesn't help, and anyone alerting on the
server process (PID …) exited with exit code 2string will silently stop matching after the upgrade.A note on the mitigations
Your first reply lists
--init/shareProcessNamespace/ a TCP probe / a longer timeout, and I think that's right — with one clarification that cost us a week.Wrapping the probe (
pg_isready … || exit 1) closes the common case but not the one reported here. If the probe is killed on its timeout, the wrapper dies first and the surviving child is the orphan, so the trailingexit 1never runs. The same applies todocker-library/healthcheck'sdocker-healthcheckscript — it ends inexit 1, but itspsqlchild exits 2 on a bad connection.--initdoes close it. Verified here: PID 1 becomes/sbin/docker-init, and both orphan shapes go from one crash-recovery cycle to zero.Not asking to revisit the packaging decision — just adding the PG18 data point and the wrapper caveat for whoever finds this thread next.
The postgres in docker hosted on K8S infrastructure with periodic health and rediness checks experiences regular restarts of the database.
Logs show '
postgres "server process <PID>exited with exit code 2"' followed by restart and recovery of the database.happening with a variety of stable postgres versions (12, 14, 16) on the same system
deployed using kubernetes, we monitor health with exec probe
as well as readiness probe
Investigation shows that the postgres db is fine - no exit 2 occurences.
Instead a regular health monitoring process that includes a
/bin/bash -c pg_isready ..causes the problem.Root cause is described here: https://www.cybertec-postgresql.com/en/docker-sudden-death-for-postgresql/ (thanks to laurentz albe)
The postgres in docker setup runs the default process (postgres) as root process of the container with PID 1.
The main postgres container is a health manager/monitor for all other spawned worker processes.
Sudden exits of worker processes lead to DB restart - remediating possible shared memory corruption.
However, orphaned other processes will get the root process as parent process (PID 1) being our postgres main entrypoint.
These processes get orphaned due to to timeout of the monitoring environment.
pg_isready is known to return exit code 2 when not able to connect.
Suggested remediation is in the referenced article: start the container using dum-init.
Example patch we deploy to remediate consists of installing the dumb-init package (using apt) and extending the entrypoint to use dumb-init as the main process (PID 1)
FROM postgres:12
RUN apt update && apt install -y dumb-init && apt clean
ENTRYPOINT ["/usr/bin/dumb-init", "docker-entrypoint.sh"]
CMD ["postgres"]