Skip to content

postgres in container with PID 1 aggregating orphaned processes leading to restarts & recovery #1349

Description

@egbert-algra

The postgres in docker hosted on K8S infrastructure with periodic health and rediness checks experiences regular restarts of the database.
Logs show 'postgres "server process <PID>exited with exit code 2" ' followed by restart and recovery of the database.

happening with a variety of stable postgres versions (12, 14, 16) on the same system

deployed using kubernetes, we monitor health with exec probe

        livenessProbe:
          exec:
            command:
            - pg_isready
            - -U
            - dmp-admin
            - -d
            - dmp-entity
          timeoutSeconds: 1

as well as readiness probe

        readinessProbe:
          exec:
            command:
            - /bin/bash
            - -c
            - pg_isready -U dmp-admin -d dmp-entity && [ ! -f /var/lib/postgresql/backup/pgdump_backup.velero.sql ]
          timeoutSeconds: 1

Investigation shows that the postgres db is fine - no exit 2 occurences.
Instead a regular health monitoring process that includes a /bin/bash -c pg_isready .. causes the problem.

Root cause is described here: https://www.cybertec-postgresql.com/en/docker-sudden-death-for-postgresql/ (thanks to laurentz albe)
The postgres in docker setup runs the default process (postgres) as root process of the container with PID 1.
The main postgres container is a health manager/monitor for all other spawned worker processes.
Sudden exits of worker processes lead to DB restart - remediating possible shared memory corruption.
However, orphaned other processes will get the root process as parent process (PID 1) being our postgres main entrypoint.
These processes get orphaned due to to timeout of the monitoring environment.
pg_isready is known to return exit code 2 when not able to connect.

Suggested remediation is in the referenced article: start the container using dum-init.

Example patch we deploy to remediate consists of installing the dumb-init package (using apt) and extending the entrypoint to use dumb-init as the main process (PID 1)

FROM postgres:12
RUN apt update && apt install -y dumb-init && apt clean
ENTRYPOINT ["/usr/bin/dumb-init", "docker-entrypoint.sh"]
CMD ["postgres"]

Activity

  1. yosifkit commented on Jul 23, 2025

    @yosifkit
    Member

    While it is regrettable that extra (crashing) processes within the container can cause PostgreSQL server itself to crash, we will not be installing dumb-init or similar in the postgres images. The reason is already in the linked article, don't run extra processes within the container. It would be best if livenessProbe and readinessProbe could be defined via a separate container, but TCP based should be enough to check that it is "up" and prevent polluting the PID namespace of the postgres container when it is PID 1. Alternatively increasing their timeout would make the failure less likely.

    This problem also could not have occurred if the psql process hadn't been started inside the container. It is good practice to consider a container running a service as a closed unit: you shouldn't start jobs or interactive sessions inside the container. I can understand the appeal of using the container's psql executable to avoid having to install the PostgreSQL client anywhere else, but it is a shortcut that you shouldn't take.

    The --init flag on docker run is the local solution to this and shareProcessNamespace should do the same thing by making the "pause" container be PID 1 for the whole pod (and thus prevent PostgreSQL from inheriting unexpected children).

  2. egbert-algra commented on Jul 28, 2025

    @egbert-algra
    Author

    shareProcessNamespace should do the same thing by making the "pause" container be PID 1 for the whole pod (and thus prevent PostgreSQL from inheriting unexpected children).

    The shareProcessNamespace K8S option does works around the primary issue - thanks for sharing that insight.
    However:

    • Today the major providers of open source postgres helm charts do not wield that option - like bitnami
    • There is security and possibly other impact of sharing process namespace with sidecar containers
    • Any self-managed K8S fabric for this standard postgres docker image will have the same potential issue lurking

    The fact that

    1. postgres depends on child process relationships for health management,
    2. spawning a postgres provided executable for health management (pg_isready) is common and best practice for health monitoring
    3. in a default container runtime environment any orphaned process now gets addded as a child
    4. postgres has no additional 'bookkeeping' on the role of specif child processes - such that any unnatural exists of inherited child processes cause instability of the main process.
      this to me suggests a fix at a fundamental container level is better than a patch / workaround in the container orchestration level.
  3. fatheus97 commented on Sep 22, 2026

    @fatheus97

    Independent confirmation, plus one thing I think is new: PostgreSQL 18 does not fix this.

    We hit this in production on 16 — roughly one crash-recovery cycle per day, and ~4.5/day during a period when a full disk was making the server slow enough to fail healthchecks constantly. It took five weeks to track down, because the victim leaves almost no trace: it isn't a backend, so it logs nothing, it has no activity string (hence no DETAIL: Failed process was running:), and it receives no signal.

    Minimal repro, no healthcheck involved:

    $ docker run -d --name t -e POSTGRES_PASSWORD=x postgres:16
    $ docker exec t sh -c '(sleep 2; exit 2) & exit 0'
    $ docker logs t
    LOG:  server process (PID 76) exited with exit code 2
    LOG:  terminating any other active server processes
    LOG:  database system was not properly shut down; automatic recovery in progress

    The same orphan exiting 1 is completely silent — CleanupBackend() whitelists 0 and 1.

    PG18 changes the label, not the outcome

    REL_18_STABLE/postmaster.c now looks the pid up before deciding, and logs "untracked child process" rather than "server process" — but the crash path is unchanged:

    if (!EXIT_STATUS_0(exitstatus) && !EXIT_STATUS_1(exitstatus))
        HandleChildCrash(pid, exitstatus, _("untracked child process"));

    Two practical consequences: upgrading to 18 doesn't help, and anyone alerting on the server process (PID …) exited with exit code 2 string will silently stop matching after the upgrade.

    A note on the mitigations

    Your first reply lists --init / shareProcessNamespace / a TCP probe / a longer timeout, and I think that's right — with one clarification that cost us a week.

    Wrapping the probe (pg_isready … || exit 1) closes the common case but not the one reported here. If the probe is killed on its timeout, the wrapper dies first and the surviving child is the orphan, so the trailing exit 1 never runs. The same applies to docker-library/healthcheck's docker-healthcheck script — it ends in exit 1, but its psql child exits 2 on a bad connection.

    --init does close it. Verified here: PID 1 becomes /sbin/docker-init, and both orphan shapes go from one crash-recovery cycle to zero.

    Not asking to revisit the packaging decision — just adding the PG18 data point and the wrapper caveat for whoever finds this thread next.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions