Building on the Docker basics covered earlier in this series, this walkthrough adds the instructions needed for a PostgreSQL image, explains why multi-stage builds keep images small, and finishes with a container that runs a single PostgreSQL instance.

Six more Dockerfile instructions

The earlier articles covered FROM, RUN, COPY, ENTRYPOINT and CMD. The following join the vocabulary:

  • ARG — a variable used only during the build. Placement matters: declared after FROM, the value lives only inside that FROM clause; declared before it, the value is global.
  • ENV — an environment variable for the container's runtime, usable by the running container.
  • VOLUME — defines a mount point for persistence. A persistent store at a given mount point can also be set up when starting the container without a volume; with VOLUME, Docker mounts a volume at that path even when none is defined explicitly.
  • EXPOSE — declares a port to be opened to the outside world. In a Dockerfile this is documentation only; the port must still be published explicitly when the container is started.
  • WORKDIR — sets the path used by all subsequent commands unless one specifies an explicit path. A missing folder is created automatically.
  • USER — containers run as root by default; this switches to a given user name or user ID.

Stages and why size matters

A stage is the part of a Dockerfile that starts at a FROM and includes every statement belonging to it. Because a Dockerfile can contain several FROM clauses, it can contain several stages — hence "multi-stage". The technique exists mainly to produce images as small as possible: a first stage might clone and compile software, pulling in packages the later running container does not need; a second stage for the running container receives only the compiled artefacts.

The CYBERTEC-pg-container project's exporter Dockerfile shows the pattern.

FROM rockylinux:9 AS builder

RUN dnf -y install --nodocs  \
    	--setopt=skip_missing_names_on_install=False \
    	git \
    	go \
    	dumb-init \
    	&& dnf -y clean all ;

RUN git clone https://github.com/prometheus-community/postgres_exporter.git && cd postgres_exporter && make build

FROM rockylinux/rockylinux:9-ubi-micro
COPY --from=builder /usr/bin/dumb-init /usr/bin/dumb-init
COPY --from=builder ./postgres_exporter/postgres_exporter /bin/postgres_exporter

COPY launcher/exporter/launch.sh /
COPY scripts/exporter/queries/ /postgres_exporter/queries

EXPOSE 9187

ENTRYPOINT ["/usr/bin/dumb-init", "--"]

CMD ["/bin/sh", "/launch.sh", "init"]
Note: For simplification purposes, the variables have been removed and replaced with fixed values.

Fetching and compiling the exporter needs extra packages such as git and go, plus others that already ship in the regular rocky9 image. None of them belong in the image that eventually runs, so the second stage starts from a much smaller base.

Simply deleting unneeded files from a single container is not a substitute, for a reason rooted in layers. Every Dockerfile instruction — FROM, RUN and so on — creates a layer, and a layer is immutable. Within one layer you can strip whatever you like, as the package-installation command below does, clearing the cache before the layer closes.

RUN dnf -y install --nodocs  \
    	--setopt=skip_missing_names_on_install=False \
    	git \
    	go \
    	dumb-init \
    	&& dnf -y clean all ;

Once removal happens in a second layer, the image grows anyway: the extra layer exists and the package remains in the earlier one.

The empty-base variant

For the PostgreSQL container a different trick is used. Rather than a second stage based on a slimmed-down image, the build starts from an empty base with no layer, then creates exactly one layer by copying everything from the preceding builder stage — already cleaned of anything unnecessary — into it. Taking the final, cleaned state and writing it in a single layer is what makes this work.

FROM rockylinux:9 AS builder

RUN dnf -y install --nodocs  \
    	--setopt=skip_missing_names_on_install=False \
    	git \
    	go \
    	dumb-init \
    	&& dnf -y clean all ;

RUN dnf -y remove git go make && dnf -y clean all ;

FROM scratch 
COPY --from=builder / /

Only particular work belongs in the last, final container. Among others, these instructions are affected:

  • ENV
  • ENTRYPOINT
  • CMD

Note too that when the final stage runs as a different user from the build stage, permissions on folders and files still need adjusting.

Building the PostgreSQL container

The goal is a simple container offering one PostgreSQL instance: install PostgreSQL from the PostgreSQL Global Development Group repositories, strip what the container does not need, build a fresh container from the content of the build stage, and start it.

Dockerfile

FROM rockylinux:9 as builder

# Install needed Repos, Packages and remove unneded Packages
RUN dnf install -y https://download.postgresql.org/pub/repos/yum/reporpms/EL-9-x86_64/pgdg-redhat-repo-latest.noarch.rpm https://dl.fedoraproject.org/pub/epel/epel-release-latest-9.noarch.rpm \
   && dnf -qy module disable postgresql \
   && dnf update -y \
   && dnf install -y postgresql17-server postgresql17 dumb-init \
   && dnf remove -y kernel kernel-core kernel-modules man-db man-pages NetworkManager\
   && dnf groupremove -y "Development Tools" "Base" "Standard" \
   && dnf autoremove -y \
   && dnf clean all \
   && mkdir -p /data \
   && chown -R postgres:postgres /data \
   && rm -rf /var/cache/dnf /usr/share/doc /usr/share/man;

# Create new Container and Copy from builder
FROM scratch
COPY --from=builder / /
COPY launch.sh /launch.sh

RUN chmod +x /launch.sh

USER postgres

ENV PGDATA=/data/pgdata \
   PATH=$PATH:/usr/pgsql-17/bin

# Start dumb-init to ensure its using PID 1
ENTRYPOINT ["/usr/bin/dumb-init", "--"]

# Starter Script
CMD ["/launch.sh"]

launcher.sh

#!/bin/bash

# Initialise DB using the initdb command
initialize_db() {
   initdb -D "$PGDATA"
   if [ $? -eq 0 ]; then
       echo "Initialization successfully completed."
   else
       echo "Error during initialization of the data directory." >&2
       exit 1
   fi
}

# Start PostgreSQL
start_postgres() {
   echo "Starting PostgreSQL Server"

   sed -i "s/logging_collector = on/logging_collector = off/g" "$PGDATA/postgresql.conf"

   exec postgres -D "$PGDATA"
}

# Check whether directory exists, if yes, start database, if no, initialise and start database
if [ -d "$PGDATA" ]; then
   start_postgres
else
   initialize_db
   start_postgres

Build and launch

$ docker build . --tag my_first_pg_container:0.0.1
   [+] Building 3.0s (11/11) FINISHED                                                                                                                                                        
   ...
    => => naming to docker.io/library/my_first_pg_container:0.0.1 

$ docker run -it docker.io/library/my_first_pg_container:0.0.1
The files belonging to this database system will be owned by user "postgres".
This user must also own the server process.

The database cluster will be initialized with locale "C".
The default database encoding has accordingly been set to "SQL_ASCII".
The default text search configuration will be set to "english".

Data page checksums are disabled.

creating directory /data/pgdata ... ok
creating subdirectories ... ok
selecting dynamic shared memory implementation ... posix
selecting default "max_connections" ... 100
selecting default "shared_buffers" ... 128MB
selecting default time zone ... UTC
creating configuration files ... ok
running bootstrap script ... ok
performing post-bootstrap initialization ... ok
syncing data to disk ... ok

initdb: warning: enabling "trust" authentication for local connections
initdb: hint: You can change this by editing pg_hba.conf or using the option -A, or --auth-local and --auth-host, the next time you run initdb.

Success. You can now start the database server using:

    pg_ctl -D /data/pgdata -l logfile start

Initialization successfully completed.
Starting PostgreSQL Server
2024-11-20 16:16:23.320 UTC [7] LOG:  starting PostgreSQL 17.1 on x86_64-pc-linux-gnu, compiled by gcc (GCC) 11.5.0 20240719 (Red Hat 11.5.0-2), 64-bit
2024-11-20 16:16:23.321 UTC [7] LOG:  listening on IPv6 address "::1", port 5432
2024-11-20 16:16:23.321 UTC [7] LOG:  listening on IPv4 address "127.0.0.1", port 5432
2024-11-20 16:16:23.333 UTC [7] LOG:  listening on Unix socket "/run/postgresql/.s.PGSQL.5432"
2024-11-20 16:16:23.346 UTC [7] LOG:  listening on Unix socket "/tmp/.s.PGSQL.5432"
2024-11-20 16:16:23.356 UTC [21] LOG:  database system was shut down at 2024-11-20 16:16:18 UTC
2024-11-20 16:16:23.371 UTC [7] LOG:  database system is ready to accept connections
2024-11-20 16:21:23.453 UTC [19] LOG:  checkpoint starting: time
2024-11-20 16:21:27.840 UTC [19] LOG:  checkpoint complete: wrote 46 buffers (0.3%); 0 WAL file(s) added, 0 removed, 0 recycled; write=4.323 s, sync=0.030 s, total=4.387 s; sync files=11, longest=0.011 s, average=0.003 s; distance=270 kB, estimate=270 kB; lsn=0/1524158, redo lsn=0/1524100

Outlook

The container built here serves as an instance for experimentation only. Productively usable PostgreSQL containers, and the proper use of containers in general, are the subject of the next part of the series.