How This Blog Ships to Production with Docker, GHCR and Caddy

·9 min read

Docker
Next.js
DevOps
CI/CD
How This Blog Ships to Production with Docker, GHCR and Caddy

This blog is a Next.js application backed by Prisma and PostgreSQL, and it ships the way I like production software to ship: as an immutable container image, through a fully automated pipeline. Every push to main builds a new image, publishes it to the GitHub Container Registry (GHCR) and rolls it out to production, typically in under five minutes and with no manual steps.

None of the individual pieces are exotic, but the details are what separate a pipeline that merely works from one that is secure, reproducible and easy to roll back. This post walks through the whole path, from git push to the page you're reading, using the real configuration files.

The big picture

Here's the request path for a page on this site:

browser
  └─ CDN / edge (DNS, TLS, caching)
       └─ Caddy reverse proxy
            └─ app container (node server.js, bound to localhost only)
                 └─ PostgreSQL

And here's the deployment path:

git push main
  └─ GitHub Actions: "Build and Push Docker Image"
       ├─ docker build (multi-stage, build secrets, GHA layer cache)
       └─ push ghcr.io/<owner>/blog-v3:<commit-sha> and :latest
  └─ GitHub Actions: "Deploy" (triggered only when the build succeeds)
       └─ write docker-compose.yml + env files → docker compose pull && up -d

The design goals:

  • One immutable image per commit, so every release is traceable to source.
  • No secrets in image layers, ever.
  • No hand-edited configuration in production; GitHub is the single source of truth.
  • Rollback in seconds by redeploying an earlier image tag, without rebuilding.

Step 1: a multi-stage Dockerfile for Next.js standalone output

Next.js has an output: "standalone" mode that traces exactly which files in node_modules the server needs and copies them into .next/standalone, together with a minimal server.js. That turns a 1 GB node_modules into a runtime image that's a fraction of the size.

// next.config.js
const nextConfig = {
    output: "standalone",
    // ...
};

The Dockerfile has four stages: base (Node + pnpm), deps (install), builder (compile) and runner (what actually ships).

FROM node:20.11.0-alpine AS base
ENV NEXT_TELEMETRY_DISABLED=1
RUN apk add --no-cache libc6-compat
RUN npm install -g [email protected]
WORKDIR /app

FROM base AS deps
COPY package*.json pnpm-lock.yaml ./
RUN pnpm install --frozen-lockfile

Copying only the manifest and lockfile before pnpm install means the dependency layer is cached until the lockfile changes. Most commits only touch application code, so this stage is a cache hit almost every time. --frozen-lockfile makes the build fail if package.json and the lockfile disagree, rather than quietly resolving new versions in CI.

Build secrets without leaking them into layers

The build step needs real configuration. next build imports modules that validate environment variables, and Prisma needs DATABASE_URL to generate its client. The naive approach is ARG DATABASE_URL, but build args are recorded in the image history: anyone who can pull the image can run docker history and read them.

BuildKit secret mounts solve this. The secret is mounted as a file for the duration of a single RUN instruction and never written to a layer:

FROM base AS builder
ARG NEXTAUTH_URL
ARG NEXT_PUBLIC_BLOG_TITLE
ENV NEXTAUTH_URL=$NEXTAUTH_URL
ENV NEXT_PUBLIC_BLOG_TITLE=$NEXT_PUBLIC_BLOG_TITLE

COPY --from=deps /app/node_modules ./node_modules
COPY . .

RUN --mount=type=secret,id=DATABASE_URL \
    --mount=type=secret,id=NEXTAUTH_SECRET \
    --mount=type=secret,id=MAILGUN_API_KEY \
    # ...one mount per secret
    sh -c 'export DATABASE_URL=$(cat /run/secrets/DATABASE_URL) && \
           export NEXTAUTH_SECRET=$(cat /run/secrets/NEXTAUTH_SECRET) && \
           export MAILGUN_API_KEY=$(cat /run/secrets/MAILGUN_API_KEY) && \
           node scripts/validate-env.mjs build && \
           pnpm run build'

Notice that two values are plain build args: NEXTAUTH_URL and NEXT_PUBLIC_BLOG_TITLE. Neither is secret, and the NEXT_PUBLIC_ one has to be present at build time anyway. Next.js inlines process.env.NEXT_PUBLIC_* references into the JavaScript bundle during next build. Changing it in production later does nothing for code that has already been compiled. I cover this in depth in validating env vars at build time and runtime.

The runtime stage

FROM base AS runner
WORKDIR /app
ENV NODE_ENV=production

RUN addgroup --system --gid 1001 nodejs && \
    adduser --system --uid 1001 nextjs

COPY --from=builder /app/.next/standalone ./
COPY --from=builder /app/.next/static ./.next/static
COPY --from=builder /app/public ./public
COPY --from=builder /app/scripts/validate-env.mjs ./scripts/validate-env.mjs
COPY --from=builder /app/node_modules/zod ./node_modules/zod

RUN chown -R nextjs:nodejs .
USER nextjs

EXPOSE 3000
ENV PORT=3000
ENV HOSTNAME="0.0.0.0"

CMD ["sh", "-c", "node ./scripts/validate-env.mjs runtime && node server.js"]

A few things here aren't obvious:

  • .next/static and public are copied separately. The standalone output deliberately leaves them out, on the assumption that you'll serve them from a CDN. If you forget these two lines, the app boots and then every CSS and JS request 404s.
  • zod is copied by hand. The env validation script runs before server.js, so the standalone file tracer never sees it import zod. Without that line the container dies on startup with Cannot find module 'zod'.
  • HOSTNAME=0.0.0.0. The standalone server binds to localhost by default, which inside a container means "unreachable from outside the container".
  • Non-root user. Cheap to do, and it limits the damage if something in the app is ever exploited.

Step 2: building and publishing in GitHub Actions

The build workflow runs on every push to main (and on pull requests, where it builds but doesn't push). The interesting part is the build step:

- name: Build Docker image
  uses: docker/build-push-action@v6
  with:
    context: .
    push: ${{ github.event_name != 'pull_request' }}
    tags: |
      ghcr.io/${{ steps.ghcr.outputs.owner }}/blog-v3:latest
      ghcr.io/${{ steps.ghcr.outputs.owner }}/blog-v3:${{ github.sha }}
    cache-from: type=gha
    cache-to: type=gha,mode=max
    build-args: |
      NEXTAUTH_URL=${{ vars.NEXTAUTH_URL }}
      NEXT_PUBLIC_BLOG_TITLE=${{ vars.NEXT_PUBLIC_BLOG_TITLE }}
    secrets: |
      DATABASE_URL=${{ secrets.DATABASE_URL }}
      NEXTAUTH_SECRET=${{ secrets.NEXTAUTH_SECRET }}
      # ...
  • Two tags per build. :latest is a convenience. :<sha> is what actually gets deployed. Because every commit has an immutable image, a rollback is just "deploy the previous SHA".
  • type=gha cache. Buildx stores layer cache in the GitHub Actions cache, so the deps stage is reused across runs. mode=max caches intermediate stages too, not just the final image.
  • Secrets vs. variables. Non-sensitive config (NEXTAUTH_URL, the blog title) lives in repository variables, which are visible in logs and easy to edit. Credentials live in secrets and are only passed through secrets:, which maps them to BuildKit secret mounts.
  • The owner name is lowercased (${GITHUB_REPOSITORY_OWNER,,} in a previous step), because GHCR image names must be lowercase and GitHub usernames often aren't.

Before building, the workflow also checks that every required variable and secret is set and fails with a readable message if one is missing. That's a lot friendlier than a Zod error from deep inside next build three minutes later.

Step 3: deploying with workflow_run

Deployment lives in a separate workflow that triggers when the build finishes:

on:
  workflow_run:
    workflows: ["Build and Push Docker Image"]
    types: [completed]
  workflow_dispatch:
    inputs:
      image_tag:
        description: "Docker image tag"
        default: "latest"

jobs:
  deploy:
    if: >-
      github.event_name == 'workflow_dispatch' ||
      (github.event.workflow_run.conclusion == 'success' &&
       github.event.workflow_run.head_branch == 'main')

types: [completed] fires for failed builds too, so the if: that checks conclusion == 'success' is essential. Without it, a broken build would trigger a deploy of whatever :latest happened to be.

The deploy job checks out the exact commit that was built (github.event.workflow_run.head_sha), connects to the production host and runs a short, idempotent script (simplified for the post):

set -eu
IMAGE_TAG="${{ github.event.inputs.image_tag || github.event.workflow_run.head_sha || 'latest' }}"
cd "$APP_DIR"

echo "$REGISTRY_TOKEN" | docker login ghcr.io -u "$OWNER" --password-stdin

# docker-compose.yml comes from the repo, so infra changes are versioned with the code
echo "$COMPOSE_B64" | base64 -d > docker-compose.yml
printf '%s\n' "APP_IMAGE=ghcr.io/$OWNER/blog-v3:$IMAGE_TAG" > .env

# runtime configuration is rewritten from GitHub secrets on every deploy
printf '%s\n' \
  "NEXTAUTH_URL=$NEXTAUTH_URL" \
  "DATABASE_URL=$DATABASE_URL" \
  "NEXTAUTH_SECRET=$NEXTAUTH_SECRET" \
  > .env.production

docker compose pull
docker compose up -d
docker image prune -f

Two things I like about this:

  1. GitHub is the single source of truth for configuration. The production env file is regenerated on every deploy, so there is no drift between what's in the secrets and what's actually running. Configuration changes go through the same audited path as code.
  2. Rollback doesn't need a rebuild. workflow_dispatch with image_tag=<older sha> redeploys a previous image in about 20 seconds.

Step 4: the Compose file

services:
  app:
    image: ${APP_IMAGE:-ghcr.io/nighthunter30/blog-v3:latest}
    container_name: blog-v3
    ports:
      - "127.0.0.1:3010:3000"
    restart: unless-stopped
    env_file:
      - .env.production
    extra_hosts:
      - "host.docker.internal:host-gateway"
    mem_limit: 512m
    healthcheck:
      test: ["CMD", "wget", "--spider", "-q", "http://127.0.0.1:3000/api/health"]
      interval: 30s
      timeout: 10s
      retries: 3
      start_period: 40s
    logging:
      driver: json-file
      options:
        max-size: "10m"
        max-file: "3"

Line by line, these are the parts that matter:

  • 127.0.0.1:3010:3000, not 3010:3000. Docker writes its own iptables rules, and a published port bypasses host firewalls like ufw. Publish on all interfaces and the app is reachable directly, around your reverse proxy and its security headers. Binding to loopback means only the proxy can reach it.
  • host.docker.internal:host-gateway. On Linux this hostname doesn't exist by default. The special host-gateway value maps it to the host's address on the Docker bridge, so DATABASE_URL can point at host.docker.internal:5432.
  • mem_limit. Explicit resource limits mean a memory leak takes down this container, not the host. It's the same discipline as requests and limits in Kubernetes.
  • Health check. /api/health is a trivial route that returns { "status": "ok" }. Compose uses it to mark the container healthy, and docker ps shows the state at a glance. The Alpine base image ships wget but not curl, hence wget --spider.
  • Log rotation. The default json-file driver never rotates. Without max-size, logs eventually fill the disk.

Step 5: connecting to PostgreSQL outside the container

The database doesn't live in the application's Compose project. Keeping stateful services separate from stateless app containers means a docker compose down -v can never take the data with it, and backups, upgrades and tuning happen independently of app releases.

If PostgreSQL runs on the Docker host itself, two settings make it reachable from containers without exposing it to the internet. First, listen on the Docker bridge address in addition to localhost:

# postgresql.conf
listen_addresses = 'localhost,172.17.0.1'

Second, allow password (SCRAM) authentication from Docker's private address range in pg_hba.conf:

host    all    all    172.16.0.0/12    scram-sha-256

The /12 matters. Compose places each project on its own network (172.18.0.0/16, 172.19.0.0/16, ...), so connections don't arrive from the default bridge's 172.17.x.x range. 172.16.0.0/12 covers all of Docker's default ranges while Postgres still only listens on private interfaces. Inside the container, DATABASE_URL simply points at host.docker.internal:5432, thanks to the host-gateway mapping in the Compose file.

Step 6: Caddy as the reverse proxy

Caddy sits in front of the container. Its configuration is refreshingly small, and it handles TLS certificates, HTTP/2 and HTTP/3 automatically:

example.com {
    encode gzip zstd
    reverse_proxy 127.0.0.1:3010

    header {
        Strict-Transport-Security "max-age=31536000; includeSubDomains"
        X-Content-Type-Options "nosniff"
        Referrer-Policy "strict-origin-when-cross-origin"
        -Server
    }

    log {
        output file /var/log/caddy/example.access.log {
            roll_size 10mb
            roll_keep 5
        }
    }
}

www.example.com {
    redir https://example.com{uri} permanent
}

Redirecting www to the apex keeps a single canonical host, which matters for SEO and for auth libraries like NextAuth that compare the request origin against NEXTAUTH_URL. If you put a CDN in front of Caddy, configure it for end-to-end encryption (on Cloudflare, "Full (strict)") so traffic is encrypted all the way to the origin.

Key takeaways

  • Tag images by commit SHA and deploy the SHA. :latest is convenient for humans; deployments should be reproducible and rollbacks instant.
  • Treat NEXT_PUBLIC_* as build-time configuration. It belongs with source code, not runtime secrets.
  • Use BuildKit secret mounts so credentials never land in image layers or history.
  • Never publish container ports on 0.0.0.0 behind a reverse proxy. Docker's iptables rules will route around your firewall.
  • Validate configuration at every edge: in CI before building, and when the container starts. A container that refuses to boot with a clear error beats one that serves 500s.
  • Version your deployment config with your code. The Compose file is checked out from the exact commit that was built, so infrastructure changes get the same review and history as application changes.

The result is a pipeline that is fast, reproducible and fully auditable: every release maps to a commit, every rollback is a single workflow run, and nothing in production is configured by hand.

Comments

Related Articles