Zero-downtime deploys without Docker: systemd and Caddy
Zero-downtime deployment on a VPS without Docker: blue-green with two systemd units, a health check and a Caddy upstream file. A copyable script inside.
You can get zero-downtime deployment on a plain VPS without Docker: run two copies of your app as systemd services on two ports, start the new release on the idle one, wait for its health check, then point Caddy at it and stop the old one. A bad release never gets traffic, and no request is dropped during the switch. This is blue-green deployment on one server.
The simplest deploy on a VPS is git pull and systemctl restart myapp. It works, but for a few seconds nothing is listening: visitors get errors, and if the new code fails to start, the site stays down until you notice. You don't need Docker or Kubernetes to fix this. This post shows the pattern by hand with systemd and Caddy, with a deploy script you can copy, then explains exactly how ox does it on every deploy, including what it does not cover.
In this guide:
- Why a restart drops requests
- Three ways to deploy without downtime, compared
- Blue-green by hand with systemd and Caddy
- How ox does it on every deploy
- What this does not cover
- Frequently asked questions
Why a restart drops requests
A restart stops the old process before the new one is ready. Between the two, the port is closed and the proxy in front answers with an error. A Django or Next.js app can take several seconds to start, longer if it loads a lot at import time. And if the new release crashes on start, there is no old process left to fall back on.
Three ways to deploy without downtime
Socket activation. systemd holds the listening socket during the restart, so connections wait instead of failing. A slow start still makes them wait, and a broken release still takes the site down.
Your app server's own graceful reload, like gunicorn's
HUPsignal or Puma's phased restart. Useful, but specific to each server and language, and a bad release can still replace a good one.Two sides, then a switch, often called blue-green. The new release starts next to the old one on another port. Only when it answers a health check does the proxy send traffic to it. Then the old one stops. A bad release never gets traffic. This works for any language.
| Approach | Requests during a slow start | A release that fails to start | Works for any language |
|---|---|---|---|
| systemd socket activation | Wait until the app is up | Site down | Only if the app accepts a passed socket |
| App server graceful reload | Served by old workers | Depends on the server | No, each server has its own |
| Blue-green with a proxy switch | Served by the old side | Old side keeps serving | Yes |
Blue-green by hand with systemd and Caddy
1. One folder per release
Build each release in its own folder, never over the running one:
/srv/myapp/releases/2026-10-08-a1b2c3/
/srv/myapp/releases/2026-10-09-d4e5f6/
/srv/myapp/side-9001 -> releases/2026-10-08-a1b2c3
/srv/myapp/side-9002 -> releases/2026-10-09-d4e5f6
Each side is a link to the release it runs, changed only while that side is stopped.
2. A systemd template unit, one instance per port
# /etc/systemd/system/[email protected]
[Unit]
Description=myapp on port %i
[Service]
User=myapp
WorkingDirectory=/srv/myapp/side-%i
EnvironmentFile=/etc/myapp.env
Environment=PORT=%i
ExecStart=/srv/myapp/side-%i/.venv/bin/gunicorn config.wsgi --bind 127.0.0.1:%i
Restart=always
myapp@9001 and myapp@9002 are now two independent services from one file. The %i is the instance name, here the port; see systemd's unit file documentation for templates and specifiers.
3. Make Caddy read the port from a file
Caddy 2 has a {file.*} placeholder that reads a file's contents when a request needs it. Use it as the upstream address, and switching becomes a file write with no reload at all. In Caddy's JSON config, the route's handler is:
{
"handler": "reverse_proxy",
"upstreams": [{ "dial": "{file./etc/myapp/upstream}" }]
}
and /etc/myapp/upstream holds one line, like 127.0.0.1:9001.
Why not edit the config and run caddy reload? A reload is careful, but it restarts Caddy's servers. On real servers under load, ox's own tests saw that restart reset connections still waiting in the accept queue. A file read per request has no such moment.
4. The deploy script
set -e
NEW=9002; OLD=9001 # whichever side is idle becomes NEW
ln -sfn releases/2026-10-09-d4e5f6 /srv/myapp/side-$NEW
systemctl start myapp@$NEW
# wait until the new side answers, or give up and leave the old side live
for i in $(seq 1 60); do
curl -fsS http://127.0.0.1:$NEW/healthz >/dev/null && break
[ "$i" = 60 ] && { systemctl stop myapp@$NEW; exit 1; }
sleep 2
done
# switch: write a temp file, then rename it over the old one
echo 127.0.0.1:$NEW > /etc/myapp/upstream.tmp
mv /etc/myapp/upstream.tmp /etc/myapp/upstream
sleep 10 # let requests on the old side finish
systemctl stop myapp@$OLD
The rename matters. mv within one filesystem replaces the file in one step, so Caddy reads either the old port or the new one, never half a line.
5. Database migrations
Run migrations after the build and before the switch. For a moment the old release runs against the new schema, so make each migration safe for both: add a column before the code uses it, and drop it in a later release. This is the one part no tool can do for you.
How ox does it on every deploy
ox runs this pattern for every app, with no setting to turn on. Each app gets two ports from ox's range, one per side. A deploy, in order:
Fetches the commit into a new release folder, installs dependencies and builds, without touching what is live.
Takes a snapshot of each PostgreSQL database that has tables, then runs your migration.
Restarts the workers your app calls over HTTP and waits for their health, then starts the app on the idle side.
Waits up to 120 seconds for the new side's health check: a 2xx or 3xx answer on your
[app] healthpath when you set one, otherwise a TCP connect.Switches: it writes the new port to the upstream file Caddy reads on every request, and the static files' release id to a second file, one right after the other. Caddy does not reload.
Stops the old side 10 seconds later, then restarts the other workers, and keeps the newest 3 releases for rollback.
If anything fails before the switch, the live release keeps serving and the run's log names the step, the cause, and a fix. If the switch itself fails, ox moves traffic back. Rollback uses the same switch with a kept release, so it needs no rebuild (how a rollback works).
The only time ox reloads Caddy is when the routes themselves change, like a new domain. For those reloads ox sets the kernel's tcp_migrate_req, which hands queued connections to the new listener instead of resetting them.
You get this by adding a server and a repository; there is nothing to configure. The push-to-deploy features show the rest of a deploy, and the [app] section of the config reference covers the health path.
What this does not cover
Background workers restart in place. Workers that don't serve HTTP, like a Celery worker, are restarted after the switch. Their queue keeps the jobs, but a job that was running may need to be retried.
Migrations are not undone. A rollback doesn't revert the schema. ox says so before you confirm, and the snapshot taken before the migration is there to restore.
Long connections close. A websocket open to the old side is closed when that side stops, so clients need to reconnect.
Both sides use memory for a moment. The server needs room for two copies of the app during the switch.
Static files can race. A request in the same instant as the switch may ask for a file only the old release has. File names with a content hash, which most build tools make, avoid it.
Key takeaways
- A plain
systemctl restartdrops requests and leaves the site down if the new release fails to start. - Blue-green on one VPS needs only two systemd instances, a health check and a proxy switch. No Docker, no Kubernetes.
- Caddy's
{file.*}placeholder makes the switch a file rename, with no config reload. - Migrations must work with both the old and the new code, because both run for a moment.
- ox runs this on every deploy and rollback with nothing to configure.
Frequently asked questions
Is this blue-green deployment?
Yes, on one server. The two sides, each a systemd instance on its own port, are the blue and green copies, and the upstream file Caddy reads is the switch. Classic blue-green uses two servers behind a load balancer; on a single VPS the idea is the same, only cheaper, as long as the server has memory for both copies.
Do I need Docker or Kubernetes for zero-downtime deploys?
No. Two systemd services and a proxy that can switch between them are enough, on any VPS. Containers are one way to run two copies side by side, but a release folder per version and a systemd template unit do the same job with less on the server, and the pattern works for any language.
Can I do this with nginx instead of Caddy?
Yes. nginx's reload starts new workers and lets the old ones finish their requests, so the same two-sides pattern works with an upstream you change and reload. ox uses Caddy because it also handles HTTPS certificates on its own, and the per-request file read means the switch needs no reload at all.
What should the health check do?
Answer quickly with a success status once the app can serve real requests: ox accepts any 2xx or 3xx on the path you set. Checking that the database connection works is useful. Calling slow outside services is not, because a payment provider having a bad minute should not block your deploy.
Does this work for Django, Next.js and FastAPI?
Yes. The switch happens in front of the app, so any language works the same way, as long as the app reads its port from $PORT. See how to deploy Django to a VPS without Docker for a full example, and the Heroku alternative guide for moving an existing app.
Next step
Try the script above on a staging server first, with a deliberately broken commit, and watch the old side keep serving. Or let ox run the same switch for you: sign up with GitHub, add a server, and push.
