Files
Yuriy PanovandClaude Opus 5 da93bbc9f3 Remove the release directory a failed deploy leaves behind
A deploy that died after mkdir -- a broken build, a missing lock file, a
health check that never went green -- left its half-finished release in
releases/. That is not just disk: it is the newest directory there, so the
KEEP=3 rotation of the next successful deploy counted it among the three to
keep and dropped a working release instead, thinning out exactly the rollback
targets the runbook tells you to use.

The release directory is now armed for cleanup the moment it is created and
disarmed once the release is live and healthy, with the removal done by the
same EXIT trap that drops the unpacked archive -- so it covers every way out:
fatal, a command failing under set -e, an aborted build. The one case left
alone is a release the swap did reach and that `current` still points at, i.e.
a first-ever deploy with nothing to roll back to: removing that would leave a
dangling symlink behind instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 22:33:54 +03:00
..

Production deployment — four EXO landings on one VPS

Four independent landings, each on its own domain root, on a single reg.ru KVM VPS. nginx serves the built SPA off disk; a small Express process per app handles only /api/* and forwards leads to amoCRM. Releases are built on the server and swapped in atomically via a current symlink.

apps.conf is the single source of truth for domains and ports — check it before provisioning.

App Domain Port amoCRM tag
fitnes fitness.exorecovery.ru 3000 fitness-landing
hotel hotel.exorecovery.ru 3001 hotel-landing
medcenterphysio physio.exorecovery.ru 3002 physio-landing
medcenterstart start.exorecovery.ru 3003 start-landing

VPS requirements

Minimum Recommended
vCPU 2 24
RAM 4 GB + 2 GB swap 48 GB
Disk 40 GB NVMe 6080 GB NVMe
Network 1 IPv4, ≥100 Mbps + IPv6
Virtualization KVM KVM

Take the KVM line, not OpenVZ/LXC — the systemd sandboxing in systemd/exo@.service and swap control need a real kernel. Moscow or SPb DC: leads are personal data of Russian citizens, so 152-ФЗ wants them processed on RU infrastructure.

Why these numbers: four Node processes ≈ 400 MB total; nginx ~30 MB; Ubuntu ~300 MB — idle well under 1 GB. The only spike is a build (~0.51 GB, one app at a time; the toolchain is native — TypeScript 7 in Go, Vite 8 on Rolldown, Tailwind's Rust oxide). Disk: ~40 MB per pruned release, 3 kept × 4 apps ≈ 1.5 GB steady.

Software

Component Version
Ubuntu 24.04 LTS (Debian 12/13 fine)
Node.js 24.x LTS — hard floor is >=22.12 from transitive deps
npm 11.x (bundled)
nginx ≥1.24 — needs http2, gzip_static, ssl, all built in
certbot python3-certbot-nginx
ufw, fail2ban, unattended-upgrades distro

No database, no Redis, no Docker, no PM2. See §14 of the plan for why not Docker.

Provisioning

1. Base OS

apt update && apt full-upgrade -y
apt install -y git curl nginx ufw fail2ban unattended-upgrades gzip unzip
timedatectl set-timezone Europe/Moscow
dpkg-reconfigure --priority=low unattended-upgrades

# Swap, if the plan doesn't provide it
fallocate -l 2G /swapfile && chmod 600 /swapfile && mkswap /swapfile && swapon /swapfile
echo '/swapfile none swap sw 0 0' >> /etc/fstab
echo 'vm.swappiness=10' > /etc/sysctl.d/99-swappiness.conf && sysctl --system

2. Node 24

curl -fsSL https://deb.nodesource.com/setup_24.x | bash -
apt install -y nodejs && node -v   # v24.x

3. User, directories, repos

adduser --system --group --home /srv/exo --shell /bin/bash exo
mkdir -p /srv/exo/{fitnes,hotel,medcenterphysio,medcenterstart}/releases
mkdir -p /etc/exo /var/www/certbot
chown -R exo:exo /srv/exo
chmod 755 /srv/exo          # www-data must traverse to reach dist/client
chown root:exo /etc/exo     # without the group, exo cannot traverse it at 750
chmod 750 /etc/exo

for app in fitnes hotel medcenterphysio medcenterstart; do
  sudo -u exo git clone --bare <REPO_URL> /srv/exo/$app/repo
done

All four clone the same repo; exo-deploy extracts one subdirectory each.

4. Env files, per app

Two files with different trust levels — see env/app.env.example and env/app.build.env.example for annotated templates.

install -m 640 -o root -g exo env/app.env.example       /etc/exo/fitnes.env
install -m 644                 env/app.build.env.example /etc/exo/fitnes.build.env
# then edit both: PORT, AMO_LONG_LIVED_TOKEN, AMO_LEAD_TAGS, VITE_SITE_URL

The split matters. VITE_SITE_URL is a build-time HTML substitution — if it is absent when vite build runs, the literal %VITE_SITE_URL% ships into the canonical tag and OG metadata (exo-deploy aborts if it detects this). Keeping the amoCRM token in the 640 runtime file means it never lands in a release directory. dotenv does not override variables already in the environment, so the release .env (VITE vars) and systemd's EnvironmentFile (secrets) coexist and the secrets win.

5. systemd

install -m 644 systemd/exo@.service /etc/systemd/system/
systemctl daemon-reload
systemctl enable exo@fitnes exo@hotel exo@medcenterphysio exo@medcenterstart

Let the deploy user restart only its own units — visudo -f /etc/sudoers.d/exo-deploy:

exo ALL=(root) NOPASSWD: /usr/bin/systemctl restart exo@fitnes, \
  /usr/bin/systemctl restart exo@hotel, \
  /usr/bin/systemctl restart exo@medcenterphysio, \
  /usr/bin/systemctl restart exo@medcenterstart

6. Firewall and log retention

ufw default deny incoming && ufw default allow outgoing
ufw allow OpenSSH && ufw allow 80/tcp && ufw allow 443/tcp && ufw enable

Ports 30003003 are never opened, and the app binds 127.0.0.1 by default anyway (HOST in server/src/config.ts). In /etc/ssh/sshd_config.d/99-hardening.conf set PasswordAuthentication no and PermitRootLogin no — but install your key on a sudo-capable non-root user first.

On an amoCRM failure the server deliberately logs the whole lead payload so a real lead is never lost. That puts names and phones in the journal, so bound it in /etc/systemd/journald.conf:

[Journal]
SystemMaxUse=500M
MaxRetentionSec=14day

7. nginx

install -m 644 nginx/snippets/*.conf /etc/nginx/snippets/
install -m 644 nginx/http-extras.conf /etc/nginx/conf.d/00-exo-http.conf

# Debian's stock nginx.conf sets some of the same http{} directives (gzip,
# server_tokens), and nginx refuses to start on a duplicate. Comment out every
# stock directive that 00-exo-http.conf now owns.
for d in $(grep -oE '^[a-z_]+' /etc/nginx/conf.d/00-exo-http.conf | sort -u); do
  sed -i -E "s@^([[:space:]]*)($d[[:space:]]+[^;]*;)@\\1# \\2@" /etc/nginx/nginx.conf
done

bin/exo-render-nginx /etc/nginx/sites-available     # after filling in apps.conf
ln -s /etc/nginx/sites-available/exo-*.conf /etc/nginx/sites-enabled/
rm -f /etc/nginx/sites-enabled/default
nginx -t && systemctl reload nginx

8. DNS and TLS

Order matters: nginx needs the port-80 vhost live before certbot can validate.

  1. Per domain at reg.ru DNS: A @ → <VPS IPv4>, A www → <VPS IPv4> (+ AAAA if available).
  2. Confirm: dig +short <domain> @77.88.8.8.
  3. Comment out the two 443 blocks in each rendered vhost (they reference certs that don't exist yet), then nginx -t && systemctl reload nginx.
  4. Issue apex + www together, per domain:
    certbot certonly --webroot -w /var/www/certbot \
      -d <domain> -d www.<domain> --agree-tos -m <admin-email> --no-eff-email
    
  5. Uncomment the 443 blocks, nginx -t && systemctl reload nginx.
  6. Make renewal reload nginx:
    printf '#!/bin/sh\nsystemctl reload nginx\n' > /etc/letsencrypt/renewal-hooks/deploy/reload-nginx
    chmod +x /etc/letsencrypt/renewal-hooks/deploy/reload-nginx
    certbot renew --dry-run
    

9. First deploy

install -m 755 bin/exo-deploy /usr/local/bin/
sudo -u exo exo-deploy medcenterphysio main   # one app
sudo -u exo exo-deploy all                    # all four, in apps.conf order

Deploying from a zip archive

exo-deploy takes its code from one of two places: the bare repo (above) or a zip archive sitting on the server. The archive path exists for code that is not in git on this box — a handover build, a contractor's snapshot, a hotfix from a laptop.

exo-deploy <app|all> [git-ref]              # from the bare repo, default ref main
exo-deploy --zip <archive.zip> <app|all>    # from an archive

The archive holds sources, not a build. Everything after the source lands in the release directory is identical in both modes — npm ci, npm run build, the %VITE_SITE_URL% check, npm prune, gzip, the atomic current swap, the health check with automatic rollback. That is the point: VITE_SITE_URL comes from /etc/exo/<app>.build.env on this server and never from whoever packed the archive.

Expected layout — the landing directories as they sit in the repo, either at the root of the archive or under one wrapping directory (what Finder's "Compress" produces):

fitnes/  hotel/  medcenterphysio/  medcenterstart/     # or bundle/fitnes/ ...
  package.json, package-lock.json, index.html,
  src/, server/, shared/, scripts/, public/, tsconfig*.json, vite.config.ts

package-lock.json is required — npm ci refuses to run without it, and exo-deploy says so up front for every app in the archive rather than failing halfway through. node_modules/, dist/, __MACOSX/ and .DS_Store are skipped during unpacking, and any .env / .env.local that travelled inside the archive is deleted before the build.

Pack it on the dev machine, from the repo root:

zip -r exo-$(date +%Y%m%d).zip fitnes hotel medcenterphysio medcenterstart \
  -x '*/node_modules/*' '*/dist/*' '*/.env' '*/.env.local' '*/.DS_Store' '*/legacy/*'

Then ship and deploy. /tmp works; the file only has to be readable by exo:

scp exo-20260101.zip <server>:/tmp/
sudo -u exo exo-deploy --zip /tmp/exo-20260101.zip all

all runs the four in apps.conf order and stops at the first failure — landings already swapped in stay on their new release, the one that failed rolls itself back. Rerun for the rest once the cause is fixed.

Each release records where it came from, so a live one can be traced long after the deploy output has scrolled away:

cat /srv/exo/<app>/current/.deploy-source
# zip exo-20260101.zip sha256=9f86d0...
# deployed 2026-01-01T09:12:44Z

Runbook

Task Command
Deploy sudo -u exo exo-deploy <app|all> [ref]
Deploy from a zip sudo -u exo exo-deploy --zip <archive.zip> <app|all>
What is live cat /srv/exo/<app>/current/.deploy-source
Rollback ln -sfn /srv/exo/<app>/releases/<older> /srv/exo/<app>/current && sudo systemctl restart exo@<app>
Tail logs journalctl -u exo@<app> -f
All leads, last hour journalctl -u 'exo@*' --since '1 hour ago' | grep '\[lead\]'
Restart all systemctl restart 'exo@*'
Check amoCRM wiring cd /srv/exo/<app>/current && npm run amo:check
Rotate amoCRM token edit /etc/exo/*.envsystemctl restart 'exo@*' → verify "amo":true
Renew certs certbot renew && systemctl reload nginx
Disk usage du -sh /srv/exo/*/releases/*

exo-deploy health-checks after the swap and rolls back automatically if the new release fails to answer /api/health. A deploy that fails anywhere — a broken build, a missing lock file, a health check that never goes green — takes its half-finished release directory with it, so the three kept releases are always three releases that actually ran and any one of them is a safe rollback target.

Monitoring

Each app answers GET /api/health with { ok, amo, pipelineId }. Point an off-box uptime checker at https://<domain>/api/health for all four, alerting on non-200 and on "amo":false — the latter is how an expired amoCRM token shows up, and systemd cannot see it. Restart=always covers crashes.

Log markers worth alerting on: [lead] amoCRM submission failed, [lead] payload was: (a recoverable lead sitting in the journal), [amo] AMO_SUBDOMAIN / AMO_LONG_LIVED_TOKEN are not set.

The amoCRM token is the most likely future outage. ~1 year TTL, fails as a 401 that visitors see as a 502 with a phone number. Calendar a rotation ~11 months out.

Backups

No database, and the code is in git, so the surface is small. Nightly tar, off the box:

  • /etc/exo/ — the amoCRM token and per-app config
  • /etc/nginx/sites-available/, /etc/nginx/snippets/, /etc/nginx/conf.d/
  • /etc/letsencrypt/
  • /etc/systemd/system/exo@.service, /etc/sudoers.d/exo-deploy

Also enable reg.ru VPS snapshots — a full-image restore beats rebuilding under pressure.

Known constraints

  • One process per app. server/src/rate-limit.ts is an in-memory fixed-window limiter, single-process by design. Horizontal scaling needs a shared store first. Not a concern at landing-page traffic; nginx limit_req is the second layer.
  • No CI. Deployment is a manual exo-deploy. A GitHub Actions job that SSHes and runs it is a natural follow-up once the flow is proven. Until then --zip is the way to deploy code that never reached this server's git remote.
  • One Metrika counter for four domains. All four landings carry the same Yandex.Metrika counter (112352796) in index.html, so reports mix the domains and every domain has to be listed in the counter's settings, or its hits get filtered.
  • The FORM_SUCCESS goal has to exist in the Metrika UI. src/lib/lead.ts fires ym(112352796, 'reachGoal', 'FORM_SUCCESS') on every successful form submission, but the goal itself is not part of the code: create it in the counter's settings as «JavaScript-событие» with the identifier FORM_SUCCESS, or the hits arrive and no report ever shows them. The honeypot answer (leadId: 0) deliberately does not count.
  • Still open before launch (not a deployment blocker): the forms show implicit 152-ФЗ consent text with no link to a published privacy policy.