Files
exodevices/deploy/README.md
T
Yuriy PanovandClaude Opus 5 da93bbc9f3 Remove the release directory a failed deploy leaves behind
A deploy that died after mkdir -- a broken build, a missing lock file, a
health check that never went green -- left its half-finished release in
releases/. That is not just disk: it is the newest directory there, so the
KEEP=3 rotation of the next successful deploy counted it among the three to
keep and dropped a working release instead, thinning out exactly the rollback
targets the runbook tells you to use.

The release directory is now armed for cleanup the moment it is created and
disarmed once the release is live and healthy, with the removal done by the
same EXIT trap that drops the unpacked archive -- so it covers every way out:
fatal, a command failing under set -e, an aborted build. The one case left
alone is a release the swap did reach and that `current` still points at, i.e.
a first-ever deploy with nothing to roll back to: removing that would leave a
dangling symlink behind instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 22:33:54 +03:00

318 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Production deployment — four EXO landings on one VPS
Four independent landings, each on its own domain root, on a single reg.ru KVM VPS.
nginx serves the built SPA off disk; a small Express process per app handles only
`/api/*` and forwards leads to amoCRM. Releases are built on the server and swapped
in atomically via a `current` symlink.
`apps.conf` is the single source of truth for domains and ports — check it before provisioning.
| App | Domain | Port | amoCRM tag |
|---|---|---|---|
| `fitnes` | `fitness.exorecovery.ru` | 3000 | `fitness-landing` |
| `hotel` | `hotel.exorecovery.ru` | 3001 | `hotel-landing` |
| `medcenterphysio` | `physio.exorecovery.ru` | 3002 | `physio-landing` |
| `medcenterstart` | `start.exorecovery.ru` | 3003 | `start-landing` |
## VPS requirements
| | Minimum | Recommended |
|---|---|---|
| vCPU | 2 | 24 |
| RAM | 4 GB + 2 GB swap | 48 GB |
| Disk | 40 GB NVMe | 6080 GB NVMe |
| Network | 1 IPv4, ≥100 Mbps | + IPv6 |
| Virtualization | **KVM** | KVM |
Take the **KVM** line, not OpenVZ/LXC — the systemd sandboxing in `systemd/exo@.service`
and swap control need a real kernel. Moscow or SPb DC: leads are personal data of
Russian citizens, so 152-ФЗ wants them processed on RU infrastructure.
Why these numbers: four Node processes ≈ 400 MB total; nginx ~30 MB; Ubuntu ~300 MB —
idle well under 1 GB. The only spike is a build (~0.51 GB, one app at a time; the
toolchain is native — TypeScript 7 in Go, Vite 8 on Rolldown, Tailwind's Rust oxide).
Disk: ~40 MB per pruned release, 3 kept × 4 apps ≈ 1.5 GB steady.
## Software
| Component | Version |
|---|---|
| Ubuntu | 24.04 LTS (Debian 12/13 fine) |
| Node.js | 24.x LTS — hard floor is `>=22.12` from transitive deps |
| npm | 11.x (bundled) |
| nginx | ≥1.24 — needs `http2`, `gzip_static`, `ssl`, all built in |
| certbot | `python3-certbot-nginx` |
| ufw, fail2ban, unattended-upgrades | distro |
No database, no Redis, no Docker, no PM2. See §14 of the plan for why not Docker.
## Provisioning
### 1. Base OS
```bash
apt update && apt full-upgrade -y
apt install -y git curl nginx ufw fail2ban unattended-upgrades gzip unzip
timedatectl set-timezone Europe/Moscow
dpkg-reconfigure --priority=low unattended-upgrades
# Swap, if the plan doesn't provide it
fallocate -l 2G /swapfile && chmod 600 /swapfile && mkswap /swapfile && swapon /swapfile
echo '/swapfile none swap sw 0 0' >> /etc/fstab
echo 'vm.swappiness=10' > /etc/sysctl.d/99-swappiness.conf && sysctl --system
```
### 2. Node 24
```bash
curl -fsSL https://deb.nodesource.com/setup_24.x | bash -
apt install -y nodejs && node -v # v24.x
```
### 3. User, directories, repos
```bash
adduser --system --group --home /srv/exo --shell /bin/bash exo
mkdir -p /srv/exo/{fitnes,hotel,medcenterphysio,medcenterstart}/releases
mkdir -p /etc/exo /var/www/certbot
chown -R exo:exo /srv/exo
chmod 755 /srv/exo # www-data must traverse to reach dist/client
chown root:exo /etc/exo # without the group, exo cannot traverse it at 750
chmod 750 /etc/exo
for app in fitnes hotel medcenterphysio medcenterstart; do
sudo -u exo git clone --bare <REPO_URL> /srv/exo/$app/repo
done
```
All four clone the same repo; `exo-deploy` extracts one subdirectory each.
### 4. Env files, per app
Two files with different trust levels — see `env/app.env.example` and
`env/app.build.env.example` for annotated templates.
```bash
install -m 640 -o root -g exo env/app.env.example /etc/exo/fitnes.env
install -m 644 env/app.build.env.example /etc/exo/fitnes.build.env
# then edit both: PORT, AMO_LONG_LIVED_TOKEN, AMO_LEAD_TAGS, VITE_SITE_URL
```
The split matters. `VITE_SITE_URL` is a **build-time HTML substitution** — if it is
absent when `vite build` runs, the literal `%VITE_SITE_URL%` ships into the canonical
tag and OG metadata (`exo-deploy` aborts if it detects this). Keeping the amoCRM token
in the 640 runtime file means it never lands in a release directory. `dotenv` does not
override variables already in the environment, so the release `.env` (VITE vars) and
systemd's `EnvironmentFile` (secrets) coexist and the secrets win.
### 5. systemd
```bash
install -m 644 systemd/exo@.service /etc/systemd/system/
systemctl daemon-reload
systemctl enable exo@fitnes exo@hotel exo@medcenterphysio exo@medcenterstart
```
Let the deploy user restart only its own units — `visudo -f /etc/sudoers.d/exo-deploy`:
```
exo ALL=(root) NOPASSWD: /usr/bin/systemctl restart exo@fitnes, \
/usr/bin/systemctl restart exo@hotel, \
/usr/bin/systemctl restart exo@medcenterphysio, \
/usr/bin/systemctl restart exo@medcenterstart
```
### 6. Firewall and log retention
```bash
ufw default deny incoming && ufw default allow outgoing
ufw allow OpenSSH && ufw allow 80/tcp && ufw allow 443/tcp && ufw enable
```
Ports 30003003 are never opened, and the app binds `127.0.0.1` by default anyway
(`HOST` in `server/src/config.ts`). In `/etc/ssh/sshd_config.d/99-hardening.conf` set
`PasswordAuthentication no` and `PermitRootLogin no` — but install your key on a
sudo-capable non-root user *first*.
On an amoCRM failure the server deliberately logs the whole lead payload so a real
lead is never lost. That puts names and phones in the journal, so bound it in
`/etc/systemd/journald.conf`:
```ini
[Journal]
SystemMaxUse=500M
MaxRetentionSec=14day
```
### 7. nginx
```bash
install -m 644 nginx/snippets/*.conf /etc/nginx/snippets/
install -m 644 nginx/http-extras.conf /etc/nginx/conf.d/00-exo-http.conf
# Debian's stock nginx.conf sets some of the same http{} directives (gzip,
# server_tokens), and nginx refuses to start on a duplicate. Comment out every
# stock directive that 00-exo-http.conf now owns.
for d in $(grep -oE '^[a-z_]+' /etc/nginx/conf.d/00-exo-http.conf | sort -u); do
sed -i -E "s@^([[:space:]]*)($d[[:space:]]+[^;]*;)@\\1# \\2@" /etc/nginx/nginx.conf
done
bin/exo-render-nginx /etc/nginx/sites-available # after filling in apps.conf
ln -s /etc/nginx/sites-available/exo-*.conf /etc/nginx/sites-enabled/
rm -f /etc/nginx/sites-enabled/default
nginx -t && systemctl reload nginx
```
### 8. DNS and TLS
Order matters: nginx needs the port-80 vhost live before certbot can validate.
1. Per domain at reg.ru DNS: `A @ → <VPS IPv4>`, `A www → <VPS IPv4>` (+ `AAAA` if available).
2. Confirm: `dig +short <domain> @77.88.8.8`.
3. Comment out the two `443` blocks in each rendered vhost (they reference certs that
don't exist yet), then `nginx -t && systemctl reload nginx`.
4. Issue apex + www together, per domain:
```bash
certbot certonly --webroot -w /var/www/certbot \
-d <domain> -d www.<domain> --agree-tos -m <admin-email> --no-eff-email
```
5. Uncomment the `443` blocks, `nginx -t && systemctl reload nginx`.
6. Make renewal reload nginx:
```bash
printf '#!/bin/sh\nsystemctl reload nginx\n' > /etc/letsencrypt/renewal-hooks/deploy/reload-nginx
chmod +x /etc/letsencrypt/renewal-hooks/deploy/reload-nginx
certbot renew --dry-run
```
### 9. First deploy
```bash
install -m 755 bin/exo-deploy /usr/local/bin/
sudo -u exo exo-deploy medcenterphysio main # one app
sudo -u exo exo-deploy all # all four, in apps.conf order
```
## Deploying from a zip archive
`exo-deploy` takes its code from one of two places: the bare repo (above) or a zip
archive sitting on the server. The archive path exists for code that is not in git on
this box — a handover build, a contractor's snapshot, a hotfix from a laptop.
```bash
exo-deploy <app|all> [git-ref] # from the bare repo, default ref main
exo-deploy --zip <archive.zip> <app|all> # from an archive
```
**The archive holds sources, not a build.** Everything after the source lands in the
release directory is identical in both modes — `npm ci`, `npm run build`, the
`%VITE_SITE_URL%` check, `npm prune`, gzip, the atomic `current` swap, the health check
with automatic rollback. That is the point: `VITE_SITE_URL` comes from
`/etc/exo/<app>.build.env` on this server and never from whoever packed the archive.
Expected layout — the landing directories as they sit in the repo, either at the root of
the archive or under one wrapping directory (what Finder's "Compress" produces):
```
fitnes/ hotel/ medcenterphysio/ medcenterstart/ # or bundle/fitnes/ ...
package.json, package-lock.json, index.html,
src/, server/, shared/, scripts/, public/, tsconfig*.json, vite.config.ts
```
`package-lock.json` is required — `npm ci` refuses to run without it, and `exo-deploy`
says so up front for every app in the archive rather than failing halfway through.
`node_modules/`, `dist/`, `__MACOSX/` and `.DS_Store` are skipped during unpacking, and
any `.env` / `.env.local` that travelled inside the archive is deleted before the build.
Pack it on the dev machine, from the repo root:
```bash
zip -r exo-$(date +%Y%m%d).zip fitnes hotel medcenterphysio medcenterstart \
-x '*/node_modules/*' '*/dist/*' '*/.env' '*/.env.local' '*/.DS_Store' '*/legacy/*'
```
Then ship and deploy. `/tmp` works; the file only has to be readable by `exo`:
```bash
scp exo-20260101.zip <server>:/tmp/
sudo -u exo exo-deploy --zip /tmp/exo-20260101.zip all
```
`all` runs the four in `apps.conf` order and **stops at the first failure** — landings
already swapped in stay on their new release, the one that failed rolls itself back.
Rerun for the rest once the cause is fixed.
Each release records where it came from, so a live one can be traced long after the
deploy output has scrolled away:
```bash
cat /srv/exo/<app>/current/.deploy-source
# zip exo-20260101.zip sha256=9f86d0...
# deployed 2026-01-01T09:12:44Z
```
## Runbook
| Task | Command |
|---|---|
| Deploy | `sudo -u exo exo-deploy <app\|all> [ref]` |
| Deploy from a zip | `sudo -u exo exo-deploy --zip <archive.zip> <app\|all>` |
| What is live | `cat /srv/exo/<app>/current/.deploy-source` |
| Rollback | `ln -sfn /srv/exo/<app>/releases/<older> /srv/exo/<app>/current && sudo systemctl restart exo@<app>` |
| Tail logs | `journalctl -u exo@<app> -f` |
| All leads, last hour | `journalctl -u 'exo@*' --since '1 hour ago' \| grep '\[lead\]'` |
| Restart all | `systemctl restart 'exo@*'` |
| Check amoCRM wiring | `cd /srv/exo/<app>/current && npm run amo:check` |
| Rotate amoCRM token | edit `/etc/exo/*.env` → `systemctl restart 'exo@*'` → verify `"amo":true` |
| Renew certs | `certbot renew && systemctl reload nginx` |
| Disk usage | `du -sh /srv/exo/*/releases/*` |
`exo-deploy` health-checks after the swap and **rolls back automatically** if the new
release fails to answer `/api/health`. A deploy that fails anywhere — a broken build, a
missing lock file, a health check that never goes green — takes its half-finished release
directory with it, so the three kept releases are always three releases that actually ran
and any one of them is a safe rollback target.
## Monitoring
Each app answers `GET /api/health` with `{ ok, amo, pipelineId }`. Point an off-box
uptime checker at `https://<domain>/api/health` for all four, alerting on non-200 **and
on `"amo":false`** — the latter is how an expired amoCRM token shows up, and systemd
cannot see it. `Restart=always` covers crashes.
Log markers worth alerting on: `[lead] amoCRM submission failed`, `[lead] payload was:`
(a recoverable lead sitting in the journal), `[amo] AMO_SUBDOMAIN / AMO_LONG_LIVED_TOKEN
are not set`.
**The amoCRM token is the most likely future outage.** ~1 year TTL, fails as a 401 that
visitors see as a 502 with a phone number. Calendar a rotation ~11 months out.
## Backups
No database, and the code is in git, so the surface is small. Nightly tar, off the box:
- `/etc/exo/` — the amoCRM token and per-app config
- `/etc/nginx/sites-available/`, `/etc/nginx/snippets/`, `/etc/nginx/conf.d/`
- `/etc/letsencrypt/`
- `/etc/systemd/system/exo@.service`, `/etc/sudoers.d/exo-deploy`
Also enable reg.ru VPS snapshots — a full-image restore beats rebuilding under pressure.
## Known constraints
- **One process per app.** `server/src/rate-limit.ts` is an in-memory fixed-window
limiter, single-process by design. Horizontal scaling needs a shared store first.
Not a concern at landing-page traffic; nginx `limit_req` is the second layer.
- **No CI.** Deployment is a manual `exo-deploy`. A GitHub Actions job that SSHes and
runs it is a natural follow-up once the flow is proven. Until then `--zip` is the way
to deploy code that never reached this server's git remote.
- **One Metrika counter for four domains.** All four landings carry the same
Yandex.Metrika counter (`112352796`) in `index.html`, so reports mix the domains and
every domain has to be listed in the counter's settings, or its hits get filtered.
- **The `FORM_SUCCESS` goal has to exist in the Metrika UI.** `src/lib/lead.ts` fires
`ym(112352796, 'reachGoal', 'FORM_SUCCESS')` on every successful form submission, but
the goal itself is not part of the code: create it in the counter's settings as
«JavaScript-событие» with the identifier `FORM_SUCCESS`, or the hits arrive and no
report ever shows them. The honeypot answer (`leadId: 0`) deliberately does not count.
- **Still open before launch** (not a deployment blocker): the forms show implicit 152-ФЗ
consent text with no link to a published privacy policy.