Payments Engine

Self-hosting

What the engine consists of, what it needs from you, and how to upgrade it without a maintenance window. It is two Node processes and one Postgres database — deliberately, so that operating it does not require operating a cluster.

Processes

ProcessCommandWhat it doesScale
Operator APInpm run start:apiServes the signed HTTP API and the health endpoints. Stateless between requests.Horizontally. Every guarantee is in the database, not in a process.
Workernpm run start:workerWebhook relay, treasury float checks, reconciliation sweeps. Each loop is isolated from the others.More than one is safe — work is claimed with FOR UPDATE SKIP LOCKED.
Consolenpm run start:webThis site and the operator console. Signs on the server so the credential never reaches a browser.Optional. It is a read surface.
Adminnpm run start:adminPlatform admin across operators. Read-only.Optional.

Environment

Configuration is read once at boot and validated as a whole. A misconfigured process refuses to start and names every variable that is wrong in one message — not the first one, and never the value.

VariableRequiredNotes
PAYMENTS_DATABASE_URLYesStandard Postgres URL. The pool sets idle_in_transaction_session_timeout so a leaked transaction cannot hold a lock forever.
PAYMENTS_PORTNo — 8080Refused rather than clamped if outside 1–65535.
PAYMENTS_OPERATORSYesid:secret, comma-separated. Held by the host process only.
PAYMENTS_MIGRATIONS_DIRNoOverrides migration discovery. Useful when the engine runs from a bundle rather than the repository.
PAYMENTS_ENVNo — developmentAppears in logs and on the console so nobody mistakes staging for production.
PAYMENTS_WORKER_PORTNo — 8081The worker serves its own health endpoint.

Platform Admin

The admin surface reads every operator's position, so it has its own sign-in rather than relying on nobody finding the port.

VariableRequiredNotes
PAYMENTS_ADMIN_USERSYesusername:$argon2id$… entries, separated by ; or a newline. Not commas — an argon2 digest contains them. With none set, nobody can sign in and the form says so.
PAYMENTS_ADMIN_SESSION_KEYYes when users existAt least 32 characters and independent from PAYMENTS_ADMIN_SECRET. Reuse is refused.
PAYMENTS_ADMIN_SECURE_COOKIESYes in productionMust be true in production. It adds Secure, which makes the cookie unusable over plain http.
PAYMENTS_ADMIN_SETTINGS_PATHNoWhere this surface keeps its own preferences. A read-only filesystem degrades to defaults rather than failing.
generating a digest
$ npm run admin:hash --workspace @payments/admin
Password for the admin account: (typed, not echoed)

alice:$argon2id$v=19$m=65536,t=3,p=4$…
The password is prompted for, never passed as an argument. An argument is in your shell history and in the process list, both readable by anyone else with a login on that machine.
Sessions are signed tokens rather than rows in a table, so there is nothing to leak — but also nothing to revoke individually. Removing a user from PAYMENTS_ADMIN_USERS and restarting invalidates their session immediately, because a session naming an account that is no longer configured is not a session.
Give every process its own operator credential. The console, the platform admin and each of your own services need separate ones. Two processes sharing a credential will eventually poll the same endpoint in the same second — identical canonical string, identical signature — and the engine admits a signature once, so one of them gets a 401 that looks like a configuration fault and is really a collision. Separate credentials also give you separate audit trails and independent revocation, which you want anyway.
env
# One per surface, not one shared between them.
PAYMENTS_OPERATOR_APP00001_SECRET=…    # your application
PAYMENTS_OPERATOR_CONSOLE01_SECRET=…   # the operator console
PAYMENTS_OPERATOR_ADMIN0001_SECRET=…   # the platform admin
Secrets never print. A configuration error names the variable, never its value, and a secret held in memory stringifies to [redacted] — so a stack trace, a log line or an accidental console.log cannot leak one. Reading it requires an explicit .reveal(), which is greppable.

Migrations

Applied at boot, inside a transaction, under pg_advisory_xact_lock — so ten instances starting simultaneously produce one migration run and nine waits. Checksums are recorded: an already-applied migration that has since been edited stops the boot rather than silently diverging from what the database actually has.

Every migration is expand-only. A column is added, backfilled and read before anything stops writing the old one, and the contract half ships in a later release. This is enforced by a lint rule in CI, and the practical consequence is that a deploy can be rolled back without a reverse data migration.

CI
$ npm run migration:lint
✓ 76 migrations, expand-then-contract respected
✓ no destructive statement outside a _contract file

Health checks

EndpointMeaningWhat to wire it to
GET /health/liveThe process is running and its event loop is responsive.Restart policy. A failure here means kill and restart.
GET /health/readyDependencies answer — the database in particular.Load balancer. A failure means take out of rotation, do not restart.
Wiring liveness to a dependency check is the classic mistake: a database blip then restarts every instance simultaneously, turning a recoverable outage into an outage plus a thundering herd.

Upgrading without a window

  1. Deploy the new version alongside the old. Expand migrations run at boot and are safe for the old code to keep running against.
  2. Shift traffic. Both versions read and write the same schema by construction.
  3. Retire the old instances.
  4. The contract migration ships in a later release, once nothing is left that reads the old shape.

Backups

The database is the system of record for everything except the chain itself. A standard pg_dump plus WAL archiving is sufficient and there is nothing else to back up — no local state files, no queue to drain, no in-memory position to lose.

Key material is the exception. Sealed signing material lives in the custody schema, so a database backup contains it in sealed form. Treat those backups exactly as you would treat the keys themselves: the sealing is what stands between a leaked dump and a drained wallet.

Observability

Every process writes one JSON object per line to stdout. Each request carries a requestId that also appears in the error envelope returned to the caller, which is how an opaque 401 is diagnosed: the caller quotes the ID, and the log says which of the four causes it actually was.

stdout
{"level":"warn","event":"auth.rejected","reason":"timestamp_outside_window",
 "operatorId":"op_abcdefgh","skewSeconds":184,"requestId":"3c1ac997-…"}

Clock skew is the most common cause in practice, and it is the one that produces intermittent failures rather than consistent ones. Run NTP on anything that signs.