The addresses on this page were filled in with this server's own address, https://qm.weekday100.com.
Queue Manager — Operations Guide
Everything you have to decide before putting this server in front of real traffic — configuration, the proxy/CDN setup that per-IP limits depend on, the API keys, and how a pass is bound to a visitor — and everything you use once it is there: the operator console, scheduled drops, and metrics.
INTEGRATION.md documents the customer-site side (script tag, /api/check). This file documents the server side. Both are served by the running server at /docs, and linked from the dashboard header. Thai translations are served at /docs/th (sources in docs/th/); the English files are authoritative.
Environment variables
Everything is optional; defaults are safe for a single-instance install on localhost, and deliberately not safe to expose unchanged.
| Variable | Default | What it does |
|---|---|---|
PORT | 8080 | Listen port. |
HOST | (unset) | Listen address. Unset binds every interface — except while ADMIN_KEY is the default, when unset means loopback (127.0.0.1) only, and a HOST that is not loopback refuses to start (see Refusals). Set it to 127.0.0.1 to answer only on loopback — worth doing when a reverse proxy on the same box is the only thing that should reach this process. |
DATA_DIR | (removed) | Removed in step 5: nothing is written to local disk. Set, it is ignored with a boot warning (DATA_DIR is set but ignored since step 5: nothing is written to local disk.), so an old manifest still deploys. Delete it from the manifest. |
ADMIN_KEY | admin-dev | Private API key: /api/admin/* (create/delete rooms, rates, flush, eject, metrics). Never distribute it. While it is the default the server listens on loopback only and refuses a non-loopback HOST — the default is printed in these docs, so it must never reach another machine. |
PUBLIC_API_KEY | (unset) | Public API key: the only thing it unlocks is vouching for an end client on /api/check?ip=&agent=. Sent as Authorization: Bearer only — ?key= is refused (401 credential_in_query). Safe to ship to every edge server / backend. Without it those calls need ADMIN_KEY, which is a privilege escalation — a warning says so. |
ADMIN_HOST | (unset) | Comma-separated hostnames the operator console and /api/admin/* answer on. Unset, a snippet-only install answers anywhere; once a room is inline an unset value means the console answers only on an address literal (127.0.0.1, the pod IP), on localhost or on the PUBLIC_URL host (QM-440), and on any other Host the console surface and /__qm/ are a bare 404 while the queue itself keeps answering, exactly as with ADMIN_HOST set (QM-428; the customer's own hostname still serves the console under /__qm/). Set this to put the console on a real name: set, it is the only place the console answers — on every other hostname (and on the bare pod address) the console surface answers 404: /, /api/admin/*, /metrics, /docs, /progress and the console's own file /public/admin.html, and on a hostname that fronts an inline room the same paths under /__qm/. It gates only the console: the queue itself (/healthz, /snippet/qm.js, /w/<room>, /api/check, /api/join, status, events, notify, the waiting page) keeps answering on any hostname, so the snippet host and a probe by pod IP work unchanged. A room may not be put inline on an ADMIN_HOST name or the PUBLIC_URL host (400 reserved_host), and the console surface on those names is never handed to a room, even one saved before that was refused. Unset, the address literals and localhost are the console's too, so their /api/admin/* and /api/v1/admin/* are never handed to a room either, even one whose targetUrl is the server's own address (QM-439); the rest of such a host stays the room's, which is how inline mode runs on a laptop. And on any host, the proxy never forwards this server's own credential: an Authorization: Bearer carrying ADMIN_KEY or an operator key is dropped before the request reaches the origin, while the app's own Authorization passes unchanged. See Two deployments. |
ADMIN_AUTH_FAIL_PER_MIN | 20 | Failed admin authentications allowed per client address per minute (token bucket). Past it, every credential from that address — the right one included — gets 429 too_many_auth_failures with Retry-After until one failure's worth has refilled. Only failures spend it. 0 disables. See Two keys, two blast radii. |
SECRET | random per boot (memory); required with SIDESTORE=pg | Root signing secret for visitor tickets, monitor links, console sessions and operator-key verifiers. By default they are signed with it directly, as in earlier releases; with TOKEN_SIGN_V2=1 each purpose signs with its own key derived from it (HKDF-SHA256) and every token names it by a kid. Both forms are always accepted. Set it if you run more than one instance, or tokens minted by one are rejected by the others. Unset, a new key is made on every boot, so a restart voids every ticket, monitor link, console session and operator key (fine for development, where nothing else survives a restart either). Required with SIDESTORE=pg: the server refuses to start without it, because a new key on every deploy voids every pass (tickets, monitor links, operator keys) that Postgres still holds. See Signing keys and rotating SECRET. |
SECRET_PREVIOUS | (unset) | The SECRET being rotated out. Tokens signed under it still verify (either form, every purpose); new tokens are always signed under SECRET. Operator keys resolved under it are re-keyed to SECRET on use. Remove it once the rotation has settled — /api/admin/health lists it under hardening while it is set. See Signing keys and rotating SECRET. |
TOKEN_SIGN_V2 | 0 | 1 signs tokens and new operator verifiers in the v2 form: a key per purpose and a kid. 0 (default) signs the legacy form, raw SECRET and no kid, which the release before per-purpose keys also reads. Both forms are accepted either way. Once you turn it on, you can't roll back below this release without losing the queue — see Signing keys and rotating SECRET. Any value other than 0/1 counts as 0 and is named in warnings. |
TRUST_PROXY | 0 | Number of reverse proxies in front of this server. The client address is taken from X-Forwarded-For counted from the right — and only when the connection comes from an address in TRUST_PROXY_IPS. 0 means the socket address is the client. |
TRUST_PROXY_IPS | (empty) | Comma-separated addresses or CIDR ranges of those proxies, IPv4 or IPv6 (10.0.0.7, 173.245.48.0/20, 2400:cb00::/32). An IPv4-mapped peer (::ffff:1.2.3.4) matches IPv4 entries. An entry that does not parse stops the server at startup, naming the entry. From any other peer, X-Forwarded-For is ignored. Required to get per-visitor throttling behind a CDN (see below). |
EDGE_SECRET | (unset) | The secret a Cloudflare request-header rule adds as X-QM-Edge (see Cloudflare edge secret). Set, every request and WebSocket upgrade without a matching header is refused 403 (edge_required), except /healthz; the client address becomes cf-connecting-ip and TRUST_PROXY/TRUST_PROXY_IPS are ignored. Comma-separated for rotation (any listed value passes). An entry shorter than 16 characters stops the server at startup. Unset with STORE=valkey or NODE_ENV=production, a startup line says bypassing requests are accepted. |
IPV6_PREFIX_BITS | 64 | What "per address" means for an IPv6 client: its first N bits (32..128). An IPv6 host is normally handed a whole /64 and may use any address in it, so every per-address limit — the join, check, notify, page and SSE budgets, the failed-login lockout, queueMaxPerIp, preQueueMaxPerIp and the traffic-classification record — is keyed on the /64, not the /128. 128 restores one key per address. IPv4 and IPv4-mapped (::ffff:1.2.3.4) clients are unaffected. Only the key changes: logs, the audit trail and the console still show the full address, and TRUST_PROXY_IPS still matches the real peer exactly or by its own CIDR. |
JOIN_LIMIT_PER_MIN | 600 | Per-address /api/join budget. 0 disables. Per instance with STORE=valkey: each instance keeps its own count, so N instances allow one client up to N× this across the fleet; Cloudflare's edge rate limiting is the fleet-wide control. |
CHECK_LIMIT_PER_MIN | 6000 | Per-address /api/check budget. 0 disables. Per instance with STORE=valkey: each instance keeps its own count, so N instances allow one client up to N× this across the fleet; Cloudflare's edge rate limiting is the fleet-wide control. |
RATE_LIMIT_MAX_KEYS | 200000 | Most distinct addresses each rate-limit bucket map holds. Past it the least recently used are evicted — never the whole map, so key churn cannot hand a throttled client a fresh budget. The ADMIN_AUTH_FAIL_PER_MIN lockout is held the same way. |
WAITING_LIMIT_PER_MIN | 120 | Per-address budget for the waiting page itself, and — in a bucket of its own — for the console page, /progress and /docs (also under /__qm/). 0 disables both. Stands down automatically for an undeclared proxy's address (that address only) — see Rate limits. Per instance with STORE=valkey: each instance keeps its own count, so N instances allow one client up to N× this across the fleet; Cloudflare's edge rate limiting is the fleet-wide control. |
SSE_PER_IP | 200 | Concurrent /events streams per address. 0 disables. Sized for a whole office or CGNAT block sharing one address; SSE_MAX_TOTAL is the cap that protects process memory. Per instance with STORE=valkey: each instance holds its own ceiling, so N instances hold up to N× this in total; Cloudflare is the fleet-wide control. |
SSE_MAX_TOTAL | 5000 | Global SSE ceiling (visitors + admins). Sized by CPU, not memory: an update that moves every visitor costs one kernel send per open stream, about 0.1 ms each on the reference box (QM-381), so 5,000 streams is ~0.5 s of CPU per update against the 1 s engine tick and 20,000 would be ~2 s. The fan-out runs in slices of 256 streams, so a large one no longer stalls the event loop, but above this cap each visitor simply gets fewer updates. A visitor refused a stream (503 sse_capacity) is not turned away: the waiting page polls /api/status every 5 s instead. Raise it only on a faster box, after measuring. Frames whose content has not changed are not re-sent, except as a keepalive to a stream that has been quiet for 4 s. Per instance with STORE=valkey: each instance holds its own ceiling, so N instances hold up to N× this in total; Cloudflare is the fleet-wide control. |
ADMIN_SSE_RESERVE | 8 | Extra stream slots above SSE_MAX_TOTAL that only a console holding a write-capable credential may use — ADMIN_KEY, or a named owner/operator key — so the dashboard can still connect when visitors have filled the global ceiling — which is exactly when you need to look at it. A monitor share link is a read-only credential and gets no exemption: it competes for SSE_MAX_TOTAL like a visitor. |
MAX_CONNECTIONS | 50000 | Soft socket ceiling, enforced by this process itself: at and above this many open connections, every NEW request is answered 503 (see qm_connections_shed_total) instead of being routed or proxied. |
CONNECTION_RESERVE | 256 | Headroom ABOVE MAX_CONNECTIONS before Node's own hard ceiling (server.maxConnections) is reached. Without it, the socket that tips the count past MAX_CONNECTIONS has nowhere left to be answered from — Node destroys it in C++ before any 503 can be sent. |
TOKEN_MAX_AGE_SEC | 86400 | How long a signed token stays verifiable at all. A visitor still waiting keeps their place past it (QM-391): once their token is half the shorter of this and QUEUE_COOKIE_MAX_AGE_SEC old (6 h with both defaults), /api/status and /events re-sign the same place with the current keys and re-set both queue cookies. Only for the browser whose qms_<room> cookie holds the place, in the room's current generation, while the place is still waiting or pre-queued; a pass, a lapsed or ejected place, or a caller matched only by fingerprint gets nothing. A token already past this age is still refused. Lowering it therefore does not cut a long wait short for a visitor whose page is open; it bounds how long a token copied out of a log stays usable. |
REFRESH_ENDS_PER_SEC | 200 | At most this many open /events streams are ended per second for their token refresh (QM-391). Streams that fall due together wait their turn, still receiving status frames, instead of all reconnecting in the same instant. The drain goes faster only when this rate would not end every waiting stream before its token lapses. |
MONITOR_TOKEN_TTL_SEC | 604800 | How long a Share monitor link stays live before it expires on its own — seven days by default. Expiry is independent of revocation: a link can be revoked earlier from the console, and an expired one cannot be re-copied. |
EMAIL_NOTIFY | (unset) | 1 (also true/yes/on) enables POST /api/notify, the turn-notification capture, and the e-mail field on the waiting page. Off by default and the page hides the control entirely when it is off, because this server records the intent and sends nothing itself — offering a channel that goes nowhere is worse than not offering one. With it off the route answers 404 email_notify_disabled. When on, an address is kept with its place in line and deleted when the ticket ends, or up to LAPSED_TTL_MS (24 h) later if its entry window lapsed: the ticket ends when it is admitted, ejected or purged, or its room is deleted; after a lapse, a visitor who comes back keeps the address on the new ticket. No entry is kept more than 24 h after it was filed, and a request against a ticket that has ended is refused (409 ticket_ended). With STORE=memory it is in this process's memory, and every restart or deploy loses all of them. With STORE=valkey it is in the shared Valkey (qm:notify:{room}), so the console on every instance sees it. Valkey RDB snapshots and provider backups can hold a copy after deletion until they rotate out. It is never written to Postgres (events or audit), the logs or /metrics, never sent by this server, and never shown to anyone unmasked (the console's visitor drill-down sees al•••@e••••••.com). Nothing in this build can export them. Turning it on raises a warning at boot and in /api/admin/health warnings, so the console badge says so too; the waiting page tells the visitor the same. |
NOTIFY_LIMIT_PER_MIN | 20 | Per-address POST /api/notify budget. A visitor sets an address once and maybe corrects it once; more than a handful a minute from one address is a script. 0 disables. Per instance with STORE=valkey: each instance keeps its own count, so N instances allow one client up to N× this across the fleet; Cloudflare's edge rate limiting is the fleet-wide control. |
NOTIFY_MAX_ENTRIES | 50000 | Hard ceiling on stored notification intents, so an address book cannot be pushed into this process until it runs out of memory. Oldest entries are evicted first. STORE=memory only: with STORE=valkey an entry needs a signed ticket, so a room's hash is bounded by the places it has handed out, and each entry goes when its ticket ends. |
QUEUE_COOKIE_MAX_AGE_SEC | 43200 | Lifetime of the queue cookies (qm_<room>, qms_<room>), counted from when they were last set. A waiting page that stays open has them re-set with its token refresh (see TOKEN_MAX_AGE_SEC); a browser that closes the page for longer than this loses the qms_<room> proof and, with it, the place. |
COOKIE_SECURE | (unset) | 1 forces Secure on cookies even when this server sees plain HTTP (TLS terminated upstream by a proxy that sends no X-Forwarded-Proto, or is not on TRUST_PROXY_IPS). Without it, cookies are still Secure on a request whose Host is the host of an https:// PUBLIC_URL. |
PROXY_TIMEOUT_MS | 30000 | Inline mode only: how long to wait for the app behind a room's proxyOrigin before answering 504. See Two deployments below. |
PROXY_ALLOW_PRIVATE | unset | 1 lets a room's proxyOrigin (and the health probe of it), and the health probe of a snippet room's targetUrl, reach loopback, RFC1918, link-local/cloud-metadata and other private addresses, which are refused by default. A refused probe reads down, with targetHealth.blocked naming the address; nothing is sent to it. Autotune makes no decision on a refused probe and holds the rate; after 3 refused in a row it undoes its own raise, once, back to the rate the room had before autotune raised it (a cut is held, and a rate an operator set is left alone). The base is kept in the room's config (written with the rate), so it survives a failover or a restart: the new leader falls back the same way. An operator's rate set while the fallback is being written is not overwritten. Set it only when the app really is on the LAN. See Two deployments below. |
PROXY_IDLE_TIMEOUT_MS | 60000 | Inline mode only: how long a response that has already started may go without a byte before both sockets are cut. A gap, not a total — a stream or a download is never cut while it is delivering. Lower it if your app has no long-lived streams. |
HEADERS_TIMEOUT_MS | 20000 | Slowloris guard. |
REQUEST_TIMEOUT_MS | 60000 | Whole-request timeout. |
PROTECTION_MODE | monitor | Server default for bot scoring: off, monitor (score and count, refuse nothing) or enforce (refuse a place in line above the threshold). Per-room protection.mode overrides it. Start on monitor and read the Security panel before you enforce anything. |
PROTECTION_WARN_AT | 40 | Score at which a request is counted as a warning. |
PROTECTION_BLOCK_AT | 70 | Score at which enforce refuses a place in line. Raising it makes the classifier more forgiving; lowering it below ~55 starts putting circumstantial-only clients in reach of a block. |
PROBE_EVERY_MS | 12000 | How often the server probes each room's target itself, on a fixed clock — measurement and autotune do not stop when a dashboard tab closes, and every operator sees the same numbers. Drives the target_down alert and the response-time figure. 0 disables probing; otherwise the minimum is 1000 — a lower value is raised to 1000 with a startup warning. Each distinct URL is probed once per interval however many rooms share it, and the probes are spread over the first 80 % of the interval. |
PROBE_TIMEOUT_MS | 8000 | How long one of those probes waits before the target counts as down. The probe only times the response headers and discards the body. A status under 500 counts as reachable; a 5xx counts as down, along with a network error or a timeout — an app answering every request with 500 is not up, and "something is listening" is not the question you are asking during an incident. |
PROTECTION_ALLOW_IPS | (empty) | Comma-separated addresses that are never scored — your own synthetic monitoring and load tests belong here. |
SHUTDOWN_GRACE_MS | 250 | How long an idle or SSE connection gets before it is cut. A socket that is mid-proxy-response (inline mode) is never cut here — only the hard ceiling below applies to it — so a deploy stops severing a customer's own page mid-transfer. |
SHUTDOWN_DRAIN_MS | 2000 (15000 when any room is inline) | Hard ceiling on the whole stop. A stuck socket must never turn a deploy into a hang. An inline deployment holds proxied responses to a longer floor by default — 2000ms was measured truncating a real proxied download on every deploy — but an operator who sets this explicitly always gets exactly that value, inline or not: with SHUTDOWN_DRAIN_MS=3000 set explicitly on an inline install and a hung origin, the process used to still drain for the full 15000ms, which SIGKILLs mid-drain under any terminationGracePeriodSeconds below that. Set it yourself if 15s is longer than your platform's grace period. |
QM_FORCE_LOCK | (removed) | Removed in step 5 with the data-directory lock. Set, it is ignored with a boot warning, like DATA_DIR. |
QM_IGNORE_CORRUPT_SNAPSHOT | (removed) | Removed in step 5 with snapshot.json. Set, it is ignored with a boot warning, like DATA_DIR. |
HISTORY_RETAIN_DAYS | 90 | How long hourly rows are kept: in Postgres with SIDESTORE=pg, in this process's memory with SIDESTORE=memory. Reports stop showing a row once it is older than this, and older rows are deleted. |
SIDESTORE | memory | Where operators, monitor links, the audit log, timeline notes, metrics buckets and hourly history are kept: memory or pg. memory keeps them in this process only, so a restart empties them (development and tests). pg keeps them in Postgres at DATABASE_URL: the schema is migrated at boot and everything is loaded before the server listens. pg needs SECRET set. With STORE=memory it assumes a single instance (another instance does not see a new monitor link or operator until it restarts); with STORE=valkey every instance re-reads operators, monitor links and timeline notes when another changes them (see RESYNC_MS). file was removed in step 5 and refuses to start by name (see Refusals); anything else refuses to start too. |
STORE | memory | Where queue state lives. memory: diskless, development only. The in-process engine writes nothing to disk, so every restart starts empty (with the demo room), and NODE_ENV=production refuses it. valkey keeps room state in Valkey at VALKEY_URL and journals every change to the events table in Postgres; it requires SIDESTORE=pg, SECRET and VALKEY_URL, and refuses to start naming each one that is missing. What works on Valkey: rooms and their config, join, status, the door check, ticks and admission, sessions and door keys, admin eject, flush and purge, URL targeting (urlPatterns, hit counters, the URL tester), autotune, stats and metrics, scheduled opens (opensAt, branding.opensAt, setOpensAt), the pre-queue (preQueueMaxPerIp) and its random draw at the open; no call answers not_supported_yet. Several instances can share one Valkey and one Postgres behind a load balancer with no sticky sessions: see Several instances (STORE=valkey). The drawn order of each open is kept in the Postgres table open_orders (about 8 bytes per pre-queue entry, one row per 20 000 entries; the open event refers to it). Like events, open_orders is not pruned yet: both grow with every open and every change for the life of the database, so size the Postgres disk for it and watch it. A room whose rebuild finds an open it cannot replay (the open_orders rows for it missing or short, or its order naming a pre-queue entry the journal never enrolled) stays fenced rather than mint tickets nobody owns: every call on it answers recovering, the rebuild is retried, and the console logs recovery rebuild <room>: rebuild <room>: … (naming the seq for missing rows) at most once a minute. To recover it, restore that room's open_orders and events rows from a Postgres backup taken after the open; the next retry rebuilds the room and unfences it. Never delete or edit its events rows to get past it. Changing IPV6_PREFIX_BITS needs the queue drained first, because the per-network keys already in Valkey were written with the old prefix. Any request (visitor or console) whose room-state call fails because Valkey is unreachable, not serving (READONLY, LOADING, MASTERDOWN, CLUSTERDOWN, TRYAGAIN) or does not answer within VALKEY_COMMAND_TIMEOUT_MS answers 503 storage_unavailable with Retry-After: 2; a command whose reply was lost is never re-sent, so a retried join cannot mint a second ticket. A Postgres failure is not this: joins and admissions answer 503 storage_unavailable when their journal commit fails (queries time out after 10 s), and other routes keep their own handling. Any other value refuses to start. |
VALKEY_URL | none | Valkey connection URL, read only with STORE=valkey (required then). rediss:// connects over TLS verified against the system CAs. The db index in the path is the one used. |
VALKEY_PREFIX | qm: | Prefix of every key the server writes in Valkey, and of its control channel (<prefix>ctl), with STORE=valkey. Two deployments on one Valkey need different prefixes. |
VALKEY_COMMAND_TIMEOUT_MS | 2000 | With STORE=valkey, how long one Valkey command may go unanswered before it fails and the request answers 503 storage_unavailable. Integer, 100–60000: a value outside the range is clamped to it and a value that is not a number falls back to 2000, each with a boot warning that names the variable. Raise it only if a slow or distant Valkey trips it under normal load; a higher value holds requests open longer while Valkey is stuck. It does not bound what each Valkey connection sends when it (re)connects (AUTH, SELECT and the INFO ready check): those may take up to 10 s (the connect timeout) before the connection is dropped and retried, so a Valkey that is up but answers slower than this value still connects. |
NODE_ENV | (unset) | production (read trimmed and in any case) is the production rule: the server refuses to start unless STORE=valkey and SIDESTORE=pg are both set, and names each one that is missing (see Refusals). It also lists an unset EDGE_SECRET under hardening (see Cloudflare edge secret) and switches off test-only hooks (QM_TEST_VALKEY_COUNT, which adds a per-request Valkey command count header, QM_TEST_DROP_CTL, which drops pub/sub messages, QM_TEST_ROOM_EVENT_DELAY_MS, which delays room events, QM_TEST_STORAGE_FAIL_FILE, which makes storage read as failing to alerts, QM_TEST_JOURNAL_FAIL_FILE, which makes this instance's journal commits fail, QM_TEST_ALERT_HOLD_FILE, which holds one alert frame past its bound, QM_TEST_HISTORY_SAMPLE_MS, which shortens the history sample period, QM_TEST_NO_TICK, which stops one instance's engine tick, QM_TEST_NOTIFY_TTL_MS, which shortens the notify intent bound, and QM_TEST_NOTIFY_FAIL_FILE, which makes notify deletes fail). |
QM_TEST_VALKEY_COUNT | unset | Test hook, never set in production: 1 with STORE=valkey adds x-qm-valkey-commands (the Valkey commands the request sent) to every response and turns off the Valkey client's auto-pipelining. |
RESYNC_MS | 5000 | With STORE=valkey, how often each instance checks that its copy of the operators and monitor links is current (the side_rev row in Postgres) and asks Valkey about the console sessions on its open admin streams. A change made on another instance normally arrives at once over Valkey pub/sub; this check repairs a lost message within this long. Operator keys, operator console sessions and monitor links are refused with 503 storage_unavailable once the copy has not been confirmed for 3× this (at least 15 s), for example while Postgres is unreachable; ADMIN_KEY still works. The same check compares the rooms revision in Valkey (<prefix>roomsrev, moved by every room create, edit and delete) and, when it moved, rebuilds this instance's copies of the room config that decide whether a request is gated: sentinel protection, the inline host index, and the last known config onBackendDown answers from while Valkey is unreachable. A copy that fails to rebuild is retried on the next check, then less often while it keeps failing (doubling up to 60 s, for example while a room is being recovered); a moved revision and a reconnect to Valkey always rebuild them all at once. Integer, 100–600000. |
QM_TEST_ROOM_EVENT_DELAY_MS | unset | Test hook, never set in production (ignored with NODE_ENV=production): holds back the Postgres commit of every room config event (upsertRoom, schedule change) by this many ms, to test that another instance's open waits for it (Notion 627). A boot warning names it when it is on. |
QM_TEST_DROP_CTL | unset | Test hook, never set in production (ignored with NODE_ENV=production): comma-separated pub/sub message kinds (acct, revoked, rooms, notes) this instance never publishes, to test the RESYNC_MS repair. A boot warning names it when it is on. |
QM_TEST_JOURNAL_FAIL_FILE | unset | Test hook, never set in production (ignored with NODE_ENV=production): with STORE=valkey, while the named file exists, this instance's events-journal commits (and its storage probe) fail, to test one instance's storage fault in a fleet. A boot warning names it when it is on. |
QM_TEST_STORAGE_FAIL_FILE | unset | Test hook, never set in production (ignored with NODE_ENV=production): while the named file exists, this instance reports its storage as failing to the alert frame, to test storage_failing across instances. A boot warning names it when it is on. |
QM_TEST_ALERT_HOLD_FILE | unset | Test hook, never set in production (ignored with NODE_ENV=production): the named file holds a point in an alert frame (evaluate, load or save); the first frame to reach that point renames the file to <file>.held.<n> and waits until that is removed, to test that a frame held past its bound saves and sends nothing. A boot warning names it when it is on. |
QM_TEST_HISTORY_SAMPLE_MS | unset | Test hook, never set in production (ignored with NODE_ENV=production): the history sampler runs every this many ms (at least 100) instead of every 15 s, to test that several instances write one history row per room-hour. A boot warning names it when it is on. |
QM_TEST_NO_TICK | unset | Test hook, never set in production (ignored with NODE_ENV=production): 1 runs no engine tick on this instance, so a test can tell which instance's tick made a change. A boot warning names it when it is on. |
QM_TEST_NOTIFY_TTL_MS | unset | Test hook, never set in production (ignored with NODE_ENV=production): bounds each notify intent to this many ms (at least 100) from when it was filed, instead of 24 h, to test that an old entry goes in a room that keeps writing. A boot warning names it when it is on. |
QM_TEST_NOTIFY_FAIL_FILE | unset | Test hook, never set in production (ignored with NODE_ENV=production): while the named file exists, deleting or clearing notify intents fails as an unreachable Valkey would, to test that eject, purge and room delete answer 503 then. A boot warning names it when it is on. |
LEADER_LEASE_MS | 15000 | With STORE=valkey, the lease that makes one instance the leader: the only one that probes targets (PROBE_EVERY_MS), runs autotune and threshold activation, and delivers alerts to ALERT_WEBHOOK_URL, so each happens once for the fleet and not once per instance. The leader renews the lease (<prefix>leader in Valkey) every third of this and stops acting as leader once half of this has passed without a renewal, before the lease can expire, so two instances never act at once; another instance takes over within about this long of the leader dying, and at once after a clean stop, which hands the lease back. The other instances still evaluate alerts and show the leader's target health and autotune notes, read from Valkey. While Valkey is unreachable no instance is leader: probes, autotune, threshold activation and alert delivery pause until it answers again, and nothing falls back to every instance acting. The one alert still sent then is valkey_unreachable (see VALKEY_ALERT_AFTER_MS); keep an external monitor on /healthz as well. Autotune and threshold-activation writes carry the leader's epoch and are refused once another instance has taken the lease, so a stale leader's late write never lands after the new leader's. leader on /healthz says whether this instance leads. With STORE=memory the one process is always the leader. Integer, 1000–600000. |
VALKEY_ALERT_AFTER_MS | 30000 | With STORE=valkey, an instance that has not reached Valkey for this long sends a valkey_unreachable alert to ALERT_WEBHOOK_URL itself, and a resolved once Valkey answers again. With Valkey down no instance leads and every other alert delivery pauses, so this is the page for that outage. Each instance that cannot reach Valkey sends its own: expect one per instance, since duplicate pages are better than none. The value used is at least one lease renewal (LEADER_LEASE_MS/3) plus the Valkey command timeout (VALKEY_COMMAND_TIMEOUT_MS) plus 1 s, because that is how often a healthy instance hears from Valkey; a lower value is raised to that with a startup warning naming both (at the defaults the floor is 5000 + 2000 + 1000 = 8000 ms). The floor grows with VALKEY_COMMAND_TIMEOUT_MS: at its 60000 maximum it is about 66 s, so a real outage pages that much later. A delivery of this alert the webhook did not take is sent again on each later check, up to 5 times in all. Integer, 1000–3600000. |
QM_INSTANCE_ID | (unset) | With STORE=valkey, this instance's display name: the instance label on every /metrics series, the instance in GET /api/admin/health and the name alerts give it. A label only: every internal identity (leases, staging keys, pub/sub sender, the counters key <prefix>ctr:<name>:<random>) is <name>:<random>, new on each boot, so instances sharing a name, or a rolling deploy that reuses it, stay correct. Unique names make metrics and alerts easier to read, and are required per scrape target when Prometheus scrapes with honor_labels: true (see Metrics and monitoring). Unset, the name is i:<random>, new on each boot. 1 to 64 characters of A-Z a-z 0-9 . _ -; anything else refuses to start. With STORE=memory it only adds the /metrics label. |
DATABASE_URL | none | Postgres connection URL, read only when SIDESTORE=pg (required then). The connection always uses TLS verified against PG_CA_FILE; any sslmode in the URL is ignored. To use a schema other than public, add options=-c search_path=<schema> to the URL; that schema is the one migrated. |
PG_CA_FILE | certs/do-ca.crt | CA certificate the Postgres server is verified against when SIDESTORE=pg. Verification is never turned off; a missing file refuses to start. |
PUBLIC_URL | (derived) | Absolute origin this server is reached at, used to build install commands, connector download URLs, waiting-room links and the copyable commands on /docs when the request's own Host is not the public one (behind a proxy that rewrites it). An https:// value also marks cookies Secure on requests for that host, whatever the socket says. With ADMIN_HOST unset its host is also where the console lives: the console answers there once a room is inline (QM-440), no room may be put inline on it (400 reserved_host), and its console surface is never handed to a room. |
ALERT_WEBHOOK_URL | (unset) | Where alerts are POSTed as JSON. Unset means nothing is delivered — conditions are still evaluated on the server and listed in Diagnostics → Alerting, so the console is the alerting. See Alerting below. |
ALERT_WAIT_SEC | 900 | Estimated wait, in seconds, at or above which a room pages. 0 disables. |
ALERT_DEPTH | 0 | Queue depth at or above which a room pages. 0 (default) disables — depth alone is not an incident; a wait nobody will sit through is. |
ALERT_STALL_SEC | 120 | People waiting and nothing let through for this long pages as critical. 0 disables. |
ALERT_FROZEN_SEC | 300 | People waiting in a room that is paused or set to 0/min for this long pages as critical — the forgotten paused room. 0 disables. |
ALERT_HOLD_SEC | 30 | How long a condition must hold before it fires. Stops a four-second spike from paging anyone. |
ALERT_REPEAT_SEC | 900 | Re-page cadence while a condition is still firing. |
ALERT_EVERY_MS | 15000 | Evaluation cadence. |
GET /api/admin/health reports the effective values, the storage state and any misconfiguration warnings. Check it after every deploy.
The dashboard renders that report behind its Diagnostics button: status, version, uptime, room and visitor counts, SSE streams, rate limits, storage, proxy, memory and every warning — plus two counters worth watching, checks for unknown rooms (a snippet pointing at a room id that does not exist here, i.e. a typo or a deleted room) and blocked return URLs (see returnOrigins below). It is a readable view of the payload rather than the raw JSON; pid, the server clock and the key-configuration block are only in the API response. The header also carries an insecure config badge for as long as any warning stands.
Behind a CDN or load balancer (read this one)
Per-IP limits key on the address this server can prove: the socket peer. Behind a CDN every visitor shares that address, so the per-IP limit silently becomes a global limit — 600 joins/min for your entire site — unless you tell the server who the proxy is:
TRUST_PROXY=1 # hops between the client and this server
TRUST_PROXY_IPS=10.0.0.7 # the address(es) those hops connect FROM
X-Forwarded-Foris believed only from a peer inTRUST_PROXY_IPS. From anyone else — including every peer whenTRUST_PROXYis set on its own — the client address is the socket address, whatever the header says. That one address is what everything uses: the throttle and SSE caps, the pre-queue identity,PROTECTION_ALLOW_IPSand the bot signals, the audit trail'sip, and theX-Forwarded-Foran inline app receives. So one machine forging 900 different values cannot buy 900 visitors' worth of joins, claim an allowlisted address, or reset its own bot score — and a CDN stays collapsed into one bucket until you addTRUST_PROXY_IPS.- With both set, each visitor gets their own bucket and the proxy is exempt from the socket-address bucket.
- Multiple hops (CDN → your LB → here):
TRUST_PROXY=2, and list the address your LB connects from. - A CDN that connects to you directly comes from a published set of ranges, not one address. List them as CIDR, IPv4 and IPv6 alike:
TRUST_PROXY_IPS=173.245.48.0/20,103.21.244.0/22,2400:cb00::/32,.... There is no built-in preset for any CDN: copy the provider's current list, and update it when they do — a range they add and you have not is a proxy whose visitors share one bucket. Every entry is checked at startup, and one that is not an address or a range (173.245.48.0/33, two entries missing their comma, a hostname) stops the server with that entry in the error: a typo in a trust list either untrusts your real proxy or trusts machines you never meant to. TRUST_PROXY_IPSis also what lets a visitor keep their pass while a CDN moves them between edge nodes — see How a pass is bound to a visitor. Forwarded identities are never accepted from a peer that is not on this list.X-Forwarded-ProtoandX-Forwarded-Hostfollow the same rule. From a peer inTRUST_PROXY_IPSthey decide whether the visitor used HTTPS (theSecureflag on every cookie, the scheme an inline app is told in its ownX-Forwarded-Proto/-Port, the URL your room's rules are matched against) and which host the absolute URLs this server hands out name (the connector manifest, the WordPress update feed, the rendered docs, whenPUBLIC_URLis unset). From anyone else — including every peer whenTRUST_PROXYis set on its own — both are ignored: the scheme is the socket's and the host isHost. A TLS-terminating proxy you have not listed therefore gets cookies withoutSecure; list it, or setCOOKIE_SECURE=1.
The same rule applies to the SSE per-IP cap.
A proxy you have not listed is named, not silent. When forwarded-address headers (X-Forwarded-For, Forwarded, CF-Connecting-IP, X-Real-IP, True-Client-IP) keep arriving from a peer that is not in TRUST_PROXY_IPS — 10 or more such requests from the same address within a minute, whether TRUST_PROXY is set or not — the server treats that peer as an undeclared proxy. Behind it, every visitor shares the proxy's one address, so queueMaxPerIp (16 places per address by default) and preQueueMaxPerIp cap the whole room at that many visitors, and the join and check budgets become one budget for the site. The server then:
- adds a warning to
warningson/api/admin/healthnaming the address and the exactTRUST_PROXY_IPS=…line to set (your existing entries plus it), setsproxy.ok: false,proxy.forwardedFromUndeclaredPeer: trueand lists the address inproxy.undeclaredPeers— the console shows it in the header badge; - logs one line, and at most one every 10 minutes while it keeps happening.
One request with a forged header does not trip it. A script sending a stream of them can, so read the address before you copy it: add it only if it is your load balancer or CDN. An address you do not recognise is a client forging the header, and listing it would let that client choose its own address. Changing TRUST_PROXY_IPS is the fix; the warning never changes whose header is believed.
Setting only one of the two is the failure mode to watch for. They are one setting in two variables, so the server refuses to let a half-configuration be silent: it prints the warning at startup, returns it in warnings on /api/admin/health with proxy.ok: false, and the dashboard header turns it into the insecure config badge (the Proxy row reads … · INCOMPLETE). A measured surge through a half-configured proxy loses roughly 94% of its joins to HTTP 429, and the behaviour is otherwise indistinguishable from the queue working correctly.
Cloudflare edge secret
In production this server is reached only through our Cloudflare zone, but its own address (the DO App Platform URL, the container) answers too. A request sent there directly skips every Cloudflare rule, cache and rate limit, and can put any value it likes in CF-Connecting-IP. EDGE_SECRET closes that door: Cloudflare adds a secret header to every request it forwards, and this server refuses any request without it.
Set it up (in this order, so nothing is refused while you do it):
- Generate a long random value, e.g.
openssl rand -hex 32. Entries shorter than 16 characters stop the server at startup: a guessable secret is worse than none, because it looks like protection. - In the Cloudflare dashboard for the zone: Rules → Transform Rules → Modify Request Header → Create rule, with Set static header
X-QM-Edge= the value. Deploy the rule. Match only this app's hostname(s) (e.g. Hostname equalsqm.weekday100.com), never All incoming requests. A zone-wide rule adds the secret to every request Cloudflare sends to every origin in the zone: your other apps, and any third-party service a hostname is CNAMEd to (a help desk, a shop platform, a status page). Any of them can log it, and anyone holding it can call this app's own address directly and be treated as Cloudflare, with aCF-Connecting-IPof their choosing. - Set
EDGE_SECRETto the same value on the app and deploy.
What it does once set:
- Every HTTP request (API, console and admin API, SSE, static files, the waiting page, and every inline gate and proxied path) and every WebSocket upgrade must carry
X-QM-Edgeequal to one listed value. Otherwise the answer is403with{"error":"edge_required","code":"edge_required"}, or a minimal403page for a browser navigation. Neither names the header. An upgrade is refused before any socket to the origin is opened. /healthzis the one exemption, because DO App Platform's health check reaches the container directly. It coversGET/HEADof exactly/healthz(a query string is fine;/HEALTHZ,//healthzor an encoded variant are not), and only on a host that fronts no inline room (on an inline host/healthzis the app's own path). The list is one constant,EDGE_EXEMPT, inserver.js.- The client address is
cf-connecting-ip, on every request that passed the check: throttle keys, SSE caps, the pre-queue identity, bot signals, the audit trail'sip, and the rightmostX-Forwarded-Forhop an inline app receives. A forgedX-Forwarded-Formoves nothing. Ifcf-connecting-ipis missing or not an address, the socket address is used and counted (edge.clientIpFallbackTotalon/api/admin/health,qm_edge_client_ip_fallback_total); behind Cloudflare that stays at 0. TRUST_PROXYandTRUST_PROXY_IPSare ignored (see Behind a CDN or load balancer). If they are still set, a warning says so.X-Forwarded-Protois believed on a request that passed the check;X-Forwarded-Hostis not, since Cloudflare passes a client's own through.- The value is never logged, audited or put in a metric, and the header is removed before any request or upgrade is forwarded to a room's
proxyOrigin. - Refusals are counted per instance:
qm_edge_refusedon/metrics(with theinstancelabel onSTORE=valkey) andedge.refusedTotalon/api/admin/health. A steady rate means something is calling the app's own address; a jump to every request means the Cloudflare rule andEDGE_SECRETno longer match.
Rotate without refusing anyone: add the new value to the list (EDGE_SECRET=old,new) and deploy; switch the Cloudflare rule to the new value; then remove the old one (EDGE_SECRET=new) and deploy again.
Unset, nothing changes from earlier releases. With STORE=valkey or NODE_ENV=production the server prints EDGE_SECRET is not set: requests that bypass Cloudflare are accepted at startup and lists it under hardening on /api/admin/health.
Two deployments: snippet and inline
There are two ways to put this queue in front of a site, and they fail differently. Choose on that, not on which is easier to install.
Snippet (default). The site loads a small script; the script asks this server whether the visitor may proceed and sends them to the waiting page if not. The queue sits beside the traffic. If this server dies, the script's check fails and the pages are served unqueued — the site stays up and the protection is gone.
Inline (proxy). DNS for the site points at this server, and every request for that host arrives here. Admitted visitors are forwarded to the app; everyone else gets the waiting page at the URL they asked for. The queue sits on the traffic. If this server dies, the site is unreachable — which is why the CDN failover page in When the queue itself is down is not optional for an inline deployment.
The three shapes side by side. With the snippet, pages come from your site and the tag in the browser asks the queue; if the queue is down, pages load unqueued. The edge connector is a snippet-mode deployment too, but a Cloudflare Worker asks the queue before your site sees the request; if the queue is down, it forwards the request to your site unqueued and marks the response X-QM-Failover. Inline, every request passes through the queue; if the queue is down, the site is unreachable unless the CDN serves a failover page.
Turn it on in the console — Edit (or + New room) → tick Serve this room inline and fill in Internal address of the site behind the queue. Clearing the tick returns the room to snippet mode, behind a confirmation that says so. Or over the API:
curl -sX POST "$BASE/api/admin/rooms" -H "Authorization: Bearer $ADMIN_KEY" \
-H 'Content-Type: application/json' -d '{
"id": "shop",
"targetUrl": "https://shop.example",
"proxyOrigin": "http://10.0.0.5:3000"
}'
targetUrlis the public address — the host visitors' browsers send, and what this server matches an incoming request against.proxyOriginis the internal one, reachable only from this server. Bare origin: scheme, host, port, nothing else. Pointing it back at the public host is refused (proxy_origin_loop) — that is an infinite loop through your CDN.- A private or internal
proxyOriginis refused unlessPROXY_ALLOW_PRIVATE=1— including the10.0.0.5in the example above. Loopback (127.0.0.0/8,::1),0.0.0.0/8, RFC1918, CGNAT100.64.0.0/10, link-local (169.254.0.0/16, which holds the169.254.169.254cloud metadata endpoint;fe80::/10), unique-localfc00::/7(holdsfd00:ec2::254),::and the IPv4-mapped forms of all of them. Whoever can edit a room could otherwise make this server fetch its own network. The check is on the address actually connected to, after DNS, on every new connection, so a hostname that starts resolving to127.0.0.1later (DNS rebinding) is refused at that point; a name with any private answer is refused outright. The server's health probe of the origin is held to the same rule. If your app really is on the LAN, setPROXY_ALLOW_PRIVATE=1. - No
proxyOriginmeans snippet mode. Nothing about an existing install changes because this feature exists. ratePerMinuteacceptsrateas an alias — that is whatGET /api/admin/roomscalls the same field, so a script that reads a room and POSTs it straight back does not have to rename anything first. Sending both is fine if they agree; disagreeing is400 conflicting_field. Any field name outside the documented set is refused (400 unknown_field), not silently dropped — a typo used to look exactly like a change that took.
What changes once a room is inline:
| Snippet | Inline | |
|---|---|---|
| Queue's own routes | /api/*, /w/<room> | under /__qm (/__qm/api/join, …) |
| Monitoring routes | /healthz, /api/version, /metrics | the same, and also under /__qm. Which one a probe gets depends on the address it uses, not on the path: reached on the customer's hostname these are the customer's paths and the visitor gate answers them 503 for ever. Probe the process directly, or use /__qm/healthz and metrics_path: /__qm/metrics. |
| Operator console | / on the queue's own host | /__qm/ on the customer's host — the same page, addressing the same prefixed routes. Bookmark that address, not /: / on that host is the shop. |
| Console sign-in | qm_console, HttpOnly, Path=/api/admin | qm_console, HttpOnly, Path=/__qm/api/admin: the customer's scripts cannot read it, and the key itself is never stored in the browser |
| Door key | in the URL (?sess=), readable by the script | qm_dk_<roomId>, HttpOnly cookie only: nothing in a query string, even under /__qm. Scoped to the room, so clearing one room's key leaves another's alone. |
Ticket token qm_<room> | readable cookie + localStorage + ?token= — this host is ours and the page's own script is the reader | HttpOnly cookie only: nothing in localStorage, nothing in a query string, even under /__qm |
| URL targeting | matches whatever the snippet reported | matches the real request URL |
returnOrigins / tagOrigins / CORS | load-bearing | unused: nothing is cross-origin |
| This server down | site up, unqueued | site down — needs CDN failover |
The prefix is reserved, not decorative: the campaign app has its own /api/*, and the queue must never shadow it. Anything on a proxied host that is not under /__qm belongs to the app.
Room ids starting with dk_ are reserved the same way (QM-533). Room dk_x's ticket cookie would be qm_dk_x, the exact name of room x's door-key cookie: the two overwrite each other in the browser, and the per-host cookie cap files the ticket under room x and expires it with x's cookies. Creating such a room answers 400 invalid_room_id. A dk_ room made before this rule keeps working and can still be edited, but it keeps the collision: recreate it under another id and move the tag over.
The rest of the reserved ids follow from the same rule (546): a room's ticket is kept under qm_<room>, as a cookie and as a localStorage key on the queue host, so no room id may make that name one the queue already uses. Refused on create with 400 invalid_room_id, case-sensitive:
- ids starting with
meta_,notify_oremail_(the waiting page's per-roomqm_meta_<room>,qm_notify_<room>,qm_email_<room>); - exactly
console,theme,lang,admin_key,autotune,autotune_cfg,wait_samples,hidden_series,selected_room,alerts,offline_attempts,origin_down_attempts.consolewould be the console sign-in cookieqm_console; the others are the storage keys of the console and the waiting, offline and origin-down pages on the same origin. A room namedadmin_keywould write its ticket where the console keeps its key.
When the server looks for a ticket among a request's cookies (/api/status, /events), qm_console is never read as one, and cookies under a reserved name are tried only after every other qm_ cookie, so eight rooms' door keys cannot crowd out the real ticket.
The ticket-token row is the same reasoning as the prefix, applied to credentials. Inline, our cookies and our localStorage land on the customer's origin, shared with every analytics tag, chat widget and ad script the site ships — so the token is kept somewhere none of them can reach, and the waiting page persists nothing. It does not need to: /api/join returns the token in its response body on every load, and /api/status and /events authenticate from the cookie when no token= is sent.
Forwarding preserves the visitor's Host and adds X-Forwarded-Host, -Proto, -Port and -For, so a Next.js app behind this builds its own absolute URLs correctly. All four describe the VISITOR's hop, never proxyOrigin: -Port is the port named in their Host, or 443/80 to match -Proto when they did not name one. WebSocket upgrades carry the same four.
The rightmost X-Forwarded-For entry is always the client address this server resolved, so an app that trusts one hop (this server) reads the truth. From a proxy in TRUST_PROXY_IPS the chain it sent is kept to the left of that entry, and its Forwarded, X-Real-IP, CF-Connecting-IP and True-Client-IP pass through. From any other peer those headers are the caller's own words: all of them are dropped and X-Forwarded-For is the socket address alone. TRUST_PROXY / TRUST_PROXY_IPS still apply: they are how this server learns the visitor's address and scheme from your CDN. The -Proto (and the default -Port) sent onward is the scheme your proxy reported only when that proxy is in TRUST_PROXY_IPS; from any other peer it is the scheme of the connection this server actually received, whatever the caller claimed.
When the app itself is the problem:
| What happened | Visitor sees | Also |
|---|---|---|
| App refused the connection / unreachable | 502 + X-QM-Origin: origin_unreachable — origin-down.html on a navigation, JSON to an XHR | counted per room; one log line a minute naming the room |
proxyOrigin is a private/internal address and PROXY_ALLOW_PRIVATE is not 1 | the same 502 + X-QM-Origin: origin_unreachable (a WebSocket gets 502 too); nothing is sent to the address | counted per room under phase blocked; the room's originErrors.reason and a log line (once a minute per room) name the refused address; the probe reads down with targetHealth.blocked saying why |
| App accepted, then never answered | 504 + X-QM-Origin: origin_timeout after PROXY_TIMEOUT_MS | same |
| App answered, then stalled mid-body | the transfer is cut — headers are already gone, so there is no status code left to send and no honest way to finish the page | counted per room under phase body_stall after PROXY_IDLE_TIMEOUT_MS; a visitor who stopped reading is counted as client_reset instead |
| Queued visitor's XHR or API call | 503 queued + Retry-After + X-QM-Queued: 1 | anything a browser renders gets the waiting page instead — see below |
This process at MAX_CONNECTIONS | 503 + Retry-After (no X-QM-Queued) — the offline page, not the waiting page | qm_connections_shed_total; /healthz answers its own 503 with ready:false too — see CONNECTION_RESERVE above |
Past MAX_CONNECTIONS a request is deliberately shed, not queued — there is no "try again in a moment and you'll be let in", the process genuinely has no room. It carries no X-QM-Queued, on purpose: a CDN failover rule that keys on 503 is meant to fire here and show its own outage page, the same as any other real outage. Without CONNECTION_RESERVE, this state was unreachable to answer at all — Node's own server.maxConnections destroyed the socket first, silently, and a hung origin pinning two sockets per stuck visitor for the whole PROXY_TIMEOUT_MS made that ceiling reachable in well under a minute during exactly the failure this product exists for.
X-QM-Queued is what stops a CDN failover rule keyed on 503 from replacing a healthy queue with the outage page, and it is also the signal a customer's app can key on to reload a page whose pass lapsed mid-session instead of failing silently. X-QM-Origin says the site behind the queue is the thing that broke.
Which of the two a request gets. The test is whether a browser will show the answer to a person, not whether the method is GET: Sec-Fetch-Mode: navigate (any method), a browsing Sec-Fetch-Dest, or — for clients that send no Fetch Metadata — Accept: text/html. A <form method="post"> submitted after the pass lapsed is a navigation, so it gets the waiting page; a fetch() asking for JSON gets JSON.
Recovering a page whose pass lapsed. The queued envelope names the room and the engine's own reason (expired, ejected, stale_generation, waiting, at_capacity, …), and Access-Control-Expose-Headers is set so the header is readable even when the call is cross-origin:
{ "error": "…", "code": "queued", "roomId": "shop", "reason": "ejected", "reload": true }
Inline there is no separate waiting-room address — the waiting page is served at the URL the visitor asked for — so re-requesting the current URL is the whole recovery:
const r = await fetch('/api/cart');
if (r.status === 503 && r.headers.get('X-QM-Queued')) { location.reload(); return; }
Without that line the site keeps rendering whatever its code does with a failed request. A visitor who is simply put back in line is told so by the waiting page itself, which names the cause rather than presenting itself as a fresh arrival.
On a hostname this server fronts no room for, the operator console is a bare 404. See ADMIN_HOST above: inline, the console answers on an address literal, localhost and the PUBLIC_URL host, or with ADMIN_HOST set on the names you list, and nowhere else. An inline room whose targetUrl is an address literal (the usual local-development setup) gets that host's pages, but with ADMIN_HOST unset never its /api/admin/*: that address is the console's too, so an admin call sent there is answered by the queue, not forwarded (QM-439). With ADMIN_HOST set that includes the customer's hostname: its /__qm/ console, /__qm/api/admin/* and /__qm/public/admin.html answer 404, and only the visitor paths under /__qm/ remain. A hostname that fronts no room is withheld the console only, ADMIN_HOST set or not (QM-428) — the snippet, /healthz, the waiting page and the visitor API still answer there, since that is the address the customer's script tag names.
This is easy to walk into (QM-454): with neither ADMIN_HOST nor PUBLIC_URL set, saving the first inline room from a console on a hostname (say qm.weekday100.com) makes that console, and every monitor link on that host, answer 404 from the next request on. The save still succeeds, and its answer carries a warning (warnings[], code: "console_host_withheld", naming the host) that the console shows after saving. While any inline room exists with both unset, /api/admin/health lists the same reason under configWarnings. The fix is to set ADMIN_HOST (or PUBLIC_URL) to that hostname and restart.
Switching the last inline room on a host back to snippet mode takes the queue's own /__qm/* surface off that host with it — including the console you are probably reading this from, if you administer the queue through the customer's domain. The console warns before that save and moves you to PUBLIC_URL afterwards; a stale /__qm/… address answers 404 {"code":"not_inline_host"} naming where the console went, rather than the generic not-found. Set PUBLIC_URL so it can name it. (With ADMIN_HOST set the console never lived on the customer's domain, so there it is a bare 404.)
None of these make this instance unhealthy: a customer origin going dark is not a reason to take the pod out of rotation, so /healthz stays 200. See Readiness. The counts are on /api/admin/health under inline — gatedTotal, bypassedTotal, originErrorsTotal and a per-room/per-phase originErrors breakdown. The whole block is absent on an install with no proxied room.
What skips the gate
Static subresources are forwarded without a queue check. This is about cost, not politeness: one page of a modern app is 30–80 subresource requests, and gating each one would run a queue check, a bot score and a cookie write per file — turning the queue into the bottleneck it was installed to prevent, while protecting nothing. Downloading a stylesheet takes nobody's place in line.
A request skips the gate only when all four hold:
- the method is
GETorHEAD; - the path ends in a static-media extension —
js mjs cjs css map, images, fonts, audio/video. Deliberately not.json,.txtor.xml, which are as often an API's answer as a file on disk; - the path as the app will receive it is a plain one: no
;, no percent-escape (%3b,%2f,%2e, even%20), no backslash, and nothing the queue's URL parser had to rewrite (a..segment). The request line is forwarded to your app verbatim, and frameworks strip or decode before they route —/api/stock;.jsis/api/stockto Tomcat and Spring — so a path like that is gated like any page. Its cost is one queue check; nobody loses a place; Sec-Fetch-Dest, when the browser sends it, is not a browsing context (document,iframe,object,embed).
The app's own API is never bypassed. POST /api/checkout is where the scarce thing happens; letting it through would mean anyone willing to skip the browser buys the item without ever waiting. Sec-Fetch-Dest is forgeable, but forging it wins nothing except a file from the list above.
The edge case to know: an app that serves capacity-heavy dynamic content from a URL ending in .js — or routes a trailing segment like /api/stock/x.css to a protected handler, or treats .js as a format suffix the way Rails does — would skip the gate for a subresource request. There is no switch for this yet (restricting the bypass to declared asset prefixes is an open product decision) — if your app does that, say so.
Streaming and WebSockets
Nothing is buffered in either direction. A streamed response starts reaching the visitor as the app produces it — a Next.js App Router shell paints at the same moment it would with no queue in front — and an upload streams through without being held in this process.
WebSocket and other connection upgrades are forwarded, and they are gated like a page, not like an asset: a socket that stays open for the length of a visit is not a subresource, and letting one through unqueued would be a way to sit inside the app without ever taking a place in line. A visitor who is still waiting gets 503 on the upgrade — no waiting page, because there is no page there; the app's own reconnect logic is left in charge of trying again. Once the handshake is through, the two sockets are joined and nothing in this process looks at the bytes again.
Upgrades on /__qm/* are refused with 501: the queue's own live updates are Server-Sent Events, which are ordinary HTTP.
Rate limits
Every throttled route answers with X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset (seconds), and adds Retry-After on 429, so a client can pace itself instead of discovering the limit by being cut off.
The waiting page is throttled too
The page is 141 KB and is the largest thing an anonymous stranger can ask this server for. It is rendered and compressed once per room and served from memory (private, no-cache + ETag, so a visitor's reload during a long wait costs a 304 and not another page), but a client that simply refuses Accept-Encoding still takes the full body every time. Measured on one laptop at 120 connections:
| bytes pulled in 10 s | other requests on the process | |
|---|---|---|
| before | 2,212 MB | 21x slower |
| after | 29 MB | 1.3x slower |
WAITING_LIMIT_PER_MIN (default 120) is roughly two page loads a second from one client — far above any real visitor. A refused navigation gets a 941-byte page that reloads itself after Retry-After; a refused XHR gets plain text. The queue is untouched: the visitor keeps their session and their place, and the status endpoint keeps answering.
It stands down when it cannot tell visitors apart. With TRUST_PROXY unset, every visitor behind a proxy looks like one client (the proxy's own address), and a per-address limit there would take the customer's whole site down — far worse than the flood. So once a peer is recognised as a proxy — forwarded-address headers from the same address 10 times within a minute, or WAITING_LIMIT_PER_MIN times if that is lower, so it is recognised before its budget runs out — the guard stands down for that address only. Every other address is still throttled. /api/admin/health then reports waitingPage.enforced: false while it stands down for any address, and the console's Proxy row says an undeclared proxy is in front. Fix the proxy settings and the guard comes back for it on restart.
One forged X-Forwarded-For does not switch the guard off: it used to, for every address and for the life of the process. A client that sends a steady stream of forged headers can still exempt its own address this way — with TRUST_PROXY unset the server cannot tell it from a proxy — but never anybody else's. Declaring your proxy (TRUST_PROXY + TRUST_PROXY_IPS) removes that too: the guard then never stands down, because it can charge each visitor. The list of exempt addresses is capped at 1024 and dropped when full; a real proxy is recognised again within a few requests.
So are the console and the docs
The console page is about half a megabyte, and /docs/* is rendered from markdown; both used to answer any client as often as it asked — on an inline host that includes /__qm/, /__qm/docs and /__qm/progress on the customer's own domain. /, /progress and /docs* now take the same guard as the waiting page: WAITING_LIMIT_PER_MIN per address, the same small self-reloading refusal, the same stand-down behind an undeclared proxy. They are charged to a separate bucket, so an operator reading the guides never spends a visitor's waiting-page allowance, and the admin API is not touched. Each guide is rendered once per file version and served from memory after that.
What churn cannot do
Each budget is a map of per-address buckets. It holds at most RATE_LIMIT_MAX_KEYS (default 200,000) and evicts the least recently used address past that. It used to clear itself instead, which gave every throttled client a fresh budget to anyone able to mint 200,000 keys — a Public API key holder naming a new ?ip= per vouched check, or a client rotating IPv6 addresses. A client that is being limited is, by definition, the most recent key in the map, so eviction never reaches it.
Vouched /api/check calls are charged to the visitor they vouch for (?ip=), by design: one connector speaks for thousands of visitors from a handful of addresses. Those buckets live in a map of their own, so however many visitors a key holder names, it cannot evict a direct caller's bucket.
Two keys, two blast radii
PUBLIC_API_KEY | ADMIN_KEY | |
|---|---|---|
GET /api/check with ?ip=&agent= | yes | yes |
/api/admin/* (rooms, rate, flush, eject, schedule, visitors, metrics, health) | no (401) | yes |
GET /metrics (Prometheus exposition) | no (401) | yes |
Set both. Distribute only the first.
Wrong keys are rationed. Each client address may fail admin authentication ADMIN_AUTH_FAIL_PER_MIN times a minute (default 20). Past that, every credential it presents — right or wrong, Bearer or SSE ticket — is answered 429 too_many_auth_failures with Retry-After, before it is checked, so a lockout never confirms a guess. A correct key, a valid monitor link and a valid key used on a route it is not scoped for (403) spend nothing; a request that presents no credential at all is not counted either. The public-key door (/api/check?ip=&agent=, /api/verify) shares the same allowance because it accepts ADMIN_KEY too: a wrong key there is charged, and a locked-out address's key is treated as absent (401, and connectors fail open). Behind a CDN the address is the forwarded client only when TRUST_PROXY_IPS names the CDN; otherwise every visitor shares the CDN's socket address and one guesser can lock the console out for a minute — another reason to set it.
No credential travels in a query string. Every admin route takes Authorization: Bearer and nothing else, and so does the Public API key on /api/check and /api/verify. ?key= used to be accepted there; it is now refused with 401 {"code":"credential_in_query"} on a vouched check, counted in qm_refused_key_checks_total like any refused key, and reported by /api/verify as "Public API key sent in the URL". Every shipped connector already sends the header. A hand-written integration that still sends ?key= fails open (its pages go unqueued) until it is moved to the header — the counter and the stderr line are how you find it. The one exception is the event stream, where EventSource cannot set a header: the console mints a single-use ticket (POST /api/admin/sse-ticket, over Bearer, scoped to whatever credential asked for it) and spends it on GET /api/admin/events?ticket=…. A ticket lives 30 seconds and is destroyed on first use, so one landing in a log is worth nothing by the time anybody reads it. Operator tooling that opens the stream must mint a ticket per connection.
The console keeps a session cookie, never the key (QM-349). The console used to keep ADMIN_KEY or an operator key in localStorage, and on an inline install the console is served from /__qm/ on the customer's origin, so any script that site ships could read it. Signing in now sends the key once, in the body of POST /api/admin/session ({"key":"…"}, header X-QM-CSRF: 1), and the answer sets qm_console: HttpOnly, SameSite=Strict, Path=/api/admin (or /__qm/api/admin on an inline host), Secure whenever the request arrived over TLS, lasting 12 hours. Its value is a signed payload naming who signed in and a session id, never the key. DELETE /api/admin/session signs out: it clears the cookie and revokes that session server-side (the cookie was already cleared in that browser). The revocation is remembered in memory with STORE=memory, so a restart forgets it, and in Valkey with STORE=valkey, so a Valkey data loss forgets it: the instances that received the sign-out's hint still refuse the session until they restart, an instance booted after the loss does not. Revoking an operator ends their sessions on the next request; changing ADMIN_KEY ends every session the old key opened. A wrong key at sign-in spends the ADMIN_AUTH_FAIL_PER_MIN allowance exactly like a wrong Bearer, and a locked-out address is refused there too. A 401 clears qm_console only when that cookie was the credential that failed: a request carrying Authorization: Bearer is judged on the Bearer alone, so a dead monitor link opened in the owner's browser does not sign the owner out (QM-460).
Revocation reaches open live streams too (QM-442). The console's live stats stream (/api/admin/events) is one long request, so checking the credential only when a request arrives was not enough. Each stream remembers what opened it: the Bearer (ADMIN_KEY, an operator key or a monitor link) or the console session cookie, carried through the SSE ticket. When a monitor link is revoked, Revoke all runs, an operator is revoked, or a console signs out, every stream opened with that credential ends at once. The admin tick checks all of them again every 2 s, which also catches expiry. The last thing the stream sends is event: revoked ({"error":"unauthorized"}), and the console shows its signed-out screen or the dead-link screen instead of reconnecting. An unspent SSE ticket whose credential was revoked is refused with 401. Removing SECRET_PREVIOUS takes a restart, and the restart closes every stream anyway.
qm_console never crosses the inline proxy (QM-427). A browser scopes a cookie by host, not port, so wherever the console and a room share a host (the console under /__qm/ on the customer's name, or both on one address) the session cookie would otherwise ride proxied requests to the app, and the app's own Set-Cookie: qm_console=… would replace the operator's session. The proxy removes qm_console from the Cookie header of every proxied request and WebSocket handshake, and drops any Set-Cookie line from the app that names it. Every other cookie passes both ways unchanged, the visitor's queue cookies included.
A change made with the cookie (any method but GET) must carry the X-QM-CSRF header, and a browser that sends Sec-Fetch-Site must say same-origin; otherwise 403 {"code":"csrf_required"}. No form can send a custom header, and a cross-origin script cannot either without a preflight this server never grants, so a page elsewhere cannot drive the console's session. Bearer callers are unchanged: a request with an Authorization header is judged on that header alone, the cookie is ignored, and no CSRF header is needed. That is also why a #monitor= link opened in an operator's browser stays read-only. A console that still has a key in localStorage from an older version moves it on first load: the key is taken out of storage, used once to sign in, and dropped.
There is a third, much smaller credential: the dashboard's Share monitor link carries a scoped viewer token in the URL fragment (#monitor=…), which no browser ever sends to a server — so it appears in no access log, no CDN log and no Referer, including the customer's own logs on an inline install. It is the read grant, not "the stats stream and nothing else" — the holder can read the live stats SSE, the room list, the metrics, the security and reports panels, GET /metrics, GET /api/admin/alerts (without a blocked probe's refusal text, the webhook address or delivery errors, each of which can name an internal address; its Send test alert button is disabled), and GET /api/admin/health (which names the process id, memory and whether the proxy settings are coherent, but not the configuration warnings, the hardening list, the proxy peer addresses, whether EDGE_SECRET is enforced or PUBLIC_API_KEY is set, Valkey's maxmemory-policy (backend.maxmemoryPolicy is left out and its incident says only that the policy is not noeviction), or the text of a storage or warehouse error, which can name a path or a host: a default ADMIN_KEY, a short SECRET or a missing EDGE_SECRET is not something to tell whoever the link was forwarded to; refused qm_return URLs appear as a count per room, never the URL, for any credential). Four things it is deliberately not shown, because the link is meant to leave the team: a room's proxyOrigin (your internal address), its origin allowlists and URL rules, the per-visitor drill-down, and the audit trail — which names every operator on your team, what each of them changed, and the IP each of them did it from. All four answer 403 insufficient_role to a monitor token. A named viewer operator key sees all of them — that is a person on your team. The share link can change nothing: any write answers 403 insufficient_role, and it can never manage operators.
And a fourth, which is the one your team should actually be using day to day: named operator keys.
Operators, roles and the audit trail
ADMIN_KEY is a shared secret with no name attached, so "who paused the room" has no answer. Mint one key per person instead, from the dashboard's Access panel or the API:
curl -sX POST https://qm.weekday100.com/api/admin/operators \
-H "Authorization: Bearer $ADMIN_KEY" \
-d '{"name":"Night ops","role":"operator"}'
# → {"operator":{"id":"535dddaa","name":"Night ops","role":"operator",...},
# "key":"qmo_535dddaa_...","shownOnce":true}
The key is returned exactly once. Only an HMAC verifier is stored, so a lost key is replaced, never recovered.
| Role | Can |
|---|---|
owner | everything ADMIN_KEY can, including minting and revoking operator keys, deleting rooms, and creating or repointing inline rooms |
operator | run the queue: rates, state, flush, eject, schedule, room settings (branding, targeting, origins, a snippet room's target URL), empty a queue |
viewer | read only: stats, metrics, security, reports, audit |
A key with insufficient role gets 403 {"code":"insufficient_role"} — not a 401, so the holder can tell "wrong key" from "not your job".
Deleting a room needs owner (or ADMIN_KEY). It is the one irreversible action here: everybody standing in the line is discarded and the room, its configuration and its schedule are gone. Emptying a queue (POST /api/admin/rooms/<id>/purge) is destructive too, but the room survives, so it stays with operator. The console hides Delete entirely for a key that may not use it, rather than letting the 403 arrive after the confirm.
Inline deployment needs owner (or ADMIN_KEY). An inline room's targetUrl names the public host this server answers for and its proxyOrigin the address that host's traffic is forwarded to, so together they decide whose traffic this server captures and where it goes. Creating an inline room, changing an inline room's targetUrl or proxyOrigin, clearing its proxyOrigin, or giving a snippet room one (which makes it inline) is refused to operator with 403 {"code":"inline_requires_owner"}, and the refusal is audited (room.create / room.update, ok: false). The refusal on an update says every other setting of the room is still the operator's to change; on a create, where there is no room yet, it says the room can still be created as a snippet room (no proxyOrigin) (QM-458). Only a change counts: a save that sends the room's current targetUrl and proxyOrigin back unchanged (the whole-body save a script or the console makes) goes through, and every other setting of an inline room stays with operator. Each accepted change is audited as room.inline, with the old and new value of each field. The console disables the inline toggle and the internal address for an operator key, and the target URL too on an inline room, and says why under the toggle.
# revoke immediately; the key stops working on the next request
curl -sX DELETE "https://qm.weekday100.com/api/admin/operators/535dddaa" \
-H "Authorization: Bearer $ADMIN_KEY"
Revoked operators are disabled, not deleted, so their past audit entries still resolve to a name. ?purge=1 is the escape hatch for one added by mistake.
Revoking an operator also closes every monitor link they shared, in the same write, so it survives a restart exactly as revoking one link does. A share link that outlived its author would be the one credential the revocation missed — and the most widely forwarded one. Links shared by other operators and by ADMIN_KEY stay open. The response carries linksRevoked, the single operator.revoke (or operator.purge) audit entry names the count (Night ops; closed 2 monitor links), and the console's confirm dialog says how many will close before you click. GET /api/admin/operators reports each operator's liveLinks.
A link is matched to its author by operator id, which is recorded when the link is minted (createdById on the link), never by name: names are free text, two operators can share one, and an operator can even be called ADMIN_KEY. Links minted before this release carry only the name and cannot be attributed, so revoking an operator leaves them open. If one of those may be in the wrong hands, revoke it from the monitor-link list, or use Revoke all; they also expire on their own after MONITOR_TOKEN_TTL_SEC (seven days by default), after which every live link is attributable.
Every state-changing admin action is appended to the audit trail (Postgres with SIDESTORE=pg) with the actor, the target, the time and the caller's address — including the server's own actions, which appear under a system actor (autotune rate changes, threshold activation flips). Field names are recorded for a room edit, not the values: the values are readable from the room, and an audit line outlives the config it describes.
Restarts are in the trail too. server.start is written when the process begins listening, and server.stop on a clean shutdown. The start line says whether the previous run stopped cleanly — the one fact that cannot be reconstructed afterwards, and the normal case on Windows, where every stop is a kill (see Stopping the server). GET /api/admin/health carries an absolute startedAt beside uptimeSec, /metrics exposes qm_process_start_time_seconds (alert on a change), and an open console raises a toast the moment the value moves under it.
curl -s "https://qm.weekday100.com/api/admin/audit?roomId=checkout&format=csv" \
-H "Authorization: Bearer $ADMIN_KEY"
There is no SSO. No SAML, no OIDC, no SCIM, no MFA. /api/admin/health reports access.sso: false so a procurement review gets a straight answer rather than a discovery.
Bot protection
Traffic asking for a place in line is scored from what this server can actually observe — a declared automation client, missing browser headers, a forged X-Forwarded-For, machine-regular arrival timing, one address holding an implausible spread of tickets — and each signal adds or subtracts points on a 0–100 scale.
PROTECTION_MODE=monitor(the default) scores and counts and refuses nothing. Run here first. The dashboard's Security panel shows the totals, which signals are firing and the recent decisions.PROTECTION_MODE=enforcerefuses a place in line abovePROTECTION_BLOCK_AT, with403 {"code":"blocked_by_protection"}and anX-QM-Protection: block score=NNheader.- Per room,
protection: {"mode":"enforce"}overrides the server default;nullinherits it.
Two properties worth knowing before you enforce:
- Enforcement refuses a ticket, never entry to your site.
/api/checkstays fail-open for everyone. The cost of a false positive is a bot-looking visitor losing their place, not you losing a customer. - A shared address cannot block itself. The scoring is calibrated so that no combination of signals a corporate NAT or CGNAT pool can produce on its own reaches the block threshold — only a client that declares itself automation gets there without help. The direct consequence is that a botnet on residential addresses driving real browsers is not caught by this; fairness against that is structural (FIFO, signed passes, one pre-queue entry per identity), not classification.
Declared crawlers (Googlebot, bingbot, Pingdom, UptimeRobot, Prometheus and the rest of the list in lib/sentinel.js) are reported and not scored — unless the same User-Agent also names an HTTP library or a scripted browser. python-requests/2.31 googlebot is scored as python-requests: a crawler's name appended to a string that already said what it is buys nothing. A UA that says only Googlebot is still taken at its word; checking it against reverse DNS is not done (it would need a network lookup per new address) and is an open product decision.
The queue cookie only counts when it is real. "Presented the queue cookie" is what keeps a returning browser away from cookieless_repeat, and it now means a qm_<room> token this server signed, for that room, still in date — any other value is treated as no cookie.
Put your own synthetic monitoring in PROTECTION_ALLOW_IPS so it is never scored. It matches the resolved client address, which is a forwarded one only when the connection comes from TRUST_PROXY_IPS — naming an allowlisted address in X-Forwarded-For from anywhere else changes nothing.
Places per address in the live line (queueMaxPerIp)
presenceSec (see Passes vs sessions) stops a client that never polls from stalling the room. It cannot stop a client that does poll from hoarding: sixty tickets spread down the line, a pass at each one's turn. That takes a ceiling on how many places one address may hold in the waiting line at once:
| Field | Default | Meaning |
|---|---|---|
queueMaxPerIp | 16 | Places one client address may hold in the waiting line. 1–100000; null for no ceiling. |
Over it, POST /api/join answers 429 {"code":"queue_identity_limit"} with Retry-After: 30 and hands back nothing — no token, no cookie. A visitor presenting a ticket they already hold is never charged; a place is freed the moment one of that address's tickets reaches the front (or is ejected). It is charged to the same address as preQueueMaxPerIp, the one thing a caller cannot mint a fresh copy of per request, and it survives a restart (it is derived from the recorded ticket owners).
The default is 16, the same as the pre-queue's. A busy live line can hold more than 16 real people behind one mobile-carrier CGNAT address at the same moment, and the ceiling refuses them, not the attacker. Raise it on a room whose audience sits behind carrier-grade NAT, or set null for no ceiling (a load test driving thousands of places from one address needs this). A room saved before the default changed, with no value stored, gets 16; one that stored null keeps no ceiling:
curl -X POST http://localhost:8080/api/v1/admin/rooms \
-H "Authorization: Bearer $ADMIN_KEY" -H 'Content-Type: application/json' \
-d '{"id":"drop","queueMaxPerIp":64}'
Threshold activation (the 3 a.m. spike)
A room can switch itself on. Set autoActivate on the room, or use the Threshold activation fields in the room editor:
{"enabled": true, "joinsPerMin": 120, "sustainSec": 60, "releaseAfterSec": 300}
The room moves bypass → active once arrivals hold at or above joinsPerMin for sustainSec, and back to bypass after releaseAfterSec of quiet. releaseAfterSec: 0 means it never stands down by itself.
- It only ever moves
bypass⇄active. A paused room is an operator decision and is never overridden. - It never stands down while anybody is still in the line.
- Both directions are annotated on the room's timeline and written to the audit trail under the
systemactor, with the number that caused the flip. While it is armed the room card carries anAUTO ≥n/MINbadge, and for 15 minutes after it acts on its own that badge turns amber and readsAUTO · SELF— so a queue the machine put up is never mistaken for one a person put up. - With
SIDESTORE=pgtimeline annotations survive a restart: they are kept in Postgres (24 h, saved about a second after each change and on clean shutdown), so the marks are still on the chart during the post-incident review. - The trigger is arrivals per minute and nothing else. A concurrency threshold is not offered because nothing meters sessions while the queue is switched off, so it could never fire — it would be a control that looks armed and is not. The number to compare against is the room card's joins/min, which is the same figure the automation reads.
This is different from AUTOTUNE, which adjusts the outflow rate of a room that is already queueing.
Historical reports
/api/admin/reports serves one row per room per hour (or per day), retained for HISTORY_RETAIN_DAYS:
curl -s "https://qm.weekday100.com/api/admin/reports?bucket=day&from=2026-08-01&format=csv" \
-H "Authorization: Bearer $ADMIN_KEY" -o last-month.csv
Columns: roomId, bucket, start, startIso, joins, passes, netUnserved, peakDepth, avgDepth, peakActive, waitP50Sec, waitP90Sec, waitMaxSec, waitAvgSec, measuredWaits, blocked, warned, samples, partial, coveredSec, missingSec. The three coverage columns are appended last, so an importer keyed on column position keeps working.
| column | what it counts |
|---|---|
joins | tickets issued |
passes | tickets promoted — permission to enter, not an entry |
netUnserved | joins - passes; null on a partial row |
peakDepth / avgDepth | queue depth, from the depth gauge |
peakActive | sessions on the protected site — the "people who got in" column |
waitP50Sec / waitP90Sec | estimated wait percentiles |
waitMaxSec / waitAvgSec | exact longest and mean wait |
measuredWaits | promotions that carried a measured join→pass duration |
blocked / warned | protection decisions; with several instances, the whole fleet's (each instance's shared counters) |
samples | sample ticks folded into the row |
Four honest labels travel with every JSON response and are worth repeating here:
waitP50Sec/waitP90Secare estimates. Each hour keeps a bounded reservoir of up to 256 waits, uniformly sampled.measuredWaits,waitAvgSecandwaitMaxSecare exact.measuredWaitsis not "people who entered the site". It was calledadmitted, and that name asked to be misread. It is the denominator underwaitAvgSecand the weight under the percentiles: how many people finished waiting in the period. A room at a concurrency ceiling of 5 that minted ten passes reportsmeasuredWaits: 10, because ten people finished queuing — five were then turned away at the door and never got in.peakActiveis the column that counts sessions on the site. See The claim window in Passes vs sessions for why those two populations differ.netUnservedis joins minus passes, not observed abandonment. Inside a single hour it also counts everyone still waiting when the clock ticked, so read it as an abandonment proxy only over a day or more.partial: truemeans the row has a hole in it. The server was not running for part of that period, sojoins,passes,measuredWaits,blockedandwarnedare lower bounds, not totals.coveredSecandmissingSecsay how much was and was not counted, andnetUnservedisnull— it cannot be derived from a partial count, and reporting it as0beside a 45,000 peak depth is how a post-mortem reaches the wrong conclusion. The console marks those rows⚠and suffixes their counters with+.
The open hour is folded in live, so a report run at 10:59 is not missing the last 59 minutes. With STORE=valkey a report Postgres cannot answer is 503 storage_unavailable (retry). It is written to the warehouse on a clean shutdown — see Stopping the server below, because on Windows the obvious way to stop a process is not one.
There is no scheduled or emailed report and no warehouse push. The CSV is the integration point.
Refusals: when the server will not start
Most misconfigurations are warnings. These five are not, because starting would be worse than stopping: the process would come up looking healthy and answer questions wrongly, or lose on the next restart what it had just accepted. Each one exits non-zero and prints the setting and the way out.
| It refuses when | Because | What to do |
|---|---|---|
NODE_ENV=production without both STORE=valkey and SIDESTORE=pg | Memory mode keeps nothing across a restart: every deploy would empty the queue, the rooms and the operators, and several instances would each run a queue of their own. The refusal names each missing setting. There is no override. | Set STORE=valkey and SIDESTORE=pg, with what they need (SECRET, VALKEY_URL, DATABASE_URL). |
SIDESTORE=file | The file side store was removed in step 5: nothing is written to local disk any more. | Use SIDESTORE=pg, or SIDESTORE=memory (the default on STORE=memory) in development. |
STORE=valkey without SIDESTORE=pg, SECRET or VALKEY_URL | The queue is journaled to Postgres, every instance must sign and verify with the same key, and there is no Valkey to connect to. The refusal names each missing setting. | Set what it names. |
SIDESTORE=pg without SECRET or DATABASE_URL | Without SECRET the key is new on every boot, and a new key on every deploy voids every pass (visitor tickets, monitor links, operator keys) that Postgres keeps. | Set SECRET (openssl rand -hex 32) and DATABASE_URL. |
ADMIN_KEY is the default admin-dev and HOST is set to anything but loopback (0.0.0.0, ::, a real address) | The default key is printed in these docs. Listening beyond loopback with it puts the surface that deletes rooms and ejects visitors on the network behind a public password. With HOST unset it does not refuse: it binds 127.0.0.1 only and says so, so node server.js on a laptop works as it always has. | Set ADMIN_KEY to a secret (openssl rand -hex 32). For local use only, HOST=127.0.0.1. |
A few single settings refuse on their own, and their rows in the environment table say so: an unknown STORE or SIDESTORE value, a missing PG_CA_FILE, a malformed QM_INSTANCE_ID. Everything else — a typo in a numeric variable, a short SECRET, thresholds in the wrong order — starts, and appears in warnings on /api/admin/health. A value that could not be applied as written is never silently substituted: the warning names the variable, what you set, and what is actually in effect.
Rolling back across the token-signing change. A release that predates per-purpose keys reads only the legacy token form. While TOKEN_SIGN_V2 is unset (the default), this release writes nothing else, so rolling back is safe. After TOKEN_SIGN_V2=1, rolling back does not refuse to start, but every v2 token is refused: visitors rejoin at the back, share links and console sessions end, and every operator whose key was minted or re-keyed under v2 gets 401 (mint them new keys). Turning TOKEN_SIGN_V2 off again stops new v2 tokens, and after the token lifetimes run out none are left. It does not convert v2 operator verifiers back, so those operators still need new keys. See Signing keys and rotating SECRET.
Signing keys and rotating SECRET
SECRET is the root of every signature the server makes. Tokens have the format base64url(JSON payload) "." base64url(HMAC-SHA256(key, body)), in one of two forms. Both forms are always accepted. TOKEN_SIGN_V2 only chooses which form this server writes.
- Legacy (default, the form every earlier release wrote): the key is
SECRETitself and the payload has nokid. - v2 (
TOKEN_SIGN_V2=1): each purpose has its own 32-byte key,K(purpose) = HKDF-SHA256(ikm = SECRET, salt = empty, info = "qm:<purpose>", 32)
for visitor (queue tickets and pre-queue handles), monitor (share links), session (the console's qm_console cookie) and operator (the stored verifier of an operator key). The payload carries kid: the lowercase hex of HKDF-SHA256(ikm = SECRET, salt = empty, info = "qm:kid", 4), eight characters that name the SECRET without revealing it. The kid picks the key a token is checked against. A kid the server does not hold is refused. Two instances agree on signatures exactly when their v2 tickets show the same kid.
What the per-purpose separation does and does not cover. A v2 token minted for one purpose never verifies as another. A legacy token carries no purpose, so it is accepted for any purpose whose payload checks pass. That is exactly as before per-purpose keys existed, and those checks (scope, room, ticket fields) are what keep a visitor ticket from opening the console. While TOKEN_SIGN_V2 is off, every token this server issues is legacy, so the separation is not yet in force. It takes effect for tokens issued after TOKEN_SIGN_V2=1, and fully once the last legacy token has expired.
Two releases. Ship this release with TOKEN_SIGN_V2 unset. It accepts both forms and writes only the legacy one, so it can be rolled back freely. Once it is settled on every instance, set TOKEN_SIGN_V2=1 everywhere. Legacy tokens already issued keep working for the rest of their life, and legacy operator verifiers are re-keyed to v2 as each operator next authenticates. For what changes if you roll back after that, see Rolling back across the token-signing change above.
Rotating. Tokens live up to TOKEN_MAX_AGE_SEC (24 h) for visitors, MONITOR_TOKEN_TTL_SEC (7 days) for share links and 12 h for console sessions. Rotate with TOKEN_SIGN_V2=1:
- Set
SECRET_PREVIOUSto the current value andSECRETto the new one, on every instance, and restart them. New tokens carry the new kid. Tokens under the old kid keep verifying, so nobody in line loses their place. - Wait out the longest life you care about: 24 h keeps every visitor ticket (the one that matters); 7 days also keeps every share link, which the console can otherwise re-copy under the new key. Operator keys have no expiry. Each operator that authenticates during the window is re-keyed to the new SECRET on that request. An operator key not used during the window stops working when you remove
SECRET_PREVIOUS, so mint that operator a new key. - Remove
SECRET_PREVIOUSand restart. Tokens under the old kid are now refused.
With TOKEN_SIGN_V2 off, the same steps still work, just without a kid. Legacy tokens name no SECRET, so each one is tried against the raw SECRET and then the raw SECRET_PREVIOUS. Nothing in a ticket tells you, or a log, which SECRET signed it, and nothing will refuse an unknown key by name: such a token just fails both tries. Visitors, share links and console sessions survive the rotation all the same. Operator verifiers are moved to the new SECRET in the form they already had (a v2 verifier is never downgraded). Prefer rotating with TOKEN_SIGN_V2=1. The kid is what lets you check every instance agrees (compare the kid in their tickets) and see when old tokens are gone.
Rotating without SECRET_PREVIOUS invalidates every ticket at once: the queue re-joins at the back (the server logs a warning when refused tickets arrive in a burst), every share link and console session ends, and every operator key stops working. A leaked SECRET is the one case where that is what you want.
Several instances (STORE=valkey)
With STORE=valkey and SIDESTORE=pg the server runs as several identical instances behind one load balancer, with no sticky sessions: a visitor or an operator may reach a different instance on every request and gets the answers one instance would give. Every instance needs the same SECRET, VALKEY_URL, VALKEY_PREFIX and DATABASE_URL; none keeps anything on its own disk. The design is in docs/adr/0001-live-state-in-valkey.md in the repository.
With STORE=memory (development only) several instances share nothing but what you give them. A shared SECRET makes a ticket's signature valid everywhere; it does not share the line. Each instance keeps its own rooms, ticket numbers and rate, and a place in line exists only on the instance that issued it, so one room spread across several makes several independent queues, each letting people through at the full rate.
- Shared. The queue (rooms, tickets, the line, admission) is in Valkey and journaled to Postgres. Every instance runs the one-second tick, and each tick credits a room the exact time since the room's shared last tick, claimed atomically, so N instances credit exactly what one ticker would crediting at all their tick instants combined: the long-run rate and the one-slot bank are unchanged. Operators, monitor links and timeline notes are written row by row to Postgres and re-read by every instance when another changes them: a hint over Valkey pub/sub (
<prefix>ctl), repaired withinRESYNC_MSif it is lost. The same check rebuilds the room config that gates requests (sentinel protection, the inline host index,onBackendDown). Console logouts, console stream tickets, theADMIN_AUTH_FAIL_PER_MINbudget, target health, the alert state and turn-notification addresses are in Valkey. An operator key, console session or monitor link whose copy this instance cannot confirm is refused with503(seeRESYNC_MS), never let through on a stale copy. Every console reads the rooms from Valkey on its 2 s tick, so a change made through one instance shows on the consoles of all of them within one tick. - One leader. Target probes, autotune, threshold activation, alert delivery and the hourly history rows run only on the instance holding the leader lease, so each happens once for the fleet (see
LEADER_LEASE_MS).leaderon/healthzsays whether this instance holds it. When the leader dies another instance takes over within aboutLEADER_LEASE_MS; a clean stop hands the lease back at once. While Valkey is unreachable no instance leads and those jobs pause, and each instance sends its ownvalkey_unreachablealert (seeVALKEY_ALERT_AFTER_MS). - Per instance. The per-address budgets (
JOIN_LIMIT_PER_MIN,CHECK_LIMIT_PER_MIN,WAITING_LIMIT_PER_MIN,NOTIFY_LIMIT_PER_MINand the rest) and the stream caps (SSE_PER_IP,SSE_MAX_TOTAL) are each instance's own: N instances allow up to N times each. Cloudflare's edge rate limiting is the fleet-wide control./metricsreports this instance's own process, with aninstancelabel (QM_INSTANCE_ID, a display name only): scrape each instance (see Metrics and monitoring). The console's counters are the sum over the live instances. - Deploys. A deploy replaces every instance, and each stop ends every stream on that instance. Nothing about the line is lost (it is in Valkey). A waiting page starts polling
/api/statusevery 5 s at once and reopens its stream: the browser retries it, and once the browser gives up the page retries after 2 s, then longer, up to 20 s. A console reconnects and gets a full frame. During a rolling deploy every instance must serve every waiting-page asset hash still being handed out (see Waiting-page assets and live streams across a deploy). - Stopping one.
POST /api/admin/shutdownstops only the instance that answered it, which behind a load balancer is whichever one the request reached (see Stopping the server).
Readiness (/healthz)
The payload carries two different judgements, and they are not the same question:
status—ok,degradedorshutting_down. "Is anything wrong that an operator should look at?" A target origin that stopped answering, or a clock that stepped, makes thisdegraded, and the same text appears inincidentson/api/admin/healthand on the console badge.ready—trueorfalse. "Should a load balancer send this instance visitors?" This is what the HTTP code follows:200whenready,503when not.
ready goes false for exactly three reasons:
- storage is degraded (on
STORE=valkey, only this instance's: see below) — the instance cannot persist a join, so it already answers503 storage_unavailableto every visitor who tries to take a ticket. It must not keep being handed a share of the surge. OnSTORE=valkeya failed events-journal commit to Postgres makesstatusdegraded(withstorage is degraded: events journal commit failed: …inincidents) and pagesstorage_failing, but leavesreadytrue: every instance shares the one Postgres, so taking them out of rotation fails the whole fleet out of the load balancer and moves no join anywhere it could be persisted. Each join still waits for its own commit and answers503 storage_unavailablewhen it fails, so no ticket is acknowledged that is not in Postgres. Storage reads healthy again once a later commit has landed and 3 ×ALERT_EVERY_MShave passed since the newest failure, so a fault that comes and goes stays failing rather than reading ok between failures;storage.writeErrorskeeps counting.storage_failingpages once two alert frames within that window saw new failures: one failed commit (or one refused join, which is its commit and its rollback's) shows as degraded but does not page; an intermittent fault, or a failure still there at the next frame's probe, does. With no traffic to prove it, each alert frame probes the journal while it is failing (the commit's insert, for one row, in a transaction rolled back), so a blip in a quiet hour recovers on its own. The exception is a fault of one instance: when this instance's journal is failing while another instance reported its storage ok in the last 3 ×ALERT_EVERY_MS, moving visitors elsewhere does get their joins persisted, soreadygoes false (storage is degraded on this instance only) and the load balancer takes only this instance out. Readiness follows a hard failure only, not the window above: it goes out when its journal has had no commit or probe land since its newest failure and two alert frames in a row each saw a new failure (a commit, or the frame's probe), and comes back at the first commit or probe that lands (the probe proves it with no traffic). A single statement timeout whose next commit lands does not take it out. An instance whose own commits are failing does not count as ok. When no other instance is ok (the shared Postgres, a fleet of one, Valkey unreachable) it stays in. shutting_down— set the instant a stop begins, before the final snapshot and before any connection is closed, so a load balancer stops sending new work while requests already in flight finish.- at the connection ceiling —
MAX_CONNECTIONSreached, so new requests are already being answered503here instead of proxied (seeCONNECTION_RESERVEin the environment table, and the surge table above). A load balancer still sending this instance its share of a surge is the opposite of what a readiness probe is for. Note thatstatusstaysokthrough this — saturation raises no incident — so "alert onstatus, page onready" below is not two ways of saying the same thing: this is the case onlyreadycatches.
A customer's origin being unreachable deliberately does not clear ready, and neither does a clock step. Both are reported as incidents; neither is a reason to deregister the pod. Wiring readiness to them turns one origin outage into a total waiting-room outage — every instance fails its probe at once and the estate empties out of the load balancer, so the component whose whole job is standing in front of a site that is down is the component that disappears. Restarting a pod does not fix somebody else's origin or this machine's clock.
With STORE=valkey the payload also carries backend: {store, valkey, recoveringRooms, ok}. valkey is up or down (this process's connection; also down for 5 s after a command timed out on a connected socket), recoveringRooms counts the rooms being rebuilt from the Postgres journal (marked here, or answered as fenced by Valkey in the last 30 s); /healthz is public, so it carries the number only and /api/admin/health names the rooms (it also carries instance, the name of the instance that answered, and fleet: {instances, partial}, see Metrics and monitoring; the public /healthz carries neither). Either makes status degraded and adds an incident on /api/admin/health; neither clears ready, because every instance shares the one Valkey and taking this one out of rotation fixes nothing. What visitors get meanwhile is onBackendDown.
leader (true or false) says whether this instance holds the leader lease and runs the fleet-wide jobs: target probes, autotune, threshold activation and alert delivery (see LEADER_LEASE_MS). In a STORE=valkey fleet exactly one instance says true, and none does while Valkey is unreachable, which pauses those jobs. It never changes ready or the HTTP code. With STORE=memory it is always true.
Point k8s readiness, your ALB target group and your uptime monitor at /healthz. Alert on status, page on ready. /api/version returns the same payload but always 200: it answers a question about the build, not about readiness, so a rollout script reading it is not tripped by a degraded instance.
Inline deployments: which address the probe uses decides what it gets. A probe that addresses the process — a pod IP, a container port, localhost — asks for a host no room fronts, so /healthz is the queue's own route and answers normally. A probe that addresses the public hostname is a visitor as far as this server is concerned, and /healthz there is the customer's path: measured on an inline host, GET /healthz answers 503 with {"code":"queued"} for ever, because a monitor never gets admitted. Wire that to an ALB target group and it deregisters every healthy instance you have.
Under the reserved prefix it is the queue's route again. On the public hostname, probe /__qm/healthz:
| Probe addresses | Path to use |
|---|---|
the process (pod IP, container port, localhost) | /healthz |
| the public hostname of an inline room | /__qm/healthz |
| the queue's own hostname (snippet deployments) | /healthz |
The same split applies to /api/version and to /metrics below.
HEAD works everywhere GET does, and answers the same status and the same Content-Length with no body — so a check that defaults to HEAD (many do), or a curl -I on a connector download, gets the truth rather than a 404. It is the same answer, so it makes the same decision: a HEAD /healthz on an inline public hostname is still 503, for the reason above.
When the queue itself is down
ready going false takes one instance out of rotation, which is the normal case and costs nobody anything. This section is about the other one: every instance unready at the same time.
If the queue is deployed inline — Users → Cloudflare → Queue → your app — then "every instance unready" means the site has no path to the app. Somebody has to decide what the visitor gets, and the decision is not the queue's to make at request time; it is a routing rule you write in advance.
The decision taken here: serve a static page, do not fail open. Failing open sends the full surge at an origin with nothing in front of it — the exact outage the queue was bought to prevent, arriving at the worst possible minute. A campaign unreachable for a few minutes is recoverable; a melted origin is not. If your campaign is worth more than that risk, route the failover pool at the app instead and accept what follows. Decide it now, not at 21:04.
public/offline.html is that page. The running server hands it out at /offline.html, and it has to end up somewhere Cloudflare can serve without this server.
It is deliberately built for that job: one self-contained file, no <link>, no <script src>, no image, no webfont — every subresource would be a request to the origin that just stopped answering, and a half-rendered apology looks like a second bug. It carries the same five languages as the waiting page and reads the same qm_lang key, so a visitor who picked Thai on the way in still gets Thai. It retries on its own with jittered backoff (20 s → 60 s, plus up to 80 % random), because everyone landed on it in the same second and a fixed interval would send the whole crowd back in one synchronised wave, hardest exactly as the queue is trying to come up.
It does not show a position, an ETA or a progress bar. The component that holds places is the one that is down, so any number there would be invented.
Wiring it up — the failover Worker (no Load Balancing subscription needed). This server builds and serves it, with its own copy of offline.html already inlined, so the deployed page cannot drift from the shipped one:
curl -fsSL http://localhost:8080/connectors/cloudflare-failover/latest.js -o qm-failover.js
npx wrangler deploy qm-failover.js --name qm-failover \
--compatibility-date 2025-01-01 --route 'YOUR-QUEUE-HOSTNAME/*'
Or take the whole wrangler project — wrangler.toml, a pre-deploy check, the README — from /connectors/cloudflare-failover/latest.tgz and keep it in your own repo. npm run deploy runs the check first. There is no room id and no API key to configure: this Worker never calls the queue's API.
Route it at the hostname the queue answers on — inline, that is your public hostname. It is not the edge gate connector (/connectors/cloudflare), which decides who waits and belongs in front of your origin in a snippet deployment. Cloudflare runs one Worker per route, so give the two different patterns.
What it acts on, and the two exceptions that make it safe:
| The queue answered | Worker does |
|---|---|
| nothing — refused, DNS, TLS, timeout | offline page, 503 |
521–527, 530 (Cloudflare could not use the origin) | offline page, 503 |
503 with X-QM-Queued | passes through |
503 without that header | offline page, 503 |
502 / 504 | passes through |
| anything else | passes through |
503 alone is never enough to act on. A perfectly healthy gate answers 503 to every poll from a visitor who is still waiting, and marks those with X-QM-Queued. Replacing them with the outage page takes a working queue off the air and swaps a live countdown for "we are down".
502/504 are not this rule's business. Those mean the queue is up and your app is down, and this server already answers them itself — see When the site behind the queue is down below. Serving offline.html there would tell the visitor entry is unavailable when in fact they are through and it is the site that is broken.
After deploying, prove it: stop the queue and load a page. The response carries X-QM-Failover: unreachable (or timeout, or http-<status>), so a curl -I tells a working failover from a coincidence. This is the one piece of the product that is never exercised in normal operation — nothing else will tell you it is broken until the night it matters.
Or a Load Balancer, if you have one: primary pool = the queue instances, health check GET /healthz (that is what it is for), failover pool = wherever you put offline.html. With STORE=memory (development only) the instances do not share a line, so the pool must pin each room to one instance — a second instance is a failover target, not extra capacity for the same room. With STORE=valkey they share it, and the primary pool is every instance (see Several instances (STORE=valkey)).
Either way — the Worker above does all five; a Load Balancer's failover pool has to be configured for them:
503, not200. A200tells crawlers and uptime monitors the page they got is the site.Retry-Afterso well-behaved clients back off on their own.Cache-Control: no-store, or Cloudflare will keep serving the apology after the queue is healthy again.- The page also carries
noindex, which is the belt to that braces. - Scope the rule to document requests. A
/favicon.icothat 404s is fine; answering it with a page of HTML is not.
The edge gate fails open, and says so
The rule above is for a queue deployed inline. The edge gate connector (/connectors/cloudflare) sits in front of your origin and asks the queue per request; when the queue is slow or failing it fails open — the visitor goes to origin unqueued — because a queue that is down must never be an outage of the site it protects. That stays. What changed (connector 1.1.0, QM-340) is that it is no longer silent:
- Every fail-open response carries
X-QM-Failover: <reason>—timeout(no answer within 2 s),unreachable(connect/DNS/TLS failed),http-<status>(any non-2xx except 429:http-503,http-401for a rejected key), orinvalid-response(a200that is neither apassnor a usablequeue). Same header and names as the failover Worker, but here it means forwarded unqueued, not offline page. An answer the queue actually gave (a pass, a redirect to the waiting room) is never marked. A Worker with no configuration or no key forwards unmarked and says so in its log instead — that is a deploy mistake, not the queue failing. - A
429is not a fail-open (connector 1.2.0, QM-379). An exhaustedCHECK_LIMIT_PER_MINis the queue up and shedding load, so the Worker sends the visitor to the waiting room (/w/<room>) instead of the origin: noX-QM-Failover, no Analytics Engine count, one log line per isolate. The Node connector (1.1.0), the WordPress plugin (1.2.0) and the browser tag (1.4.0) do the same. It shows up on the server asqm_rate_limited_checks_total; a sustained increase means visitors the room may have had space for are waiting — raiseCHECK_LIMIT_PER_MIN. - Each one is counted in Workers Analytics Engine when a dataset is bound as
QM_ANALYTICS. Without the binding nothing is counted and nothing breaks.# wrangler.toml of the edge Worker [[analytics_engine_datasets]] binding = "QM_ANALYTICS" dataset = "qm_failover"
Each data point is blob1 = room id, blob2 = reason, double1 = 1, index1 = room id.
Alerting on it. The queue server cannot count checks it never received, so the alert has to read the edge's numbers, not the server's. Run this against the Analytics Engine SQL API from whatever already pages you (a cron, your monitoring system's HTTP check), every few minutes, with an API token that has Account Analytics: Read:
curl -s "https://api.cloudflare.com/client/v4/accounts/$CF_ACCOUNT_ID/analytics_engine/sql" \
-H "Authorization: Bearer $CF_API_TOKEN" --data "
SELECT blob1 AS room, blob2 AS reason, SUM(_sample_interval) AS fail_opens
FROM qm_failover
WHERE timestamp > NOW() - INTERVAL '5' MINUTE
GROUP BY room, reason
ORDER BY fail_opens DESC
FORMAT JSON"
Alert when any row comes back: every one is a visitor who reached the origin without being queued. SUM(_sample_interval), not COUNT(), because Analytics Engine samples at high volume and records the rate it kept. timeout/unreachable/http-5xx means the queue is down or drowning — see /healthz above. For a spot check without the dataset: curl -sI https://shop.example.com/checkout | grep -i x-qm-failover.
Why this and not the Worker reporting its counts back to the queue when it recovers: a Worker's memory is per isolate, isolates are evicted freely and there are many per colo, so a replayed count is lossy exactly during a long outage; it could not report a 401 at all (the key is what is failing); and it would add a write endpoint to the server. Analytics Engine is written at the moment of failure, independently of the queue, and aggregated across every colo.
When the backend is down (onBackendDown)
With STORE=valkey, a room's live state is in Valkey. When Valkey cannot be reached (connection lost, command timeout, failover) or a room is being rebuilt from the Postgres journal after Valkey lost or rolled back its state, the API answers 503 storage_unavailable with Retry-After: /api/check, /api/join, /api/status and the rest, snippet and inline alike. A room that is recovering never answers no_room.
Operator changes around a rebuild. A rebuild replays what the Postgres journal held when it read it. A change that ran in Valkey but reached the journal only after that read is not in the rebuilt room. Creating or editing a room, changing its schedule (opensAt) or deleting it waits for its journal entry, so such a change answers 503 storage_unavailable, although it ran, when the rebuild left it out (or the room is fenced when it lands): retry it once the room answers again, and it applies to the rebuilt room. Pausing, resuming or changing a room's rate (by an operator, autotune or threshold activation) is journaled in the background and answers 200 at once, so one made in the seconds before Valkey lost the room can be undone by the rebuild without an error: after a recovery, check the room's state and rate in the console and set them again if they are not what you left.
Valkey must run with maxmemory-policy noeviction (REQUIRED). Under an evicting policy Valkey can drop part of a room — its admissions, the passed set, ticket ownership — while the room's main hash survives. Nothing detects that partial state, and a visitor whose admission was evicted is sent back to the queue or let through wrongly. With noeviction a full Valkey refuses writes instead, which the API answers as 503. At boot each process reads the policy from INFO memory (CONFIG is disabled on DO managed Valkey), logs a loud warning when it is not noeviction, and reports it as backend.maxmemoryPolicy and an incident on /api/admin/health (a share link sees the incident with the policy withheld, and no maxmemoryPolicy).
An inline room also decides what its own site's visitors get, with the room setting onBackendDown:
| Value | Browser navigation | API / XHR / fetch |
|---|---|---|
closed (default) | the room's waiting page in retry mode: 503 + Retry-After: 5, no place in line taken, a "temporarily unavailable" line, and the page reloads itself; the gate decides again on the reload | 503 storage_unavailable JSON, Retry-After: 5 |
open | forwarded to proxyOrigin as if the room were in bypass: no queue check, no cookie | forwarded the same way |
open lets the full surge through to the origin for as long as the backend is down, so choose it only for a site that would rather be slow than closed.
While Valkey is down the gate cannot read the room from it, so it uses the last config this process saw for that host (the inline host index, refreshed on every room change). A host this process has never seen as inline is not gated at all (it is routed as a non-inline request). A process that boots while Valkey is unreachable still starts (it logs valkey: unreachable at boot and keeps reconnecting; /healthz reports valkey: "down") and builds the inline host index from the room configs in Postgres (the rooms table, deleted rooms left out), logging inline host index built from Postgres, so onBackendDown applies to every inline room at once: a closed room's host answers the retry page, never a 404 or the origin. inline.indexFrom on /api/admin/health says postgres meanwhile; once Valkey answers, the rooms check replaces the index with Valkey's copy (valkey). If Postgres cannot be read either, the process logs Valkey and Postgres are both unreachable and gates nothing (its inline hosts answer 404) until one answers: the Postgres read is retried every 1 s, doubling up to 30 s. A WebSocket upgrade on an inline host follows the same setting while the backend is down: closed answers 503 with Retry-After: 5, open forwards the handshake to the origin. Only three kinds of connect error refuse to start, each named without printing the URL: credentials Valkey refuses (WRONGPASS, NOAUTH, NOPERM, ACL errors, an invalid password), a host name that does not resolve (ENOTFOUND), and a TLS certificate that does not verify. Every other connect error boots degraded and keeps reconnecting: refused, timed out or reset, the host or network down, a Valkey that is loading, has lost its master or asks to try again (LOADING, MASTERDOWN, TRYAGAIN), and any error this list does not recognise, such as ERR max number of clients reached or a TLS proxy that is only half up (EPROTO, ERR_SSL_*). An unrecognised one is logged once at boot as a WARNING naming its code or first words. A VALKEY_URL that points at a reachable address with nothing listening therefore boots degraded, so after every deploy alert on backend.valkeyEverConnected on /healthz: it is false until this process has reached Valkey once, and stays true after. It is reported only and never makes the instance unready. Each instance decides from its own last-seen config: a change to onBackendDown (or to the room) made on another instance just before the outage may not have reached this one, so for the length of the outage two instances can answer the same host differently. Before planned Valkey maintenance, make the onBackendDown change at least RESYNC_MS (default 5 s) ahead, so every instance holds the new value. After the first failure that means Valkey is unreachable, each process answers the backend-down reply at once for 5 s (while the connection is not back, or one request at a time probes a Valkey that timed out) instead of waiting for a command timeout per request. Counted per mode in qm_proxy_backend_down_total{mode} and in inline.backendDownOpenTotal / backendDownClosedTotal on /api/admin/health; one log line per room per minute.
Set it like any room field: POST /api/admin/rooms with {"id": "shop", "onBackendDown": "open"}; any other value is 400 invalid_on_backend_down. It is stored with STORE=memory too, where it has no effect: the in-process engine has no backend to lose.
When the site behind the queue is down
The opposite outage, and inline the more likely one: this server is healthy and proxyOrigin is not answering. Nothing needs to be wired up for it — the queue handles it — but you need to recognise it.
| queue down | site down | |
|---|---|---|
| Status | 503 (from your CDN rule) | 502 unreachable, 504 timed out |
| Header | — | X-QM-Origin: origin_unreachable|origin_timeout |
| Page | public/offline.html, served by the CDN | public/origin-down.html, served by this server |
| It says | entry is unavailable, you are not in a queue | you are through, the site is not answering |
| Chip | red: the queue system itself is failing | amber: hold on, your place is safe |
Neither page says paused. That word is the operator's: on the waiting page it means a person chose to hold the line, and a visitor who has seen both surfaces must be able to tell an outage from that choice.
A visitor who navigates gets the page; a background fetch() gets JSON with the same code, because a document handed to an XHR fails as if the app were buggy. Both carry X-QM-Origin, so one header answers "which side is broken" whatever shape the request was.
origin-down.html follows the same rules as the failover page — one file, no subresources, five languages, jittered retry (15 s → 60 s) — with one difference that matters: it never suggests the visitor has lost their place, because they have not. A retry is a plain reload against a pass they still hold.
Where you will actually see it. The console does not make you go looking:
- The room card carries the line "Site not answering — N visitors got an error page" and its health chip flips to
site · failing. That chip outranks the probe, because a HEAD on/can succeed while every real page 502s. - The header incident badge leads with
⚠ origin down — <room>even when another room has a deeper queue. A snippet room's dead target hurts people who will be let through; a dead inline origin is hurting everyone already through it, and itswaitingcan honestly read 0 while the whole site is down. - Diagnostics adds two rows: Inline rooms (gated vs bypassed, with the ratio) and Origin errors (the total, broken down per room and per failure). Visitor disconnects are listed separately and never added to that total — a closed laptop is not the site failing.
The numbers behind all of it: inline.originErrorsTotal and inline.originErrors in GET /api/admin/health, and in /metrics:
| Series | Alert on |
|---|---|
qm_proxy_origin_errors_total{room,phase} | any sustained increase — the site behind that room is failing. connect = refused/unreachable, timeout = accepted and never answered, body_stall = answered and then died mid-response, upgrade = a WebSocket handshake failed. client_reset is the visitor's own socket and is excluded from the health total on purpose. |
qm_proxy_gated_total / qm_proxy_bypassed_total | a ratio far from ~1:40 — the asset rule matches everything or nothing |
qm_proxy_inline_rooms | a drop you did not make |
qm_join_rate_limited_total | increase — visitors are being refused a place in line |
qm_waiting_page_rate_limited_total | sustained increase — something is pulling the waiting page far faster than a browser would |
qm_open_connections / qm_max_connections | the first approaching the second |
qm_connections_refused_total | any increase — you are at the socket ceiling and shedding |
qm_process_resident_bytes | growth that does not level off |
The card, the badge and the counters all say the same thing on purpose: the queue is up, the site is not, and nothing on the rate dial will help.
Stopping the server
A clean stop marks the process not-ready, stops accepting new connections, folds the open hour into the warehouse, hands the leader lease back (STORE=valkey), makes the events journal's last commit to Postgres (STORE=valkey), saves the timeline annotations and records server.stop in the audit trail. A kill does none of that: the hour in progress is lost, and with STORE=valkey the other instances wait for the lease to expire (LEADER_LEASE_MS) before one of them leads. With STORE=memory a stop of either kind loses the queue, the rooms and the operators: nothing is kept across a restart.
Idle and SSE connections get SHUTDOWN_GRACE_MS to finish, then are cut — SSE streams never end on their own. A socket carrying an inline proxied response is left alone at that point, since unlike SSE it does end on its own; it is only bound by SHUTDOWN_DRAIN_MS, the hard ceiling on the whole stop. An inline deployment with no SHUTDOWN_DRAIN_MS set defaults to a 15s floor instead of the plain 2s default, since 2s was measured truncating a real proxied download mid-transfer on every deploy — but an explicit SHUTDOWN_DRAIN_MS always wins, inline or not: it no longer gets silently raised to the 15s floor.
The audit line is honest about what happened: if the final checkpoint did not land — on STORE=valkey, an events-journal commit to Postgres that failed — it records shutdown … WITH STORAGE FAILING and says the changes since the last good commit are not in the events journal, rather than writing the word "clean" over a checkpoint that never happened.
| Platform | Clean stop |
|---|---|
| Linux / macOS, foreground | Ctrl-C, or kill <pid> (SIGTERM) |
| Linux / macOS, service | systemctl stop, docker stop — both send SIGTERM |
| Windows | Ctrl-C in the console window, or the admin API below |
Windows has no other clean stop. taskkill /PID <pid> without /F answers This process can only be terminated forcefully (with /F option), and /F is a hard kill the handler never sees. So for anything not running in a console you have your hands on — a service wrapper, a scheduled task, a remote box — use:
curl -sX POST http://localhost:8080/api/admin/shutdown -H "Authorization: Bearer $ADMIN_KEY"
It answers {"ok":true,"stopping":true,"uptimeSec":N} and then exits through the same path SIGTERM takes. It needs the owner role (or ADMIN_KEY) — the same bar as deleting a room, and deliberately out of reach of an operator key that runs the queue.
It stops the one process that answered. With several instances behind a load balancer that is whichever instance the request reached; the others keep serving. To stop a particular instance, send it to that instance's own address, or stop it through the platform.
With STORE=valkey a hard kill is still survivable: the line is in Valkey and every acknowledged join is in the Postgres journal, so no visitor loses their place. What is lost is the part of the current hour that had not been written yet.
Deploy, backup and recovery
Production runs STORE=valkey + SIDESTORE=pg (NODE_ENV=production refuses anything else), and no instance keeps anything on its own disk: the line is in Valkey, and everything that has to survive is in Postgres. The design (rolling deploys, Valkey and Postgres backups, RPO 1 s / RTO 5 min, capacity figures and how they were measured) is in docs/adr/0002-deploy-backup-recovery.md in the repository, as amended for step 5.
Deploy. The instances are interchangeable, so replace them one at a time behind the load balancer: start a new one, wait for its /healthz to answer 200, then stop an old one cleanly (see Stopping the server). Nobody loses their place. Each stop ends the streams on that instance, and the waiting pages reconnect to another (see Deploys under Several instances (STORE=valkey)). Two versions run side by side meanwhile, so every instance must serve every waiting-page asset hash still being handed out (see the next section). With SIDESTORE=pg the schema is migrated at boot. Deploy away from a scheduled drop.
Rollback. Deploy the previous version the same way. A rollback past the release that added kid to tokens re-queues every visitor who joined after the upgrade, because the older build cannot verify those tokens (see Rolling back across the token-signing change under Refusals).
Backup set.
| What | Why it matters |
|---|---|
Postgres (the database at DATABASE_URL) | The durable record. The events journal, open_orders and the room_snapshots checkpoints hold every room's queue: it is what a room is rebuilt from. The same database holds operators and monitor links, the audit trail, timeline notes, the 5 s metrics and the hourly history. Back it up with your provider's point-in-time recovery, or pg_dump. |
| Valkey | The live line. When Valkey loses or rolls back a room's state, the room is rebuilt from the Postgres journal (see When the backend is down), so a Valkey backup shortens a recovery but is not what makes one possible. Run it with maxmemory-policy noeviction. |
SECRET (and SECRET_PREVIOUS during a rotation) | Keep it in your secret manager, never in a data backup. Lose it and every ticket, monitor link, console session and operator key stops working at once. |
Restoring Postgres. Restore it together with an empty Valkey (or a new VALKEY_PREFIX): every room is then missing from Valkey, and each is rebuilt from the restored journal as it was at the restore point. Every join since that point is lost. Do not restore Postgres under a Valkey that still holds the newer state: a room that is further along in Valkey than in its journal reads as events not yet committed, and is not rebuilt. To check a backup before you need it, restore it into a new database and start a test instance on it with a VALKEY_PREFIX of its own, then compare waiting per room on /api/admin/rooms.
With STORE=memory there is nothing to back up: the server keeps nothing across a restart, and a deploy empties the queue, the rooms and the operators. That is why production refuses it.
What a backup cannot bring back. Turn-notification e-mail addresses (EMAIL_NOTIFY) are deleted when the ticket ends, or up to LAPSED_TTL_MS (24 h) later if its entry window lapsed. They are never written to Postgres. With STORE=memory every restart or deploy drops them all; with STORE=valkey they are in the shared Valkey until then. Valkey RDB snapshots and provider backups can hold a copy after deletion until they rotate out. Nothing in this build sends them or can export them. The ADR describes what production does instead.
Waiting-page assets and live streams across a deploy
Waiting-page assets. The waiting page's style and main script are not inline. They are served at a content-hashed address, /assets/waiting-<hash>.css and /assets/waiting-<hash>.js, where <hash> is the first 22 characters of the base64url SHA-256 of the file. On an inline host they are under /__qm/assets/…. They are sent with Cache-Control: public, max-age=31536000, immutable, so a browser or CDN keeps each one for a year. A new build that changes public/waiting.html gets new hashes. A process serves only the hashes of its own build. An old hash is a 404, and a page that asked for it renders with no style and no script.
So the rule for more than one instance behind one address, and for a rolling deploy: every instance must serve every asset hash that any instance is still handing out. A page rendered by an old instance asks for old hashes, and the load balancer may send that request to a new instance. Ways to meet it:
- Deploy all instances at once (this build's stop-then-start deploy does this).
- Pin a visitor to one instance for the length of a page load (not possible behind a load balancer with no sticky sessions, such as DO App Platform's).
- Keep the previous build's asset files reachable, for example at a CDN or object store in front of
/assets/, until no old instance is left.
A CDN that already cached an old hash covers most requests, but a cold edge does not, so do not rely on it.
Visitor streams. An open waiting-page stream (/events) is ended by the server when the token it stands for is due to be re-signed (see TOKEN_MAX_AGE_SEC). The server first sends event: refresh with data {}, and the page reopens the stream itself. The new stream's response re-signs the place and re-sets the cookies. The due time is TOKEN_REFRESH_AFTER_MS after the token was issued: half the shorter of TOKEN_MAX_AGE_SEC and QUEUE_COOKIE_MAX_AGE_SEC. Each stream is ended up to 10% of that window later than the due time. The extra delay is fixed per place (a hash of the place), so a burst of joins that share one issue time does not reconnect all at once. A port must send the same frame and use the same rule. A deploy ends every stream, and pages reconnect by themselves.
Console stream. /api/admin/events sends a full stats frame on connect and again on the next 2 s tick. After that, a tick sends only a delta: {"delta":true,"rooms":[changed rooms],"removed":[ids],…} plus the same totals as a full frame. If rooms changed order (a room deleted and created again under the same id), that tick sends a full frame instead. A console that holds a delta with no full frame under it reconnects, and that is the resync. A deploy ends the stream, so the console gets a full frame again.
How a pass is bound to a visitor
The pass token travels in the URL (?qm_token=), so it is written into every origin and CDN access log. Anything that treats "possession of the token" as "is the visitor" hands your queue to whoever reads those logs. Three layers stop that:
- Ticket ownership. A ticket is recorded against the client that queued for it. Presenting somebody else's token gets you a fresh ticket at the back of the line — never their place, and never their pass.
- A browser-bound session cookie (
qms_<roomId>,HttpOnly,SameSite=Lax, first-party to this server, only its hash is stored). This is the ownership proof, in the same spirit as Queue-it's session cookie: it never appears in a URL or a log, it survives the visitor changing network, and the holder can use it to take their pass back from anyone who managed to claim it. - Client fingerprints for browsers that block cookies: a strict one (socket peer + forwarded identity +
User-Agent) and a second tier without the socket peer, so a CDN answering from a different edge node does not lock a visitor out of their own pass.
The second tier is built entirely out of request headers, and those headers — client address, User-Agent — sit on the same access-log line as the token. So it is only computed for a request that arrived from a proxy you named in TRUST_PROXY_IPS. Without that allowlist there is no way to tell a CDN edge node from an attacker replaying a log line, and the strict tier (which pins the socket peer, the one thing a sender cannot choose) is the only identity.
Consequences worth knowing:
- A visitor who blocks cookies and changes network address mid-queue is treated as a new visitor and starts at the back — the same as a visitor who cleared their storage.
- On a deployment behind a CDN with
TRUST_PROXY_IPSunset, a cookie-blocking visitor whose requests move between edge nodes will have to rejoin. SetTRUST_PROXY_IPS(/api/admin/healthwarns you if you have not).
What this stores, and what belongs in your privacy notice
Per ticket the server keeps a truncated SHA-256 digest of the client address, the socket peer, the forwarded chain and the User-Agent. The events journal in Postgres (STORE=valkey) additionally records the client address in the clear, on every join. That is all of it: no third-party cookie, no cross-site identifier, no profile, and nothing leaves your server.
It is still personal data under GDPR, because an IP address is — so it needs naming in your privacy notice, with your own counsel deciding the lawful basis and the retention period (which is however long you keep the Postgres events journal and its backups: the journal is not pruned yet). The purpose is narrow and worth stating plainly: the pass token travels in a URL and therefore lands in access logs, and this binding is the only thing stopping whoever reads one of those lines from taking the visitor's place. The customer-facing FAQ in the Integration guide links here.
Passes vs sessions
Two different clocks, per room:
| Field | Default | Meaning |
|---|---|---|
passedTtlSec | 600 | How long a promoted visitor has to walk through the door. Unused after that, the pass expires and they rejoin. |
sessionTtlSec | 1800 | How long they stay inside once admitted. Idle timeout: every /api/check slides it forward (Queue-it's extendCookieValidity). |
sessionMaxSec | null (4 h) | Hard ceiling on a single visit, however active, counted from walking in. null means 14400 s, or sessionTtlSec if that is longer; set a number to change it. |
A session holds exactly one maxConcurrent slot no matter how many times its token is replayed, and releases it when the visitor goes quiet for sessionTtlSec (or is ejected).
The claim window: a third clock, and the one that surprises people.
An unclaimed pass holds a maxConcurrent slot for 120 s (or passedTtlSec if that is shorter), not for the pass's whole life. Otherwise one visitor who was promoted and closed the tab would wedge a slot for the full ten minutes. Past the window the pass is still valid — its holder can still walk in — but it no longer counts against the ceiling, and its holder has to acquire a slot at the door like anybody else.
The clocks in order. Promotion issues the pass, an entry grant valid for passedTtlSec (600 s). For at most its first 120 s (the claim window: shorter if passedTtlSec is, and sooner still when the holder is absent, see presenceSec below) the unused pass holds a maxConcurrent slot. Walking in starts the session: every check slides it forward, and it ends after sessionTtlSec (1800 s) without one, or sessionMaxSec after walking in, whichever comes first. A pass still unused when the grant ends expires, and its holder rejoins; a session already running is not cut short by the grant ending.
The consequence is worth stating plainly, because it looks alarming and is not:
A room with maxConcurrent: 5 can have ten people holding valid, unexpired passes. Five own slots; five gave theirs back. If all ten arrive, exactly five get in and five are turned away — the ceiling is never breached. Bounced pass-holders are admitted strictly oldest-ticket-first as slots free, and keep their place while they keep polling.
Those five are real people and they are not in waiting — they already have a pass. The room card shows them under CAPACITY as +5 holding passes, and the API reports them per room:
| Field | Meaning |
|---|---|
occupancy | Slots in use: active sessions + passes still inside their claim window. |
passesOutstanding | Valid passes past the claim window, owning no slot. These people will be turned away if they arrive now. |
atDoor | The subset of them being turned away right now (knocking within the last 30 s). |
claimWindowSec | The window this room is actually using. |
passesOutstanding climbing while occupancy sits at the ceiling is the normal shape of a capacity-bound drop. passesOutstanding climbing while occupancy is below the ceiling means passes are being minted faster than they are being walked — usually promotion outpacing a slow target page. It also climbs when the line is full of ghosts, below.
Absent holders do not hold slots (presenceSec)
Before this rule, one client could freeze a room without ever looking at it. Measured on a room at 600/min with maxConcurrent: 5: one address sent 200 joins with no cookies and never polled. All 200 got a place; the next real visitor got position 201 and 15 seconds later was still at 196, with occupancy 5, active 0 — every slot held by an unclaimed ghost pass for its 120 s claim window, five at a time. Sixty ghosts froze that room for about 24 minutes.
A pass now reserves a slot only for a holder who is demonstrably there:
| Field | Default | Meaning |
|---|---|---|
presenceSec | 30 | How long a visitor may go unseen and still have a slot held for them. null switches the rule off for the room. |
"Seen" means the visitor joined, polled /api/status, or is holding the /events stream (the waiting page does one of these every 5 s, so an open page is never judged absent). Two consequences:
- At their turn, a ticket whose holder has been unseen for
presenceSecis still promoted, in order — nobody loses their place — but its pass owns no slot and costs the room's rate nothing. It is exactly a pass past its claim window: if the holder comes back insidepassedTtlSecthey are through, and take a slot at the door, oldest ticket first, like any other bounced pass-holder. A phone that was locked while the line moved loses nothing but the reservation. - After promotion, a pass whose holder was never told they are through (no status call since) gives its slot back once they have been unseen for
presenceSec, instead of holding it for the whole claim window. A visitor who was told keeps the slot for the full window, because they are now on their way to your page and there is nothing left to watch.
The same 200-ghost room now lets the real visitor through about presenceSec after the burst; the ghosts' passes show up in passesOutstanding, not in occupancy. A restart treats every ticket as seen at boot (presence is not persisted), so a crash can cost nobody a slot and can give a ghost at most one more presenceSec. Shortening the claim window for everybody was the rejected alternative: it bounces real visitors on every slow target page and still lets a burst of fresh ghosts hold every slot for the shorter window.
Two origin lists, and why they are not one (returnOrigins, tagOrigins)
They used to be a single field, which meant one edit decided two things that fail in opposite directions:
| Field | Answers | Getting it wrong |
|---|---|---|
returnOrigins | Where a promoted visitor may be sent | An open redirect on your own domain that hands out pass tokens. Loud, security-owned, kept tight. |
tagOrigins | Which pages may read a door check (the browser CORS allowlist) | That page is served completely unqueued, silently: the browser discards an answer with no Access-Control-Allow-Origin, the snippet fails open, and the room still looks healthy here. |
The room's own targetUrl origin is always on both. tagOrigins defaults to null, which means "use returnOrigins" — every room configured before the split behaves exactly as it did, and you only set it when the two genuinely differ (a tag on a marketing domain that visitors are never returned to, or a narrow return policy that must not widen just to unblock an install).
Set them independently:
curl -X POST http://localhost:8080/api/v1/admin/rooms \
-H "Authorization: Bearer $ADMIN_KEY" -H 'Content-Type: application/json' \
-d '{"id":"checkout",
"returnOrigins":["https://www.shop.example"],
"tagOrigins":["https://www.shop.example","https://landing.brand.example"]}'
Refusals of the second kind are counted per (room, origin) and shown in Diagnostics and in qm_refused_origin_checks_total. Non-zero means a page carrying your tag is live and unprotected right now. The console's install verifier asks about the origin as well as the URL, and its one-click fix writes tagOrigins only — unblocking a tag is not a decision to widen where visitors may be redirected.
Where a visitor is allowed to land (returnOrigins)
A promoted visitor is sent back to the page they came from, which reaches the waiting room as a qm_return URL — an attacker-supplied string. The server refuses every return URL not covered by the room's policy:
- the origin of the room's own
targetUrl, always; plus - any origin listed in the room's
returnOrigins—https://shop.example, orhttps://*.shop.exampleto cover subdomains. The dot in the wildcard is required, so the pattern never matchesevilshop.example, and the scheme and port must match too.
Relative URLs, non-http(s) schemes and URLs carrying embedded credentials (https://user:pass@…) are refused outright. A refused return is not an error for the visitor — they simply land on the room's target instead — so this fails quietly by design. It is counted, and Diagnostics shows the tally as blocked return URLs.
Watch that counter. A missing return origin looks exactly like an attack: if checkout lives on https://shop.example but the room's target is https://www.shop.example, every real visitor is silently dropped on the wrong page. Non-zero means either you have a domain to add, or somebody is trying to use your waiting room as an open redirect.
Scheduled drops
A room can be given a doors-open instant. Before it there is no line at all: arrivals go into a pre-queue and get an opaque handle instead of a ticket number, and at the open instant the whole pre-queue is shuffled with a crypto-grade random draw and becomes the FIFO line.
curl -X POST http://localhost:8080/api/v1/admin/rooms/drop/schedule \
-H "Authorization: Bearer $ADMIN_KEY" -H 'Content-Type: application/json' \
-d '{"opensAt":"2026-09-01T09:00:00Z","preQueueMaxPerIp":16}'
| Field | Meaning |
|---|---|
opensAt | null (no schedule — doors are open), epoch milliseconds, or an ISO datetime. |
preQueueMaxPerIp | Entries one identity may hold in the pre-queue. Default 16; null for unlimited. |
Each field is applied only if you send it. Adjusting the cap mid-drop — -d '{"preQueueMaxPerIp":4}' — changes the cap and leaves the drop scheduled. Cancelling a drop takes an explicit {"opensAt": null}, so a routine tweak can never silently throw the doors open on a queue that was waiting for a draw. A body with neither field is refused (400 nothing_to_change) rather than accepted as a no-op.
Why it works this way:
- Arriving early buys nothing. Clicking the moment the link is published and clicking a minute before the drop draw from the same hat. That is the entire point: a FIFO line that opens hours early just moves the stampede to whenever you published the link.
- The draw is round-robin across identities. The second entry one identity holds is drawn strictly behind the first entry of every other identity, so extra entries can never crowd out first-time arrivals.
preQueueMaxPerIpis a fairness floor, not an anti-fraud control. A whole office or a CGNAT block shares one address, which is why the default is generous. Raise it if your audience sits behind carrier-grade NAT; lowering it punishes shared addresses long before it inconveniences a determined attacker.- The permutation is recorded, not recomputed. A restart across the open instant replays the order that was drawn instead of drawing a second one.
While the doors are shut, POST /api/join answers state:"scheduled", preQueued:true, a preCount, and opensAt/now so the page can count down against the server's clock. position and ahead are null, because neither exists yet.
Note the two views of a scheduled room. GET /api/admin/rooms reports the room's configured state (active, paused, bypass) alongside opensAt and a preQueue count; the visitor wire reports the effective state, scheduled, until the doors open. The dashboard composes the two into a SCHEDULED badge, an "opens …" time, and an in pre-queue counter in place of waiting.
The operator console
The dashboard at / is where an operator lives during an incident. Everything on it is live — a stats frame arrives every 2 s over SSE — and every control applies immediately. There is no save step.
Per room
| Control | What it does |
|---|---|
Active / Paused / Bypass | active queues normally; paused freezes automatic outflow and holds everyone's place; bypass switches the queue off and lets everyone through. Pausing states its consequence before you confirm. |
| Rate slider / number | ratePerMinute, the smooth per-second outflow. Applies on the next slice, not the next minute. |
Let through n | Promote n visitors now, on top of the rate — including on a paused room, which is the one way anybody gets in while paused. Deliberate: releasing a few people during a pause is a normal operator action. It releases the front of the line in ticket order — there is no way to pick a particular visitor. The confirm says which of the two applies. |
Autotune | Drives the rate from the target site's measured response time instead of your guess; the gear opens its bounds. |
Install | The script tag for this room, ready to paste — or, for a room with a proxyOrigin, the inline deployment checklist instead: where DNS points, the internal address, the reserved /__qm prefix and the failover rule. An inline room is never shown a snippet; there is no page to paste one into. Rooms deployed inline also carry an inline chip on the card. |
Waiting page | Opens /w/<room> — exactly what a visitor sees. |
Empty queue | Discards the entire waiting backlog without admitting anybody — for a room that has outlived its event. Nobody is promoted, so the chart gains no spike that never happened; purged tickets read as expired and have to join again. A scheduled room's pre-queue is part of the backlog and is emptied too: those visitors are told their pre-queue entry closed and join again. The room, its settings and its history are kept. Irreversible, and the confirm says so. |
Edit / Delete | Room settings (target URL, TTLs, branding, tagOrigins, returnOrigins, schedule) and removal. |
Reading a room card
CAPACITYalways has a denominator.12/25is twelve of twenty-five slots in use;0/∞is a room with no concurrency ceiling at all. (A bare0under the word CAPACITY, beside a neighbour reading25/25, says the opposite of what it means.) The number counts live sessions on the protected site plus passes issued inside the claim window and not yet used — see Passes vs sessions.EST. WAITis compact and never truncated:40s,3m,1h20m, and past ten hours just109h, because at that point the minutes are not information. A leading≥means the figure is a lower bound, not an estimate — the room is at its ceiling and nothing has come out for a minute, so raising the rate will not help. Hover it for which case you are in.∞under EST. WAIT is a room letting nobody through — paused, or active at 0/min — with people still in it. It is waiting ÷ 0 at the room's current setting, not a forecast: it goes finite the moment you resume or raise the rate, and the tooltip says which of the two it is waiting on. The visitor's own page words the same gap as "no estimate", because a visitor cannot act on the rate and "unbounded" would be a promise about their future; you can, so the card shows you the arithmetic. The state itself is on the badge, once.- The header's WORST EST. WAIT leaves those
∞rooms out, and says so. The top-line number is the worst wait across the rooms that are moving, so one paused room cannot pin the whole estate at∞. Whenever a room with people in it is letting nobody through, the line under the number readsexcludes 3 at ∞: the header and a card reading∞are then both right, and the header is telling you which rooms it is not counting. When every room with a queue is stopped the number itself is∞and the line reads3 not admitting. - A scheduled room counts down. Under the room name,
Doors open in 4m 12s · 12:13 PM, amber inside the last five minutes. The line forms before then — visitors can join and hold a place — it just does not move.
Across rooms
- Chart — queue depth, joins/min and passes/min, plus two wait lines that are deliberately not the same number: est. wait is computed (depth ÷ measured outflow) while actual wait p50/p90 are measured server-side from real join→pass durations. When they disagree, believe the measured one.
- The rate scale can split. Normally joins/min and passes/min share one amber
/minscale inside the left axis, because the gap between the two lines is the rate the queue is growing at. When one is several times the other — a surge at 1,000 joins/min against 20 passes/min — the smaller line would be pressed flat onto the axis, and passes/min is the line you steer by: nudging the dial from 20 to 30 would move it a fraction of a pixel. Past a 4× ratio each line gets its own scale, its ticks printed inside the left axis in its own colour (amber for joins, green for passes), and both legend entries change from(rate scale)to(own scale). While it says own scale, the lines crossing does not mean the numbers crossed — hover for both figures. It reverts on its own once the two are comparable again. - Events — vertical marks for everything that changed the room: rate moves, state flips, threshold activation, autotune, flushes, purges, ejects and protection blocks. Hover one for the note and its clock time; marks landing on the same pixel are collapsed and the tooltip lists all of them. Toggle the series off in the legend when a busy timeline gets in the way.
- Export CSV — the full retained history (1 h), percentiles included.
- URL targeting — the rules deciding which pages this room protects, each with how many door checks it has matched and when it last matched. Change them in
Edit; they take effect on the next check, with no deploy. A rule reading never matched during a live sale is the visible symptom of a room pointed at the wrong pages. The box below it is a dry run: paste a URL, see which rule would win, before you save anything. Warnings appear here when a rule names an origin the room will not return visitors to, and when checks are arriving with no page URL at all (an integration older than qm.js 1.2.0, which cannot be scoped and is therefore queued as if unscoped). - Visitor drill-down — look up either a ticket number or the
Refcode the visitor is reading off their own waiting page: state, position, time waited, which clock is running (pass expires in before they enter the site, session expires in once they are on it), whether they are on the site right now, and whether the ticket is bound to a client. The waiting page shows both the ref and the raw ticket number, so whichever one the caller reads out is one you can paste. (A ref is a short digest, so a very large room can produce a collision; the lookup then lists the candidate tickets instead of guessing.)Ejectrevokes the pass, drops any live session and sends them to the back of the line. It is a two-step confirm; the arming survives the 2 s refresh, and is dropped if you look up somebody else.
state is one of:
| State | Meaning |
|---|---|
waiting | In the line. position and est. wait are theirs. |
passed | Holds a pass that would get them in right now, or is already on the site (on site, with the session clock). |
holding | Shown as holding pass. A valid pass that gave its slot back (claim window lapsed) while the room is full: turned away at the door until a slot frees, oldest ticket first. These are the room card's holding passes. pass expires in is still running; door position appears while they are actually knocking. Not expired, and Eject voids the pass. |
expired | The engine is finished with the ticket: the pass has expired and no session is live (or the line was emptied before their turn). They have to join again. |
ejected | An operator ejected them. |
prequeued | A PQ- ref: in a scheduled room's pre-queue, before the draw. After the draw the same PQ- ref opens the ticket it was drawn into, marked drawn from pre-queue (same room generation only). |
unknown | That ticket number was never issued in this room. There is no position or wait to report. |
A ref also names the room generation it was issued in. Deleting a room and creating it again under the same id restarts ticket numbers at 1, so a ref read out by a caller from before the reset is not resolved to whoever holds that number today: the lookup says which ticket of which earlier generation it was (previous_generation), and that the caller has to join again. The last 3 generations are checked, over their first 50,000 tickets.
The wait field is labelled by tense, and the tense is the point. waiting is a number still climbing — time since they joined the line. waited is the finished journey, frozen at the instant they were promoted; it does not grow while you read it, so a ticket you look up an hour after the incident still answers how long did this caller wait? with the wait, not the hour. For a room with opensAt, the clock starts when the visitor joined the pre-queue, not when the doors opened — the hour they spent holding is real waiting and every number here says so, including the Actual wait series on the chart and the percentiles in Reports. joined is that instant. passed is when they were let through; it stays after the pass itself expires. A session admitted by an older version has no record of it once the pass expires, and shows passed: unknown rather than a guess.
After a reclaim (the holder takes their session back from a second tab or device), passed and waited still read the original promotion, not the re-grant the reclaim issues.
When the engine has finished with the ticket (the pass has expired and no session is live), state becomes expired, and joined, passed and waited are kept. That covers a visitor who was promoted and never came in, too. The last 50,000 finished tickets per room are kept, oldest dropped first. An eject forgets them straight away.
- Share monitor — a read-only link for stakeholders. It carries a scoped viewer token, not
ADMIN_KEY, so it can observe but never change. The console asks the server what the signed-in credential may do and hides every mutating control accordingly, so the read-only view survives the#monitorfragment being lost in a paste — and a namedviewerkey gets the same treatment. The header names which credential you are signed in with.
Share monitor shows a live read-only link, never minting on open (Mint a new link creates one), in a selectable field — copy it from there, or with the button beside it. Minting a second link never revokes the first, so the dialog says so and offers the existing link first. Copy on a row in the list below puts an existing link back in front of you rather than creating another, so the number of live links stays the number of people who are meant to have one. It is audited as monitor_token.reveal, separately from monitor_token.mint, because who else was sent this is a different question from who made it. A revoked or expired link cannot be re-copied: that would be a way to resurrect exactly the credential somebody just took back. Links expire on their own after MONITOR_TOKEN_TTL_SEC — seven days by default; revoking one is separate and takes effect immediately. Revoking an operator closes every link that operator shared (see Operators, roles and the audit trail).
- Alerts — desktop notifications for surge start, target-URL failure and stalled outflow.
- Diagnostics — the health report described above.
Where the controls live. The header bar holds only what you reach for while a queue is moving: Alerts, Share monitor, Security, Diagnostics, and + New room. Everything you set up between incidents is one click away under Console — Reports, Access, Docs and Sign out. The menu is drawn over the page rather than in it, so opening it never moves a button you were aiming at. On a read-only credential the bar and the menu both shrink to what that credential can actually read, which for a shared monitor link is Security and Reports.
The console never stores ADMIN_KEY: signing in trades it once for an HttpOnly qm_console session cookie (POST /api/admin/session, see the QM-349 note above), and Sign out (under Console) sends DELETE /api/admin/session, which clears the cookie and revokes the session server-side.
Running it from the keyboard
A busy console is well over a hundred tab stops, and a room card is about eighteen of them, so reaching the sixth room by tabbing is not a real option.
- Skip links. The first
Tabon the page offers Skip to rooms and Skip to chart. They are off-screen until focused. - Skip past a room. The first stop inside every room card is **Skip past name to the next room**, so moving down the list costs one
Enterper room instead of eighteen tabs. The last card's skip goes to the chart. The target is worked out when you press it, so it follows the live card order. - The rate dial.
←/→move it by 10/min,PageUp/PageDownby 50/min, andShift+←/→by 1/min. The number box beside it is still the way to type an exact figure. (The underlying step is 1, so a rate the server sends is always shown exactly — only the keys move in jumps.) - The Console menu.
Enteror↓opens it,↑/↓move through it and wrap,Home/Endjump to the ends,Escapecloses it and puts focus back on the button.Tableaves and closes it rather than cycling inside. - Dialogs return you where you were. Closing a dialog — by
Escape,Close, or clicking outside — puts focus back on the control that opened it, including a confirm opened from inside a panel: the firstEscapereturns to the control in the panel, the second to the button that opened the panel. If the room a dialog belonged to has been deleted meanwhile, focus lands on that room's card heading, or on the room list — never on nothing.
Metrics and monitoring
| Endpoint | Format | Auth |
|---|---|---|
GET /metrics | Prometheus text exposition | ADMIN_KEY |
GET /api/v1/admin/metrics?roomId=&range= | JSON time series for charts | ADMIN_KEY |
GET /api/admin/health | JSON health, config and warnings (incl. absolute startedAt) | ADMIN_KEY |
POST /api/admin/shutdown | Clean stop of the instance that answers — the only one Windows has off-console | owner / ADMIN_KEY |
GET /healthz | Readiness: 200 while ready, 503 when storage is failing (STORE=memory; on STORE=valkey only when this instance fails while another is ok), the process is draining, or this instance is at MAX_CONNECTIONS | none |
GET /api/version | The same payload, always 200 | none |
/metrics is not public — it answers 401 without the key, because room ids and live queue depths are commercially sensitive. Give Prometheus a bearer token:
scrape_configs:
- job_name: queue-manager
metrics_path: /metrics # inline, on the public hostname: /__qm/metrics
authorization: { type: Bearer, credentials: "<ADMIN_KEY>" }
static_configs: [{ targets: ["qm.weekday100.com"] }]
metrics_path is spelled out because Prometheus defaults it to /metrics, and on the public hostname of an inline room that path belongs to the customer's site: the visitor gate answers the scrape 503 and the job goes up=0 while the server is perfectly healthy. A scrape target is a host and port with no room for a prefix, so the path is the only place to say it — see the probe table under Readiness. Scraping the process directly (pod IP, container port) needs no change.
Per-room series are labelled {room="…"}: qm_room_waiting, qm_room_active_sessions, qm_room_pending_entry, qm_room_occupancy, qm_room_max_concurrent, qm_room_rate_per_minute, qm_room_passed_last_minute, qm_room_oldest_wait_seconds, qm_room_state. Process-wide: qm_rooms, qm_sse_subscribers, qm_uptime_seconds, qm_shutting_down, and the counters worth alerting on — qm_storage_write_errors_total, qm_unknown_room_checks_total, qm_blocked_return_urls_total, qm_refused_key_checks_all_total, qm_rate_limited_checks_all_total and qm_rejected_tokens_all_total.
Several instances (STORE=valkey). Every series also carries instance="<QM_INSTANCE_ID>", and each instance reports only its own process: its own counters, connections and SSE streams (the room gauges read the shared Valkey, so every instance reports the same queue). Scrape each instance and sum in Prometheus. Prometheus sets an instance target label of its own, so with the default honor_labels: false this one arrives as exported_instance; set honor_labels: true on the job to keep the server's name. With honor_labels: true the name must be unique per scrape target: give each instance its own QM_INSTANCE_ID, or leave it unset (i:<random> is new on each boot). Two targets that report the same name write the same series, and Prometheus drops or mixes their samples. Where the platform exposes one address for the whole fleet (DO App Platform), a scrape reaches whichever instance the load balancer picks. The console is the fleet view there: each instance writes its origin-error, protection and diagnostic counters to Valkey every 5 s (<prefix>ctr:<id>, expiring after 15 s), and the console's room cards and GET /api/admin/health show the sum over the live instances. A stopped instance's share drops out of that sum within 15 s. Health's fleet: {instances, partial} says how many instances the sum covers, and partial: true when the other instances' counters could not be read from Valkey (the sum is then this instance's own, plus what it last read within 15 s). Connection and SSE counts in health stay this instance's own, beside its own limits.
Alert on the rate of that last one. A trickle of refused tickets is stale bookmarks; a step change means everyone in a room was ejected at once, which is what a regenerated signing key or a moved clock does — and it is the one failure with no other symptom anywhere. Those visitors are not shown an error; they are quietly put at the back of the line.
The one alert you should not skip is storage: a server that cannot append to its event log must stop handing out tickets, and it does (503 storage_unavailable), but you want to hear about it before your visitors do.
Alerting (paging someone who is not looking at the dashboard)
The console's Alerts button raises desktop notifications in that tab. That is useful while somebody is watching and worthless at 3 a.m. — it dies with the tab and exists once per operator laptop.
Everything below runs inside the queue process. When the process is down, wedged or unreachable, nothing here fires — no webhook, no resolved, nothing in the panel. The only thing that pages you then is an external monitor on /healthz (see Readiness); set one up, because it is the only alert that survives the failure it reports.
The conditions below are evaluated on the server's own clock every ALERT_EVERY_MS whether or not a webhook is configured, and whatever is firing is listed under Diagnostics → Alerting → Firing now either way. With nothing configured that panel is the alerting: the conditions are real and nobody is being paged for them, which the panel says in as many words. Several of them — queue_frozen, outflow_stalled — depend on hold times no dashboard replicates, so this is the only place they appear.
Set ALERT_WEBHOOK_URL and the server also POSTs JSON:
{
"source": "queue-manager",
"event": "outflow_stalled",
"status": "firing",
"severity": "critical",
"roomId": "shop",
"message": "Shop: 412 waiting and nobody let through in the last minute",
"ts": 1730000000000,
"since": 1729999880000,
"text": "[queue-manager] CRITICAL: Shop: 412 waiting and nobody let through in the last minute"
}
text is there so a Slack incoming webhook renders something readable with no transform; the structured fields are for PagerDuty Events v2, Opsgenie and anything else that parses. Events:
event | Severity | Fires when |
|---|---|---|
storage_failing | critical | The event log cannot be written, so joins are being refused. On STORE=valkey this is events-journal commits to Postgres failing in two alert frames within 3 × ALERT_EVERY_MS (one failed commit does not page), until a later commit has landed and that window has passed since the newest failure. With several instances the message names the instance(s) whose storage is failing. |
valkey_unreachable | critical | With STORE=valkey: this instance has not reached Valkey for VALKEY_ALERT_AFTER_MS. Sent by each such instance itself, not by the leader. |
target_down | critical | The server's own probe stops reaching a room's target site. |
outflow_stalled | critical | People are waiting and nobody has been let through for ALERT_STALL_SEC. The room is trying and failing. |
queue_frozen | critical | People are waiting and the room is paused or set to 0/min for ALERT_FROZEN_SEC — the forgotten paused room. The room is not trying, because a human said so and may not have come back. |
queue_deep | warning | Depth at or above ALERT_DEPTH (off by default). |
wait_high | warning | Estimated wait at or above ALERT_WAIT_SEC. |
Every condition must hold for ALERT_HOLD_SEC before it fires, re-pages every ALERT_REPEAT_SEC while it is still true, and sends a "status": "resolved" event when it clears — an alert channel that never says "it is over" is one operators learn to ignore.
With STORE=valkey and several instances, every instance evaluates the conditions, so Firing now and GET /api/admin/alerts are the same on any of them, and only the leader (see LEADER_LEASE_MS) POSTs to the webhook. The state that decides a delivery (when a condition started, when it was paged, when it was last re-paged) is kept in Valkey (<prefix>alerts:state), so when the leader dies mid-incident the next one neither pages the open incident again nor forgets to send its resolved. Do not delete <prefix>leader:epoch by hand: the stored state is fenced on it, so if <prefix>alerts:state survives, the epoch restarts below the one saved with that state and every later save of the alert state is refused until the epoch passes the stored one again (or <prefix>alerts:state is deleted too). If you must reset the epoch, delete both keys together (an alert that is still firing is then paged again). Storage is each instance's own: every instance reports its storage status to Valkey (<prefix>alerts:storage) each frame, the leader pages storage_failing when any instance reported failing within the last 3 × ALERT_EVERY_MS and names it (instance on /api/admin/health is each instance's name), and resolves it only once every instance reports healthy. While Valkey is unreachable nothing is delivered except valkey_unreachable, which each instance that cannot reach Valkey sends itself after VALKEY_ALERT_AFTER_MS (so expect one page per instance); an external monitor on /healthz is still the alert that survives the queue itself failing. The delivery log, sent and failed are the answering instance's own: the leader's list is the one with the deliveries in it.
ts is when this page was decided and since when the condition began, which a firing and its resolved share. A resolve can reach the webhook before its firing while the firing is still in flight (ALERT_EVERY_MS below the 5 s webhook timeout, or a receiver that accepts late), so order pages by ts and pair them by event, roomId and since.
GET /api/admin/alerts reports the configuration, what is firing and the last 20 deliveries with their HTTP results. POST /api/admin/alerts (write grant) fires a test event; the console does this from Diagnostics → Alerting → Send test alert, which is the only way to know the path works before you need it. The button is always in that panel — with no ALERT_WEBHOOK_URL set it is there but inert, and says so, rather than not existing at all for the one reader who is following these instructions because nothing is configured yet. It is disabled, saying why, for a credential without the write grant (a share link or a named viewer key), which the server would answer with 403.
Durability
With STORE=valkey the durable record is the events journal in Postgres. Valkey holds the live line, and a room whose Valkey state is lost or rolled back is rebuilt from the journal (see When the backend is down). With STORE=memory nothing is written anywhere: every wait below is met at once, a restart starts empty, and the 503 answers below cannot happen.
- Joins are acknowledged only after the join event is committed to the journal. A failed commit makes
/api/joinanswer503 storage_unavailablerather than hand out a signed ticket nothing recorded, and the ticket (or pre-queue entry) it minted is taken back: it does not stay in the line, does not count againstqueueMaxPerIp/preQueueMaxPerIp, and does not come back in a rebuild. - The same holds for everything else a restart must not take back (QM-390): the pre-queue draw, an eject, a purge and a room's definition.
/api/statusand/eventsanswer only once the draw they reveal is committed, and eject, purge, room create/update/delete, schedule and flush answer200only once their event is. If the commit fails they answer503 storage_unavailable(/eventsholds the frame instead). Retrying an eject or purge is safe. When nothing is pending the check costs one map lookup. - Admissions at the door are held to the same rule (QM-431).
/api/check(direct and vouched) answers apassthat admits somebody only once the room's admissions are committed, and the inline gate forwards the page or the WebSocket upgrade only then. If the commit fails the answer is503 storage_unavailable(a refused upgrade inline), and the admission is taken back: the pass is pending again with the slot and claim window it had, so the retry is admitted on a fresh door key instead of being toldsession_elsewhere. The snippet and the connectors fail open on that503, as on any 5xx. The wait is per room and separate from the one/api/statusand/eventsuse, so they never wait on the door. - A refused join or admission can still come back as an orphan if the process dies before its compensating event is committed, or that commit fails too. An orphaned ticket lapses like a closed tab: it counts against
queueMaxPerIpuntil it reaches the head of the line, is promoted without a capacity slot once its holder has been unseen forpresenceSec, and expires with its pass. An orphaned pre-queue entry counts againstpreQueueMaxPerIpuntil the open and is then drawn into such a ticket. An orphaned admission holds a slot until the session idles out, and its visitor, whose door key it does not know, is toldsession_elsewhereand rejoins unless their queue cookie reclaims it. - Ticket ownership, admissions and ejections are all in the journal, so a
kill -9cannot turn a leaked token into a free ride. GET /api/admin/healthexposesstorage.ok,storage.writeErrors(onSTORE=valkey, the journal commits that failed) and the last error; alert on it, or on thestorage_failingalert.- Checkpoints keep a rebuild short. One instance at a time folds each busy room's journal into a
room_snapshotsrow, so a rebuild replays the latest checkpoint and the events after it instead of the room's whole life. A checkpoint is read from Postgres only, never from Valkey, so it always agrees with the journal. - The journal is guarded per event against a rollback: a rebuild that meets an event type this build does not know, or an event stamped with a newer format (
lv, absent means 1), fails instead of skipping it, and the room stays fenced until a build that can read it runs the rebuild. For whoever changes the format: a new kind of event gets a new type; changing what an existing type means stamps those events with the nextlv(EVENT_LOG_VERSIONinlib/engine.js). - The hourly history is one row per room and hour in Postgres. The leader rewrites the open hour on every sample, and rows older than
HISTORY_RETAIN_DAYSare deleted. A warehouse that cannot be read at boot is logged ([history]) and refuses the boot rather than start on empty state and overwrite it.
Clocks
Outflow pacing uses a monotonic clock, so an NTP correction cannot stall or burst promotions; recorded timestamps use the wall clock but never move backwards, which is what keeps the log and the prune order consistent.
What a clock step does cost is tickets. Every ticket carries the instant it was issued and is refused if it was issued more than one minute in the future or is older than TOKEN_MAX_AGE_SEC (24 hours by default). That check is what stops a ticket from last week being replayed, and it is not negotiable. So a host whose clock moves (a VM booting with a bad RTC before NTP steps it, a machine resumed from a snapshot, a container inheriting a wrong clock) refuses every ticket in the queue at once, and each of those visitors silently rejoins at the back of the line.
The server cannot prevent that, so it names it: a step of more than five seconds is logged, listed in incidents for the following hour, and sets status to degraded on /healthz. The probe still answers 200: taking the pod out of rotation does not fix the host's clock (see Readiness). Watch qm_rejected_tokens_all_total for the size of it. Keep NTP running on the host, and prefer a slew to a step.
No public-internet dependency
Every surface this server hands out — the console, the waiting page and /docs — renders completely with no request to any origin but this one. There is no CDN, no analytics, no external stylesheet and no external script. You can run the whole product on an airgapped network, behind a corporate proxy that blocks everything unlisted, or during the DNS outage that is the reason somebody opened the console in the first place, and it will look exactly the same.
That mattered enough to stop loading the typeface from Google Fonts. Geist and Geist Mono are checked into public/fonts/ and served by this process at /fonts/<file>.woff2:
| File | Face | Subset |
|---|---|---|
geist-latin.woff2 | Geist, variable 400–700 | latin |
geist-latin-ext.woff2 | Geist, variable 400–700 | latin-ext |
geist-mono-latin.woff2 | Geist Mono, variable 400–600 | latin |
geist-mono-latin-ext.woff2 | Geist Mono, variable 400–600 | latin-ext |
- About 82 KB total, and a page only fetches the subsets it actually shows — an English waiting page pulls two files, not four.
- Served
Cache-Control: public, max-age=604800, so reloading a dashboard mid-incident does not refetch them. They are the one asset here that is pure content: identical in every deployment and carrying nothing about yours. - Sent uncompressed on purpose. woff2 is Brotli internally; gzipping it again costs CPU and adds bytes.
- Latin only. Thai, and every other non-Latin script, falls through to the system face — as it did on the CDN, whose subsets did not cover Thai either. The waiting page in Thai renders correctly; it just renders in the visitor's system font.
- Licensed SIL OFL 1.1 (Vercel / basement.studio). The licence ships beside the files in
public/fonts/LICENSE.txt— keep it there, redistribution requires it.
To update them, replace the four files, keep the names, and re-run the tests: test/surfaces.test.js fetches each one over HTTP and checks the wOF2 magic and the byte count, so a truncated or mistyped file fails the suite instead of silently falling back to Arial in production.
Running the HTTP tests
Every HTTP suite starts its server through startServer() in test/_harness.js, which runs node server.js with PORT, HOST=127.0.0.1, ADMIN_KEY, SECRET, the limit knobs set to 0 (off), plus whatever the suite adds (see the environment table above). Each server gets a temporary directory of its own (dataDir). Nothing is written there since step 5; it names the server's Valkey run below, so a restart on the same one sees the same state.
On the Valkey store. With STORE=valkey in the test process's environment (and TEST_VALKEY_URL, TEST_DATABASE_URL set), startServer() runs every server on STORE=valkey + SIDESTORE=pg: the test Valkey under a key prefix of its own and a Postgres schema of its own per dataDir, both deleted after the file. It never uses VALKEY_URL or DATABASE_URL from the environment, and refuses Valkey db 0 and a database whose name does not end in _test:
STORE=valkey npx vitest run test/conformance-http.test.js
Which suites pass on Valkey today, and why the others do not yet, is in docs/superpowers/plans/2026-09-25-valkey-step3a-http.md. The tests whose point is state surviving a restart (test/_valkey-only.js) run only on Valkey, since memory mode keeps nothing across a restart. Without TEST_VALKEY_URL and TEST_DATABASE_URL they are skipped, with one warning per file that names each skipped test; with QM_REQUIRE_VALKEY=1 (CI) they fail instead.
Against another server command. QM_SERVER_CMD (test-only, read by test/_harness.js) replaces node server.js with another command, to run the same HTTP suites against another implementation or build. It is split on whitespace (double quotes group a token) and spawned with no shell, with the same environment as above. It must be the server process itself, not a wrapper, because suites kill it and read its exit code, and it must print a line containing listening to stdout once it accepts connections. Unset, the harness runs node server.js from this repo.
test:conformance runs only the files that talk to the server over HTTP and nothing else: no require('../lib/...'), no assertions on its stderr, no decoding of ticket numbers out of a token. These files are: abuse, admin-host-gate, conformance-http, contract, edge-connector, failover, forwarded-proto-trust, forwarded-trust, identity, origin-down, prequeue-forwarded, proxy, snippet, stream-generation, surfaces, timeline and visitor-support. A file with even one white-box block stays out, so a white-box test goes in a sibling <name>.internal.test.js (as the ones split out of proxy, abuse, visitor-support, failover, identity and edge-connector did). The list is the test:conformance script in package.json; a new file that meets the bar goes there.