Queue Manager docs

The addresses on this page were filled in with this server's own address, https://qm.weekday100.com.

Queue Manager — Operations Guide

Everything you have to decide before putting this server in front of real traffic — configuration, the proxy/CDN setup that per-IP limits depend on, the API keys, and how a pass is bound to a visitor — and everything you use once it is there: the operator console, scheduled drops, and metrics.

INTEGRATION.md documents the customer-site side (script tag, /api/check). This file documents the server side. Both are served by the running server at /docs, and linked from the dashboard header. Thai translations are served at /docs/th (sources in docs/th/); the English files are authoritative.


Environment variables

Everything is optional; defaults are safe for a single-instance install on localhost, and deliberately not safe to expose unchanged.

VariableDefaultWhat it does
PORT8080Listen port.
HOST(unset)Listen address. Unset binds every interface — except while ADMIN_KEY is the default, when unset means loopback (127.0.0.1) only, and a HOST that is not loopback refuses to start (see Refusals). Set it to 127.0.0.1 to answer only on loopback — worth doing when a reverse proxy on the same box is the only thing that should reach this process.
DATA_DIR(removed)Removed in step 5: nothing is written to local disk. Set, it is ignored with a boot warning (DATA_DIR is set but ignored since step 5: nothing is written to local disk.), so an old manifest still deploys. Delete it from the manifest.
ADMIN_KEYadmin-devPrivate API key: /api/admin/* (create/delete rooms, rates, flush, eject, metrics). Never distribute it. While it is the default the server listens on loopback only and refuses a non-loopback HOST — the default is printed in these docs, so it must never reach another machine.
PUBLIC_API_KEY(unset)Public API key: the only thing it unlocks is vouching for an end client on /api/check?ip=&agent=. Sent as Authorization: Bearer only — ?key= is refused (401 credential_in_query). Safe to ship to every edge server / backend. Without it those calls need ADMIN_KEY, which is a privilege escalation — a warning says so.
ADMIN_HOST(unset)Comma-separated hostnames the operator console and /api/admin/* answer on. Unset, a snippet-only install answers anywhere; once a room is inline an unset value means the console answers only on an address literal (127.0.0.1, the pod IP), on localhost or on the PUBLIC_URL host (QM-440), and on any other Host the console surface and /__qm/ are a bare 404 while the queue itself keeps answering, exactly as with ADMIN_HOST set (QM-428; the customer's own hostname still serves the console under /__qm/). Set this to put the console on a real name: set, it is the only place the console answers — on every other hostname (and on the bare pod address) the console surface answers 404: /, /api/admin/*, /metrics, /docs, /progress and the console's own file /public/admin.html, and on a hostname that fronts an inline room the same paths under /__qm/. It gates only the console: the queue itself (/healthz, /snippet/qm.js, /w/<room>, /api/check, /api/join, status, events, notify, the waiting page) keeps answering on any hostname, so the snippet host and a probe by pod IP work unchanged. A room may not be put inline on an ADMIN_HOST name or the PUBLIC_URL host (400 reserved_host), and the console surface on those names is never handed to a room, even one saved before that was refused. Unset, the address literals and localhost are the console's too, so their /api/admin/* and /api/v1/admin/* are never handed to a room either, even one whose targetUrl is the server's own address (QM-439); the rest of such a host stays the room's, which is how inline mode runs on a laptop. And on any host, the proxy never forwards this server's own credential: an Authorization: Bearer carrying ADMIN_KEY or an operator key is dropped before the request reaches the origin, while the app's own Authorization passes unchanged. See Two deployments.
ADMIN_AUTH_FAIL_PER_MIN20Failed admin authentications allowed per client address per minute (token bucket). Past it, every credential from that address — the right one included — gets 429 too_many_auth_failures with Retry-After until one failure's worth has refilled. Only failures spend it. 0 disables. See Two keys, two blast radii.
SECRETrandom per boot (memory); required with SIDESTORE=pgRoot signing secret for visitor tickets, monitor links, console sessions and operator-key verifiers. By default they are signed with it directly, as in earlier releases; with TOKEN_SIGN_V2=1 each purpose signs with its own key derived from it (HKDF-SHA256) and every token names it by a kid. Both forms are always accepted. Set it if you run more than one instance, or tokens minted by one are rejected by the others. Unset, a new key is made on every boot, so a restart voids every ticket, monitor link, console session and operator key (fine for development, where nothing else survives a restart either). Required with SIDESTORE=pg: the server refuses to start without it, because a new key on every deploy voids every pass (tickets, monitor links, operator keys) that Postgres still holds. See Signing keys and rotating SECRET.
SECRET_PREVIOUS(unset)The SECRET being rotated out. Tokens signed under it still verify (either form, every purpose); new tokens are always signed under SECRET. Operator keys resolved under it are re-keyed to SECRET on use. Remove it once the rotation has settled — /api/admin/health lists it under hardening while it is set. See Signing keys and rotating SECRET.
TOKEN_SIGN_V201 signs tokens and new operator verifiers in the v2 form: a key per purpose and a kid. 0 (default) signs the legacy form, raw SECRET and no kid, which the release before per-purpose keys also reads. Both forms are accepted either way. Once you turn it on, you can't roll back below this release without losing the queue — see Signing keys and rotating SECRET. Any value other than 0/1 counts as 0 and is named in warnings.
TRUST_PROXY0Number of reverse proxies in front of this server. The client address is taken from X-Forwarded-For counted from the right — and only when the connection comes from an address in TRUST_PROXY_IPS. 0 means the socket address is the client.
TRUST_PROXY_IPS(empty)Comma-separated addresses or CIDR ranges of those proxies, IPv4 or IPv6 (10.0.0.7, 173.245.48.0/20, 2400:cb00::/32). An IPv4-mapped peer (::ffff:1.2.3.4) matches IPv4 entries. An entry that does not parse stops the server at startup, naming the entry. From any other peer, X-Forwarded-For is ignored. Required to get per-visitor throttling behind a CDN (see below).
EDGE_SECRET(unset)The secret a Cloudflare request-header rule adds as X-QM-Edge (see Cloudflare edge secret). Set, every request and WebSocket upgrade without a matching header is refused 403 (edge_required), except /healthz; the client address becomes cf-connecting-ip and TRUST_PROXY/TRUST_PROXY_IPS are ignored. Comma-separated for rotation (any listed value passes). An entry shorter than 16 characters stops the server at startup. Unset with STORE=valkey or NODE_ENV=production, a startup line says bypassing requests are accepted.
IPV6_PREFIX_BITS64What "per address" means for an IPv6 client: its first N bits (32..128). An IPv6 host is normally handed a whole /64 and may use any address in it, so every per-address limit — the join, check, notify, page and SSE budgets, the failed-login lockout, queueMaxPerIp, preQueueMaxPerIp and the traffic-classification record — is keyed on the /64, not the /128. 128 restores one key per address. IPv4 and IPv4-mapped (::ffff:1.2.3.4) clients are unaffected. Only the key changes: logs, the audit trail and the console still show the full address, and TRUST_PROXY_IPS still matches the real peer exactly or by its own CIDR.
JOIN_LIMIT_PER_MIN600Per-address /api/join budget. 0 disables. Per instance with STORE=valkey: each instance keeps its own count, so N instances allow one client up to N× this across the fleet; Cloudflare's edge rate limiting is the fleet-wide control.
CHECK_LIMIT_PER_MIN6000Per-address /api/check budget. 0 disables. Per instance with STORE=valkey: each instance keeps its own count, so N instances allow one client up to N× this across the fleet; Cloudflare's edge rate limiting is the fleet-wide control.
RATE_LIMIT_MAX_KEYS200000Most distinct addresses each rate-limit bucket map holds. Past it the least recently used are evicted — never the whole map, so key churn cannot hand a throttled client a fresh budget. The ADMIN_AUTH_FAIL_PER_MIN lockout is held the same way.
WAITING_LIMIT_PER_MIN120Per-address budget for the waiting page itself, and — in a bucket of its own — for the console page, /progress and /docs (also under /__qm/). 0 disables both. Stands down automatically for an undeclared proxy's address (that address only) — see Rate limits. Per instance with STORE=valkey: each instance keeps its own count, so N instances allow one client up to N× this across the fleet; Cloudflare's edge rate limiting is the fleet-wide control.
SSE_PER_IP200Concurrent /events streams per address. 0 disables. Sized for a whole office or CGNAT block sharing one address; SSE_MAX_TOTAL is the cap that protects process memory. Per instance with STORE=valkey: each instance holds its own ceiling, so N instances hold up to N× this in total; Cloudflare is the fleet-wide control.
SSE_MAX_TOTAL5000Global SSE ceiling (visitors + admins). Sized by CPU, not memory: an update that moves every visitor costs one kernel send per open stream, about 0.1 ms each on the reference box (QM-381), so 5,000 streams is ~0.5 s of CPU per update against the 1 s engine tick and 20,000 would be ~2 s. The fan-out runs in slices of 256 streams, so a large one no longer stalls the event loop, but above this cap each visitor simply gets fewer updates. A visitor refused a stream (503 sse_capacity) is not turned away: the waiting page polls /api/status every 5 s instead. Raise it only on a faster box, after measuring. Frames whose content has not changed are not re-sent, except as a keepalive to a stream that has been quiet for 4 s. Per instance with STORE=valkey: each instance holds its own ceiling, so N instances hold up to N× this in total; Cloudflare is the fleet-wide control.
ADMIN_SSE_RESERVE8Extra stream slots above SSE_MAX_TOTAL that only a console holding a write-capable credential may use — ADMIN_KEY, or a named owner/operator key — so the dashboard can still connect when visitors have filled the global ceiling — which is exactly when you need to look at it. A monitor share link is a read-only credential and gets no exemption: it competes for SSE_MAX_TOTAL like a visitor.
MAX_CONNECTIONS50000Soft socket ceiling, enforced by this process itself: at and above this many open connections, every NEW request is answered 503 (see qm_connections_shed_total) instead of being routed or proxied.
CONNECTION_RESERVE256Headroom ABOVE MAX_CONNECTIONS before Node's own hard ceiling (server.maxConnections) is reached. Without it, the socket that tips the count past MAX_CONNECTIONS has nowhere left to be answered from — Node destroys it in C++ before any 503 can be sent.
TOKEN_MAX_AGE_SEC86400How long a signed token stays verifiable at all. A visitor still waiting keeps their place past it (QM-391): once their token is half the shorter of this and QUEUE_COOKIE_MAX_AGE_SEC old (6 h with both defaults), /api/status and /events re-sign the same place with the current keys and re-set both queue cookies. Only for the browser whose qms_<room> cookie holds the place, in the room's current generation, while the place is still waiting or pre-queued; a pass, a lapsed or ejected place, or a caller matched only by fingerprint gets nothing. A token already past this age is still refused. Lowering it therefore does not cut a long wait short for a visitor whose page is open; it bounds how long a token copied out of a log stays usable.
REFRESH_ENDS_PER_SEC200At most this many open /events streams are ended per second for their token refresh (QM-391). Streams that fall due together wait their turn, still receiving status frames, instead of all reconnecting in the same instant. The drain goes faster only when this rate would not end every waiting stream before its token lapses.
MONITOR_TOKEN_TTL_SEC604800How long a Share monitor link stays live before it expires on its own — seven days by default. Expiry is independent of revocation: a link can be revoked earlier from the console, and an expired one cannot be re-copied.
EMAIL_NOTIFY(unset)1 (also true/yes/on) enables POST /api/notify, the turn-notification capture, and the e-mail field on the waiting page. Off by default and the page hides the control entirely when it is off, because this server records the intent and sends nothing itself — offering a channel that goes nowhere is worse than not offering one. With it off the route answers 404 email_notify_disabled. When on, an address is kept with its place in line and deleted when the ticket ends, or up to LAPSED_TTL_MS (24 h) later if its entry window lapsed: the ticket ends when it is admitted, ejected or purged, or its room is deleted; after a lapse, a visitor who comes back keeps the address on the new ticket. No entry is kept more than 24 h after it was filed, and a request against a ticket that has ended is refused (409 ticket_ended). With STORE=memory it is in this process's memory, and every restart or deploy loses all of them. With STORE=valkey it is in the shared Valkey (qm:notify:{room}), so the console on every instance sees it. Valkey RDB snapshots and provider backups can hold a copy after deletion until they rotate out. It is never written to Postgres (events or audit), the logs or /metrics, never sent by this server, and never shown to anyone unmasked (the console's visitor drill-down sees al•••@e••••••.com). Nothing in this build can export them. Turning it on raises a warning at boot and in /api/admin/health warnings, so the console badge says so too; the waiting page tells the visitor the same.
NOTIFY_LIMIT_PER_MIN20Per-address POST /api/notify budget. A visitor sets an address once and maybe corrects it once; more than a handful a minute from one address is a script. 0 disables. Per instance with STORE=valkey: each instance keeps its own count, so N instances allow one client up to N× this across the fleet; Cloudflare's edge rate limiting is the fleet-wide control.
NOTIFY_MAX_ENTRIES50000Hard ceiling on stored notification intents, so an address book cannot be pushed into this process until it runs out of memory. Oldest entries are evicted first. STORE=memory only: with STORE=valkey an entry needs a signed ticket, so a room's hash is bounded by the places it has handed out, and each entry goes when its ticket ends.
QUEUE_COOKIE_MAX_AGE_SEC43200Lifetime of the queue cookies (qm_<room>, qms_<room>), counted from when they were last set. A waiting page that stays open has them re-set with its token refresh (see TOKEN_MAX_AGE_SEC); a browser that closes the page for longer than this loses the qms_<room> proof and, with it, the place.
COOKIE_SECURE(unset)1 forces Secure on cookies even when this server sees plain HTTP (TLS terminated upstream by a proxy that sends no X-Forwarded-Proto, or is not on TRUST_PROXY_IPS). Without it, cookies are still Secure on a request whose Host is the host of an https:// PUBLIC_URL.
PROXY_TIMEOUT_MS30000Inline mode only: how long to wait for the app behind a room's proxyOrigin before answering 504. See Two deployments below.
PROXY_ALLOW_PRIVATEunset1 lets a room's proxyOrigin (and the health probe of it), and the health probe of a snippet room's targetUrl, reach loopback, RFC1918, link-local/cloud-metadata and other private addresses, which are refused by default. A refused probe reads down, with targetHealth.blocked naming the address; nothing is sent to it. Autotune makes no decision on a refused probe and holds the rate; after 3 refused in a row it undoes its own raise, once, back to the rate the room had before autotune raised it (a cut is held, and a rate an operator set is left alone). The base is kept in the room's config (written with the rate), so it survives a failover or a restart: the new leader falls back the same way. An operator's rate set while the fallback is being written is not overwritten. Set it only when the app really is on the LAN. See Two deployments below.
PROXY_IDLE_TIMEOUT_MS60000Inline mode only: how long a response that has already started may go without a byte before both sockets are cut. A gap, not a total — a stream or a download is never cut while it is delivering. Lower it if your app has no long-lived streams.
HEADERS_TIMEOUT_MS20000Slowloris guard.
REQUEST_TIMEOUT_MS60000Whole-request timeout.
PROTECTION_MODEmonitorServer default for bot scoring: off, monitor (score and count, refuse nothing) or enforce (refuse a place in line above the threshold). Per-room protection.mode overrides it. Start on monitor and read the Security panel before you enforce anything.
PROTECTION_WARN_AT40Score at which a request is counted as a warning.
PROTECTION_BLOCK_AT70Score at which enforce refuses a place in line. Raising it makes the classifier more forgiving; lowering it below ~55 starts putting circumstantial-only clients in reach of a block.
PROBE_EVERY_MS12000How often the server probes each room's target itself, on a fixed clock — measurement and autotune do not stop when a dashboard tab closes, and every operator sees the same numbers. Drives the target_down alert and the response-time figure. 0 disables probing; otherwise the minimum is 1000 — a lower value is raised to 1000 with a startup warning. Each distinct URL is probed once per interval however many rooms share it, and the probes are spread over the first 80 % of the interval.
PROBE_TIMEOUT_MS8000How long one of those probes waits before the target counts as down. The probe only times the response headers and discards the body. A status under 500 counts as reachable; a 5xx counts as down, along with a network error or a timeout — an app answering every request with 500 is not up, and "something is listening" is not the question you are asking during an incident.
PROTECTION_ALLOW_IPS(empty)Comma-separated addresses that are never scored — your own synthetic monitoring and load tests belong here.
SHUTDOWN_GRACE_MS250How long an idle or SSE connection gets before it is cut. A socket that is mid-proxy-response (inline mode) is never cut here — only the hard ceiling below applies to it — so a deploy stops severing a customer's own page mid-transfer.
SHUTDOWN_DRAIN_MS2000 (15000 when any room is inline)Hard ceiling on the whole stop. A stuck socket must never turn a deploy into a hang. An inline deployment holds proxied responses to a longer floor by default — 2000ms was measured truncating a real proxied download on every deploy — but an operator who sets this explicitly always gets exactly that value, inline or not: with SHUTDOWN_DRAIN_MS=3000 set explicitly on an inline install and a hung origin, the process used to still drain for the full 15000ms, which SIGKILLs mid-drain under any terminationGracePeriodSeconds below that. Set it yourself if 15s is longer than your platform's grace period.
QM_FORCE_LOCK(removed)Removed in step 5 with the data-directory lock. Set, it is ignored with a boot warning, like DATA_DIR.
QM_IGNORE_CORRUPT_SNAPSHOT(removed)Removed in step 5 with snapshot.json. Set, it is ignored with a boot warning, like DATA_DIR.
HISTORY_RETAIN_DAYS90How long hourly rows are kept: in Postgres with SIDESTORE=pg, in this process's memory with SIDESTORE=memory. Reports stop showing a row once it is older than this, and older rows are deleted.
SIDESTOREmemoryWhere operators, monitor links, the audit log, timeline notes, metrics buckets and hourly history are kept: memory or pg. memory keeps them in this process only, so a restart empties them (development and tests). pg keeps them in Postgres at DATABASE_URL: the schema is migrated at boot and everything is loaded before the server listens. pg needs SECRET set. With STORE=memory it assumes a single instance (another instance does not see a new monitor link or operator until it restarts); with STORE=valkey every instance re-reads operators, monitor links and timeline notes when another changes them (see RESYNC_MS). file was removed in step 5 and refuses to start by name (see Refusals); anything else refuses to start too.
STOREmemoryWhere queue state lives. memory: diskless, development only. The in-process engine writes nothing to disk, so every restart starts empty (with the demo room), and NODE_ENV=production refuses it. valkey keeps room state in Valkey at VALKEY_URL and journals every change to the events table in Postgres; it requires SIDESTORE=pg, SECRET and VALKEY_URL, and refuses to start naming each one that is missing. What works on Valkey: rooms and their config, join, status, the door check, ticks and admission, sessions and door keys, admin eject, flush and purge, URL targeting (urlPatterns, hit counters, the URL tester), autotune, stats and metrics, scheduled opens (opensAt, branding.opensAt, setOpensAt), the pre-queue (preQueueMaxPerIp) and its random draw at the open; no call answers not_supported_yet. Several instances can share one Valkey and one Postgres behind a load balancer with no sticky sessions: see Several instances (STORE=valkey). The drawn order of each open is kept in the Postgres table open_orders (about 8 bytes per pre-queue entry, one row per 20 000 entries; the open event refers to it). Like events, open_orders is not pruned yet: both grow with every open and every change for the life of the database, so size the Postgres disk for it and watch it. A room whose rebuild finds an open it cannot replay (the open_orders rows for it missing or short, or its order naming a pre-queue entry the journal never enrolled) stays fenced rather than mint tickets nobody owns: every call on it answers recovering, the rebuild is retried, and the console logs recovery rebuild <room>: rebuild <room>: … (naming the seq for missing rows) at most once a minute. To recover it, restore that room's open_orders and events rows from a Postgres backup taken after the open; the next retry rebuilds the room and unfences it. Never delete or edit its events rows to get past it. Changing IPV6_PREFIX_BITS needs the queue drained first, because the per-network keys already in Valkey were written with the old prefix. Any request (visitor or console) whose room-state call fails because Valkey is unreachable, not serving (READONLY, LOADING, MASTERDOWN, CLUSTERDOWN, TRYAGAIN) or does not answer within VALKEY_COMMAND_TIMEOUT_MS answers 503 storage_unavailable with Retry-After: 2; a command whose reply was lost is never re-sent, so a retried join cannot mint a second ticket. A Postgres failure is not this: joins and admissions answer 503 storage_unavailable when their journal commit fails (queries time out after 10 s), and other routes keep their own handling. Any other value refuses to start.
VALKEY_URLnoneValkey connection URL, read only with STORE=valkey (required then). rediss:// connects over TLS verified against the system CAs. The db index in the path is the one used.
VALKEY_PREFIXqm:Prefix of every key the server writes in Valkey, and of its control channel (<prefix>ctl), with STORE=valkey. Two deployments on one Valkey need different prefixes.
VALKEY_COMMAND_TIMEOUT_MS2000With STORE=valkey, how long one Valkey command may go unanswered before it fails and the request answers 503 storage_unavailable. Integer, 100–60000: a value outside the range is clamped to it and a value that is not a number falls back to 2000, each with a boot warning that names the variable. Raise it only if a slow or distant Valkey trips it under normal load; a higher value holds requests open longer while Valkey is stuck. It does not bound what each Valkey connection sends when it (re)connects (AUTH, SELECT and the INFO ready check): those may take up to 10 s (the connect timeout) before the connection is dropped and retried, so a Valkey that is up but answers slower than this value still connects.
NODE_ENV(unset)production (read trimmed and in any case) is the production rule: the server refuses to start unless STORE=valkey and SIDESTORE=pg are both set, and names each one that is missing (see Refusals). It also lists an unset EDGE_SECRET under hardening (see Cloudflare edge secret) and switches off test-only hooks (QM_TEST_VALKEY_COUNT, which adds a per-request Valkey command count header, QM_TEST_DROP_CTL, which drops pub/sub messages, QM_TEST_ROOM_EVENT_DELAY_MS, which delays room events, QM_TEST_STORAGE_FAIL_FILE, which makes storage read as failing to alerts, QM_TEST_JOURNAL_FAIL_FILE, which makes this instance's journal commits fail, QM_TEST_ALERT_HOLD_FILE, which holds one alert frame past its bound, QM_TEST_HISTORY_SAMPLE_MS, which shortens the history sample period, QM_TEST_NO_TICK, which stops one instance's engine tick, QM_TEST_NOTIFY_TTL_MS, which shortens the notify intent bound, and QM_TEST_NOTIFY_FAIL_FILE, which makes notify deletes fail).
QM_TEST_VALKEY_COUNTunsetTest hook, never set in production: 1 with STORE=valkey adds x-qm-valkey-commands (the Valkey commands the request sent) to every response and turns off the Valkey client's auto-pipelining.
RESYNC_MS5000With STORE=valkey, how often each instance checks that its copy of the operators and monitor links is current (the side_rev row in Postgres) and asks Valkey about the console sessions on its open admin streams. A change made on another instance normally arrives at once over Valkey pub/sub; this check repairs a lost message within this long. Operator keys, operator console sessions and monitor links are refused with 503 storage_unavailable once the copy has not been confirmed for 3× this (at least 15 s), for example while Postgres is unreachable; ADMIN_KEY still works. The same check compares the rooms revision in Valkey (<prefix>roomsrev, moved by every room create, edit and delete) and, when it moved, rebuilds this instance's copies of the room config that decide whether a request is gated: sentinel protection, the inline host index, and the last known config onBackendDown answers from while Valkey is unreachable. A copy that fails to rebuild is retried on the next check, then less often while it keeps failing (doubling up to 60 s, for example while a room is being recovered); a moved revision and a reconnect to Valkey always rebuild them all at once. Integer, 100–600000.
QM_TEST_ROOM_EVENT_DELAY_MSunsetTest hook, never set in production (ignored with NODE_ENV=production): holds back the Postgres commit of every room config event (upsertRoom, schedule change) by this many ms, to test that another instance's open waits for it (Notion 627). A boot warning names it when it is on.
QM_TEST_DROP_CTLunsetTest hook, never set in production (ignored with NODE_ENV=production): comma-separated pub/sub message kinds (acct, revoked, rooms, notes) this instance never publishes, to test the RESYNC_MS repair. A boot warning names it when it is on.
QM_TEST_JOURNAL_FAIL_FILEunsetTest hook, never set in production (ignored with NODE_ENV=production): with STORE=valkey, while the named file exists, this instance's events-journal commits (and its storage probe) fail, to test one instance's storage fault in a fleet. A boot warning names it when it is on.
QM_TEST_STORAGE_FAIL_FILEunsetTest hook, never set in production (ignored with NODE_ENV=production): while the named file exists, this instance reports its storage as failing to the alert frame, to test storage_failing across instances. A boot warning names it when it is on.
QM_TEST_ALERT_HOLD_FILEunsetTest hook, never set in production (ignored with NODE_ENV=production): the named file holds a point in an alert frame (evaluate, load or save); the first frame to reach that point renames the file to <file>.held.<n> and waits until that is removed, to test that a frame held past its bound saves and sends nothing. A boot warning names it when it is on.
QM_TEST_HISTORY_SAMPLE_MSunsetTest hook, never set in production (ignored with NODE_ENV=production): the history sampler runs every this many ms (at least 100) instead of every 15 s, to test that several instances write one history row per room-hour. A boot warning names it when it is on.
QM_TEST_NO_TICKunsetTest hook, never set in production (ignored with NODE_ENV=production): 1 runs no engine tick on this instance, so a test can tell which instance's tick made a change. A boot warning names it when it is on.
QM_TEST_NOTIFY_TTL_MSunsetTest hook, never set in production (ignored with NODE_ENV=production): bounds each notify intent to this many ms (at least 100) from when it was filed, instead of 24 h, to test that an old entry goes in a room that keeps writing. A boot warning names it when it is on.
QM_TEST_NOTIFY_FAIL_FILEunsetTest hook, never set in production (ignored with NODE_ENV=production): while the named file exists, deleting or clearing notify intents fails as an unreachable Valkey would, to test that eject, purge and room delete answer 503 then. A boot warning names it when it is on.
LEADER_LEASE_MS15000With STORE=valkey, the lease that makes one instance the leader: the only one that probes targets (PROBE_EVERY_MS), runs autotune and threshold activation, and delivers alerts to ALERT_WEBHOOK_URL, so each happens once for the fleet and not once per instance. The leader renews the lease (<prefix>leader in Valkey) every third of this and stops acting as leader once half of this has passed without a renewal, before the lease can expire, so two instances never act at once; another instance takes over within about this long of the leader dying, and at once after a clean stop, which hands the lease back. The other instances still evaluate alerts and show the leader's target health and autotune notes, read from Valkey. While Valkey is unreachable no instance is leader: probes, autotune, threshold activation and alert delivery pause until it answers again, and nothing falls back to every instance acting. The one alert still sent then is valkey_unreachable (see VALKEY_ALERT_AFTER_MS); keep an external monitor on /healthz as well. Autotune and threshold-activation writes carry the leader's epoch and are refused once another instance has taken the lease, so a stale leader's late write never lands after the new leader's. leader on /healthz says whether this instance leads. With STORE=memory the one process is always the leader. Integer, 1000–600000.
VALKEY_ALERT_AFTER_MS30000With STORE=valkey, an instance that has not reached Valkey for this long sends a valkey_unreachable alert to ALERT_WEBHOOK_URL itself, and a resolved once Valkey answers again. With Valkey down no instance leads and every other alert delivery pauses, so this is the page for that outage. Each instance that cannot reach Valkey sends its own: expect one per instance, since duplicate pages are better than none. The value used is at least one lease renewal (LEADER_LEASE_MS/3) plus the Valkey command timeout (VALKEY_COMMAND_TIMEOUT_MS) plus 1 s, because that is how often a healthy instance hears from Valkey; a lower value is raised to that with a startup warning naming both (at the defaults the floor is 5000 + 2000 + 1000 = 8000 ms). The floor grows with VALKEY_COMMAND_TIMEOUT_MS: at its 60000 maximum it is about 66 s, so a real outage pages that much later. A delivery of this alert the webhook did not take is sent again on each later check, up to 5 times in all. Integer, 1000–3600000.
QM_INSTANCE_ID(unset)With STORE=valkey, this instance's display name: the instance label on every /metrics series, the instance in GET /api/admin/health and the name alerts give it. A label only: every internal identity (leases, staging keys, pub/sub sender, the counters key <prefix>ctr:<name>:<random>) is <name>:<random>, new on each boot, so instances sharing a name, or a rolling deploy that reuses it, stay correct. Unique names make metrics and alerts easier to read, and are required per scrape target when Prometheus scrapes with honor_labels: true (see Metrics and monitoring). Unset, the name is i:<random>, new on each boot. 1 to 64 characters of A-Z a-z 0-9 . _ -; anything else refuses to start. With STORE=memory it only adds the /metrics label.
DATABASE_URLnonePostgres connection URL, read only when SIDESTORE=pg (required then). The connection always uses TLS verified against PG_CA_FILE; any sslmode in the URL is ignored. To use a schema other than public, add options=-c search_path=<schema> to the URL; that schema is the one migrated.
PG_CA_FILEcerts/do-ca.crtCA certificate the Postgres server is verified against when SIDESTORE=pg. Verification is never turned off; a missing file refuses to start.
PUBLIC_URL(derived)Absolute origin this server is reached at, used to build install commands, connector download URLs, waiting-room links and the copyable commands on /docs when the request's own Host is not the public one (behind a proxy that rewrites it). An https:// value also marks cookies Secure on requests for that host, whatever the socket says. With ADMIN_HOST unset its host is also where the console lives: the console answers there once a room is inline (QM-440), no room may be put inline on it (400 reserved_host), and its console surface is never handed to a room.
ALERT_WEBHOOK_URL(unset)Where alerts are POSTed as JSON. Unset means nothing is delivered — conditions are still evaluated on the server and listed in Diagnostics → Alerting, so the console is the alerting. See Alerting below.
ALERT_WAIT_SEC900Estimated wait, in seconds, at or above which a room pages. 0 disables.
ALERT_DEPTH0Queue depth at or above which a room pages. 0 (default) disables — depth alone is not an incident; a wait nobody will sit through is.
ALERT_STALL_SEC120People waiting and nothing let through for this long pages as critical. 0 disables.
ALERT_FROZEN_SEC300People waiting in a room that is paused or set to 0/min for this long pages as critical — the forgotten paused room. 0 disables.
ALERT_HOLD_SEC30How long a condition must hold before it fires. Stops a four-second spike from paging anyone.
ALERT_REPEAT_SEC900Re-page cadence while a condition is still firing.
ALERT_EVERY_MS15000Evaluation cadence.

GET /api/admin/health reports the effective values, the storage state and any misconfiguration warnings. Check it after every deploy.

The dashboard renders that report behind its Diagnostics button: status, version, uptime, room and visitor counts, SSE streams, rate limits, storage, proxy, memory and every warning — plus two counters worth watching, checks for unknown rooms (a snippet pointing at a room id that does not exist here, i.e. a typo or a deleted room) and blocked return URLs (see returnOrigins below). It is a readable view of the payload rather than the raw JSON; pid, the server clock and the key-configuration block are only in the API response. The header also carries an insecure config badge for as long as any warning stands.

Behind a CDN or load balancer (read this one)

Per-IP limits key on the address this server can prove: the socket peer. Behind a CDN every visitor shares that address, so the per-IP limit silently becomes a global limit — 600 joins/min for your entire site — unless you tell the server who the proxy is:

TRUST_PROXY=1                 # hops between the client and this server
TRUST_PROXY_IPS=10.0.0.7      # the address(es) those hops connect FROM

The same rule applies to the SSE per-IP cap.

A proxy you have not listed is named, not silent. When forwarded-address headers (X-Forwarded-For, Forwarded, CF-Connecting-IP, X-Real-IP, True-Client-IP) keep arriving from a peer that is not in TRUST_PROXY_IPS — 10 or more such requests from the same address within a minute, whether TRUST_PROXY is set or not — the server treats that peer as an undeclared proxy. Behind it, every visitor shares the proxy's one address, so queueMaxPerIp (16 places per address by default) and preQueueMaxPerIp cap the whole room at that many visitors, and the join and check budgets become one budget for the site. The server then:

One request with a forged header does not trip it. A script sending a stream of them can, so read the address before you copy it: add it only if it is your load balancer or CDN. An address you do not recognise is a client forging the header, and listing it would let that client choose its own address. Changing TRUST_PROXY_IPS is the fix; the warning never changes whose header is believed.

Setting only one of the two is the failure mode to watch for. They are one setting in two variables, so the server refuses to let a half-configuration be silent: it prints the warning at startup, returns it in warnings on /api/admin/health with proxy.ok: false, and the dashboard header turns it into the insecure config badge (the Proxy row reads … · INCOMPLETE). A measured surge through a half-configured proxy loses roughly 94% of its joins to HTTP 429, and the behaviour is otherwise indistinguishable from the queue working correctly.

Cloudflare edge secret

In production this server is reached only through our Cloudflare zone, but its own address (the DO App Platform URL, the container) answers too. A request sent there directly skips every Cloudflare rule, cache and rate limit, and can put any value it likes in CF-Connecting-IP. EDGE_SECRET closes that door: Cloudflare adds a secret header to every request it forwards, and this server refuses any request without it.

Set it up (in this order, so nothing is refused while you do it):

  1. Generate a long random value, e.g. openssl rand -hex 32. Entries shorter than 16 characters stop the server at startup: a guessable secret is worse than none, because it looks like protection.
  2. In the Cloudflare dashboard for the zone: Rules → Transform Rules → Modify Request Header → Create rule, with Set static header X-QM-Edge = the value. Deploy the rule. Match only this app's hostname(s) (e.g. Hostname equals qm.weekday100.com), never All incoming requests. A zone-wide rule adds the secret to every request Cloudflare sends to every origin in the zone: your other apps, and any third-party service a hostname is CNAMEd to (a help desk, a shop platform, a status page). Any of them can log it, and anyone holding it can call this app's own address directly and be treated as Cloudflare, with a CF-Connecting-IP of their choosing.
  3. Set EDGE_SECRET to the same value on the app and deploy.

What it does once set:

Rotate without refusing anyone: add the new value to the list (EDGE_SECRET=old,new) and deploy; switch the Cloudflare rule to the new value; then remove the old one (EDGE_SECRET=new) and deploy again.

Unset, nothing changes from earlier releases. With STORE=valkey or NODE_ENV=production the server prints EDGE_SECRET is not set: requests that bypass Cloudflare are accepted at startup and lists it under hardening on /api/admin/health.

Two deployments: snippet and inline

There are two ways to put this queue in front of a site, and they fail differently. Choose on that, not on which is easier to install.

Snippet (default). The site loads a small script; the script asks this server whether the visitor may proceed and sends them to the waiting page if not. The queue sits beside the traffic. If this server dies, the script's check fails and the pages are served unqueued — the site stays up and the protection is gone.

Inline (proxy). DNS for the site points at this server, and every request for that host arrives here. Admitted visitors are forwarded to the app; everyone else gets the waiting page at the URL they asked for. The queue sits on the traffic. If this server dies, the site is unreachable — which is why the CDN failover page in When the queue itself is down is not optional for an inline deployment.

Solid: page traffic. Dashed: asks the queue. Snippet: a tag in the page Browser Your site Queue Queue down: pages load unqueued. Edge: a Cloudflare Worker Browser Worker Your site Queue Queue down: forwarded to your site unqueued, marked X-QM-Failover. Inline: DNS points at the queue Browser Queue Your site Queue down: the site is unreachable, unless the CDN serves a failover page.

The three shapes side by side. With the snippet, pages come from your site and the tag in the browser asks the queue; if the queue is down, pages load unqueued. The edge connector is a snippet-mode deployment too, but a Cloudflare Worker asks the queue before your site sees the request; if the queue is down, it forwards the request to your site unqueued and marks the response X-QM-Failover. Inline, every request passes through the queue; if the queue is down, the site is unreachable unless the CDN serves a failover page.

Turn it on in the console — Edit (or + New room) → tick Serve this room inline and fill in Internal address of the site behind the queue. Clearing the tick returns the room to snippet mode, behind a confirmation that says so. Or over the API:

curl -sX POST "$BASE/api/admin/rooms" -H "Authorization: Bearer $ADMIN_KEY" \
  -H 'Content-Type: application/json' -d '{
    "id": "shop",
    "targetUrl": "https://shop.example",
    "proxyOrigin": "http://10.0.0.5:3000"
  }'

What changes once a room is inline:

SnippetInline
Queue's own routes/api/*, /w/<room>under /__qm (/__qm/api/join, …)
Monitoring routes/healthz, /api/version, /metricsthe same, and also under /__qm. Which one a probe gets depends on the address it uses, not on the path: reached on the customer's hostname these are the customer's paths and the visitor gate answers them 503 for ever. Probe the process directly, or use /__qm/healthz and metrics_path: /__qm/metrics.
Operator console/ on the queue's own host/__qm/ on the customer's host — the same page, addressing the same prefixed routes. Bookmark that address, not /: / on that host is the shop.
Console sign-inqm_console, HttpOnly, Path=/api/adminqm_console, HttpOnly, Path=/__qm/api/admin: the customer's scripts cannot read it, and the key itself is never stored in the browser
Door keyin the URL (?sess=), readable by the scriptqm_dk_<roomId>, HttpOnly cookie only: nothing in a query string, even under /__qm. Scoped to the room, so clearing one room's key leaves another's alone.
Ticket token qm_<room>readable cookie + localStorage + ?token= — this host is ours and the page's own script is the readerHttpOnly cookie only: nothing in localStorage, nothing in a query string, even under /__qm
URL targetingmatches whatever the snippet reportedmatches the real request URL
returnOrigins / tagOrigins / CORSload-bearingunused: nothing is cross-origin
This server downsite up, unqueuedsite down — needs CDN failover

The prefix is reserved, not decorative: the campaign app has its own /api/*, and the queue must never shadow it. Anything on a proxied host that is not under /__qm belongs to the app.

Room ids starting with dk_ are reserved the same way (QM-533). Room dk_x's ticket cookie would be qm_dk_x, the exact name of room x's door-key cookie: the two overwrite each other in the browser, and the per-host cookie cap files the ticket under room x and expires it with x's cookies. Creating such a room answers 400 invalid_room_id. A dk_ room made before this rule keeps working and can still be edited, but it keeps the collision: recreate it under another id and move the tag over.

The rest of the reserved ids follow from the same rule (546): a room's ticket is kept under qm_<room>, as a cookie and as a localStorage key on the queue host, so no room id may make that name one the queue already uses. Refused on create with 400 invalid_room_id, case-sensitive:

When the server looks for a ticket among a request's cookies (/api/status, /events), qm_console is never read as one, and cookies under a reserved name are tried only after every other qm_ cookie, so eight rooms' door keys cannot crowd out the real ticket.

The ticket-token row is the same reasoning as the prefix, applied to credentials. Inline, our cookies and our localStorage land on the customer's origin, shared with every analytics tag, chat widget and ad script the site ships — so the token is kept somewhere none of them can reach, and the waiting page persists nothing. It does not need to: /api/join returns the token in its response body on every load, and /api/status and /events authenticate from the cookie when no token= is sent.

Forwarding preserves the visitor's Host and adds X-Forwarded-Host, -Proto, -Port and -For, so a Next.js app behind this builds its own absolute URLs correctly. All four describe the VISITOR's hop, never proxyOrigin: -Port is the port named in their Host, or 443/80 to match -Proto when they did not name one. WebSocket upgrades carry the same four.

The rightmost X-Forwarded-For entry is always the client address this server resolved, so an app that trusts one hop (this server) reads the truth. From a proxy in TRUST_PROXY_IPS the chain it sent is kept to the left of that entry, and its Forwarded, X-Real-IP, CF-Connecting-IP and True-Client-IP pass through. From any other peer those headers are the caller's own words: all of them are dropped and X-Forwarded-For is the socket address alone. TRUST_PROXY / TRUST_PROXY_IPS still apply: they are how this server learns the visitor's address and scheme from your CDN. The -Proto (and the default -Port) sent onward is the scheme your proxy reported only when that proxy is in TRUST_PROXY_IPS; from any other peer it is the scheme of the connection this server actually received, whatever the caller claimed.

When the app itself is the problem:

What happenedVisitor seesAlso
App refused the connection / unreachable502 + X-QM-Origin: origin_unreachable — origin-down.html on a navigation, JSON to an XHRcounted per room; one log line a minute naming the room
proxyOrigin is a private/internal address and PROXY_ALLOW_PRIVATE is not 1the same 502 + X-QM-Origin: origin_unreachable (a WebSocket gets 502 too); nothing is sent to the addresscounted per room under phase blocked; the room's originErrors.reason and a log line (once a minute per room) name the refused address; the probe reads down with targetHealth.blocked saying why
App accepted, then never answered504 + X-QM-Origin: origin_timeout after PROXY_TIMEOUT_MSsame
App answered, then stalled mid-bodythe transfer is cut — headers are already gone, so there is no status code left to send and no honest way to finish the pagecounted per room under phase body_stall after PROXY_IDLE_TIMEOUT_MS; a visitor who stopped reading is counted as client_reset instead
Queued visitor's XHR or API call503 queued + Retry-After + X-QM-Queued: 1anything a browser renders gets the waiting page instead — see below
This process at MAX_CONNECTIONS503 + Retry-After (no X-QM-Queued) — the offline page, not the waiting pageqm_connections_shed_total; /healthz answers its own 503 with ready:false too — see CONNECTION_RESERVE above

Past MAX_CONNECTIONS a request is deliberately shed, not queued — there is no "try again in a moment and you'll be let in", the process genuinely has no room. It carries no X-QM-Queued, on purpose: a CDN failover rule that keys on 503 is meant to fire here and show its own outage page, the same as any other real outage. Without CONNECTION_RESERVE, this state was unreachable to answer at all — Node's own server.maxConnections destroyed the socket first, silently, and a hung origin pinning two sockets per stuck visitor for the whole PROXY_TIMEOUT_MS made that ceiling reachable in well under a minute during exactly the failure this product exists for.

X-QM-Queued is what stops a CDN failover rule keyed on 503 from replacing a healthy queue with the outage page, and it is also the signal a customer's app can key on to reload a page whose pass lapsed mid-session instead of failing silently. X-QM-Origin says the site behind the queue is the thing that broke.

Which of the two a request gets. The test is whether a browser will show the answer to a person, not whether the method is GET: Sec-Fetch-Mode: navigate (any method), a browsing Sec-Fetch-Dest, or — for clients that send no Fetch Metadata — Accept: text/html. A <form method="post"> submitted after the pass lapsed is a navigation, so it gets the waiting page; a fetch() asking for JSON gets JSON.

Recovering a page whose pass lapsed. The queued envelope names the room and the engine's own reason (expired, ejected, stale_generation, waiting, at_capacity, …), and Access-Control-Expose-Headers is set so the header is readable even when the call is cross-origin:

{ "error": "…", "code": "queued", "roomId": "shop", "reason": "ejected", "reload": true }

Inline there is no separate waiting-room address — the waiting page is served at the URL the visitor asked for — so re-requesting the current URL is the whole recovery:

const r = await fetch('/api/cart');
if (r.status === 503 && r.headers.get('X-QM-Queued')) { location.reload(); return; }

Without that line the site keeps rendering whatever its code does with a failed request. A visitor who is simply put back in line is told so by the waiting page itself, which names the cause rather than presenting itself as a fresh arrival.

On a hostname this server fronts no room for, the operator console is a bare 404. See ADMIN_HOST above: inline, the console answers on an address literal, localhost and the PUBLIC_URL host, or with ADMIN_HOST set on the names you list, and nowhere else. An inline room whose targetUrl is an address literal (the usual local-development setup) gets that host's pages, but with ADMIN_HOST unset never its /api/admin/*: that address is the console's too, so an admin call sent there is answered by the queue, not forwarded (QM-439). With ADMIN_HOST set that includes the customer's hostname: its /__qm/ console, /__qm/api/admin/* and /__qm/public/admin.html answer 404, and only the visitor paths under /__qm/ remain. A hostname that fronts no room is withheld the console only, ADMIN_HOST set or not (QM-428) — the snippet, /healthz, the waiting page and the visitor API still answer there, since that is the address the customer's script tag names.

This is easy to walk into (QM-454): with neither ADMIN_HOST nor PUBLIC_URL set, saving the first inline room from a console on a hostname (say qm.weekday100.com) makes that console, and every monitor link on that host, answer 404 from the next request on. The save still succeeds, and its answer carries a warning (warnings[], code: "console_host_withheld", naming the host) that the console shows after saving. While any inline room exists with both unset, /api/admin/health lists the same reason under configWarnings. The fix is to set ADMIN_HOST (or PUBLIC_URL) to that hostname and restart.

Switching the last inline room on a host back to snippet mode takes the queue's own /__qm/* surface off that host with it — including the console you are probably reading this from, if you administer the queue through the customer's domain. The console warns before that save and moves you to PUBLIC_URL afterwards; a stale /__qm/… address answers 404 {"code":"not_inline_host"} naming where the console went, rather than the generic not-found. Set PUBLIC_URL so it can name it. (With ADMIN_HOST set the console never lived on the customer's domain, so there it is a bare 404.)

None of these make this instance unhealthy: a customer origin going dark is not a reason to take the pod out of rotation, so /healthz stays 200. See Readiness. The counts are on /api/admin/health under inline — gatedTotal, bypassedTotal, originErrorsTotal and a per-room/per-phase originErrors breakdown. The whole block is absent on an install with no proxied room.

What skips the gate

Static subresources are forwarded without a queue check. This is about cost, not politeness: one page of a modern app is 30–80 subresource requests, and gating each one would run a queue check, a bot score and a cookie write per file — turning the queue into the bottleneck it was installed to prevent, while protecting nothing. Downloading a stylesheet takes nobody's place in line.

A request skips the gate only when all four hold:

The app's own API is never bypassed. POST /api/checkout is where the scarce thing happens; letting it through would mean anyone willing to skip the browser buys the item without ever waiting. Sec-Fetch-Dest is forgeable, but forging it wins nothing except a file from the list above.

The edge case to know: an app that serves capacity-heavy dynamic content from a URL ending in .js — or routes a trailing segment like /api/stock/x.css to a protected handler, or treats .js as a format suffix the way Rails does — would skip the gate for a subresource request. There is no switch for this yet (restricting the bypass to declared asset prefixes is an open product decision) — if your app does that, say so.

Streaming and WebSockets

Nothing is buffered in either direction. A streamed response starts reaching the visitor as the app produces it — a Next.js App Router shell paints at the same moment it would with no queue in front — and an upload streams through without being held in this process.

WebSocket and other connection upgrades are forwarded, and they are gated like a page, not like an asset: a socket that stays open for the length of a visit is not a subresource, and letting one through unqueued would be a way to sit inside the app without ever taking a place in line. A visitor who is still waiting gets 503 on the upgrade — no waiting page, because there is no page there; the app's own reconnect logic is left in charge of trying again. Once the handshake is through, the two sockets are joined and nothing in this process looks at the bytes again.

Upgrades on /__qm/* are refused with 501: the queue's own live updates are Server-Sent Events, which are ordinary HTTP.

Rate limits

Every throttled route answers with X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset (seconds), and adds Retry-After on 429, so a client can pace itself instead of discovering the limit by being cut off.

The waiting page is throttled too

The page is 141 KB and is the largest thing an anonymous stranger can ask this server for. It is rendered and compressed once per room and served from memory (private, no-cache + ETag, so a visitor's reload during a long wait costs a 304 and not another page), but a client that simply refuses Accept-Encoding still takes the full body every time. Measured on one laptop at 120 connections:

bytes pulled in 10 sother requests on the process
before2,212 MB21x slower
after29 MB1.3x slower

WAITING_LIMIT_PER_MIN (default 120) is roughly two page loads a second from one client — far above any real visitor. A refused navigation gets a 941-byte page that reloads itself after Retry-After; a refused XHR gets plain text. The queue is untouched: the visitor keeps their session and their place, and the status endpoint keeps answering.

It stands down when it cannot tell visitors apart. With TRUST_PROXY unset, every visitor behind a proxy looks like one client (the proxy's own address), and a per-address limit there would take the customer's whole site down — far worse than the flood. So once a peer is recognised as a proxy — forwarded-address headers from the same address 10 times within a minute, or WAITING_LIMIT_PER_MIN times if that is lower, so it is recognised before its budget runs out — the guard stands down for that address only. Every other address is still throttled. /api/admin/health then reports waitingPage.enforced: false while it stands down for any address, and the console's Proxy row says an undeclared proxy is in front. Fix the proxy settings and the guard comes back for it on restart.

One forged X-Forwarded-For does not switch the guard off: it used to, for every address and for the life of the process. A client that sends a steady stream of forged headers can still exempt its own address this way — with TRUST_PROXY unset the server cannot tell it from a proxy — but never anybody else's. Declaring your proxy (TRUST_PROXY + TRUST_PROXY_IPS) removes that too: the guard then never stands down, because it can charge each visitor. The list of exempt addresses is capped at 1024 and dropped when full; a real proxy is recognised again within a few requests.

So are the console and the docs

The console page is about half a megabyte, and /docs/* is rendered from markdown; both used to answer any client as often as it asked — on an inline host that includes /__qm/, /__qm/docs and /__qm/progress on the customer's own domain. /, /progress and /docs* now take the same guard as the waiting page: WAITING_LIMIT_PER_MIN per address, the same small self-reloading refusal, the same stand-down behind an undeclared proxy. They are charged to a separate bucket, so an operator reading the guides never spends a visitor's waiting-page allowance, and the admin API is not touched. Each guide is rendered once per file version and served from memory after that.

What churn cannot do

Each budget is a map of per-address buckets. It holds at most RATE_LIMIT_MAX_KEYS (default 200,000) and evicts the least recently used address past that. It used to clear itself instead, which gave every throttled client a fresh budget to anyone able to mint 200,000 keys — a Public API key holder naming a new ?ip= per vouched check, or a client rotating IPv6 addresses. A client that is being limited is, by definition, the most recent key in the map, so eviction never reaches it.

Vouched /api/check calls are charged to the visitor they vouch for (?ip=), by design: one connector speaks for thousands of visitors from a handful of addresses. Those buckets live in a map of their own, so however many visitors a key holder names, it cannot evict a direct caller's bucket.

Two keys, two blast radii

PUBLIC_API_KEYADMIN_KEY
GET /api/check with ?ip=&agent=yesyes
/api/admin/* (rooms, rate, flush, eject, schedule, visitors, metrics, health)no (401)yes
GET /metrics (Prometheus exposition)no (401)yes

Set both. Distribute only the first.

Wrong keys are rationed. Each client address may fail admin authentication ADMIN_AUTH_FAIL_PER_MIN times a minute (default 20). Past that, every credential it presents — right or wrong, Bearer or SSE ticket — is answered 429 too_many_auth_failures with Retry-After, before it is checked, so a lockout never confirms a guess. A correct key, a valid monitor link and a valid key used on a route it is not scoped for (403) spend nothing; a request that presents no credential at all is not counted either. The public-key door (/api/check?ip=&agent=, /api/verify) shares the same allowance because it accepts ADMIN_KEY too: a wrong key there is charged, and a locked-out address's key is treated as absent (401, and connectors fail open). Behind a CDN the address is the forwarded client only when TRUST_PROXY_IPS names the CDN; otherwise every visitor shares the CDN's socket address and one guesser can lock the console out for a minute — another reason to set it.

No credential travels in a query string. Every admin route takes Authorization: Bearer and nothing else, and so does the Public API key on /api/check and /api/verify. ?key= used to be accepted there; it is now refused with 401 {"code":"credential_in_query"} on a vouched check, counted in qm_refused_key_checks_total like any refused key, and reported by /api/verify as "Public API key sent in the URL". Every shipped connector already sends the header. A hand-written integration that still sends ?key= fails open (its pages go unqueued) until it is moved to the header — the counter and the stderr line are how you find it. The one exception is the event stream, where EventSource cannot set a header: the console mints a single-use ticket (POST /api/admin/sse-ticket, over Bearer, scoped to whatever credential asked for it) and spends it on GET /api/admin/events?ticket=…. A ticket lives 30 seconds and is destroyed on first use, so one landing in a log is worth nothing by the time anybody reads it. Operator tooling that opens the stream must mint a ticket per connection.

The console keeps a session cookie, never the key (QM-349). The console used to keep ADMIN_KEY or an operator key in localStorage, and on an inline install the console is served from /__qm/ on the customer's origin, so any script that site ships could read it. Signing in now sends the key once, in the body of POST /api/admin/session ({"key":"…"}, header X-QM-CSRF: 1), and the answer sets qm_console: HttpOnly, SameSite=Strict, Path=/api/admin (or /__qm/api/admin on an inline host), Secure whenever the request arrived over TLS, lasting 12 hours. Its value is a signed payload naming who signed in and a session id, never the key. DELETE /api/admin/session signs out: it clears the cookie and revokes that session server-side (the cookie was already cleared in that browser). The revocation is remembered in memory with STORE=memory, so a restart forgets it, and in Valkey with STORE=valkey, so a Valkey data loss forgets it: the instances that received the sign-out's hint still refuse the session until they restart, an instance booted after the loss does not. Revoking an operator ends their sessions on the next request; changing ADMIN_KEY ends every session the old key opened. A wrong key at sign-in spends the ADMIN_AUTH_FAIL_PER_MIN allowance exactly like a wrong Bearer, and a locked-out address is refused there too. A 401 clears qm_console only when that cookie was the credential that failed: a request carrying Authorization: Bearer is judged on the Bearer alone, so a dead monitor link opened in the owner's browser does not sign the owner out (QM-460).

Revocation reaches open live streams too (QM-442). The console's live stats stream (/api/admin/events) is one long request, so checking the credential only when a request arrives was not enough. Each stream remembers what opened it: the Bearer (ADMIN_KEY, an operator key or a monitor link) or the console session cookie, carried through the SSE ticket. When a monitor link is revoked, Revoke all runs, an operator is revoked, or a console signs out, every stream opened with that credential ends at once. The admin tick checks all of them again every 2 s, which also catches expiry. The last thing the stream sends is event: revoked ({"error":"unauthorized"}), and the console shows its signed-out screen or the dead-link screen instead of reconnecting. An unspent SSE ticket whose credential was revoked is refused with 401. Removing SECRET_PREVIOUS takes a restart, and the restart closes every stream anyway.

qm_console never crosses the inline proxy (QM-427). A browser scopes a cookie by host, not port, so wherever the console and a room share a host (the console under /__qm/ on the customer's name, or both on one address) the session cookie would otherwise ride proxied requests to the app, and the app's own Set-Cookie: qm_console=… would replace the operator's session. The proxy removes qm_console from the Cookie header of every proxied request and WebSocket handshake, and drops any Set-Cookie line from the app that names it. Every other cookie passes both ways unchanged, the visitor's queue cookies included.

A change made with the cookie (any method but GET) must carry the X-QM-CSRF header, and a browser that sends Sec-Fetch-Site must say same-origin; otherwise 403 {"code":"csrf_required"}. No form can send a custom header, and a cross-origin script cannot either without a preflight this server never grants, so a page elsewhere cannot drive the console's session. Bearer callers are unchanged: a request with an Authorization header is judged on that header alone, the cookie is ignored, and no CSRF header is needed. That is also why a #monitor= link opened in an operator's browser stays read-only. A console that still has a key in localStorage from an older version moves it on first load: the key is taken out of storage, used once to sign in, and dropped.

There is a third, much smaller credential: the dashboard's Share monitor link carries a scoped viewer token in the URL fragment (#monitor=…), which no browser ever sends to a server — so it appears in no access log, no CDN log and no Referer, including the customer's own logs on an inline install. It is the read grant, not "the stats stream and nothing else" — the holder can read the live stats SSE, the room list, the metrics, the security and reports panels, GET /metrics, GET /api/admin/alerts (without a blocked probe's refusal text, the webhook address or delivery errors, each of which can name an internal address; its Send test alert button is disabled), and GET /api/admin/health (which names the process id, memory and whether the proxy settings are coherent, but not the configuration warnings, the hardening list, the proxy peer addresses, whether EDGE_SECRET is enforced or PUBLIC_API_KEY is set, Valkey's maxmemory-policy (backend.maxmemoryPolicy is left out and its incident says only that the policy is not noeviction), or the text of a storage or warehouse error, which can name a path or a host: a default ADMIN_KEY, a short SECRET or a missing EDGE_SECRET is not something to tell whoever the link was forwarded to; refused qm_return URLs appear as a count per room, never the URL, for any credential). Four things it is deliberately not shown, because the link is meant to leave the team: a room's proxyOrigin (your internal address), its origin allowlists and URL rules, the per-visitor drill-down, and the audit trail — which names every operator on your team, what each of them changed, and the IP each of them did it from. All four answer 403 insufficient_role to a monitor token. A named viewer operator key sees all of them — that is a person on your team. The share link can change nothing: any write answers 403 insufficient_role, and it can never manage operators.

And a fourth, which is the one your team should actually be using day to day: named operator keys.

Operators, roles and the audit trail

ADMIN_KEY is a shared secret with no name attached, so "who paused the room" has no answer. Mint one key per person instead, from the dashboard's Access panel or the API:

curl -sX POST https://qm.weekday100.com/api/admin/operators \
  -H "Authorization: Bearer $ADMIN_KEY" \
  -d '{"name":"Night ops","role":"operator"}'
# → {"operator":{"id":"535dddaa","name":"Night ops","role":"operator",...},
#    "key":"qmo_535dddaa_...","shownOnce":true}

The key is returned exactly once. Only an HMAC verifier is stored, so a lost key is replaced, never recovered.

RoleCan
ownereverything ADMIN_KEY can, including minting and revoking operator keys, deleting rooms, and creating or repointing inline rooms
operatorrun the queue: rates, state, flush, eject, schedule, room settings (branding, targeting, origins, a snippet room's target URL), empty a queue
viewerread only: stats, metrics, security, reports, audit

A key with insufficient role gets 403 {"code":"insufficient_role"} — not a 401, so the holder can tell "wrong key" from "not your job".

Deleting a room needs owner (or ADMIN_KEY). It is the one irreversible action here: everybody standing in the line is discarded and the room, its configuration and its schedule are gone. Emptying a queue (POST /api/admin/rooms/<id>/purge) is destructive too, but the room survives, so it stays with operator. The console hides Delete entirely for a key that may not use it, rather than letting the 403 arrive after the confirm.

Inline deployment needs owner (or ADMIN_KEY). An inline room's targetUrl names the public host this server answers for and its proxyOrigin the address that host's traffic is forwarded to, so together they decide whose traffic this server captures and where it goes. Creating an inline room, changing an inline room's targetUrl or proxyOrigin, clearing its proxyOrigin, or giving a snippet room one (which makes it inline) is refused to operator with 403 {"code":"inline_requires_owner"}, and the refusal is audited (room.create / room.update, ok: false). The refusal on an update says every other setting of the room is still the operator's to change; on a create, where there is no room yet, it says the room can still be created as a snippet room (no proxyOrigin) (QM-458). Only a change counts: a save that sends the room's current targetUrl and proxyOrigin back unchanged (the whole-body save a script or the console makes) goes through, and every other setting of an inline room stays with operator. Each accepted change is audited as room.inline, with the old and new value of each field. The console disables the inline toggle and the internal address for an operator key, and the target URL too on an inline room, and says why under the toggle.

# revoke immediately; the key stops working on the next request
curl -sX DELETE "https://qm.weekday100.com/api/admin/operators/535dddaa" \
  -H "Authorization: Bearer $ADMIN_KEY"

Revoked operators are disabled, not deleted, so their past audit entries still resolve to a name. ?purge=1 is the escape hatch for one added by mistake.

Revoking an operator also closes every monitor link they shared, in the same write, so it survives a restart exactly as revoking one link does. A share link that outlived its author would be the one credential the revocation missed — and the most widely forwarded one. Links shared by other operators and by ADMIN_KEY stay open. The response carries linksRevoked, the single operator.revoke (or operator.purge) audit entry names the count (Night ops; closed 2 monitor links), and the console's confirm dialog says how many will close before you click. GET /api/admin/operators reports each operator's liveLinks.

A link is matched to its author by operator id, which is recorded when the link is minted (createdById on the link), never by name: names are free text, two operators can share one, and an operator can even be called ADMIN_KEY. Links minted before this release carry only the name and cannot be attributed, so revoking an operator leaves them open. If one of those may be in the wrong hands, revoke it from the monitor-link list, or use Revoke all; they also expire on their own after MONITOR_TOKEN_TTL_SEC (seven days by default), after which every live link is attributable.

Every state-changing admin action is appended to the audit trail (Postgres with SIDESTORE=pg) with the actor, the target, the time and the caller's address — including the server's own actions, which appear under a system actor (autotune rate changes, threshold activation flips). Field names are recorded for a room edit, not the values: the values are readable from the room, and an audit line outlives the config it describes.

Restarts are in the trail too. server.start is written when the process begins listening, and server.stop on a clean shutdown. The start line says whether the previous run stopped cleanly — the one fact that cannot be reconstructed afterwards, and the normal case on Windows, where every stop is a kill (see Stopping the server). GET /api/admin/health carries an absolute startedAt beside uptimeSec, /metrics exposes qm_process_start_time_seconds (alert on a change), and an open console raises a toast the moment the value moves under it.

curl -s "https://qm.weekday100.com/api/admin/audit?roomId=checkout&format=csv" \
  -H "Authorization: Bearer $ADMIN_KEY"

There is no SSO. No SAML, no OIDC, no SCIM, no MFA. /api/admin/health reports access.sso: false so a procurement review gets a straight answer rather than a discovery.

Bot protection

Traffic asking for a place in line is scored from what this server can actually observe — a declared automation client, missing browser headers, a forged X-Forwarded-For, machine-regular arrival timing, one address holding an implausible spread of tickets — and each signal adds or subtracts points on a 0–100 scale.

Two properties worth knowing before you enforce:

  1. Enforcement refuses a ticket, never entry to your site. /api/check stays fail-open for everyone. The cost of a false positive is a bot-looking visitor losing their place, not you losing a customer.
  2. A shared address cannot block itself. The scoring is calibrated so that no combination of signals a corporate NAT or CGNAT pool can produce on its own reaches the block threshold — only a client that declares itself automation gets there without help. The direct consequence is that a botnet on residential addresses driving real browsers is not caught by this; fairness against that is structural (FIFO, signed passes, one pre-queue entry per identity), not classification.

Declared crawlers (Googlebot, bingbot, Pingdom, UptimeRobot, Prometheus and the rest of the list in lib/sentinel.js) are reported and not scored — unless the same User-Agent also names an HTTP library or a scripted browser. python-requests/2.31 googlebot is scored as python-requests: a crawler's name appended to a string that already said what it is buys nothing. A UA that says only Googlebot is still taken at its word; checking it against reverse DNS is not done (it would need a network lookup per new address) and is an open product decision.

The queue cookie only counts when it is real. "Presented the queue cookie" is what keeps a returning browser away from cookieless_repeat, and it now means a qm_<room> token this server signed, for that room, still in date — any other value is treated as no cookie.

Put your own synthetic monitoring in PROTECTION_ALLOW_IPS so it is never scored. It matches the resolved client address, which is a forwarded one only when the connection comes from TRUST_PROXY_IPS — naming an allowlisted address in X-Forwarded-For from anywhere else changes nothing.

Places per address in the live line (queueMaxPerIp)

presenceSec (see Passes vs sessions) stops a client that never polls from stalling the room. It cannot stop a client that does poll from hoarding: sixty tickets spread down the line, a pass at each one's turn. That takes a ceiling on how many places one address may hold in the waiting line at once:

FieldDefaultMeaning
queueMaxPerIp16Places one client address may hold in the waiting line. 1–100000; null for no ceiling.

Over it, POST /api/join answers 429 {"code":"queue_identity_limit"} with Retry-After: 30 and hands back nothing — no token, no cookie. A visitor presenting a ticket they already hold is never charged; a place is freed the moment one of that address's tickets reaches the front (or is ejected). It is charged to the same address as preQueueMaxPerIp, the one thing a caller cannot mint a fresh copy of per request, and it survives a restart (it is derived from the recorded ticket owners).

The default is 16, the same as the pre-queue's. A busy live line can hold more than 16 real people behind one mobile-carrier CGNAT address at the same moment, and the ceiling refuses them, not the attacker. Raise it on a room whose audience sits behind carrier-grade NAT, or set null for no ceiling (a load test driving thousands of places from one address needs this). A room saved before the default changed, with no value stored, gets 16; one that stored null keeps no ceiling:

curl -X POST http://localhost:8080/api/v1/admin/rooms \
     -H "Authorization: Bearer $ADMIN_KEY" -H 'Content-Type: application/json' \
     -d '{"id":"drop","queueMaxPerIp":64}'

Threshold activation (the 3 a.m. spike)

A room can switch itself on. Set autoActivate on the room, or use the Threshold activation fields in the room editor:

{"enabled": true, "joinsPerMin": 120, "sustainSec": 60, "releaseAfterSec": 300}

The room moves bypass → active once arrivals hold at or above joinsPerMin for sustainSec, and back to bypass after releaseAfterSec of quiet. releaseAfterSec: 0 means it never stands down by itself.

This is different from AUTOTUNE, which adjusts the outflow rate of a room that is already queueing.

Historical reports

/api/admin/reports serves one row per room per hour (or per day), retained for HISTORY_RETAIN_DAYS:

curl -s "https://qm.weekday100.com/api/admin/reports?bucket=day&from=2026-08-01&format=csv" \
  -H "Authorization: Bearer $ADMIN_KEY" -o last-month.csv

Columns: roomId, bucket, start, startIso, joins, passes, netUnserved, peakDepth, avgDepth, peakActive, waitP50Sec, waitP90Sec, waitMaxSec, waitAvgSec, measuredWaits, blocked, warned, samples, partial, coveredSec, missingSec. The three coverage columns are appended last, so an importer keyed on column position keeps working.

columnwhat it counts
joinstickets issued
passestickets promoted — permission to enter, not an entry
netUnservedjoins - passes; null on a partial row
peakDepth / avgDepthqueue depth, from the depth gauge
peakActivesessions on the protected site — the "people who got in" column
waitP50Sec / waitP90Secestimated wait percentiles
waitMaxSec / waitAvgSecexact longest and mean wait
measuredWaitspromotions that carried a measured join→pass duration
blocked / warnedprotection decisions; with several instances, the whole fleet's (each instance's shared counters)
samplessample ticks folded into the row

Four honest labels travel with every JSON response and are worth repeating here:

The open hour is folded in live, so a report run at 10:59 is not missing the last 59 minutes. With STORE=valkey a report Postgres cannot answer is 503 storage_unavailable (retry). It is written to the warehouse on a clean shutdown — see Stopping the server below, because on Windows the obvious way to stop a process is not one.

There is no scheduled or emailed report and no warehouse push. The CSV is the integration point.

Refusals: when the server will not start

Most misconfigurations are warnings. These five are not, because starting would be worse than stopping: the process would come up looking healthy and answer questions wrongly, or lose on the next restart what it had just accepted. Each one exits non-zero and prints the setting and the way out.

It refuses whenBecauseWhat to do
NODE_ENV=production without both STORE=valkey and SIDESTORE=pgMemory mode keeps nothing across a restart: every deploy would empty the queue, the rooms and the operators, and several instances would each run a queue of their own. The refusal names each missing setting. There is no override.Set STORE=valkey and SIDESTORE=pg, with what they need (SECRET, VALKEY_URL, DATABASE_URL).
SIDESTORE=fileThe file side store was removed in step 5: nothing is written to local disk any more.Use SIDESTORE=pg, or SIDESTORE=memory (the default on STORE=memory) in development.
STORE=valkey without SIDESTORE=pg, SECRET or VALKEY_URLThe queue is journaled to Postgres, every instance must sign and verify with the same key, and there is no Valkey to connect to. The refusal names each missing setting.Set what it names.
SIDESTORE=pg without SECRET or DATABASE_URLWithout SECRET the key is new on every boot, and a new key on every deploy voids every pass (visitor tickets, monitor links, operator keys) that Postgres keeps.Set SECRET (openssl rand -hex 32) and DATABASE_URL.
ADMIN_KEY is the default admin-dev and HOST is set to anything but loopback (0.0.0.0, ::, a real address)The default key is printed in these docs. Listening beyond loopback with it puts the surface that deletes rooms and ejects visitors on the network behind a public password. With HOST unset it does not refuse: it binds 127.0.0.1 only and says so, so node server.js on a laptop works as it always has.Set ADMIN_KEY to a secret (openssl rand -hex 32). For local use only, HOST=127.0.0.1.

A few single settings refuse on their own, and their rows in the environment table say so: an unknown STORE or SIDESTORE value, a missing PG_CA_FILE, a malformed QM_INSTANCE_ID. Everything else — a typo in a numeric variable, a short SECRET, thresholds in the wrong order — starts, and appears in warnings on /api/admin/health. A value that could not be applied as written is never silently substituted: the warning names the variable, what you set, and what is actually in effect.

Rolling back across the token-signing change. A release that predates per-purpose keys reads only the legacy token form. While TOKEN_SIGN_V2 is unset (the default), this release writes nothing else, so rolling back is safe. After TOKEN_SIGN_V2=1, rolling back does not refuse to start, but every v2 token is refused: visitors rejoin at the back, share links and console sessions end, and every operator whose key was minted or re-keyed under v2 gets 401 (mint them new keys). Turning TOKEN_SIGN_V2 off again stops new v2 tokens, and after the token lifetimes run out none are left. It does not convert v2 operator verifiers back, so those operators still need new keys. See Signing keys and rotating SECRET.

Signing keys and rotating SECRET

SECRET is the root of every signature the server makes. Tokens have the format base64url(JSON payload) "." base64url(HMAC-SHA256(key, body)), in one of two forms. Both forms are always accepted. TOKEN_SIGN_V2 only chooses which form this server writes.

for visitor (queue tickets and pre-queue handles), monitor (share links), session (the console's qm_console cookie) and operator (the stored verifier of an operator key). The payload carries kid: the lowercase hex of HKDF-SHA256(ikm = SECRET, salt = empty, info = "qm:kid", 4), eight characters that name the SECRET without revealing it. The kid picks the key a token is checked against. A kid the server does not hold is refused. Two instances agree on signatures exactly when their v2 tickets show the same kid.

What the per-purpose separation does and does not cover. A v2 token minted for one purpose never verifies as another. A legacy token carries no purpose, so it is accepted for any purpose whose payload checks pass. That is exactly as before per-purpose keys existed, and those checks (scope, room, ticket fields) are what keep a visitor ticket from opening the console. While TOKEN_SIGN_V2 is off, every token this server issues is legacy, so the separation is not yet in force. It takes effect for tokens issued after TOKEN_SIGN_V2=1, and fully once the last legacy token has expired.

Two releases. Ship this release with TOKEN_SIGN_V2 unset. It accepts both forms and writes only the legacy one, so it can be rolled back freely. Once it is settled on every instance, set TOKEN_SIGN_V2=1 everywhere. Legacy tokens already issued keep working for the rest of their life, and legacy operator verifiers are re-keyed to v2 as each operator next authenticates. For what changes if you roll back after that, see Rolling back across the token-signing change above.

Rotating. Tokens live up to TOKEN_MAX_AGE_SEC (24 h) for visitors, MONITOR_TOKEN_TTL_SEC (7 days) for share links and 12 h for console sessions. Rotate with TOKEN_SIGN_V2=1:

  1. Set SECRET_PREVIOUS to the current value and SECRET to the new one, on every instance, and restart them. New tokens carry the new kid. Tokens under the old kid keep verifying, so nobody in line loses their place.
  2. Wait out the longest life you care about: 24 h keeps every visitor ticket (the one that matters); 7 days also keeps every share link, which the console can otherwise re-copy under the new key. Operator keys have no expiry. Each operator that authenticates during the window is re-keyed to the new SECRET on that request. An operator key not used during the window stops working when you remove SECRET_PREVIOUS, so mint that operator a new key.
  3. Remove SECRET_PREVIOUS and restart. Tokens under the old kid are now refused.

With TOKEN_SIGN_V2 off, the same steps still work, just without a kid. Legacy tokens name no SECRET, so each one is tried against the raw SECRET and then the raw SECRET_PREVIOUS. Nothing in a ticket tells you, or a log, which SECRET signed it, and nothing will refuse an unknown key by name: such a token just fails both tries. Visitors, share links and console sessions survive the rotation all the same. Operator verifiers are moved to the new SECRET in the form they already had (a v2 verifier is never downgraded). Prefer rotating with TOKEN_SIGN_V2=1. The kid is what lets you check every instance agrees (compare the kid in their tickets) and see when old tokens are gone.

Rotating without SECRET_PREVIOUS invalidates every ticket at once: the queue re-joins at the back (the server logs a warning when refused tickets arrive in a burst), every share link and console session ends, and every operator key stops working. A leaked SECRET is the one case where that is what you want.

Several instances (STORE=valkey)

With STORE=valkey and SIDESTORE=pg the server runs as several identical instances behind one load balancer, with no sticky sessions: a visitor or an operator may reach a different instance on every request and gets the answers one instance would give. Every instance needs the same SECRET, VALKEY_URL, VALKEY_PREFIX and DATABASE_URL; none keeps anything on its own disk. The design is in docs/adr/0001-live-state-in-valkey.md in the repository.

With STORE=memory (development only) several instances share nothing but what you give them. A shared SECRET makes a ticket's signature valid everywhere; it does not share the line. Each instance keeps its own rooms, ticket numbers and rate, and a place in line exists only on the instance that issued it, so one room spread across several makes several independent queues, each letting people through at the full rate.

Readiness (/healthz)

The payload carries two different judgements, and they are not the same question:

ready goes false for exactly three reasons:

A customer's origin being unreachable deliberately does not clear ready, and neither does a clock step. Both are reported as incidents; neither is a reason to deregister the pod. Wiring readiness to them turns one origin outage into a total waiting-room outage — every instance fails its probe at once and the estate empties out of the load balancer, so the component whose whole job is standing in front of a site that is down is the component that disappears. Restarting a pod does not fix somebody else's origin or this machine's clock.

With STORE=valkey the payload also carries backend: {store, valkey, recoveringRooms, ok}. valkey is up or down (this process's connection; also down for 5 s after a command timed out on a connected socket), recoveringRooms counts the rooms being rebuilt from the Postgres journal (marked here, or answered as fenced by Valkey in the last 30 s); /healthz is public, so it carries the number only and /api/admin/health names the rooms (it also carries instance, the name of the instance that answered, and fleet: {instances, partial}, see Metrics and monitoring; the public /healthz carries neither). Either makes status degraded and adds an incident on /api/admin/health; neither clears ready, because every instance shares the one Valkey and taking this one out of rotation fixes nothing. What visitors get meanwhile is onBackendDown.

leader (true or false) says whether this instance holds the leader lease and runs the fleet-wide jobs: target probes, autotune, threshold activation and alert delivery (see LEADER_LEASE_MS). In a STORE=valkey fleet exactly one instance says true, and none does while Valkey is unreachable, which pauses those jobs. It never changes ready or the HTTP code. With STORE=memory it is always true.

Point k8s readiness, your ALB target group and your uptime monitor at /healthz. Alert on status, page on ready. /api/version returns the same payload but always 200: it answers a question about the build, not about readiness, so a rollout script reading it is not tripped by a degraded instance.

Inline deployments: which address the probe uses decides what it gets. A probe that addresses the process — a pod IP, a container port, localhost — asks for a host no room fronts, so /healthz is the queue's own route and answers normally. A probe that addresses the public hostname is a visitor as far as this server is concerned, and /healthz there is the customer's path: measured on an inline host, GET /healthz answers 503 with {"code":"queued"} for ever, because a monitor never gets admitted. Wire that to an ALB target group and it deregisters every healthy instance you have.

Under the reserved prefix it is the queue's route again. On the public hostname, probe /__qm/healthz:

Probe addressesPath to use
the process (pod IP, container port, localhost)/healthz
the public hostname of an inline room/__qm/healthz
the queue's own hostname (snippet deployments)/healthz

The same split applies to /api/version and to /metrics below.

HEAD works everywhere GET does, and answers the same status and the same Content-Length with no body — so a check that defaults to HEAD (many do), or a curl -I on a connector download, gets the truth rather than a 404. It is the same answer, so it makes the same decision: a HEAD /healthz on an inline public hostname is still 503, for the reason above.

When the queue itself is down

ready going false takes one instance out of rotation, which is the normal case and costs nobody anything. This section is about the other one: every instance unready at the same time.

If the queue is deployed inline — Users → Cloudflare → Queue → your app — then "every instance unready" means the site has no path to the app. Somebody has to decide what the visitor gets, and the decision is not the queue's to make at request time; it is a routing rule you write in advance.

The decision taken here: serve a static page, do not fail open. Failing open sends the full surge at an origin with nothing in front of it — the exact outage the queue was bought to prevent, arriving at the worst possible minute. A campaign unreachable for a few minutes is recoverable; a melted origin is not. If your campaign is worth more than that risk, route the failover pool at the app instead and accept what follows. Decide it now, not at 21:04.

public/offline.html is that page. The running server hands it out at /offline.html, and it has to end up somewhere Cloudflare can serve without this server.

It is deliberately built for that job: one self-contained file, no <link>, no <script src>, no image, no webfont — every subresource would be a request to the origin that just stopped answering, and a half-rendered apology looks like a second bug. It carries the same five languages as the waiting page and reads the same qm_lang key, so a visitor who picked Thai on the way in still gets Thai. It retries on its own with jittered backoff (20 s → 60 s, plus up to 80 % random), because everyone landed on it in the same second and a fixed interval would send the whole crowd back in one synchronised wave, hardest exactly as the queue is trying to come up.

It does not show a position, an ETA or a progress bar. The component that holds places is the one that is down, so any number there would be invented.

Wiring it up — the failover Worker (no Load Balancing subscription needed). This server builds and serves it, with its own copy of offline.html already inlined, so the deployed page cannot drift from the shipped one:

curl -fsSL http://localhost:8080/connectors/cloudflare-failover/latest.js -o qm-failover.js
npx wrangler deploy qm-failover.js --name qm-failover \
  --compatibility-date 2025-01-01 --route 'YOUR-QUEUE-HOSTNAME/*'

Or take the whole wrangler project — wrangler.toml, a pre-deploy check, the README — from /connectors/cloudflare-failover/latest.tgz and keep it in your own repo. npm run deploy runs the check first. There is no room id and no API key to configure: this Worker never calls the queue's API.

Route it at the hostname the queue answers on — inline, that is your public hostname. It is not the edge gate connector (/connectors/cloudflare), which decides who waits and belongs in front of your origin in a snippet deployment. Cloudflare runs one Worker per route, so give the two different patterns.

What it acts on, and the two exceptions that make it safe:

The queue answeredWorker does
nothing — refused, DNS, TLS, timeoutoffline page, 503
521–527, 530 (Cloudflare could not use the origin)offline page, 503
503 with X-QM-Queuedpasses through
503 without that headeroffline page, 503
502 / 504passes through
anything elsepasses through

503 alone is never enough to act on. A perfectly healthy gate answers 503 to every poll from a visitor who is still waiting, and marks those with X-QM-Queued. Replacing them with the outage page takes a working queue off the air and swaps a live countdown for "we are down".

502/504 are not this rule's business. Those mean the queue is up and your app is down, and this server already answers them itself — see When the site behind the queue is down below. Serving offline.html there would tell the visitor entry is unavailable when in fact they are through and it is the site that is broken.

After deploying, prove it: stop the queue and load a page. The response carries X-QM-Failover: unreachable (or timeout, or http-<status>), so a curl -I tells a working failover from a coincidence. This is the one piece of the product that is never exercised in normal operation — nothing else will tell you it is broken until the night it matters.

Or a Load Balancer, if you have one: primary pool = the queue instances, health check GET /healthz (that is what it is for), failover pool = wherever you put offline.html. With STORE=memory (development only) the instances do not share a line, so the pool must pin each room to one instance — a second instance is a failover target, not extra capacity for the same room. With STORE=valkey they share it, and the primary pool is every instance (see Several instances (STORE=valkey)).

Either way — the Worker above does all five; a Load Balancer's failover pool has to be configured for them:

The edge gate fails open, and says so

The rule above is for a queue deployed inline. The edge gate connector (/connectors/cloudflare) sits in front of your origin and asks the queue per request; when the queue is slow or failing it fails open — the visitor goes to origin unqueued — because a queue that is down must never be an outage of the site it protects. That stays. What changed (connector 1.1.0, QM-340) is that it is no longer silent:

Each data point is blob1 = room id, blob2 = reason, double1 = 1, index1 = room id.

Alerting on it. The queue server cannot count checks it never received, so the alert has to read the edge's numbers, not the server's. Run this against the Analytics Engine SQL API from whatever already pages you (a cron, your monitoring system's HTTP check), every few minutes, with an API token that has Account Analytics: Read:

curl -s "https://api.cloudflare.com/client/v4/accounts/$CF_ACCOUNT_ID/analytics_engine/sql" \
  -H "Authorization: Bearer $CF_API_TOKEN" --data "
    SELECT blob1 AS room, blob2 AS reason, SUM(_sample_interval) AS fail_opens
    FROM qm_failover
    WHERE timestamp > NOW() - INTERVAL '5' MINUTE
    GROUP BY room, reason
    ORDER BY fail_opens DESC
    FORMAT JSON"

Alert when any row comes back: every one is a visitor who reached the origin without being queued. SUM(_sample_interval), not COUNT(), because Analytics Engine samples at high volume and records the rate it kept. timeout/unreachable/http-5xx means the queue is down or drowning — see /healthz above. For a spot check without the dataset: curl -sI https://shop.example.com/checkout | grep -i x-qm-failover.

Why this and not the Worker reporting its counts back to the queue when it recovers: a Worker's memory is per isolate, isolates are evicted freely and there are many per colo, so a replayed count is lossy exactly during a long outage; it could not report a 401 at all (the key is what is failing); and it would add a write endpoint to the server. Analytics Engine is written at the moment of failure, independently of the queue, and aggregated across every colo.

When the backend is down (onBackendDown)

With STORE=valkey, a room's live state is in Valkey. When Valkey cannot be reached (connection lost, command timeout, failover) or a room is being rebuilt from the Postgres journal after Valkey lost or rolled back its state, the API answers 503 storage_unavailable with Retry-After: /api/check, /api/join, /api/status and the rest, snippet and inline alike. A room that is recovering never answers no_room.

Operator changes around a rebuild. A rebuild replays what the Postgres journal held when it read it. A change that ran in Valkey but reached the journal only after that read is not in the rebuilt room. Creating or editing a room, changing its schedule (opensAt) or deleting it waits for its journal entry, so such a change answers 503 storage_unavailable, although it ran, when the rebuild left it out (or the room is fenced when it lands): retry it once the room answers again, and it applies to the rebuilt room. Pausing, resuming or changing a room's rate (by an operator, autotune or threshold activation) is journaled in the background and answers 200 at once, so one made in the seconds before Valkey lost the room can be undone by the rebuild without an error: after a recovery, check the room's state and rate in the console and set them again if they are not what you left.

Valkey must run with maxmemory-policy noeviction (REQUIRED). Under an evicting policy Valkey can drop part of a room — its admissions, the passed set, ticket ownership — while the room's main hash survives. Nothing detects that partial state, and a visitor whose admission was evicted is sent back to the queue or let through wrongly. With noeviction a full Valkey refuses writes instead, which the API answers as 503. At boot each process reads the policy from INFO memory (CONFIG is disabled on DO managed Valkey), logs a loud warning when it is not noeviction, and reports it as backend.maxmemoryPolicy and an incident on /api/admin/health (a share link sees the incident with the policy withheld, and no maxmemoryPolicy).

An inline room also decides what its own site's visitors get, with the room setting onBackendDown:

ValueBrowser navigationAPI / XHR / fetch
closed (default)the room's waiting page in retry mode: 503 + Retry-After: 5, no place in line taken, a "temporarily unavailable" line, and the page reloads itself; the gate decides again on the reload503 storage_unavailable JSON, Retry-After: 5
openforwarded to proxyOrigin as if the room were in bypass: no queue check, no cookieforwarded the same way

open lets the full surge through to the origin for as long as the backend is down, so choose it only for a site that would rather be slow than closed.

While Valkey is down the gate cannot read the room from it, so it uses the last config this process saw for that host (the inline host index, refreshed on every room change). A host this process has never seen as inline is not gated at all (it is routed as a non-inline request). A process that boots while Valkey is unreachable still starts (it logs valkey: unreachable at boot and keeps reconnecting; /healthz reports valkey: "down") and builds the inline host index from the room configs in Postgres (the rooms table, deleted rooms left out), logging inline host index built from Postgres, so onBackendDown applies to every inline room at once: a closed room's host answers the retry page, never a 404 or the origin. inline.indexFrom on /api/admin/health says postgres meanwhile; once Valkey answers, the rooms check replaces the index with Valkey's copy (valkey). If Postgres cannot be read either, the process logs Valkey and Postgres are both unreachable and gates nothing (its inline hosts answer 404) until one answers: the Postgres read is retried every 1 s, doubling up to 30 s. A WebSocket upgrade on an inline host follows the same setting while the backend is down: closed answers 503 with Retry-After: 5, open forwards the handshake to the origin. Only three kinds of connect error refuse to start, each named without printing the URL: credentials Valkey refuses (WRONGPASS, NOAUTH, NOPERM, ACL errors, an invalid password), a host name that does not resolve (ENOTFOUND), and a TLS certificate that does not verify. Every other connect error boots degraded and keeps reconnecting: refused, timed out or reset, the host or network down, a Valkey that is loading, has lost its master or asks to try again (LOADING, MASTERDOWN, TRYAGAIN), and any error this list does not recognise, such as ERR max number of clients reached or a TLS proxy that is only half up (EPROTO, ERR_SSL_*). An unrecognised one is logged once at boot as a WARNING naming its code or first words. A VALKEY_URL that points at a reachable address with nothing listening therefore boots degraded, so after every deploy alert on backend.valkeyEverConnected on /healthz: it is false until this process has reached Valkey once, and stays true after. It is reported only and never makes the instance unready. Each instance decides from its own last-seen config: a change to onBackendDown (or to the room) made on another instance just before the outage may not have reached this one, so for the length of the outage two instances can answer the same host differently. Before planned Valkey maintenance, make the onBackendDown change at least RESYNC_MS (default 5 s) ahead, so every instance holds the new value. After the first failure that means Valkey is unreachable, each process answers the backend-down reply at once for 5 s (while the connection is not back, or one request at a time probes a Valkey that timed out) instead of waiting for a command timeout per request. Counted per mode in qm_proxy_backend_down_total{mode} and in inline.backendDownOpenTotal / backendDownClosedTotal on /api/admin/health; one log line per room per minute.

Set it like any room field: POST /api/admin/rooms with {"id": "shop", "onBackendDown": "open"}; any other value is 400 invalid_on_backend_down. It is stored with STORE=memory too, where it has no effect: the in-process engine has no backend to lose.

When the site behind the queue is down

The opposite outage, and inline the more likely one: this server is healthy and proxyOrigin is not answering. Nothing needs to be wired up for it — the queue handles it — but you need to recognise it.

queue downsite down
Status503 (from your CDN rule)502 unreachable, 504 timed out
Header—X-QM-Origin: origin_unreachable|origin_timeout
Pagepublic/offline.html, served by the CDNpublic/origin-down.html, served by this server
It saysentry is unavailable, you are not in a queueyou are through, the site is not answering
Chipred: the queue system itself is failingamber: hold on, your place is safe

Neither page says paused. That word is the operator's: on the waiting page it means a person chose to hold the line, and a visitor who has seen both surfaces must be able to tell an outage from that choice.

A visitor who navigates gets the page; a background fetch() gets JSON with the same code, because a document handed to an XHR fails as if the app were buggy. Both carry X-QM-Origin, so one header answers "which side is broken" whatever shape the request was.

origin-down.html follows the same rules as the failover page — one file, no subresources, five languages, jittered retry (15 s → 60 s) — with one difference that matters: it never suggests the visitor has lost their place, because they have not. A retry is a plain reload against a pass they still hold.

Where you will actually see it. The console does not make you go looking:

The numbers behind all of it: inline.originErrorsTotal and inline.originErrors in GET /api/admin/health, and in /metrics:

SeriesAlert on
qm_proxy_origin_errors_total{room,phase}any sustained increase — the site behind that room is failing. connect = refused/unreachable, timeout = accepted and never answered, body_stall = answered and then died mid-response, upgrade = a WebSocket handshake failed. client_reset is the visitor's own socket and is excluded from the health total on purpose.
qm_proxy_gated_total / qm_proxy_bypassed_totala ratio far from ~1:40 — the asset rule matches everything or nothing
qm_proxy_inline_roomsa drop you did not make
qm_join_rate_limited_totalincrease — visitors are being refused a place in line
qm_waiting_page_rate_limited_totalsustained increase — something is pulling the waiting page far faster than a browser would
qm_open_connections / qm_max_connectionsthe first approaching the second
qm_connections_refused_totalany increase — you are at the socket ceiling and shedding
qm_process_resident_bytesgrowth that does not level off

The card, the badge and the counters all say the same thing on purpose: the queue is up, the site is not, and nothing on the rate dial will help.

Stopping the server

A clean stop marks the process not-ready, stops accepting new connections, folds the open hour into the warehouse, hands the leader lease back (STORE=valkey), makes the events journal's last commit to Postgres (STORE=valkey), saves the timeline annotations and records server.stop in the audit trail. A kill does none of that: the hour in progress is lost, and with STORE=valkey the other instances wait for the lease to expire (LEADER_LEASE_MS) before one of them leads. With STORE=memory a stop of either kind loses the queue, the rooms and the operators: nothing is kept across a restart.

Idle and SSE connections get SHUTDOWN_GRACE_MS to finish, then are cut — SSE streams never end on their own. A socket carrying an inline proxied response is left alone at that point, since unlike SSE it does end on its own; it is only bound by SHUTDOWN_DRAIN_MS, the hard ceiling on the whole stop. An inline deployment with no SHUTDOWN_DRAIN_MS set defaults to a 15s floor instead of the plain 2s default, since 2s was measured truncating a real proxied download mid-transfer on every deploy — but an explicit SHUTDOWN_DRAIN_MS always wins, inline or not: it no longer gets silently raised to the 15s floor.

The audit line is honest about what happened: if the final checkpoint did not land — on STORE=valkey, an events-journal commit to Postgres that failed — it records shutdown … WITH STORAGE FAILING and says the changes since the last good commit are not in the events journal, rather than writing the word "clean" over a checkpoint that never happened.

PlatformClean stop
Linux / macOS, foregroundCtrl-C, or kill <pid> (SIGTERM)
Linux / macOS, servicesystemctl stop, docker stop — both send SIGTERM
WindowsCtrl-C in the console window, or the admin API below

Windows has no other clean stop. taskkill /PID <pid> without /F answers This process can only be terminated forcefully (with /F option), and /F is a hard kill the handler never sees. So for anything not running in a console you have your hands on — a service wrapper, a scheduled task, a remote box — use:

curl -sX POST http://localhost:8080/api/admin/shutdown -H "Authorization: Bearer $ADMIN_KEY"

It answers {"ok":true,"stopping":true,"uptimeSec":N} and then exits through the same path SIGTERM takes. It needs the owner role (or ADMIN_KEY) — the same bar as deleting a room, and deliberately out of reach of an operator key that runs the queue.

It stops the one process that answered. With several instances behind a load balancer that is whichever instance the request reached; the others keep serving. To stop a particular instance, send it to that instance's own address, or stop it through the platform.

With STORE=valkey a hard kill is still survivable: the line is in Valkey and every acknowledged join is in the Postgres journal, so no visitor loses their place. What is lost is the part of the current hour that had not been written yet.

Deploy, backup and recovery

Production runs STORE=valkey + SIDESTORE=pg (NODE_ENV=production refuses anything else), and no instance keeps anything on its own disk: the line is in Valkey, and everything that has to survive is in Postgres. The design (rolling deploys, Valkey and Postgres backups, RPO 1 s / RTO 5 min, capacity figures and how they were measured) is in docs/adr/0002-deploy-backup-recovery.md in the repository, as amended for step 5.

Deploy. The instances are interchangeable, so replace them one at a time behind the load balancer: start a new one, wait for its /healthz to answer 200, then stop an old one cleanly (see Stopping the server). Nobody loses their place. Each stop ends the streams on that instance, and the waiting pages reconnect to another (see Deploys under Several instances (STORE=valkey)). Two versions run side by side meanwhile, so every instance must serve every waiting-page asset hash still being handed out (see the next section). With SIDESTORE=pg the schema is migrated at boot. Deploy away from a scheduled drop.

Rollback. Deploy the previous version the same way. A rollback past the release that added kid to tokens re-queues every visitor who joined after the upgrade, because the older build cannot verify those tokens (see Rolling back across the token-signing change under Refusals).

Backup set.

WhatWhy it matters
Postgres (the database at DATABASE_URL)The durable record. The events journal, open_orders and the room_snapshots checkpoints hold every room's queue: it is what a room is rebuilt from. The same database holds operators and monitor links, the audit trail, timeline notes, the 5 s metrics and the hourly history. Back it up with your provider's point-in-time recovery, or pg_dump.
ValkeyThe live line. When Valkey loses or rolls back a room's state, the room is rebuilt from the Postgres journal (see When the backend is down), so a Valkey backup shortens a recovery but is not what makes one possible. Run it with maxmemory-policy noeviction.
SECRET (and SECRET_PREVIOUS during a rotation)Keep it in your secret manager, never in a data backup. Lose it and every ticket, monitor link, console session and operator key stops working at once.

Restoring Postgres. Restore it together with an empty Valkey (or a new VALKEY_PREFIX): every room is then missing from Valkey, and each is rebuilt from the restored journal as it was at the restore point. Every join since that point is lost. Do not restore Postgres under a Valkey that still holds the newer state: a room that is further along in Valkey than in its journal reads as events not yet committed, and is not rebuilt. To check a backup before you need it, restore it into a new database and start a test instance on it with a VALKEY_PREFIX of its own, then compare waiting per room on /api/admin/rooms.

With STORE=memory there is nothing to back up: the server keeps nothing across a restart, and a deploy empties the queue, the rooms and the operators. That is why production refuses it.

What a backup cannot bring back. Turn-notification e-mail addresses (EMAIL_NOTIFY) are deleted when the ticket ends, or up to LAPSED_TTL_MS (24 h) later if its entry window lapsed. They are never written to Postgres. With STORE=memory every restart or deploy drops them all; with STORE=valkey they are in the shared Valkey until then. Valkey RDB snapshots and provider backups can hold a copy after deletion until they rotate out. Nothing in this build sends them or can export them. The ADR describes what production does instead.

Waiting-page assets and live streams across a deploy

Waiting-page assets. The waiting page's style and main script are not inline. They are served at a content-hashed address, /assets/waiting-<hash>.css and /assets/waiting-<hash>.js, where <hash> is the first 22 characters of the base64url SHA-256 of the file. On an inline host they are under /__qm/assets/…. They are sent with Cache-Control: public, max-age=31536000, immutable, so a browser or CDN keeps each one for a year. A new build that changes public/waiting.html gets new hashes. A process serves only the hashes of its own build. An old hash is a 404, and a page that asked for it renders with no style and no script.

So the rule for more than one instance behind one address, and for a rolling deploy: every instance must serve every asset hash that any instance is still handing out. A page rendered by an old instance asks for old hashes, and the load balancer may send that request to a new instance. Ways to meet it:

A CDN that already cached an old hash covers most requests, but a cold edge does not, so do not rely on it.

Visitor streams. An open waiting-page stream (/events) is ended by the server when the token it stands for is due to be re-signed (see TOKEN_MAX_AGE_SEC). The server first sends event: refresh with data {}, and the page reopens the stream itself. The new stream's response re-signs the place and re-sets the cookies. The due time is TOKEN_REFRESH_AFTER_MS after the token was issued: half the shorter of TOKEN_MAX_AGE_SEC and QUEUE_COOKIE_MAX_AGE_SEC. Each stream is ended up to 10% of that window later than the due time. The extra delay is fixed per place (a hash of the place), so a burst of joins that share one issue time does not reconnect all at once. A port must send the same frame and use the same rule. A deploy ends every stream, and pages reconnect by themselves.

Console stream. /api/admin/events sends a full stats frame on connect and again on the next 2 s tick. After that, a tick sends only a delta: {"delta":true,"rooms":[changed rooms],"removed":[ids],…} plus the same totals as a full frame. If rooms changed order (a room deleted and created again under the same id), that tick sends a full frame instead. A console that holds a delta with no full frame under it reconnects, and that is the resync. A deploy ends the stream, so the console gets a full frame again.

How a pass is bound to a visitor

The pass token travels in the URL (?qm_token=), so it is written into every origin and CDN access log. Anything that treats "possession of the token" as "is the visitor" hands your queue to whoever reads those logs. Three layers stop that:

  1. Ticket ownership. A ticket is recorded against the client that queued for it. Presenting somebody else's token gets you a fresh ticket at the back of the line — never their place, and never their pass.
  2. A browser-bound session cookie (qms_<roomId>, HttpOnly, SameSite=Lax, first-party to this server, only its hash is stored). This is the ownership proof, in the same spirit as Queue-it's session cookie: it never appears in a URL or a log, it survives the visitor changing network, and the holder can use it to take their pass back from anyone who managed to claim it.
  3. Client fingerprints for browsers that block cookies: a strict one (socket peer + forwarded identity + User-Agent) and a second tier without the socket peer, so a CDN answering from a different edge node does not lock a visitor out of their own pass.

The second tier is built entirely out of request headers, and those headers — client address, User-Agent — sit on the same access-log line as the token. So it is only computed for a request that arrived from a proxy you named in TRUST_PROXY_IPS. Without that allowlist there is no way to tell a CDN edge node from an attacker replaying a log line, and the strict tier (which pins the socket peer, the one thing a sender cannot choose) is the only identity.

Consequences worth knowing:

What this stores, and what belongs in your privacy notice

Per ticket the server keeps a truncated SHA-256 digest of the client address, the socket peer, the forwarded chain and the User-Agent. The events journal in Postgres (STORE=valkey) additionally records the client address in the clear, on every join. That is all of it: no third-party cookie, no cross-site identifier, no profile, and nothing leaves your server.

It is still personal data under GDPR, because an IP address is — so it needs naming in your privacy notice, with your own counsel deciding the lawful basis and the retention period (which is however long you keep the Postgres events journal and its backups: the journal is not pruned yet). The purpose is narrow and worth stating plainly: the pass token travels in a URL and therefore lands in access logs, and this binding is the only thing stopping whoever reads one of those lines from taking the visitor's place. The customer-facing FAQ in the Integration guide links here.

Passes vs sessions

Two different clocks, per room:

FieldDefaultMeaning
passedTtlSec600How long a promoted visitor has to walk through the door. Unused after that, the pass expires and they rejoin.
sessionTtlSec1800How long they stay inside once admitted. Idle timeout: every /api/check slides it forward (Queue-it's extendCookieValidity).
sessionMaxSecnull (4 h)Hard ceiling on a single visit, however active, counted from walking in. null means 14400 s, or sessionTtlSec if that is longer; set a number to change it.

A session holds exactly one maxConcurrent slot no matter how many times its token is replayed, and releases it when the visitor goes quiet for sessionTtlSec (or is ejected).

The claim window: a third clock, and the one that surprises people.

An unclaimed pass holds a maxConcurrent slot for 120 s (or passedTtlSec if that is shorter), not for the pass's whole life. Otherwise one visitor who was promoted and closed the tab would wedge a slot for the full ten minutes. Past the window the pass is still valid — its holder can still walk in — but it no longer counts against the ceiling, and its holder has to acquire a slot at the door like anybody else.

Entry grant: passedTtlSec (600 s) Claim window: holds a slot, max 120 s Session: every check slides it; it ends after sessionTtlSec (1800 s) idle sessionMaxSec: hard cap from walk-in (unset: 4 h, never below sessionTtlSec) promoted walks in grant ends Not to scale. A pass unused when the grant ends expires, and the visitor rejoins. A live session outlasts the grant.

The clocks in order. Promotion issues the pass, an entry grant valid for passedTtlSec (600 s). For at most its first 120 s (the claim window: shorter if passedTtlSec is, and sooner still when the holder is absent, see presenceSec below) the unused pass holds a maxConcurrent slot. Walking in starts the session: every check slides it forward, and it ends after sessionTtlSec (1800 s) without one, or sessionMaxSec after walking in, whichever comes first. A pass still unused when the grant ends expires, and its holder rejoins; a session already running is not cut short by the grant ending.

The consequence is worth stating plainly, because it looks alarming and is not:

A room with maxConcurrent: 5 can have ten people holding valid, unexpired passes. Five own slots; five gave theirs back. If all ten arrive, exactly five get in and five are turned away — the ceiling is never breached. Bounced pass-holders are admitted strictly oldest-ticket-first as slots free, and keep their place while they keep polling.

Those five are real people and they are not in waiting — they already have a pass. The room card shows them under CAPACITY as +5 holding passes, and the API reports them per room:

FieldMeaning
occupancySlots in use: active sessions + passes still inside their claim window.
passesOutstandingValid passes past the claim window, owning no slot. These people will be turned away if they arrive now.
atDoorThe subset of them being turned away right now (knocking within the last 30 s).
claimWindowSecThe window this room is actually using.

passesOutstanding climbing while occupancy sits at the ceiling is the normal shape of a capacity-bound drop. passesOutstanding climbing while occupancy is below the ceiling means passes are being minted faster than they are being walked — usually promotion outpacing a slow target page. It also climbs when the line is full of ghosts, below.

Absent holders do not hold slots (presenceSec)

Before this rule, one client could freeze a room without ever looking at it. Measured on a room at 600/min with maxConcurrent: 5: one address sent 200 joins with no cookies and never polled. All 200 got a place; the next real visitor got position 201 and 15 seconds later was still at 196, with occupancy 5, active 0 — every slot held by an unclaimed ghost pass for its 120 s claim window, five at a time. Sixty ghosts froze that room for about 24 minutes.

A pass now reserves a slot only for a holder who is demonstrably there:

FieldDefaultMeaning
presenceSec30How long a visitor may go unseen and still have a slot held for them. null switches the rule off for the room.

"Seen" means the visitor joined, polled /api/status, or is holding the /events stream (the waiting page does one of these every 5 s, so an open page is never judged absent). Two consequences:

The same 200-ghost room now lets the real visitor through about presenceSec after the burst; the ghosts' passes show up in passesOutstanding, not in occupancy. A restart treats every ticket as seen at boot (presence is not persisted), so a crash can cost nobody a slot and can give a ghost at most one more presenceSec. Shortening the claim window for everybody was the rejected alternative: it bounces real visitors on every slow target page and still lets a burst of fresh ghosts hold every slot for the shorter window.

Two origin lists, and why they are not one (returnOrigins, tagOrigins)

They used to be a single field, which meant one edit decided two things that fail in opposite directions:

FieldAnswersGetting it wrong
returnOriginsWhere a promoted visitor may be sentAn open redirect on your own domain that hands out pass tokens. Loud, security-owned, kept tight.
tagOriginsWhich pages may read a door check (the browser CORS allowlist)That page is served completely unqueued, silently: the browser discards an answer with no Access-Control-Allow-Origin, the snippet fails open, and the room still looks healthy here.

The room's own targetUrl origin is always on both. tagOrigins defaults to null, which means "use returnOrigins" — every room configured before the split behaves exactly as it did, and you only set it when the two genuinely differ (a tag on a marketing domain that visitors are never returned to, or a narrow return policy that must not widen just to unblock an install).

Set them independently:

curl -X POST http://localhost:8080/api/v1/admin/rooms \
     -H "Authorization: Bearer $ADMIN_KEY" -H 'Content-Type: application/json' \
     -d '{"id":"checkout",
          "returnOrigins":["https://www.shop.example"],
          "tagOrigins":["https://www.shop.example","https://landing.brand.example"]}'

Refusals of the second kind are counted per (room, origin) and shown in Diagnostics and in qm_refused_origin_checks_total. Non-zero means a page carrying your tag is live and unprotected right now. The console's install verifier asks about the origin as well as the URL, and its one-click fix writes tagOrigins only — unblocking a tag is not a decision to widen where visitors may be redirected.

Where a visitor is allowed to land (returnOrigins)

A promoted visitor is sent back to the page they came from, which reaches the waiting room as a qm_return URL — an attacker-supplied string. The server refuses every return URL not covered by the room's policy:

Relative URLs, non-http(s) schemes and URLs carrying embedded credentials (https://user:pass@…) are refused outright. A refused return is not an error for the visitor — they simply land on the room's target instead — so this fails quietly by design. It is counted, and Diagnostics shows the tally as blocked return URLs.

Watch that counter. A missing return origin looks exactly like an attack: if checkout lives on https://shop.example but the room's target is https://www.shop.example, every real visitor is silently dropped on the wrong page. Non-zero means either you have a domain to add, or somebody is trying to use your waiting room as an open redirect.

Scheduled drops

A room can be given a doors-open instant. Before it there is no line at all: arrivals go into a pre-queue and get an opaque handle instead of a ticket number, and at the open instant the whole pre-queue is shuffled with a crypto-grade random draw and becomes the FIFO line.

curl -X POST http://localhost:8080/api/v1/admin/rooms/drop/schedule \
     -H "Authorization: Bearer $ADMIN_KEY" -H 'Content-Type: application/json' \
     -d '{"opensAt":"2026-09-01T09:00:00Z","preQueueMaxPerIp":16}'
FieldMeaning
opensAtnull (no schedule — doors are open), epoch milliseconds, or an ISO datetime.
preQueueMaxPerIpEntries one identity may hold in the pre-queue. Default 16; null for unlimited.

Each field is applied only if you send it. Adjusting the cap mid-drop — -d '{"preQueueMaxPerIp":4}' — changes the cap and leaves the drop scheduled. Cancelling a drop takes an explicit {"opensAt": null}, so a routine tweak can never silently throw the doors open on a queue that was waiting for a draw. A body with neither field is refused (400 nothing_to_change) rather than accepted as a no-op.

Why it works this way:

While the doors are shut, POST /api/join answers state:"scheduled", preQueued:true, a preCount, and opensAt/now so the page can count down against the server's clock. position and ahead are null, because neither exists yet.

Note the two views of a scheduled room. GET /api/admin/rooms reports the room's configured state (active, paused, bypass) alongside opensAt and a preQueue count; the visitor wire reports the effective state, scheduled, until the doors open. The dashboard composes the two into a SCHEDULED badge, an "opens …" time, and an in pre-queue counter in place of waiting.

The operator console

The dashboard at / is where an operator lives during an incident. Everything on it is live — a stats frame arrives every 2 s over SSE — and every control applies immediately. There is no save step.

Per room

ControlWhat it does
Active / Paused / Bypassactive queues normally; paused freezes automatic outflow and holds everyone's place; bypass switches the queue off and lets everyone through. Pausing states its consequence before you confirm.
Rate slider / numberratePerMinute, the smooth per-second outflow. Applies on the next slice, not the next minute.
Let through nPromote n visitors now, on top of the rate — including on a paused room, which is the one way anybody gets in while paused. Deliberate: releasing a few people during a pause is a normal operator action. It releases the front of the line in ticket order — there is no way to pick a particular visitor. The confirm says which of the two applies.
AutotuneDrives the rate from the target site's measured response time instead of your guess; the gear opens its bounds.
InstallThe script tag for this room, ready to paste — or, for a room with a proxyOrigin, the inline deployment checklist instead: where DNS points, the internal address, the reserved /__qm prefix and the failover rule. An inline room is never shown a snippet; there is no page to paste one into. Rooms deployed inline also carry an inline chip on the card.
Waiting pageOpens /w/<room> — exactly what a visitor sees.
Empty queueDiscards the entire waiting backlog without admitting anybody — for a room that has outlived its event. Nobody is promoted, so the chart gains no spike that never happened; purged tickets read as expired and have to join again. A scheduled room's pre-queue is part of the backlog and is emptied too: those visitors are told their pre-queue entry closed and join again. The room, its settings and its history are kept. Irreversible, and the confirm says so.
Edit / DeleteRoom settings (target URL, TTLs, branding, tagOrigins, returnOrigins, schedule) and removal.

Reading a room card

Across rooms

state is one of:

StateMeaning
waitingIn the line. position and est. wait are theirs.
passedHolds a pass that would get them in right now, or is already on the site (on site, with the session clock).
holdingShown as holding pass. A valid pass that gave its slot back (claim window lapsed) while the room is full: turned away at the door until a slot frees, oldest ticket first. These are the room card's holding passes. pass expires in is still running; door position appears while they are actually knocking. Not expired, and Eject voids the pass.
expiredThe engine is finished with the ticket: the pass has expired and no session is live (or the line was emptied before their turn). They have to join again.
ejectedAn operator ejected them.
prequeuedA PQ- ref: in a scheduled room's pre-queue, before the draw. After the draw the same PQ- ref opens the ticket it was drawn into, marked drawn from pre-queue (same room generation only).
unknownThat ticket number was never issued in this room. There is no position or wait to report.

A ref also names the room generation it was issued in. Deleting a room and creating it again under the same id restarts ticket numbers at 1, so a ref read out by a caller from before the reset is not resolved to whoever holds that number today: the lookup says which ticket of which earlier generation it was (previous_generation), and that the caller has to join again. The last 3 generations are checked, over their first 50,000 tickets.

The wait field is labelled by tense, and the tense is the point. waiting is a number still climbing — time since they joined the line. waited is the finished journey, frozen at the instant they were promoted; it does not grow while you read it, so a ticket you look up an hour after the incident still answers how long did this caller wait? with the wait, not the hour. For a room with opensAt, the clock starts when the visitor joined the pre-queue, not when the doors opened — the hour they spent holding is real waiting and every number here says so, including the Actual wait series on the chart and the percentiles in Reports. joined is that instant. passed is when they were let through; it stays after the pass itself expires. A session admitted by an older version has no record of it once the pass expires, and shows passed: unknown rather than a guess.

After a reclaim (the holder takes their session back from a second tab or device), passed and waited still read the original promotion, not the re-grant the reclaim issues.

When the engine has finished with the ticket (the pass has expired and no session is live), state becomes expired, and joined, passed and waited are kept. That covers a visitor who was promoted and never came in, too. The last 50,000 finished tickets per room are kept, oldest dropped first. An eject forgets them straight away.

Share monitor shows a live read-only link, never minting on open (Mint a new link creates one), in a selectable field — copy it from there, or with the button beside it. Minting a second link never revokes the first, so the dialog says so and offers the existing link first. Copy on a row in the list below puts an existing link back in front of you rather than creating another, so the number of live links stays the number of people who are meant to have one. It is audited as monitor_token.reveal, separately from monitor_token.mint, because who else was sent this is a different question from who made it. A revoked or expired link cannot be re-copied: that would be a way to resurrect exactly the credential somebody just took back. Links expire on their own after MONITOR_TOKEN_TTL_SEC — seven days by default; revoking one is separate and takes effect immediately. Revoking an operator closes every link that operator shared (see Operators, roles and the audit trail).

Where the controls live. The header bar holds only what you reach for while a queue is moving: Alerts, Share monitor, Security, Diagnostics, and + New room. Everything you set up between incidents is one click away under Console — Reports, Access, Docs and Sign out. The menu is drawn over the page rather than in it, so opening it never moves a button you were aiming at. On a read-only credential the bar and the menu both shrink to what that credential can actually read, which for a shared monitor link is Security and Reports.

The console never stores ADMIN_KEY: signing in trades it once for an HttpOnly qm_console session cookie (POST /api/admin/session, see the QM-349 note above), and Sign out (under Console) sends DELETE /api/admin/session, which clears the cookie and revokes the session server-side.

Running it from the keyboard

A busy console is well over a hundred tab stops, and a room card is about eighteen of them, so reaching the sixth room by tabbing is not a real option.

Metrics and monitoring

EndpointFormatAuth
GET /metricsPrometheus text expositionADMIN_KEY
GET /api/v1/admin/metrics?roomId=&range=JSON time series for chartsADMIN_KEY
GET /api/admin/healthJSON health, config and warnings (incl. absolute startedAt)ADMIN_KEY
POST /api/admin/shutdownClean stop of the instance that answers — the only one Windows has off-consoleowner / ADMIN_KEY
GET /healthzReadiness: 200 while ready, 503 when storage is failing (STORE=memory; on STORE=valkey only when this instance fails while another is ok), the process is draining, or this instance is at MAX_CONNECTIONSnone
GET /api/versionThe same payload, always 200none

/metrics is not public — it answers 401 without the key, because room ids and live queue depths are commercially sensitive. Give Prometheus a bearer token:

scrape_configs:
  - job_name: queue-manager
    metrics_path: /metrics          # inline, on the public hostname: /__qm/metrics
    authorization: { type: Bearer, credentials: "<ADMIN_KEY>" }
    static_configs: [{ targets: ["qm.weekday100.com"] }]

metrics_path is spelled out because Prometheus defaults it to /metrics, and on the public hostname of an inline room that path belongs to the customer's site: the visitor gate answers the scrape 503 and the job goes up=0 while the server is perfectly healthy. A scrape target is a host and port with no room for a prefix, so the path is the only place to say it — see the probe table under Readiness. Scraping the process directly (pod IP, container port) needs no change.

Per-room series are labelled {room="…"}: qm_room_waiting, qm_room_active_sessions, qm_room_pending_entry, qm_room_occupancy, qm_room_max_concurrent, qm_room_rate_per_minute, qm_room_passed_last_minute, qm_room_oldest_wait_seconds, qm_room_state. Process-wide: qm_rooms, qm_sse_subscribers, qm_uptime_seconds, qm_shutting_down, and the counters worth alerting on — qm_storage_write_errors_total, qm_unknown_room_checks_total, qm_blocked_return_urls_total, qm_refused_key_checks_all_total, qm_rate_limited_checks_all_total and qm_rejected_tokens_all_total.

Several instances (STORE=valkey). Every series also carries instance="<QM_INSTANCE_ID>", and each instance reports only its own process: its own counters, connections and SSE streams (the room gauges read the shared Valkey, so every instance reports the same queue). Scrape each instance and sum in Prometheus. Prometheus sets an instance target label of its own, so with the default honor_labels: false this one arrives as exported_instance; set honor_labels: true on the job to keep the server's name. With honor_labels: true the name must be unique per scrape target: give each instance its own QM_INSTANCE_ID, or leave it unset (i:<random> is new on each boot). Two targets that report the same name write the same series, and Prometheus drops or mixes their samples. Where the platform exposes one address for the whole fleet (DO App Platform), a scrape reaches whichever instance the load balancer picks. The console is the fleet view there: each instance writes its origin-error, protection and diagnostic counters to Valkey every 5 s (<prefix>ctr:<id>, expiring after 15 s), and the console's room cards and GET /api/admin/health show the sum over the live instances. A stopped instance's share drops out of that sum within 15 s. Health's fleet: {instances, partial} says how many instances the sum covers, and partial: true when the other instances' counters could not be read from Valkey (the sum is then this instance's own, plus what it last read within 15 s). Connection and SSE counts in health stay this instance's own, beside its own limits.

Alert on the rate of that last one. A trickle of refused tickets is stale bookmarks; a step change means everyone in a room was ejected at once, which is what a regenerated signing key or a moved clock does — and it is the one failure with no other symptom anywhere. Those visitors are not shown an error; they are quietly put at the back of the line.

The one alert you should not skip is storage: a server that cannot append to its event log must stop handing out tickets, and it does (503 storage_unavailable), but you want to hear about it before your visitors do.

Alerting (paging someone who is not looking at the dashboard)

The console's Alerts button raises desktop notifications in that tab. That is useful while somebody is watching and worthless at 3 a.m. — it dies with the tab and exists once per operator laptop.

Everything below runs inside the queue process. When the process is down, wedged or unreachable, nothing here fires — no webhook, no resolved, nothing in the panel. The only thing that pages you then is an external monitor on /healthz (see Readiness); set one up, because it is the only alert that survives the failure it reports.

The conditions below are evaluated on the server's own clock every ALERT_EVERY_MS whether or not a webhook is configured, and whatever is firing is listed under Diagnostics → Alerting → Firing now either way. With nothing configured that panel is the alerting: the conditions are real and nobody is being paged for them, which the panel says in as many words. Several of them — queue_frozen, outflow_stalled — depend on hold times no dashboard replicates, so this is the only place they appear.

Set ALERT_WEBHOOK_URL and the server also POSTs JSON:

{
  "source": "queue-manager",
  "event": "outflow_stalled",
  "status": "firing",
  "severity": "critical",
  "roomId": "shop",
  "message": "Shop: 412 waiting and nobody let through in the last minute",
  "ts": 1730000000000,
  "since": 1729999880000,
  "text": "[queue-manager] CRITICAL: Shop: 412 waiting and nobody let through in the last minute"
}

text is there so a Slack incoming webhook renders something readable with no transform; the structured fields are for PagerDuty Events v2, Opsgenie and anything else that parses. Events:

eventSeverityFires when
storage_failingcriticalThe event log cannot be written, so joins are being refused. On STORE=valkey this is events-journal commits to Postgres failing in two alert frames within 3 × ALERT_EVERY_MS (one failed commit does not page), until a later commit has landed and that window has passed since the newest failure. With several instances the message names the instance(s) whose storage is failing.
valkey_unreachablecriticalWith STORE=valkey: this instance has not reached Valkey for VALKEY_ALERT_AFTER_MS. Sent by each such instance itself, not by the leader.
target_downcriticalThe server's own probe stops reaching a room's target site.
outflow_stalledcriticalPeople are waiting and nobody has been let through for ALERT_STALL_SEC. The room is trying and failing.
queue_frozencriticalPeople are waiting and the room is paused or set to 0/min for ALERT_FROZEN_SEC — the forgotten paused room. The room is not trying, because a human said so and may not have come back.
queue_deepwarningDepth at or above ALERT_DEPTH (off by default).
wait_highwarningEstimated wait at or above ALERT_WAIT_SEC.

Every condition must hold for ALERT_HOLD_SEC before it fires, re-pages every ALERT_REPEAT_SEC while it is still true, and sends a "status": "resolved" event when it clears — an alert channel that never says "it is over" is one operators learn to ignore.

With STORE=valkey and several instances, every instance evaluates the conditions, so Firing now and GET /api/admin/alerts are the same on any of them, and only the leader (see LEADER_LEASE_MS) POSTs to the webhook. The state that decides a delivery (when a condition started, when it was paged, when it was last re-paged) is kept in Valkey (<prefix>alerts:state), so when the leader dies mid-incident the next one neither pages the open incident again nor forgets to send its resolved. Do not delete <prefix>leader:epoch by hand: the stored state is fenced on it, so if <prefix>alerts:state survives, the epoch restarts below the one saved with that state and every later save of the alert state is refused until the epoch passes the stored one again (or <prefix>alerts:state is deleted too). If you must reset the epoch, delete both keys together (an alert that is still firing is then paged again). Storage is each instance's own: every instance reports its storage status to Valkey (<prefix>alerts:storage) each frame, the leader pages storage_failing when any instance reported failing within the last 3 × ALERT_EVERY_MS and names it (instance on /api/admin/health is each instance's name), and resolves it only once every instance reports healthy. While Valkey is unreachable nothing is delivered except valkey_unreachable, which each instance that cannot reach Valkey sends itself after VALKEY_ALERT_AFTER_MS (so expect one page per instance); an external monitor on /healthz is still the alert that survives the queue itself failing. The delivery log, sent and failed are the answering instance's own: the leader's list is the one with the deliveries in it.

ts is when this page was decided and since when the condition began, which a firing and its resolved share. A resolve can reach the webhook before its firing while the firing is still in flight (ALERT_EVERY_MS below the 5 s webhook timeout, or a receiver that accepts late), so order pages by ts and pair them by event, roomId and since.

GET /api/admin/alerts reports the configuration, what is firing and the last 20 deliveries with their HTTP results. POST /api/admin/alerts (write grant) fires a test event; the console does this from Diagnostics → Alerting → Send test alert, which is the only way to know the path works before you need it. The button is always in that panel — with no ALERT_WEBHOOK_URL set it is there but inert, and says so, rather than not existing at all for the one reader who is following these instructions because nothing is configured yet. It is disabled, saying why, for a credential without the write grant (a share link or a named viewer key), which the server would answer with 403.

Durability

With STORE=valkey the durable record is the events journal in Postgres. Valkey holds the live line, and a room whose Valkey state is lost or rolled back is rebuilt from the journal (see When the backend is down). With STORE=memory nothing is written anywhere: every wait below is met at once, a restart starts empty, and the 503 answers below cannot happen.

Clocks

Outflow pacing uses a monotonic clock, so an NTP correction cannot stall or burst promotions; recorded timestamps use the wall clock but never move backwards, which is what keeps the log and the prune order consistent.

What a clock step does cost is tickets. Every ticket carries the instant it was issued and is refused if it was issued more than one minute in the future or is older than TOKEN_MAX_AGE_SEC (24 hours by default). That check is what stops a ticket from last week being replayed, and it is not negotiable. So a host whose clock moves (a VM booting with a bad RTC before NTP steps it, a machine resumed from a snapshot, a container inheriting a wrong clock) refuses every ticket in the queue at once, and each of those visitors silently rejoins at the back of the line.

The server cannot prevent that, so it names it: a step of more than five seconds is logged, listed in incidents for the following hour, and sets status to degraded on /healthz. The probe still answers 200: taking the pod out of rotation does not fix the host's clock (see Readiness). Watch qm_rejected_tokens_all_total for the size of it. Keep NTP running on the host, and prefer a slew to a step.

No public-internet dependency

Every surface this server hands out — the console, the waiting page and /docs — renders completely with no request to any origin but this one. There is no CDN, no analytics, no external stylesheet and no external script. You can run the whole product on an airgapped network, behind a corporate proxy that blocks everything unlisted, or during the DNS outage that is the reason somebody opened the console in the first place, and it will look exactly the same.

That mattered enough to stop loading the typeface from Google Fonts. Geist and Geist Mono are checked into public/fonts/ and served by this process at /fonts/<file>.woff2:

FileFaceSubset
geist-latin.woff2Geist, variable 400–700latin
geist-latin-ext.woff2Geist, variable 400–700latin-ext
geist-mono-latin.woff2Geist Mono, variable 400–600latin
geist-mono-latin-ext.woff2Geist Mono, variable 400–600latin-ext

To update them, replace the four files, keep the names, and re-run the tests: test/surfaces.test.js fetches each one over HTTP and checks the wOF2 magic and the byte count, so a truncated or mistyped file fails the suite instead of silently falling back to Arial in production.

Running the HTTP tests

Every HTTP suite starts its server through startServer() in test/_harness.js, which runs node server.js with PORT, HOST=127.0.0.1, ADMIN_KEY, SECRET, the limit knobs set to 0 (off), plus whatever the suite adds (see the environment table above). Each server gets a temporary directory of its own (dataDir). Nothing is written there since step 5; it names the server's Valkey run below, so a restart on the same one sees the same state.

On the Valkey store. With STORE=valkey in the test process's environment (and TEST_VALKEY_URL, TEST_DATABASE_URL set), startServer() runs every server on STORE=valkey + SIDESTORE=pg: the test Valkey under a key prefix of its own and a Postgres schema of its own per dataDir, both deleted after the file. It never uses VALKEY_URL or DATABASE_URL from the environment, and refuses Valkey db 0 and a database whose name does not end in _test:

STORE=valkey npx vitest run test/conformance-http.test.js

Which suites pass on Valkey today, and why the others do not yet, is in docs/superpowers/plans/2026-09-25-valkey-step3a-http.md. The tests whose point is state surviving a restart (test/_valkey-only.js) run only on Valkey, since memory mode keeps nothing across a restart. Without TEST_VALKEY_URL and TEST_DATABASE_URL they are skipped, with one warning per file that names each skipped test; with QM_REQUIRE_VALKEY=1 (CI) they fail instead.

Against another server command. QM_SERVER_CMD (test-only, read by test/_harness.js) replaces node server.js with another command, to run the same HTTP suites against another implementation or build. It is split on whitespace (double quotes group a token) and spawned with no shell, with the same environment as above. It must be the server process itself, not a wrapper, because suites kill it and read its exit code, and it must print a line containing listening to stdout once it accepts connections. Unset, the harness runs node server.js from this repo.

test:conformance runs only the files that talk to the server over HTTP and nothing else: no require('../lib/...'), no assertions on its stderr, no decoding of ticket numbers out of a token. These files are: abuse, admin-host-gate, conformance-http, contract, edge-connector, failover, forwarded-proto-trust, forwarded-trust, identity, origin-down, prequeue-forwarded, proxy, snippet, stream-generation, surfaces, timeline and visitor-support. A file with even one white-box block stays out, so a white-box test goes in a sibling <name>.internal.test.js (as the ones split out of proxy, abuse, visitor-support, failover, identity and edge-connector did). The list is the test:conformance script in package.json; a new file that meets the bar goes there.