DreamLake

Relay config

Niche / dev-only path

The relay is opt-in and serves a single niche: HPC clusters where compute nodes have no outbound internet but a sibling host (typically a login node) does. The default production setup — daemons talking directly to the Lakeshore control plane on Heroku — does not use a relay. If your compute nodes can reach the control plane URL on outbound 443, skip this page entirely.

The relay is now a subcommand of the nymph daemon binary. There is no separate Python service to install. A relay is a host that has both (a) network reach to the control plane, and (b) shared filesystem visibility into compute nodes that cannot open outbound TCP. The relay watches a directory of named pipes (FIFOs) created by compute-node daemons, reads framed messages out of them, forwards each frame to the control plane over HTTPS, and writes the response back into a sibling FIFO the daemon reads.

                  cluster boundary (egress firewall)
       ┌────────────────────────────────────────────────────┐
       │                                                    │
       │  compute-1  ── req.fifo / resp.fifo ──┐            │
       │  compute-2  ── req.fifo / resp.fifo ──┼─▶ [nymph relay] ─┼─▶ control plane
       │  compute-3  ── req.fifo / resp.fifo ──┘   (login node)  │
       │      (shared filesystem)                                │
       └────────────────────────────────────────────────────┘

The relay is a dumb byte-mover: it does not decode the msgpack payloads, does not buffer beyond a single frame, does not retry, and does not track per-daemon identity. The control plane authenticates each daemon via tokens carried inside the forwarded payload; the relay only carries its own bearer token (used for the relay → control-plane channel).

Why FIFOs instead of HTTPS

Earlier versions of the relay were an HTTPS listener that daemons dialed at the cluster-internal address of a login node. That worked but required TLS keys on the login node and a daemon-side trust anchor for those keys, and it pretended the relay was the control plane (with bespoke per-daemon auth flowing through the relay's listener). The FIFO model is simpler:

  • No TLS material on the relay host. Auth is one bearer token in a file.
  • No firewall rules. Communication is entirely over the shared filesystem.
  • Daemons treat the relay as an asynchronous transport, not a fake control plane. The daemon's own bearer token is embedded inside the forwarded body and is verified at the real control plane.

How daemons reach the relay

Compute-node daemons configure their transport to write into the shared-filesystem FIFOs instead of dialing the control plane directly. The exact daemon-side configuration knob will live under [server] in udf-daemon.toml once it lands; the wire shape is already pinned (see Frame format below).

Running the relay

bash
nymph relay \
  --prefix    /shared/lakeshore/pipes \
  --upstream  https://api.lakeshore.dreamlake.ai \
  --token-file /etc/lakeshore/relay-token

Flags (also accepted via environment variables):

FlagEnv varDefaultNotes
--prefixLAKESHORE_RELAY_PREFIXrequiredDirectory the relay watches. Each daemon owns a subdirectory <prefix>/<machine_id>/ containing req.fifo and resp.fifo.
--upstreamLAKESHORE_RELAY_UPSTREAMrequiredReal control-plane base URL (no trailing slash).
--token-fileLAKESHORE_RELAY_TOKEN_FILEnoneFile holding the bearer token attached to every forwarded request. Loaded once at startup.
--tls-insecureLAKESHORE_RELAY_TLS_INSECUREfalseDisable upstream TLS verification. Local dev only; never in production.

Pipe convention

For each daemon talking through this relay:

<prefix>/<machine_id>/req.fifo      daemon writes  — relay reads
<prefix>/<machine_id>/resp.fifo     relay writes   — daemon reads

<machine_id> is arbitrary — the relay never validates it. It just gives each daemon its own pair of pipes so multiple daemons can share one relay process without head-of-line blocking each other. Create each pair with mkfifo; setup scripts on the cluster typically own that.

Frame format

Each request and response on the wire is a length-prefixed frame:

┌─────────────────────┬─────────────────────────┐
│  4-byte BE u32 len  │  `len` bytes of payload │
└─────────────────────┴─────────────────────────┘

For request frames the payload is a msgpack map of shape:

{
  "endpoint": "/v1/daemon/poll",   # which control-plane endpoint
  "body":      <bytes>,            # pre-encoded msgpack body
}

body is a msgpack bin field carrying the already-encoded body the daemon would have POSTed directly to the control plane. The relay decodes only the outer envelope; the inner body is forwarded verbatim as the HTTP request body, never re-encoded.

Response frames carry the upstream HTTP response body as raw bytes (still length-prefixed, but no envelope) — the daemon decodes those bytes as if it had received the response directly.

The 4-byte length is unsigned big-endian (network byte order). The relay enforces a 256 MiB hard ceiling to refuse obviously-malformed frames.

Auth model

The relay carries one bearer token (loaded from --token-file) which it attaches as Authorization: Bearer … on every forwarded request. Per-daemon authentication is not the relay's job — daemons embed their own tokens inside the msgpack body and the real control plane verifies them. This means:

  • Rotating a daemon's token requires no relay restart.
  • A compromised relay token does not, by itself, let an attacker impersonate a specific daemon — they would also need the daemon's token from inside the body.

Watching

The relay uses the notify crate (FSEvents on macOS, inotify on Linux) to detect req.fifo paths appearing under <prefix>. At startup it also scans the directory once so daemons that created their FIFOs before the relay process started are picked up immediately. For each req.fifo it sees, the relay spawns a tokio task that:

  1. Opens req.fifo (and the sibling resp.fifo) in non-blocking read+write mode, so open(2) returns immediately even if the daemon-side reader/writer hasn't connected yet.
  2. Polls the request fd for readability via tokio's reactor, reads one length-prefixed frame.
  3. POSTs the inner body to the upstream control plane.
  4. Polls the response fd for writability, writes the upstream response back as a length-prefixed frame.
  5. Loops.

Because the I/O is non-blocking and driven by the tokio reactor, a relay shutdown (SIGTERM / SIGINT) terminates all pump tasks cleanly — there are no blocked syscalls holding worker threads.

Operational notes

  • Stateless. A crash interrupts in-flight long-polls; daemons reconnect and the next frame they write opens the new pipe. Restart-safe, no recovery file.
  • One per cluster. Typically one relay per cluster, running on the login node or a dedicated egress host. Multiple relays are fine too — daemons hash to subdirectories, so two relays sharing the same prefix would race on the same pipes; partition the prefix instead.
  • No TLS material on the relay. The relay terminates no TLS; it only originates outbound HTTPS to the upstream. Daemon-side TLS, if any, is moot because daemons do not dial out.
  • mkfifo ownership. Whatever script provisions a daemon also creates its pipe pair. The relay does not auto-create pipes — it only watches and forwards.

See also

  • Architecture — where the relay fits between the Runtime and Fabric layers.
  • Daemon config — set url to the real control plane even when going through a relay.
  • Daemon protocol — the wire shape of the bodies the relay forwards.