DreamLake

Daemons

A daemon is one nymph process registered with the control plane — one Worker row in the database. It is a single static Rust binary; the same binary and the same protocol run on a laptop, an EC2 instance, a SLURM allocation, and a Kubernetes node. What differs is only where it runs and how it got there.

For the CLI verbs that manage daemons, see Daemon lifecycle. For the wire format, see Daemon protocol.

Getting one onto a host

PathCommandBest for
Hand installinstall.shA box you are already sitting on.
SSH bootstraplakeshore daemon install <alias>A host you can ssh to but the control plane cannot reach.
Provider launchlakeshore daemon launchA VM the control plane provisions for you (EC2 / GCE / SLURM / Kube).

Once registered, the control plane cannot tell the three apart.

`lakeshore worker start` is a different thing

lakeshore worker start supervises the native Python queue worker (python -m dreamlake.lakeshore.daemon) in the foreground on your own machine. It does not run nymph. Use it for local development against the Python SDK; use the paths above for fleet daemons.

Boot sequence

read config ─▶ detect host ─▶ [introspect server] ─▶ hello ─▶ poll loop ─▶ drain ─▶ exit
                                                                  │
                                                    backoff ◀─────┘ (on error)
  1. Config. Read udf-daemon.toml once — from --config, else /etc/dreamlake/udf-daemon.toml, else $HOME/.config/dreamlake/udf-daemon.toml, else built-in defaults. It is never re-read.
  2. Host detection. Probe the environment once on a blocking thread and derive extra tags and runners (see Host detection).
  3. Introspection. If [introspect] enabled = true, bind the loopback status server — deliberately before /hello, so a daemon that cannot reach the control plane is still inspectable. A bind failure logs a warning and is not fatal.
  4. Hello. POST /v1/daemon/hello with identity, tags, capabilities, versions, and (with an identity configured) the public key. A reject in the response aborts startup.
  5. Poll loop. Long-poll, dispatch, repeat.
  6. Drain. On SIGTERM / SIGINT, stop polling and let in-flight work finish. See Shutdown.

Before /hello returns a canonical id, the daemon uses a provisional worker id of {machine_id}-{ULID} — or {hostname}-{ULID} when [daemon] machine_id is "auto" or empty.

The poll loop

Every iteration races three signals: the cancellation token, an internal wake notification fired the instant any exec completes, and the poll itself. The wake path exists because the control plane pops at most one queued exec per poll — without it, back-to-back execs would stall a full wait_max_s each.

Concurrency. Runs and execs share one task set, bounded by a semaphore sized to [runtime] max_invocations (default 100; 0 means unbounded). When zero permits are free the daemon stops long-polling entirely and waits for one to be released, so it never claims work it cannot start.

Backoff. A failed poll waits backoff + rand(0,1) · min(backoff, 5) seconds — full jitter, so a fleet that lost the control plane together does not come back in lockstep — then doubles up to [poll] retry_max_s. A successful poll resets to retry_min_s. lakeshore daemon reset queues a reset_backoff command to collapse a long backoff without a restart.

Idle policy

[runtime] keep_alive_s decides what happens when there is nothing to do. The idle clock starts only when a poll returned zero commands and nothing is in flight; a completing invocation restarts it.

ValueBehaviour
-1Pool mode. Never idle-exits. (Default for hand installs.)
0Exit as soon as the idle clock starts — effectively single-job.
NExit after N continuous idle seconds.

An idle exit routes through the same graceful-drain path SIGTERM uses.

Shutdown and drain

SIGTERM or SIGINT fires one cancellation token: the poll loop stops issuing polls, but in-flight work keeps running. The daemon then waits up to [shutdown] grace_s (default 60) for natural completion. If the grace expires, a second token fires; each in-flight dispatch holds a child of it, and the runners SIGTERM their children and return killed. Every outcome is still acked before exit.

Runners

The runner kind comes from run_config.runner, falling back to [runtime] default_runner. An unknown kind does not silently fall back — it fails the invocation with unknown runner kind: {other}.

KindIsolationNotes
processNone — a child process on the hostDefault. Works anywhere Python is installed.
subprocessNoneAn accepted alias for process.
dockerNamespaces + cgroups (--runtime=runc)Requires run_config.image.
gvisorUser-space kernel (--runtime=runsc)A 26-line wrapper around the docker runner.
slurmWhatever the cluster gives the jobsbatch + squeue polling.
kubePodkubectl apply + phase polling.

A mock runner exists only under the test-mock-runner cargo feature.

Shared workdir contract

process, docker, gvisor, slurm, and kube all follow the same per-invocation contract:

  • The daemon creates {[runtime] workdir}/{invocation_id} (default root /var/lib/dreamlake/udf).
  • The daemon writes RunBody.args_blob to args.msgpack in that directory.
  • The child writes result.msgpack on success, or error.json with { type, message, traceback } on failure. error.json is surfaced as "<type>: <message>".
  • On success the workdir is deleted. On failure it is left on disk so you can inspect it.

Container runners bind-mount the workdir at /workspace.

process

Spawns python3 -m dreamlake.lakeshore.daemon.entry with cwd set to the per-invocation workdir, stdin null, stdout/stderr inherited into the daemon log, and kill_on_drop. Environment added to the child: DREAMLAKE_INVOCATION_ID, DREAMLAKE_WORKDIR, plus DREAMLAKE_PARENT_INVOCATION_ID / DREAMLAKE_ROOT_INVOCATION_ID when set.

Two escape hatches on RunBody:

  • script — when set, the runner executes bash -c <script> instead of the Python entry. Success is decided by exit status alone; there is no result.msgpack contract and the outcome blob is empty. This is how setup commands are executed.
  • as_user — wraps the spawn as runuser -u <as_user> -- <program> …. The username is a separate argv element, never interpolated into a shell string.

docker and gvisor

gvisor is docker with --runtime=runsc; everything else — workdir layout, image resolution, env propagation, registry auth, cancellation — is identical. It needs gVisor installed and registered as a docker runtime on the host.

The argv is built as:

docker run --rm -i --runtime=<runc|runsc> --cidfile <workdir>/container.cid \
  --network=<network, default host> \
  -v <host_workdir>:/workspace [extra -v …] \
  [caller -e K=V …] \
  -e DREAMLAKE_INVOCATION_ID=… -e DREAMLAKE_WORKDIR=/workspace \
  [raw run_config.docker.args …] [--user <uid>:<gid>] \
  <image> python -m dreamlake.lakeshore.daemon.entry

Note the in-container interpreter is python, while the process runner uses python3 on the host.

run_config.image is required — without it the invocation fails immediately with { type: "ImageRequired" }. Recognised sub-keys under run_config.docker: network, env, volumes ("host:cont" strings), args (raw docker flags), and registry_auth. Anything else is ignored.

--user defaults to the daemon's own uid:gid unless the caller already passed --user / -u in run_config.docker.args.

registry_auth is { server, username, password | password_env } — password_env wins over a literal password. Login runs docker login <server> --username <u> --password-stdin with the secret on stdin, never in argv or the environment, and the password is redacted from captured stderr. Any auth failure returns { type: "RegistryAuthFailed" } and the run is not attempted.

Cancellation captures the container id via --cidfile, runs docker stop <cid> (SIGTERM, then SIGKILL after 10 s), and kills the local docker client. --rm plus kill_on_drop prevent leaked containers.

slurm

Writes job.sbatch into the workdir, runs sbatch, parses the id out of Submitted batch job 12345, then polls squeue -j <id> -h -o '%T' every run_config.slurm.poll_interval_s (default 5). scancel <id> on cancel. CD maps to success; F / CA / TO map to failure; empty squeue output means the job left the queue, so the result / error files decide.

Config keys: partition, time (default "01:00:00"), gres, nodes, ntasks-per-node, cpus-per-task, mem, poll_interval_s.

kube

Renders a Pod manifest, applies it with kubectl … apply -f - on stdin, polls kubectl … get pod <name> -o jsonpath='{.status.phase}' every 5 s, and kubectl delete pod <name> on cancel. Config keys: image (default python:3.12-slim), namespace (default default), context, gpu, cpu, memory, poll_interval_s.

kube workdir sharing is hostPath

The per-invocation workdir is mounted into the Pod as a hostPath volume, which assumes the daemon and the scheduling node share a filesystem. That holds for single-node kind / minikube / k3d, or for a daemon running as a DaemonSet co-located with the workload. Multi-node clusters need an RWX PVC, which the source flags as a TODO.

Host detection

At boot the daemon probes its environment once and merges the result into the tags and runner list it announces on /hello. Tags are emitted as key=value strings.

ContextDetection signalTags addedRunner added
SLURMsinfo on PATH and /etc/slurm/slurm.conf existshost=slurm, slurm:cluster=<ClusterName>slurm
Kubekubectl on PATH and kubectl version --client=false --request-timeout=2s --output=json exits 0host=kube, kube:context=<ctx>kube
AWSIMDSv1 GET http://169.254.169.254/latest/dynamic/instance-identity/documenthost=aws, aws:instance-type=<t>, region=<r>—
GCPGET http://metadata.google.internal/…/machine-type with Metadata-Flavor: Googlehost=gcp, gcp:machine-type=<t>—
Localfallback when no host tag firedhost=local—

Merge rules: config tags win on a key= collision and detected tags are appended (deduped by key); the configured [runtime] runners list is authoritative and detected runners are only appended if not already present.

Setup commands

A setup command pushes a bash script at an already-running daemon, delivered through the poll channel. It is the supported way to install docker, gVisor, NVIDIA drivers, or anything else a host needs after the daemon is up.

bash
lakeshore daemon setup <worker-id> \
  --script ./host-setup/docker.sh --sudo --refresh-capabilities

The script content (not the path) is appended to the Worker row's pendingSetupCommands FIFO. The daemon pops one entry per poll onto a serial queue drained by a dedicated task, so the poll loop is never blocked by a slow install. --sudo sets as_user = "root"; --refresh-capabilities re-runs host detection after exit 0 and ships the new capability snapshot on the next poll, which flips the Worker from setting_up back to active.

Two-phase provisioning. lakeshore daemon launch accepts two kinds of script, and the split matters:

PhaseFieldWhenDebuggability
Bootstrapbootstrap_scriptCloud-init, as root, before nymph startsCloud-init log on the VM only. No control-plane visibility.
Setupsetup_commands[]After register, FIFO over the poll channelLive; the control plane records the outcome; re-runnable.

Keep bootstrap_script to the bare minimum nymph needs to start. Anything that can fail interestingly belongs in setup, where you can see it fail and retry without re-launching the VM.

host-setup/base.shbash
#!/usr/bin/env bash
# Cloud-init bootstrap. Runs as root, BEFORE nymph starts. Keep it tiny.
set -euo pipefail
apt-get update -y
apt-get install -y --no-install-recommends ca-certificates curl gnupg
host-setup/docker.shbash
#!/usr/bin/env bash
# Post-register setup. Must be idempotent — it can be redelivered.
set -euo pipefail

if ! command -v docker >/dev/null; then
  apt-get update -y
  apt-get install -y docker.io
fi

systemctl enable --now docker
usermod -aG docker "$(id -un)" || true
docker --version
Idempotency is required, not advisory

The daemon sends cursor: null on every poll — there is no command de-duplication on the wire. A reconnect can redeliver a command the daemon already ran. Probe before installing, use apt-get install -y, prefer systemctl enable --now, and read-modify-write configs rather than blind-append.

Mounts

The Mount registry is metadata only on both sides today. The control plane stores mount rows and validates their config shapes, but nothing attaches them into a job's sandbox.

On the daemon side there is a content-addressed blob cache module (nymph/src/mounts.rs): SHA-256-keyed, two-level directory fan-out, atomic rename(2) inserts, mtime-based LRU eviction against [mounts] cache_max_bytes (default 20 GiB under /var/lib/dreamlake/cache). It has no call sites yet — it is infrastructure waiting for the mount pipeline, not a shipping feature.

Known gaps

  • No poll cursor. PollRequest.cursor is hard-coded null.
  • No stdout tailing for runs. InvocationStatus.stdout_tail is always empty. Streaming output exists for exec only, via LogSink.
  • No per-invocation cancel. The protocol has no cancel command; the control plane returns 409 when you try to kill a running invocation.
  • Mounts are unwired, as above.
  • startup (compose) is a no-op. The lakeshore.yaml daemon-group startup: field is parsed by the CLI and forwarded on the launch body, but neither the control plane nor nymph reads it today.

Read next