DreamLake

Daemon lifecycle

A daemon is one nymph process registered with the control plane — one Worker row. This page is the operator's map of lakeshore daemon …: how to bring one online, poke at it while it runs, upgrade it, and get rid of it.

For what the daemon does on the host, see Daemons; for the wire format, Daemon protocol.

Bring one online

SSH bootstrap

bash
export LAKESHORE_ADMIN_TOKEN=…
lakeshore daemon install bos14

<ssh-alias> is anything ssh <alias> accepts. The CLI mints a per-daemon token, SSHes in, drops the binary and config under $HOME, starts the process detached, and (by default) waits up to 60 s for the Worker row to appear. Full mechanics on Bootstrap a daemon on a remote host.

FlagDefaultMeaning
--label <name>the SSH aliasLabel prefix; a fresh launch ULID is appended.
--tag <key:value>(none)Capability tag. Repeatable.
--runner <name>processRunner kinds advertised; the first becomes the default. Repeatable.
--keep-alive-s <n>-1Idle policy: -1 pool, 0 single-job, N seconds.
--nymph-base-url <url>the R2 release bucketOverride the binary download base URL.
--binary <path>(download)rsync a local ELF binary instead of downloading. Mutually exclusive with --nymph-base-url.
--admin-token <token>$LAKESHORE_ADMIN_TOKENAdmin token used to mint the per-daemon token.
--no-wait(waits 60 s)Skip the registration poll.
--slurmoffLaunch under sbatch instead of setsid nohup.
--slurm-partition <name>(none)#SBATCH --partition= value; only meaningful with --slurm.
bash
# GPU host with tags and a 10-minute idle window
lakeshore daemon install h100-rig \
  --label gpu-rig --tag gpu:h100 --tag gpu-count:8 \
  --runner docker --runner process --keep-alive-s 600

# SLURM login node
lakeshore daemon install bos14-login --slurm --slurm-partition xeon-g6-volta
SLURM lifetime is the allocation's lifetime

--slurm runs nymph inside an sbatch job, so the daemon dies when the wall-clock runs out. Re-run daemon install <alias> --slurm for a fresh allocation — there is no auto-renewal.

Each install prints a launch ULID and a separate log path. The label is <prefix>-<launchId>; under --slurm, the job name is lakeshore-nymph-<launchId>. See remote installation for config paths and commands to match a worker to its job and logs.

Launch through a provider

When the control plane can reach the host it is provisioning, let it do the work:

bash
lakeshore daemon launch \
  --provider aws-prod \
  --name gpu-rig-east \
  --capacity '{"instance_type": "g5.2xlarge"}' \
  --tag gpu:h100 --queue training \
  --keep-alive 300 --wait

The control plane mints a bootstrap token, renders a cloud-init script (binary into /opt/dreamlake/bin/nymph, config at /etc/dreamlake/udf-daemon.toml, a dreamlake-nymph.service systemd unit with Restart=always), and calls the provider's launcher. Any setup_commands land on the Worker row as a FIFO, so the initial state is setting_up when the queue is non-empty and active otherwise.

FlagMeaning
--provider <name>Provider registered on the control plane.
--name <text>Friendly name — sets the daemon label and the cloud Name tag. With --count > 1 it becomes a prefix (name-00, name-01, …). Defaults to <provider>-<short-ulid>.
--count <n>Launch N daemons in parallel.
--concurrency <n>Max in-flight launches when --count > 1. Default 5 (the AWS RunInstances rate limit); retries on throttle.
--label <text>Worker-row label; --name is usually what you want.
--capacity <json>JSON object passed verbatim to the launcher.
--tag <key:value>Capability tag. Repeatable. Replaces config tags when supplied.
--queue <name>Queue the daemon joins on first /hello. Repeatable. Empty means the default queue.
--runner <name>Runner kinds advertised; the first becomes the default. Repeatable.
--keep-alive <seconds>Idle policy: 0 single-job, N seconds, -1 pool.
--bootstrap-script <path>Cloud-init bash, runs as root before nymph. Content is read locally and sent inline.
--setup-script <path>Post-register setup script. Repeatable; appends to the config list.
--no-configSkip .lakeshore discovery.
--no-setupDrop every setup script, from config and flags alike.
--waitBlock until the daemon registers.

Parallel launches are safe — each call mints its own bootstrap token, so 20 concurrent launches produce 20 distinct Workers.

bash
seq 0 19 | xargs -P 20 -I {} lakeshore daemon launch --name fleet-{}
# or, equivalently:
lakeshore daemon launch --name fleet --count 20

Project defaults — .lakeshore

lakeshore daemon launch reads defaults from a .lakeshore YAML file at the project root, layered with a gitignored .lakeshore.local.

.lakeshoreyaml
daemon:
  provider: my-ec2-pro
  runners: [docker, gvisor]
  tags: [example:docker-gvisor]
  queues: [training]
  keep_alive_s: 1800
  capacity: { instance_type: g5.2xlarge }
  bootstrap_script: ./host-setup/base.sh
  setup_scripts:
    - ./host-setup/docker.sh
    - ./host-setup/gvisor.sh

Discovery walks up from the current directory; the first level where either file exists wins, and the walk stops at the git root (a directory containing .git) or the filesystem root. Script paths resolve relative to the config file's directory, not your cwd, and the CLI ships the script text on the wire.

Precedence is CLI flag > .lakeshore.local > .lakeshore > built-in default. Arrays replace rather than concatenate — the one exception is --setup-script, which appends.

Every field under daemon: is optional: provider, name, runners, tags, queues (legacy synonym lanes), keep_alive_s, capacity, bootstrap_script, setup_scripts. Unknown keys are ignored.

Inspect

bash
lakeshore daemon list                 # every daemon in the namespace
lakeshore daemon list --json
lakeshore daemon list --watch 5       # live-refresh every 5s
lakeshore daemon show <id>            # one row: versions, tags, runners, queues
lakeshore daemon show <id> --json

Both accept --watch [seconds] (default 2; Ctrl-C exits).

A worker whose lastSeenAt is older than 60 s reads back as stale — that is derived at read time and never stored. The persisted states are pending, setting_up, joining, active, and gone.

lakeshore daemon launch-log <id> fetches the cloud serial-console output for a control-plane-launched daemon (EC2 only today) — the tool of choice when a daemon never reached its first /hello. Flags: --json, --tail <n>, --no-latest.

There are also two local-only verbs: lakeshore daemon status (is a daemon running on this machine?) and lakeshore daemon stop (--timeout <seconds>, --force).

Run something on a daemon

bash
lakeshore daemon exec <daemon> nvidia-smi
lakeshore exec --queue training uname -a       # let the CP pick a worker

Everything after the daemon argument is the remote command — no -- separator, but CLI-side flags must come first. The CLI posts the command, then long-polls the control plane until the daemon ships back stdout/stderr/exit code.

FlagMeaning
--idTreat <daemon> as a literal worker id, skipping name lookup.
--timeout <seconds>SIGTERM after N seconds, then SIGKILL 5 s later.
--env <K=V>Extra environment variable, merged on top of the daemon env. Repeatable.
--workdir <path>Override the working directory.
--stdin <path|->Pipe a file (or the CLI's own stdin) to the command.
--wait <seconds>Total CLI wait budget. Default 600.
--jsonEmit the raw envelope instead of streaming text.

lakeshore exec (no daemon subcommand) is the same surface with the daemon argument dropped: it pops an interactive picker, or skips the picker entirely when you pass --queue <name> and lets the control plane route to an eligible member.

Push setup at a running daemon

bash
lakeshore daemon setup <id> --script ./host-setup/docker.sh \
  --sudo --refresh-capabilities --name docker

--script <path> is required; the file's contents go on the wire. --sudo runs it as root, --refresh-capabilities re-probes the host afterwards so a freshly installed runner is advertised, --name <label> is a human label, --json emits the envelope. See Daemons → Setup commands for the idempotency requirements.

Pause, unstick, upgrade

bash
lakeshore daemon hibernate <id> --until +2h     # or an ISO-8601 timestamp
lakeshore daemon reset <id>                     # collapse a long poll backoff
lakeshore daemon update <id> --version 0.1.10

Hibernate accepts a relative duration (+30s, +5m, +1h, +3d) or an ISO-8601 timestamp; a deadline in the past is a 400. The daemon stops polling until the deadline, then re-enters the poll loop from scratch.

Reset queues a reset_backoff command so a daemon stuck in exponential backoff retries at retry_min_s on its next poll — cheaper than a kill-and-relaunch.

Update queues an over-the-air binary update. With ota-self-replace baked in (the default build), the daemon exec()s into the new binary — no manual restart.

FlagMeaning
--version <v>Target version. Omitted → the namespace's target_nymph_version.
--all / --prefix <t>Fan out to every daemon, or every daemon whose label starts with <t>.
--yesSkip the confirmation prompt required by --all / --prefix on a TTY.
--url <u>Override the download URL (the control plane normally derives it).
--sha256 <h>Override the checksum (normally fetched from the <url>.sha256 sidecar).
--add-queue <name>Add the daemon to a queue. Repeatable. Errors if the queue does not exist.
--remove-queue <name>Remove the daemon from a queue. Repeatable.

Queue-only edits skip the OTA path entirely and do not need --version.

The fleet-wide target is set separately:

bash
lakeshore daemon target-version                 # show it
lakeshore daemon set-target-version 0.1.10      # set it
lakeshore daemon set-target-version --clear

Daemons with [runtime] auto_update = true converge on it on their own; everyone else waits for an explicit daemon update.

Kill

bash
lakeshore daemon kill <id>
lakeshore daemon kill <id> --no-terminate       # leave the cloud instance up
lakeshore daemon kill --prefix fleet --yes      # bulk kill by label prefix

By default the control plane also terminates the cloud instance using the provider credentials it holds; --no-terminate removes only the Worker row. --prefix is mutually exclusive with a positional id and prompts for confirmation unless you pass --yes.

Killing the row does not reach into the daemon process itself. A pool daemon whose row disappeared keeps polling until you stop the process yourself (ssh HOST pkill nymph) or its host goes away.

Cleanup

bash
lakeshore daemon cleanup                          # default 24h, prompts
lakeshore daemon cleanup --older-than 2h --dry-run
lakeshore daemon cleanup --older-than 5d --yes
lakeshore daemon cleanup --older-than 6h --json

Deletes every Worker row whose lastSeenAt is older than the threshold.

InputParsed as
30m30 minutes
2h2 hours
5d5 days
2424 hours (a bare number is hours)

The default is 24h and the server cap is 168h (one week) — anything above that, or zero/negative, is rejected with a 400. Without --dry-run and without --yes the CLI prompts before deleting.

Server side this is POST /v1/namespaces/:ns/workers/cleanup with { olderThanHours, dryRun }, returning { count, deleted: [...] }.

Common flows

Clean up after an experiment.

bash
lakeshore daemon list | grep <label>
lakeshore daemon kill <id>
lakeshore daemon cleanup --older-than 1d

Spot-check before pruning.

bash
lakeshore daemon cleanup --older-than 6h --dry-run --json | jq '.deleted'

A daemon launched but never registered.

bash
lakeshore daemon launch-log <id> --tail 100

Read next