DreamLake

Bootstrap a daemon on a remote host

The most common question after the control plane is up: I have a host I want jobs to run on, but the control plane can't reach it. How do I get a daemon there?

The daemon never needs to be reachable — it long-polls outbound to the control plane. Anywhere you can ssh from your laptop, you can drop a daemon and walk away.

When to use this

  • HPC login node behind a VPN your laptop is on but your hosted control plane is not.
  • On-prem box with no public IP.
  • A subnet reachable only from a jump host.
  • Quick experiments — a spare laptop, a colleague's dev box, a hand-started instance.

If the control plane can reach the host — a normal EC2 or GCE instance you provision through a Provider — use lakeshore daemon launch instead. That path gets you a systemd unit, cloud-init bootstrap, and serial-console logs.

The command

bash
export LAKESHORE_ADMIN_TOKEN=…
lakeshore daemon install bos14

The full flag surface lives on Daemon lifecycle → SSH bootstrap. Exit codes: 0 success, 1 the ssh/rsync/mint step failed or the wait timed out, 2 a configuration error (no admin token, no saved auth).

Release compatibility

The launch-specific paths below require the unique-launch installer fix. Older versions print no launch ID, share ~/.local/share/nymph.log, and may write flat TOML that nymph cannot parse. Upgrade before running multiple installs against the same host. This change does not stop older processes or move their files.

What it does

Laptop side, step 1 — mint a token. POST /v1/namespaces/:ns/tokens with your LAKESHORE_ADMIN_TOKEN as the bearer. The token name is daemon-<sanitized-alias>-<launchId>, so re-installing the same host never collides. The plaintext is captured in memory and never written locally.

Laptop side, step 2 (only with --binary) — stage a local build. The path is validated as a real, readable file whose first four bytes are the ELF magic \x7fELF (a Mach-O build fails here with a clear error), then rsync -av --progress copies it to <alias>:/tmp/nymph-<launchId>.

Remote side. The CLI spawns ssh -T -o BatchMode=yes <alias> bash -s and writes a script to its stdin. That script:

  1. Records the SSH user (whoami) — everything installs under that account's $HOME; nothing runs as root.
  2. Creates a private ~/.local/share/nymph/launches/<launchId>/ directory with mode 0700. A repeated ID is refused rather than overwriting files.
  3. Best-effort sudo -n usermod -aG docker "$INSTALL_USER", so the daemon can reach the docker socket without sudo. Never fails the install if sudo or the group is missing.
  4. Detects the target triple: uname -s → unknown-linux-gnu / apple-darwin, uname -m → x86_64 / aarch64. Anything else is a hard error.
  5. Downloads <base>/dreamlake/nymph/latest/nymph-<triple> into that directory as nymph and marks it executable — or, with --binary, moves /tmp/nymph-<launchId> there instead.
  6. Writes the minted token beside it as token, with mode 0600.
  7. Writes sectioned udf-daemon.toml beside the binary and token.
  8. Changes into the launch directory and starts ./nymph --config udf-daemon.toml through setsid nohup, with stdin disconnected and output redirected to ~/.local/share/nymph-<launchId>.log. With --slurm, sbatch receives --job-name=lakeshore-nymph-<launchId>, the same log path, and --chdir pointing to the private launch directory.

Laptop side, step 3 — wait. Unless you passed --no-wait, the CLI polls GET /v1/namespaces/:ns/workers every 2 s for up to 60 s, looking for a row whose label matches and whose state is active.

Nothing on the remote host needs a package install, a systemd unit, or root. The binary is self-contained and everything lives under $HOME.

The generated config uses nymph's [server], [daemon], and [runtime] sections. The worker label is <label-prefix>-<launchId>; the prefix defaults to the SSH alias, or comes from --label. Nymph preserves this explicit label when registering, so the launch ID reaches the worker row unchanged.

~/.local/share/nymph/launches/<launchId>/udf-daemon.tomltoml
[server]
url = "https://api.lakeshore.dreamlake.ai"
token_file = "token"
namespace = "default"

[daemon]
label = "bos14-<launchId>"
tags = ["installed-from-mac"]

[runtime]
runners = ["process"]
default_runner = "process"
keep_alive_s = -1

token_file is relative to the process working directory. Both generated launch commands set that directory; a manual restart must do the same:

bash
cd ~/.local/share/nymph/launches/<launchId>
./nymph --config udf-daemon.toml

See Daemon configuration for every key.

Bootstrap stays minimal — install runtimes afterwards

The SSH heredoc does the absolute minimum to get nymph polling. To add docker, gVisor, an NVIDIA driver, or anything else the daemon should advertise as a runner, push them as post-register setup scripts:

bash
lakeshore daemon setup <worker-id> \
  --script ./host-setup/docker.sh --sudo --refresh-capabilities

The daemon picks each one up on its next long-poll, runs it, and — with --refresh-capabilities — ships back a re-probed capability list so the scheduler starts routing docker work to it.

Baking runtime installs into the SSH heredoc would be large, hard to debug (no live log stream until nymph is up), and impossible to retry without re-running the whole bootstrap. Get nymph up first; install runtimes from the live control-plane channel. Same two-phase model as daemon launch.

SLURM

On a login node you usually do not want a long-lived process on the login node itself. --slurm wraps the launch in sbatch so the daemon runs as a regular cluster job: it survives login-node restarts, inherits the partition's resource limits, and counts against fair-share like anything else.

bash
lakeshore daemon install bos14-login --slurm --slurm-partition xeon-g6-volta

One ULID is minted per install, including installs without --slurm. The CLI prints it and the log path. For an allocation, use that ID to join all three views:

SurfaceName
Worker label<alias-or-label>-<launchId>
Slurm job namelakeshore-nymph-<launchId>
Log~/.local/share/nymph-<launchId>.log
bash
# On the cluster; copy the ID printed by the installer.
launch_id=01K4G8AX000000000000000001
squeue --name="lakeshore-nymph-$launch_id" -o '%.18i %.48j %.10T'
tail -f "$HOME/.local/share/nymph-$launch_id.log"

# On the machine with CLI auth:
lakeshore daemon list | grep "$launch_id"

Two queued allocations keep separate binaries, configs, and tokens. Starting the second install cannot change the first job's label while it waits. A new install creates another daemon; it does not replace or stop an old one.

The daemon dies when the allocation's wall-clock expires. Re-run to start a fresh allocation. Launch directories and logs remain for inspection; remove a launch's files only after its process or Slurm job has stopped.

Troubleshooting

SymptomLikely cause
daemon install needs LAKESHORE_ADMIN_TOKENNo admin token in the environment. Export it or pass --admin-token.
Exits 2 before SSHingNo saved control-plane auth. Run lakeshore auth login --server <url> first.
Never registers within 60 sssh <host> tail ~/.local/share/nymph-<launchId>.log. Check the server URL, outbound HTTPS, and token. A Slurm job may still be pending in the queue.
unsupported arch / unsupported OSThe host is not one of the published targets. Cross-compile locally and use --binary.
--binary path does not look like an ELF executableYou built for macOS. Cross-compile for x86_64-unknown-linux-gnu (or the host's triple).
rsync exits non-zerorsync is not installed on one side, or the SSH alias does not resolve for rsync's transport.

Read next