Bootstrap a daemon on a remote host
The most common question after the control plane is up: I have a host I want jobs to run on, but the control plane can't reach it. How do I get a daemon there?
The daemon never needs to be reachable — it long-polls outbound to
the control plane. Anywhere you can ssh from your laptop, you can drop
a daemon and walk away.
When to use this
- HPC login node behind a VPN your laptop is on but your hosted control plane is not.
- On-prem box with no public IP.
- A subnet reachable only from a jump host.
- Quick experiments — a spare laptop, a colleague's dev box, a hand-started instance.
If the control plane can reach the host — a normal EC2 or GCE instance
you provision through a Provider — use
lakeshore daemon launch
instead. That path gets you a systemd unit, cloud-init bootstrap, and
serial-console logs.
The command
The full flag surface lives on
Daemon lifecycle → SSH bootstrap.
Exit codes: 0 success, 1 the ssh/rsync/mint step failed or the wait
timed out, 2 a configuration error (no admin token, no saved auth).
The launch-specific paths below require the
unique-launch installer fix.
Older versions print no launch ID, share ~/.local/share/nymph.log, and may
write flat TOML that nymph cannot parse. Upgrade before running multiple
installs against the same host. This change does not stop older processes or
move their files.
What it does
Laptop side, step 1 — mint a token. POST /v1/namespaces/:ns/tokens
with your LAKESHORE_ADMIN_TOKEN as the bearer. The token name is
daemon-<sanitized-alias>-<launchId>, so re-installing the same host
never collides. The plaintext is captured in memory and never written
locally.
Laptop side, step 2 (only with --binary) — stage a local build.
The path is validated as a real, readable file whose first four bytes are
the ELF magic \x7fELF (a Mach-O build fails here with a clear error),
then rsync -av --progress copies it to <alias>:/tmp/nymph-<launchId>.
Remote side. The CLI spawns ssh -T -o BatchMode=yes <alias> bash -s
and writes a script to its stdin. That script:
- Records the SSH user (
whoami) — everything installs under that account's$HOME; nothing runs as root. - Creates a private
~/.local/share/nymph/launches/<launchId>/directory with mode0700. A repeated ID is refused rather than overwriting files. - Best-effort
sudo -n usermod -aG docker "$INSTALL_USER", so the daemon can reach the docker socket without sudo. Never fails the install if sudo or the group is missing. - Detects the target triple:
uname -s→unknown-linux-gnu/apple-darwin,uname -m→x86_64/aarch64. Anything else is a hard error. - Downloads
<base>/dreamlake/nymph/latest/nymph-<triple>into that directory asnymphand marks it executable — or, with--binary, moves/tmp/nymph-<launchId>there instead. - Writes the minted token beside it as
token, with mode0600. - Writes sectioned
udf-daemon.tomlbeside the binary and token. - Changes into the launch directory and starts
./nymph --config udf-daemon.tomlthroughsetsid nohup, with stdin disconnected and output redirected to~/.local/share/nymph-<launchId>.log. With--slurm,sbatchreceives--job-name=lakeshore-nymph-<launchId>, the same log path, and--chdirpointing to the private launch directory.
Laptop side, step 3 — wait. Unless you passed --no-wait, the CLI
polls GET /v1/namespaces/:ns/workers every 2 s for up to 60 s, looking
for a row whose label matches and whose state is active.
Nothing on the remote host needs a package install, a systemd unit, or
root. The binary is self-contained and everything lives under $HOME.
The generated config uses nymph's [server], [daemon], and [runtime]
sections. The worker label is <label-prefix>-<launchId>; the prefix defaults
to the SSH alias, or comes from --label. Nymph preserves this explicit label
when registering, so the launch ID reaches the worker row unchanged.
token_file is relative to the process working directory. Both generated
launch commands set that directory; a manual restart must do the same:
See Daemon configuration for every key.
Bootstrap stays minimal — install runtimes afterwards
The SSH heredoc does the absolute minimum to get nymph polling. To add
docker, gVisor, an NVIDIA driver, or anything else the daemon should
advertise as a runner, push them as post-register setup scripts:
The daemon picks each one up on its next long-poll, runs it, and — with
--refresh-capabilities — ships back a re-probed capability list so the
scheduler starts routing docker work to it.
Baking runtime installs into the SSH heredoc would be large, hard to
debug (no live log stream until nymph is up), and impossible to retry
without re-running the whole bootstrap. Get nymph up first; install
runtimes from the live control-plane channel. Same two-phase model as
daemon launch.
SLURM
On a login node you usually do not want a long-lived process on the login
node itself. --slurm wraps the launch in sbatch so the daemon runs as
a regular cluster job: it survives login-node restarts, inherits the
partition's resource limits, and counts against fair-share like anything
else.
One ULID is minted per install, including installs without --slurm.
The CLI prints it and the log path. For an allocation, use that ID to join
all three views:
| Surface | Name |
|---|---|
| Worker label | <alias-or-label>-<launchId> |
| Slurm job name | lakeshore-nymph-<launchId> |
| Log | ~/.local/share/nymph-<launchId>.log |
Two queued allocations keep separate binaries, configs, and tokens. Starting the second install cannot change the first job's label while it waits. A new install creates another daemon; it does not replace or stop an old one.
The daemon dies when the allocation's wall-clock expires. Re-run to start a fresh allocation. Launch directories and logs remain for inspection; remove a launch's files only after its process or Slurm job has stopped.
Troubleshooting
| Symptom | Likely cause |
|---|---|
daemon install needs LAKESHORE_ADMIN_TOKEN | No admin token in the environment. Export it or pass --admin-token. |
Exits 2 before SSHing | No saved control-plane auth. Run lakeshore auth login --server <url> first. |
| Never registers within 60 s | ssh <host> tail ~/.local/share/nymph-<launchId>.log. Check the server URL, outbound HTTPS, and token. A Slurm job may still be pending in the queue. |
unsupported arch / unsupported OS | The host is not one of the published targets. Cross-compile locally and use --binary. |
--binary path does not look like an ELF executable | You built for macOS. Cross-compile for x86_64-unknown-linux-gnu (or the host's triple). |
rsync exits non-zero | rsync is not installed on one side, or the SSH alias does not resolve for rsync's transport. |
Read next
- Daemon lifecycle — every
lakeshore daemonverb with flags. - Daemons — runners, host detection, setup commands.
- Installation — the standalone
install.shpath. - Daemon configuration — the
udf-daemon.tomlschema.