# Bootstrap a daemon on a remote host

The most common question after the control plane is up: *I have a host I
want jobs to run on, but the control plane can't reach it. How do I get a
daemon there?*

The daemon never needs to be reachable — it long-polls **outbound** to
the control plane. Anywhere you can `ssh` from your laptop, you can drop
a daemon and walk away.

## When to use this

- **HPC login node** behind a VPN your laptop is on but your hosted
  control plane is not.
- **On-prem box** with no public IP.
- **A subnet reachable only from a jump host.**
- **Quick experiments** — a spare laptop, a colleague's dev box, a
  hand-started instance.

If the control plane *can* reach the host — a normal EC2 or GCE instance
you provision through a Provider — use
[`lakeshore daemon launch`](/nymph/daemon-lifecycle.md#launch-through-a-provider)
instead. That path gets you a systemd unit, cloud-init bootstrap, and
serial-console logs.

## The command

```bash
export LAKESHORE_ADMIN_TOKEN=…
lakeshore daemon install bos14
```

The full flag surface lives on
[Daemon lifecycle → SSH bootstrap](/nymph/daemon-lifecycle.md#ssh-bootstrap).
Exit codes: `0` success, `1` the ssh/rsync/mint step failed or the wait
timed out, `2` a configuration error (no admin token, no saved auth).

> **Warning:** The launch-specific paths below require the
> [unique-launch installer fix](https://github.com/dreamlake-ai/lakeshore/pull/25).
> Older versions print no launch ID, share `~/.local/share/nymph.log`, and may
> write flat TOML that nymph cannot parse. Upgrade before running multiple
> installs against the same host. This change does not stop older processes or
> move their files.

## What it does

**Laptop side, step 1 — mint a token.** `POST /v1/namespaces/:ns/tokens`
with your `LAKESHORE_ADMIN_TOKEN` as the bearer. The token name is
`daemon-<sanitized-alias>-<launchId>`, so re-installing the same host
never collides. The plaintext is captured in memory and never written
locally.

**Laptop side, step 2 (only with `--binary`) — stage a local build.**
The path is validated as a real, readable file whose first four bytes are
the ELF magic `\x7fELF` (a Mach-O build fails here with a clear error),
then `rsync -av --progress` copies it to `<alias>:/tmp/nymph-<launchId>`.

**Remote side.** The CLI spawns `ssh -T -o BatchMode=yes <alias> bash -s`
and writes a script to its stdin. That script:

1. Records the SSH user (`whoami`) — everything installs under that
   account's `$HOME`; nothing runs as root.
2. Creates a private `~/.local/share/nymph/launches/<launchId>/` directory
   with mode `0700`. A repeated ID is refused rather than overwriting files.
3. Best-effort `sudo -n usermod -aG docker "$INSTALL_USER"`, so the
   daemon can reach the docker socket without sudo. Never fails the
   install if sudo or the group is missing.
4. Detects the target triple: `uname -s` → `unknown-linux-gnu` /
   `apple-darwin`, `uname -m` → `x86_64` / `aarch64`. Anything else is a
   hard error.
5. Downloads `<base>/dreamlake/nymph/latest/nymph-<triple>` into that
   directory as `nymph` and marks it executable — or, with `--binary`,
   moves `/tmp/nymph-<launchId>` there instead.
6. Writes the minted token beside it as `token`, with mode `0600`.
7. Writes sectioned `udf-daemon.toml` beside the binary and token.
8. Changes into the launch directory and starts `./nymph --config udf-daemon.toml`
   through `setsid nohup`, with stdin disconnected and output redirected to
   `~/.local/share/nymph-<launchId>.log`. With `--slurm`, `sbatch` receives
   `--job-name=lakeshore-nymph-<launchId>`, the same log path, and
   `--chdir` pointing to the private launch directory.

**Laptop side, step 3 — wait.** Unless you passed `--no-wait`, the CLI
polls `GET /v1/namespaces/:ns/workers` every 2 s for up to 60 s, looking
for a row whose label matches **and** whose state is `active`.

Nothing on the remote host needs a package install, a systemd unit, or
root. The binary is self-contained and everything lives under `$HOME`.

The generated config uses nymph's `[server]`, `[daemon]`, and `[runtime]`
sections. The worker label is `<label-prefix>-<launchId>`; the prefix defaults
to the SSH alias, or comes from `--label`. Nymph preserves this explicit label
when registering, so the launch ID reaches the worker row unchanged.

```toml file="~/.local/share/nymph/launches/<launchId>/udf-daemon.toml"
[server]
url = "https://api.lakeshore.dreamlake.ai"
token_file = "token"
namespace = "default"

[daemon]
label = "bos14-<launchId>"
tags = ["installed-from-mac"]

[runtime]
runners = ["process"]
default_runner = "process"
keep_alive_s = -1
```

`token_file` is relative to the process working directory. Both generated
launch commands set that directory; a manual restart must do the same:

```bash
cd ~/.local/share/nymph/launches/<launchId>
./nymph --config udf-daemon.toml
```

See [Daemon configuration](/api/configuration.md) for every key.

## Bootstrap stays minimal — install runtimes afterwards

The SSH heredoc does the absolute minimum to get `nymph` polling. To add
docker, gVisor, an NVIDIA driver, or anything else the daemon should
advertise as a runner, push them as **post-register** setup scripts:

```bash
lakeshore daemon setup <worker-id> \
  --script ./host-setup/docker.sh --sudo --refresh-capabilities
```

The daemon picks each one up on its next long-poll, runs it, and — with
`--refresh-capabilities` — ships back a re-probed capability list so the
scheduler starts routing docker work to it.

Baking runtime installs into the SSH heredoc would be large, hard to
debug (no live log stream until nymph is up), and impossible to retry
without re-running the whole bootstrap. Get nymph up first; install
runtimes from the live control-plane channel. Same two-phase model as
[`daemon launch`](/nymph/daemons.md#setup-commands).

## SLURM

On a login node you usually do not want a long-lived process on the login
node itself. `--slurm` wraps the launch in `sbatch` so the daemon runs as
a regular cluster job: it survives login-node restarts, inherits the
partition's resource limits, and counts against fair-share like anything
else.

```bash
lakeshore daemon install bos14-login --slurm --slurm-partition xeon-g6-volta
```

One ULID is minted per install, including installs without `--slurm`.
The CLI prints it and the log path. For an allocation, use that ID to join
all three views:

| Surface | Name |
| --- | --- |
| Worker label | `<alias-or-label>-<launchId>` |
| Slurm job name | `lakeshore-nymph-<launchId>` |
| Log | `~/.local/share/nymph-<launchId>.log` |

```bash
# On the cluster; copy the ID printed by the installer.
launch_id=01K4G8AX000000000000000001
squeue --name="lakeshore-nymph-$launch_id" -o '%.18i %.48j %.10T'
tail -f "$HOME/.local/share/nymph-$launch_id.log"

# On the machine with CLI auth:
lakeshore daemon list | grep "$launch_id"
```

Two queued allocations keep separate binaries, configs, and tokens. Starting
the second install cannot change the first job's label while it waits.
A new install creates another daemon; it does not replace or stop an old one.

The daemon dies when the allocation's wall-clock expires. Re-run to start a
fresh allocation. Launch directories and logs remain for inspection; remove a
launch's files only after its process or Slurm job has stopped.

## Troubleshooting

| Symptom                                    | Likely cause                                                                                          |
| ------------------------------------------ | ----------------------------------------------------------------------------------------------------- |
| `daemon install` needs `LAKESHORE_ADMIN_TOKEN` | No admin token in the environment. Export it or pass `--admin-token`.                              |
| Exits `2` before SSHing                    | No saved control-plane auth. Run `lakeshore auth login --server <url>` first.                          |
| Never registers within 60 s                | `ssh <host> tail ~/.local/share/nymph-<launchId>.log`. Check the server URL, outbound HTTPS, and token. A Slurm job may still be pending in the queue. |
| `unsupported arch` / `unsupported OS`      | The host is not one of the published targets. Cross-compile locally and use `--binary`.                |
| `--binary path does not look like an ELF executable` | You built for macOS. Cross-compile for `x86_64-unknown-linux-gnu` (or the host's triple).      |
| `rsync` exits non-zero                     | `rsync` is not installed on one side, or the SSH alias does not resolve for rsync's transport.          |

## Read next

- [Daemon lifecycle](/nymph/daemon-lifecycle.md) — every `lakeshore daemon`
  verb with flags.
- [Daemons](/nymph/daemons.md) — runners, host detection, setup commands.
- [Installation](/nymph/installation.md) — the standalone `install.sh` path.
- [Daemon configuration](/api/configuration.md) — the `udf-daemon.toml`
  schema.
