# Daemons

A **daemon** is one `nymph` process registered with the control plane —
one Worker row in the database. It is a single static Rust binary; the
same binary and the same protocol run on a laptop, an EC2 instance, a
SLURM allocation, and a Kubernetes node. What differs is only *where* it
runs and *how* it got there.

For the CLI verbs that manage daemons, see
[Daemon lifecycle](/nymph/daemon-lifecycle.md). For the wire format, see
[Daemon protocol](/nymph/protocol.md).

## Getting one onto a host

| Path                              | Command                                       | Best for                                  |
| --------------------------------- | --------------------------------------------- | ----------------------------------------- |
| Hand install                      | [`install.sh`](/nymph/installation.md)           | A box you are already sitting on.         |
| SSH bootstrap                     | [`lakeshore daemon install <alias>`](/nymph/install-remote-daemon.md) | A host you can `ssh` to but the control plane cannot reach. |
| Provider launch                   | [`lakeshore daemon launch`](/nymph/daemon-lifecycle.md#launch-through-a-provider) | A VM the control plane provisions for you (EC2 / GCE / SLURM / Kube). |

Once registered, the control plane cannot tell the three apart.

> **Note:** `lakeshore worker start` supervises the **native Python queue worker**
> (`python -m dreamlake.lakeshore.daemon`) in the foreground on your own
> machine. It does not run nymph. Use it for local development against the
> Python SDK; use the paths above for fleet daemons.

## Boot sequence

```text
read config ─▶ detect host ─▶ [introspect server] ─▶ hello ─▶ poll loop ─▶ drain ─▶ exit
                                                                  │
                                                    backoff ◀─────┘ (on error)
```

1. **Config.** Read `udf-daemon.toml` once — from `--config`, else
   `/etc/dreamlake/udf-daemon.toml`, else
   `$HOME/.config/dreamlake/udf-daemon.toml`, else built-in defaults. It
   is never re-read.
2. **Host detection.** Probe the environment once on a blocking thread
   and derive extra tags and runners (see
   [Host detection](#host-detection)).
3. **Introspection.** If `[introspect] enabled = true`, bind the
   loopback status server — deliberately *before* `/hello`, so a daemon
   that cannot reach the control plane is still inspectable. A bind
   failure logs a warning and is not fatal.
4. **Hello.** `POST /v1/daemon/hello` with identity, tags, capabilities,
   versions, and (with an identity configured) the public key. A `reject`
   in the response aborts startup.
5. **Poll loop.** Long-poll, dispatch, repeat.
6. **Drain.** On `SIGTERM` / `SIGINT`, stop polling and let in-flight
   work finish. See [Shutdown](#shutdown-and-drain).

Before `/hello` returns a canonical id, the daemon uses a provisional
worker id of `{machine_id}-{ULID}` — or `{hostname}-{ULID}` when
`[daemon] machine_id` is `"auto"` or empty.

## The poll loop

Every iteration races three signals: the cancellation token, an internal
wake notification fired the instant any exec completes, and the poll
itself. The wake path exists because the control plane pops at most one
queued exec per poll — without it, back-to-back execs would stall a full
`wait_max_s` each.

**Concurrency.** Runs and execs share one task set, bounded by a
semaphore sized to `[runtime] max_invocations` (default 100; `0` means
unbounded). When zero permits are free the daemon stops long-polling
entirely and waits for one to be released, so it never claims work it
cannot start.

**Backoff.** A failed poll waits `backoff + rand(0,1) · min(backoff, 5)`
seconds — full jitter, so a fleet that lost the control plane together
does not come back in lockstep — then doubles up to `[poll] retry_max_s`.
A successful poll resets to `retry_min_s`. `lakeshore daemon reset`
queues a `reset_backoff` command to collapse a long backoff without a
restart.

### Idle policy

`[runtime] keep_alive_s` decides what happens when there is nothing to
do. The idle clock starts only when a poll returned zero commands *and*
nothing is in flight; a completing invocation restarts it.

| Value | Behaviour                                                       |
| ----- | --------------------------------------------------------------- |
| `-1`  | Pool mode. Never idle-exits. (Default for hand installs.)        |
| `0`   | Exit as soon as the idle clock starts — effectively single-job.  |
| `N`   | Exit after `N` continuous idle seconds.                          |

An idle exit routes through the same graceful-drain path `SIGTERM` uses.

### Shutdown and drain

`SIGTERM` or `SIGINT` fires one cancellation token: the poll loop stops
issuing polls, but in-flight work keeps running. The daemon then waits up
to `[shutdown] grace_s` (default 60) for natural completion. If the grace
expires, a **second** token fires; each in-flight dispatch holds a child
of it, and the runners `SIGTERM` their children and return `killed`.
Every outcome is still acked before exit.

## Runners

The runner kind comes from `run_config.runner`, falling back to
`[runtime] default_runner`. An unknown kind does **not** silently fall
back — it fails the invocation with `unknown runner kind: {other}`.

| Kind         | Isolation                                     | Notes                                              |
| ------------ | --------------------------------------------- | -------------------------------------------------- |
| `process`    | None — a child process on the host            | Default. Works anywhere Python is installed.       |
| `subprocess` | None                                          | An accepted **alias** for `process`.               |
| `docker`     | Namespaces + cgroups (`--runtime=runc`)       | Requires `run_config.image`.                        |
| `gvisor`     | User-space kernel (`--runtime=runsc`)         | A 26-line wrapper around the docker runner.        |
| `slurm`      | Whatever the cluster gives the job            | `sbatch` + `squeue` polling.                        |
| `kube`       | Pod                                           | `kubectl apply` + phase polling.                    |

A `mock` runner exists only under the `test-mock-runner` cargo feature.

### Shared workdir contract

`process`, `docker`, `gvisor`, `slurm`, and `kube` all follow the same
per-invocation contract:

- The daemon creates `{[runtime] workdir}/{invocation_id}` (default root
  `/var/lib/dreamlake/udf`).
- The daemon writes `RunBody.args_blob` to `args.msgpack` in that
  directory.
- The child writes `result.msgpack` on success, or `error.json` with
  `{ type, message, traceback }` on failure. `error.json` is surfaced as
  `"<type>: <message>"`.
- On success the workdir is deleted. **On failure it is left on disk** so
  you can inspect it.

Container runners bind-mount the workdir at `/workspace`.

### `process`

Spawns `python3 -m dreamlake.lakeshore.daemon.entry` with `cwd` set to
the per-invocation workdir, stdin null, stdout/stderr inherited into the
daemon log, and `kill_on_drop`. Environment added to the child:
`DREAMLAKE_INVOCATION_ID`, `DREAMLAKE_WORKDIR`, plus
`DREAMLAKE_PARENT_INVOCATION_ID` / `DREAMLAKE_ROOT_INVOCATION_ID` when
set.

Two escape hatches on `RunBody`:

- **`script`** — when set, the runner executes `bash -c <script>` instead
  of the Python entry. Success is decided by exit status alone; there is
  no `result.msgpack` contract and the outcome blob is empty. This is how
  `setup` commands are executed.
- **`as_user`** — wraps the spawn as `runuser -u <as_user> -- <program> …`.
  The username is a separate argv element, never interpolated into a
  shell string.

### `docker` and `gvisor`

`gvisor` is `docker` with `--runtime=runsc`; everything else — workdir
layout, image resolution, env propagation, registry auth, cancellation —
is identical. It needs gVisor installed and registered as a docker
runtime on the host.

The argv is built as:

```text
docker run --rm -i --runtime=<runc|runsc> --cidfile <workdir>/container.cid \
  --network=<network, default host> \
  -v <host_workdir>:/workspace [extra -v …] \
  [caller -e K=V …] \
  -e DREAMLAKE_INVOCATION_ID=… -e DREAMLAKE_WORKDIR=/workspace \
  [raw run_config.docker.args …] [--user <uid>:<gid>] \
  <image> python -m dreamlake.lakeshore.daemon.entry
```

Note the in-container interpreter is `python`, while the `process` runner
uses `python3` on the host.

`run_config.image` is required — without it the invocation fails
immediately with `{ type: "ImageRequired" }`. Recognised sub-keys under
`run_config.docker`: `network`, `env`, `volumes` (`"host:cont"` strings),
`args` (raw docker flags), and `registry_auth`. Anything else is ignored.

`--user` defaults to the daemon's own uid:gid *unless* the caller already
passed `--user` / `-u` in `run_config.docker.args`.

`registry_auth` is `{ server, username, password | password_env }` —
`password_env` wins over a literal `password`. Login runs
`docker login <server> --username <u> --password-stdin` with the secret
on **stdin**, never in argv or the environment, and the password is
redacted from captured stderr. Any auth failure returns
`{ type: "RegistryAuthFailed" }` and the run is not attempted.

Cancellation captures the container id via `--cidfile`, runs
`docker stop <cid>` (SIGTERM, then SIGKILL after 10 s), and kills the
local docker client. `--rm` plus `kill_on_drop` prevent leaked
containers.

### `slurm`

Writes `job.sbatch` into the workdir, runs `sbatch`, parses the id out of
`Submitted batch job 12345`, then polls `squeue -j <id> -h -o '%T'` every
`run_config.slurm.poll_interval_s` (default 5). `scancel <id>` on cancel.
`CD` maps to success; `F` / `CA` / `TO` map to failure; empty `squeue`
output means the job left the queue, so the result / error files decide.

Config keys: `partition`, `time` (default `"01:00:00"`), `gres`, `nodes`,
`ntasks-per-node`, `cpus-per-task`, `mem`, `poll_interval_s`.

### `kube`

Renders a Pod manifest, applies it with `kubectl … apply -f -` on stdin,
polls `kubectl … get pod <name> -o jsonpath='{.status.phase}'` every 5 s,
and `kubectl delete pod <name>` on cancel. Config keys: `image` (default
`python:3.12-slim`), `namespace` (default `default`), `context`, `gpu`,
`cpu`, `memory`, `poll_interval_s`.

> **Warning:** The per-invocation workdir is mounted into the Pod as a **hostPath**
> volume, which assumes the daemon and the scheduling node share a
> filesystem. That holds for single-node kind / minikube / k3d, or for a
> daemon running as a DaemonSet co-located with the workload. Multi-node
> clusters need an RWX PVC, which the source flags as a TODO.

## Host detection

At boot the daemon probes its environment once and merges the result into
the tags and runner list it announces on `/hello`. Tags are emitted as
`key=value` strings.

| Context   | Detection signal                                                                                  | Tags added                                             | Runner added |
| --------- | ------------------------------------------------------------------------------------------------- | ------------------------------------------------------ | ------------ |
| SLURM     | `sinfo` on `PATH` **and** `/etc/slurm/slurm.conf` exists                                           | `host=slurm`, `slurm:cluster=`            | `slurm`      |
| Kube      | `kubectl` on `PATH` **and** `kubectl version --client=false --request-timeout=2s --output=json` exits 0 | `host=kube`, `kube:context=<ctx>`                 | `kube`       |
| AWS       | IMDSv1 `GET http://169.254.169.254/latest/dynamic/instance-identity/document`                      | `host=aws`, `aws:instance-type=<t>`, `region=<r>`      | —            |
| GCP       | `GET http://metadata.google.internal/…/machine-type` with `Metadata-Flavor: Google`                 | `host=gcp`, `gcp:machine-type=<t>`                     | —            |
| Local     | fallback when no host tag fired                                                                    | `host=local`                                            | —            |

Merge rules: config tags **win** on a `key=` collision and detected tags
are appended (deduped by key); the configured `[runtime] runners` list is
authoritative and detected runners are only appended if not already
present.

## Setup commands

A `setup` command pushes a bash script at an already-running daemon,
delivered through the poll channel. It is the supported way to install
docker, gVisor, NVIDIA drivers, or anything else a host needs *after* the
daemon is up.

```bash
lakeshore daemon setup <worker-id> \
  --script ./host-setup/docker.sh --sudo --refresh-capabilities
```

The script content (not the path) is appended to the Worker row's
`pendingSetupCommands` FIFO. The daemon pops one entry per poll onto a
serial queue drained by a dedicated task, so the poll loop is never
blocked by a slow install. `--sudo` sets `as_user = "root"`;
`--refresh-capabilities` re-runs host detection after exit 0 and ships
the new capability snapshot on the next poll, which flips the Worker from
`setting_up` back to `active`.

**Two-phase provisioning.** `lakeshore daemon launch` accepts *two* kinds
of script, and the split matters:

| Phase         | Field              | When                                        | Debuggability                                     |
| ------------- | ------------------ | ------------------------------------------- | ------------------------------------------------- |
| **Bootstrap** | `bootstrap_script` | Cloud-init, as root, **before** nymph starts | Cloud-init log on the VM only. No control-plane visibility. |
| **Setup**     | `setup_commands[]` | After register, FIFO over the poll channel  | Live; the control plane records the outcome; re-runnable. |

Keep `bootstrap_script` to the bare minimum nymph needs to start.
Anything that can fail interestingly belongs in setup, where you can see
it fail and retry without re-launching the VM.

```bash file="host-setup/base.sh"
#!/usr/bin/env bash
# Cloud-init bootstrap. Runs as root, BEFORE nymph starts. Keep it tiny.
set -euo pipefail
apt-get update -y
apt-get install -y --no-install-recommends ca-certificates curl gnupg
```

```bash file="host-setup/docker.sh"
#!/usr/bin/env bash
# Post-register setup. Must be idempotent — it can be redelivered.
set -euo pipefail

if ! command -v docker >/dev/null; then
  apt-get update -y
  apt-get install -y docker.io
fi

systemctl enable --now docker
usermod -aG docker "$(id -un)" || true
docker --version
```

> **Warning:** The daemon sends `cursor: null` on every poll — there is no command
> de-duplication on the wire. A reconnect can redeliver a command the
> daemon already ran. Probe before installing, use `apt-get install -y`,
> prefer `systemctl enable --now`, and read-modify-write configs rather
> than blind-append.

## Mounts

The [`Mount` registry](/get-started/mounts.md) is **metadata only** on both
sides today. The control plane stores mount rows and validates their
config shapes, but nothing attaches them into a job's sandbox.

On the daemon side there is a content-addressed blob cache module
(`nymph/src/mounts.rs`): SHA-256-keyed, two-level directory fan-out,
atomic `rename(2)` inserts, mtime-based LRU eviction against
`[mounts] cache_max_bytes` (default 20 GiB under
`/var/lib/dreamlake/cache`). It has **no call sites yet** — it is
infrastructure waiting for the mount pipeline, not a shipping feature.

## Known gaps

- **No poll cursor.** `PollRequest.cursor` is hard-coded `null`.
- **No stdout tailing for runs.** `InvocationStatus.stdout_tail` is
  always empty. Streaming output exists for `exec` only, via `LogSink`.
- **No per-invocation cancel.** The protocol has no cancel command; the
  control plane returns 409 when you try to kill a *running* invocation.
- **Mounts are unwired**, as above.
- **`startup` (compose) is a no-op.** The `lakeshore.yaml` daemon-group
  `startup:` field is parsed by the CLI and forwarded on the launch body,
  but neither the control plane nor nymph reads it today.

## Read next

- [Daemon lifecycle](/nymph/daemon-lifecycle.md) — install, launch, list,
  exec, update, kill, cleanup.
- [Daemon protocol](/nymph/protocol.md) — the exact wire contract.
- [Daemon configuration](/api/configuration.md) — the `udf-daemon.toml` schema.
- [Telemetry and introspection](/nymph/telemetry.md) — logs and the local
  status endpoint.
- [Compose](/get-started/compose.md) — declaring daemon groups in
  `lakeshore.yaml`.
