# Compose

A `lakeshore.yaml` at the project root declares the cluster you want to
exist: which queues, which daemon groups, how many of each. One
`lakeshore up` reconciles the actual state to it; `lakeshore down` tears
it back down. Docker-compose for fleets of daemons.

The five verbs are **top-level commands**, not subcommands of a
`compose` group — there is no `lakeshore compose …`.

## When you want this

You have a script in your repo that does this:

```bash
lakeshore queues add cpu-queue --kind fifo
lakeshore queues add gpu-t4    --kind priority
lakeshore daemon launch --queue cpu-queue --provider ec2-cpu --count 5 --name cpu --wait
lakeshore daemon launch --queue gpu-t4    --provider ec2-gpu --count 5 --name gpu --wait
# ... work ...
lakeshore daemon kill --prefix cpu- --yes
lakeshore daemon kill --prefix gpu- --yes
lakeshore queues rm cpu-queue
lakeshore queues rm gpu-t4
```

That sequence is the literal serialization of a document you should
have. With compose it collapses to:

```bash
lakeshore up
# ... work ...
lakeshore down --terminate
```

And the document lives in your repo, version-controlled with the code
that depends on it.

## A minimal `lakeshore.yaml`

```yaml file="lakeshore.yaml"
version: 1
project: hello-cluster
provider: ec2-prod            # project-level default; groups can override

queues:
  cpu-queue:
    kind: fifo
    description: "CPU pool — batch + smoke"
  gpu-t4:
    kind: priority

daemons:
  cpu:                        # label prefix → cpu-0, cpu-1, …
    queues: [cpu-queue]
    count: 5
    instance_type: t3.medium
    keep_alive_s: 600
  gpu:
    queues: [gpu-t4]
    count: 5
    instance_type: g4dn.xlarge
    keep_alive_s: 1800
```

`version: 1` and a slug-shaped `project:` are required, as is a
`daemons:` block with at least one group. The file is looked up as
`./lakeshore.yaml` from the current directory — there is **no upward
walk** in v1.

## Schema

### Top level

| Key | Required | Meaning |
| --- | --- | --- |
| `version` | yes | Must be exactly `1`. |
| `project` | yes | Slug (`^[a-z0-9][a-z0-9-]*$`) tagged onto every Worker row this file spawns. |
| `namespace` | no | Overrides the namespace from `auth.yml`. |
| `provider` | no | Project-level default provider. |
| `queues` | no | Mapping of name → queue spec. |
| `daemons` | **yes** | Mapping of group name → group spec. At least one. |

### Queue spec

| Key | Meaning |
| --- | --- |
| `kind` | `fifo`, `priority`, `boltzmann`, or `filo`. |
| `description` | Free text. |
| `elasticity` | `{ kind?, min?, max?, threshold? }` — and only those four. |
| `daemon_template` | `{ instance_type?, image_id?, runner? }`. |

Elasticity kinds in YAML are **underscored**: `fixed`, `fully_elastic`,
`pool_with_threshold`, `max_count`. (The CLI's `queues add
--elasticity` flag takes the hyphenated spelling; the YAML validator
does not.)

> **Warning:** A queue spec is parsed for `kind`, `description`, `elasticity`, and
> `daemon_template` only, and inside `elasticity` only for `kind`, `min`,
> `max`, and `threshold`. An `admission:` block, or `pool` / `match` /
> `steal_delay_s` / `cost_weight` inside `elasticity`, is silently
> discarded on the way to the control plane rather than rejected. Set
> those by PATCHing the queue over HTTP — see
> [Elasticity](/get-started/elasticity.md#setting-the-knobs).

### Daemon group spec

`queues` (non-empty list of slugs) and `count` (positive integer, no
default) are both required. Everything else is optional:

| Key | Type | Meaning |
| --- | --- | --- |
| `provider` | string | Overrides the project default. |
| `instance_type` | string | e.g. `t3.medium`, `g4dn.xlarge`. |
| `image_id` | string | AMI / image id. |
| `runner` | string | `process`, `docker`, `gvisor`, … |
| `keep_alive_s` | number | Idle-exit seconds. |
| `setup` | **string[]** | Bash snippets run as setup commands on the host after the daemon registers. |
| `startup` | string | Python-environment startup script. |
| `python` | string | Python version, e.g. `"3.12"`. |
| `mounts` | list | `{ storage, mount_path, readonly? }` entries. |

Any other key inside a daemon group is a **hard error** listing the
allowed set — including `mode:`, which is not part of the compose
schema. Unknown *top-level* keys, by contrast, are silently ignored.

## The verbs

```bash
lakeshore up                    # converge to desired state (idempotent)
lakeshore up gpu                # only the `gpu` group
lakeshore up --wait             # block until every spawned daemon is active
lakeshore up --plan             # dry-run — print what would change

lakeshore ps                    # list daemons owned by this compose
lakeshore status                # current vs desired drift, read-only

lakeshore down                  # drain (default) — instances stay alive
lakeshore down --terminate      # also terminate cloud instances
lakeshore down --prune-queues   # delete declared queues that have zero members
lakeshore down --plan           # dry-run

lakeshore logs <group|label>            # multiplexed cloud-console output
lakeshore logs gpu --follow --poll-ms 5000
```

All five read `./lakeshore.yaml`; pass `-f <path>` to point elsewhere.
`up`, `down`, `ps`, and `status` all take `--json`.

## Project isolation

Each file is tagged with its `project:` slug, which is written to the
`composeProject` field on every Worker row it spawns. The compose verbs
scope to that slug, so daemons launched ad-hoc via `lakeshore daemon
launch` — which leave `composeProject` null — are invisible to `ps`,
`status`, and `down`. Same isolation pattern as docker-compose's
`com.docker.compose.project` label.

## Runtime configuration

Three fields shape the execution environment on each group:

```yaml file="lakeshore.yaml"
daemons:
  gpu:
    queues: [training]
    count: 4
    instance_type: g4dn.xlarge
    python: "3.12"
    setup:
      - apt-get update -y && apt-get install -y python3.12-venv
      - pip install torch torchvision --index-url https://download.pytorch.org/whl/cu121
    startup: "pip install -q transformers datasets"
    keep_alive_s: 1800
```

| Field | Runs |
| --- | --- |
| `setup` | On the **host**, once after the daemon registers. Delivered as setup commands, one per poll, drained FIFO. |
| `startup` | Inside the Python execution environment. |
| `python` | Target Python version for UDF execution. |

`setup` must be a **list** of strings. A single YAML block scalar is a
validation error — split the commands into list entries.

See [Daemons](/nymph/daemons.md) for how setup commands are delivered.

## What it does not do

- **Autoscaling.** `count` is a hard number, not a target. Reactive
  scaling is declared on the queue — see
  [Elasticity](/get-started/elasticity.md).
- **Live drift reconciliation.** No background process watches and
  converges. `up` runs only when you invoke it, and state is derived
  from the control plane every run (there is no lock file).
- **Cross-namespace composes.** One file, one namespace.

## See also

- [Compose a cluster](/get-started/tutorials/compose.md) — the hands-on
  walkthrough.
- [`/dev/compose`](/dev/compose) — the design doc: schema, resolution
  order, rollout phases.
- [Queues](/get-started/queues.md) — what `queues:` blocks declare.
- [Daemons](/nymph/daemons.md) — what each `daemons:` entry reconciles to.
