# Compose a cluster

> **Example code:** [`08-queues-cpu-vs-gpu`](https://github.com/dreamlake-ai/lakeshore-examples/tree/main/08-queues-cpu-vs-gpu) and [`14-compose-cluster`](https://github.com/dreamlake-ai/lakeshore-examples/tree/main/14-compose-cluster)

Declare the cluster you want — queues, daemon groups, instance types,
counts — in a single `lakeshore.yaml`. One verb reconciles; another
tears it down. The mental model is docker-compose for daemon fleets.

## Prerequisites

| Requirement | Why |
| --- | --- |
| CLI installed (`npm i -g @dreamlake/lakeshore`) | `up` / `down` / `ps` / `status` / `logs` live here. |
| Authenticated (`lakeshore auth login --server <url>`) | The CLI needs a control plane to reconcile against. |
| A provider registered (`lakeshore providers add … --launcher EC2`) | Daemons need somewhere to launch. |

## Step 1 — Write a `lakeshore.yaml`

Create the file in the directory you will run `lakeshore up` from — the
lookup is `./lakeshore.yaml` with no upward walk.

```yaml file="lakeshore.yaml"
version: 1
project: my-first-cluster
provider: ec2-prod

queues:
  work:
    kind: fifo
    description: "General work queue"

daemons:
  worker:
    queues: [work]
    count: 2
    instance_type: t3.medium
    keep_alive_s: 600
```

Every field maps to an existing primitive:

- `project` is written to `composeProject` on every Worker row this file
  spawns, so the compose verbs scope to just this cluster.
- `queues` declares queues to create; creation is idempotent.
- `daemons` declares groups. Each group gets labels `<name>-0` through
  `<name>-(count-1)`.

`version: 1`, `project`, and a `daemons` block with at least one group
are all required. `count` is required per group and has no default.

## Step 2 — Bring it up

```bash
lakeshore up --wait
```

The CLI reads the file, creates any missing queues, then launches the
delta of daemons. `--wait` blocks until each one reaches `state=active`.

Run it again and nothing happens — desired already matches actual. `up`
is idempotent, and state is re-derived from the control plane on every
run (there is no lock file).

Preview first with `lakeshore up --plan`, which prints the plan and takes
no side effects.

## Step 3 — Inspect with `ps` and `status`

```bash
lakeshore ps          # every daemon owned by this compose
lakeshore status      # desired vs actual drift, read-only
```

Both take `--json`. Ad-hoc daemons launched with `lakeshore daemon
launch` have a null `composeProject` and never show up here.

## Step 4 — Scale

Edit `count: 2` to `count: 4` and re-run:

```bash
lakeshore up --wait
```

It sees two existing daemons, wants four, and launches the delta of two.
The originals are untouched. To reconcile only one group:

```bash
lakeshore up worker
```

## Step 5 — Tear it down

```bash
lakeshore down                 # drain — instances stay alive
lakeshore down --terminate     # also terminate the cloud instances
```

Draining is the default: `down` removes the Worker rows this project
owns without killing the machines. `--terminate` kills the EC2 / GCE /
Kube instances too. Ad-hoc daemons are untouched either way.

Optionally prune the queues:

```bash
lakeshore down --terminate --prune-queues
```

`--prune-queues` only deletes a declared queue when it has zero remaining
members, so a queue another project also joined stays alive.

## Multi-queue example

A CPU tier and a GPU tier:

```yaml file="lakeshore.yaml"
version: 1
project: training-cluster
provider: ec2-prod

queues:
  cpu-pool:
    kind: fifo
    description: "CPU preprocessing"
  gpu-t4:
    kind: priority
    description: "T4 GPUs for training"

daemons:
  cpu:
    queues: [cpu-pool]
    count: 5
    instance_type: t3.medium
    keep_alive_s: 600
  gpu:
    queues: [gpu-t4]
    count: 3
    instance_type: g4dn.xlarge
    image_id: ami-012ba162b9cd2729c
    keep_alive_s: 1800
```

```bash
lakeshore up --wait          # 2 queues + 8 daemons
lakeshore up gpu             # re-reconcile only the gpu group
lakeshore up --plan          # dry-run
lakeshore down --terminate
```

## Setup, startup, and Python version

`setup` is a **list** of bash snippets run on the host once after the
daemon registers — they arrive as setup commands and drain FIFO, one per
poll. `startup` is a Python-environment script; `python` names the target
version.

```yaml file="lakeshore.yaml"
version: 1
project: ml-training
provider: ec2-prod

queues:
  training:
    kind: fifo

daemons:
  gpu:
    queues: [training]
    count: 2
    instance_type: g4dn.xlarge
    image_id: ami-012ba162b9cd2729c
    python: "3.12"
    setup:
      - apt-get update -y && apt-get install -y python3-pip
      - pip3 install --break-system-packages dreamlake-lakeshore torch torchvision --index-url https://download.pytorch.org/whl/cu121
    startup: "pip install -q transformers datasets"
```

> **Warning:** A single YAML block scalar under `setup:` fails validation — split it
> into list entries. Any key inside a daemon group other than `queues`,
> `count`, `provider`, `instance_type`, `image_id`, `runner`,
> `keep_alive_s`, `setup`, `startup`, `python`, and `mounts` is a hard
> error naming the allowed set. In particular there is **no `mode:` key**
> and **no `tags:` key** on a daemon group.

## Mounting storage into a group

```yaml file="lakeshore.yaml"
daemons:
  train:
    queues: [training]
    count: 4
    instance_type: g5.2xlarge
    mounts:
      - storage: checkpoints
        mount_path: /mnt/checkpoints
      - storage: datasets
        mount_path: /mnt/data
        readonly: true
```

`storage` names a [Storage](/get-started/storages.md) entry (optionally
`<namespace>/<name>`); `mount_path` and `readonly` are the other two
keys.

## Declaring elasticity

Elasticity lives on the queue, not on the daemon group. In YAML the
kinds are **underscored**:

```yaml file="lakeshore.yaml"
queues:
  training:
    kind: fifo
    elasticity:
      kind: pool_with_threshold
      min: 8
      max: 20
      threshold: 0.8
```

Only `kind`, `min`, `max`, and `threshold` survive the loader —
`pool`, `match`, `steal_delay_s`, `cost_weight`, and any `admission:`
block are silently dropped. Set those by PATCHing the queue over HTTP;
see [Elasticity](/get-started/elasticity.md#setting-the-knobs). And
remember the controller is off unless the control plane runs with
`ELASTICITY_ENABLED=true`.

## Verb reference

| Command | What it does |
| --- | --- |
| `lakeshore up [group]` | Reconcile to desired state (idempotent). |
| `lakeshore up --plan` | Dry-run — print what would change. |
| `lakeshore up --wait` | Block until every spawned daemon is active. |
| `lakeshore down` | Remove this project's workers; leave instances alive. |
| `lakeshore down --terminate` | Also terminate the cloud instances. |
| `lakeshore down --prune-queues` | Delete declared queues that have no members left. |
| `lakeshore ps` | List daemons owned by this compose. |
| `lakeshore status` | Desired vs actual drift (read-only). |
| `lakeshore logs <group\|label>` | Multiplexed cloud-console output. `--follow`, `--poll-ms <n>`. |

All read `./lakeshore.yaml`; override with `-f <path>`. `up`, `down`,
`ps`, and `status` accept `--json`.

## Read next

- [Compose](/get-started/compose.md) — the concept page and full schema.
- [`/dev/compose`](/dev/compose) — the design doc.
- [Queues](/get-started/queues.md) — the primitive compose wraps.
- [Fan out](/get-started/tutorials/fan-out.md) — put work on the queue you
  just created.
