# 12 · Compose a cluster declaratively

Write one YAML, get a working fleet. `up` reconciles toward the declared
state and is idempotent — re-running with a delta applies only the delta.

> **Warning:** There is no `lakeshore compose` command. It is `lakeshore up`, `down`,
> `ps`, `status`, and `logs`, even though the source lives in
> `src/cli/compose/`.

## Exercises

- `lakeshore up` — parse `lakeshore.yaml`, diff against the control
  plane, apply.
- The daemon-group abstraction — N daemons sharing config and queue
  membership, labelled `<group>-0`, `<group>-1`, …
- `lakeshore ps` / `status` — fleet state scoped by the `project` slug.
- `lakeshore down` — drain (or terminate) the fleet.
- Reconciliation: state is derived from the control plane on every run.
  There is no lock file in v1.

## Requires

- A provider that can provision — EC2, GCE, or Kube. Its credentials
  registered as secrets and referenced from the provider kwargs.
- A `lakeshore.yaml` in the current directory. Lookup is cwd-only; there
  is no upward walk, though `-f, --file <path>` overrides it.

## The file

```yaml file="lakeshore.yaml"
version: 1
project: 08-queues-tutorial
provider: aws-us-east

queues:
  cpu-lane:
    kind: fifo
    description: "CPU pool — batch + smoke"
  gpu-t4:
    kind: priority
    description: "T4 GPUs (g4dn.xlarge)"

daemons:
  cpu:
    queues: [cpu-lane]
    count: 5
    instance_type: t3.medium
    keep_alive_s: 600
    setup:
      - apt-get update -y && apt-get install -y python3-pip
      - pip3 install --break-system-packages dreamlake-lakeshore
  gpu:
    queues: [gpu-t4]
    count: 5
    instance_type: g4dn.xlarge
    image_id: ami-012ba162b9cd2729c
    keep_alive_s: 1800
```

`version: 1`, `project`, and a non-empty `daemons` mapping are required.
`project` and every queue and group name must match
`^[a-z0-9][a-z0-9-]*$`. Inside a daemon group, `queues` and `count` are
required — `count` has no default on purpose, because a silent `1` is
more footgun than convenience — and any key outside the allowed set is a
hard error naming the allowed keys.

Queue `elasticity.kind` in YAML is **underscored**: `fixed`,
`fully_elastic`, `pool_with_threshold`, `max_count`. (The
`queues add --elasticity` CLI flag spells the same values with hyphens.
The two surfaces are not interchangeable.)

## Run

```bash
git clone --recurse-submodules \
  https://github.com/dreamlake-ai/lakeshore-workspace.git
export LAKESHORE_WORKSPACE=~/lakeshore-workspace
export LAKESHORE_URL=http://localhost:8080

cd $LAKESHORE_WORKSPACE/lakeshore-examples/08-queues-cpu-vs-gpu

lakeshore up --plan          # dry run — prints the diff, no side effects
lakeshore up --wait          # apply, block until every daemon is active
lakeshore ps                 # fleet state
lakeshore up cpu             # reconcile one group only
lakeshore down --terminate   # default is drain; --terminate kills instances
```

## Expected output

`--plan`:

```text
Plan for project '08-queues-tutorial':
  queue cpu-lane  (would create)
  queue gpu-t4  (would create)
  daemons cpu  0 → 5  (would launch: cpu-0, cpu-1, cpu-2, cpu-3, cpu-4)
  daemons gpu  0 → 5  (would launch: gpu-0, gpu-1, gpu-2, gpu-3, gpu-4)
```

Applying:

```text
✓ project '08-queues-tutorial': 10 daemon(s) launched, 2 queue(s) created
  + queue cpu-lane
  + queue gpu-t4
  + daemon cpu-0  (6712a…)
  …
```

A second `up` with nothing changed reports `(no change)` per group and
`already present` per queue.

`down`:

```text
✓ project '08-queues-tutorial': 10 worker(s) deleted, 10 cloud instance(s) terminated
  - cpu-0 (6712a…)
  …
```

`--prune-queues` also removes the declared queues; without it they are
kept and the reason is printed.

## How scoping works

Every daemon `up` spawns is stamped with `compose_project: <project>`.
`ps`, `status`, and `down` list workers with
`?compose_project=<project>` and match group membership by the
`<group>-` label prefix. Two projects in the same namespace never see
each other's daemons.

## If it fails

| Symptom | Likely cause |
| ------- | ------------ |
| Parse error naming a key path | The loader validates strictly. The message includes the offending path — e.g. `daemons.gpu.count — required, positive integer (got null)`. |
| `up` reports failures per daemon | The provider could not provision. Each failure line carries the launcher's error; `providers instances <name>` usually shows why. |
| Queue created but no daemons | A crash mid-reconcile. `down` cleans up; if it does not, `lakeshore daemon kill --prefix <group>-` and `lakeshore queues rm <name>`. |
| Re-running `up` re-creates things | A reconciliation bug — file an issue with the `--plan` output from both runs. |

## Status

Manual.

## Next

→ [13 · Cloud provider launch](/admin/happy-paths/13-cloud-provider-launch.md)
