DreamLake

12 · Compose a cluster declaratively

Write one YAML, get a working fleet. up reconciles toward the declared state and is idempotent — re-running with a delta applies only the delta.

The verbs are top level

There is no lakeshore compose command. It is lakeshore up, down, ps, status, and logs, even though the source lives in src/cli/compose/.

Exercises

  • lakeshore up — parse lakeshore.yaml, diff against the control plane, apply.
  • The daemon-group abstraction — N daemons sharing config and queue membership, labelled <group>-0, <group>-1, …
  • lakeshore ps / status — fleet state scoped by the project slug.
  • lakeshore down — drain (or terminate) the fleet.
  • Reconciliation: state is derived from the control plane on every run. There is no lock file in v1.

Requires

  • A provider that can provision — EC2, GCE, or Kube. Its credentials registered as secrets and referenced from the provider kwargs.
  • A lakeshore.yaml in the current directory. Lookup is cwd-only; there is no upward walk, though -f, --file <path> overrides it.

The file

lakeshore.yamlyaml
version: 1
project: 08-queues-tutorial
provider: aws-us-east

queues:
  cpu-lane:
    kind: fifo
    description: "CPU pool — batch + smoke"
  gpu-t4:
    kind: priority
    description: "T4 GPUs (g4dn.xlarge)"

daemons:
  cpu:
    queues: [cpu-lane]
    count: 5
    instance_type: t3.medium
    keep_alive_s: 600
    setup:
      - apt-get update -y && apt-get install -y python3-pip
      - pip3 install --break-system-packages dreamlake-lakeshore
  gpu:
    queues: [gpu-t4]
    count: 5
    instance_type: g4dn.xlarge
    image_id: ami-012ba162b9cd2729c
    keep_alive_s: 1800

version: 1, project, and a non-empty daemons mapping are required. project and every queue and group name must match ^[a-z0-9][a-z0-9-]*$. Inside a daemon group, queues and count are required — count has no default on purpose, because a silent 1 is more footgun than convenience — and any key outside the allowed set is a hard error naming the allowed keys.

Queue elasticity.kind in YAML is underscored: fixed, fully_elastic, pool_with_threshold, max_count. (The queues add --elasticity CLI flag spells the same values with hyphens. The two surfaces are not interchangeable.)

Run

bash
git clone --recurse-submodules \
  https://github.com/dreamlake-ai/lakeshore-workspace.git
export LAKESHORE_WORKSPACE=~/lakeshore-workspace
export LAKESHORE_URL=http://localhost:8080

cd $LAKESHORE_WORKSPACE/lakeshore-examples/08-queues-cpu-vs-gpu

lakeshore up --plan          # dry run — prints the diff, no side effects
lakeshore up --wait          # apply, block until every daemon is active
lakeshore ps                 # fleet state
lakeshore up cpu             # reconcile one group only
lakeshore down --terminate   # default is drain; --terminate kills instances

Expected output

--plan:

Plan for project '08-queues-tutorial':
  queue cpu-lane  (would create)
  queue gpu-t4  (would create)
  daemons cpu  0 → 5  (would launch: cpu-0, cpu-1, cpu-2, cpu-3, cpu-4)
  daemons gpu  0 → 5  (would launch: gpu-0, gpu-1, gpu-2, gpu-3, gpu-4)

Applying:

✓ project '08-queues-tutorial': 10 daemon(s) launched, 2 queue(s) created
  + queue cpu-lane
  + queue gpu-t4
  + daemon cpu-0  (6712a…)
  …

A second up with nothing changed reports (no change) per group and already present per queue.

down:

✓ project '08-queues-tutorial': 10 worker(s) deleted, 10 cloud instance(s) terminated
  - cpu-0 (6712a…)
  …

--prune-queues also removes the declared queues; without it they are kept and the reason is printed.

How scoping works

Every daemon up spawns is stamped with compose_project: <project>. ps, status, and down list workers with ?compose_project=<project> and match group membership by the <group>- label prefix. Two projects in the same namespace never see each other's daemons.

If it fails

SymptomLikely cause
Parse error naming a key pathThe loader validates strictly. The message includes the offending path — e.g. daemons.gpu.count — required, positive integer (got null).
up reports failures per daemonThe provider could not provision. Each failure line carries the launcher's error; providers instances <name> usually shows why.
Queue created but no daemonsA crash mid-reconcile. down cleans up; if it does not, lakeshore daemon kill --prefix <group>- and lakeshore queues rm <name>.
Re-running up re-creates thingsA reconciliation bug — file an issue with the --plan output from both runs.

Status

Manual.

Next

→ 13 · Cloud provider launch