DreamLake

Compose a cluster

Example code: 08-queues-cpu-vs-gpu and 14-compose-cluster

Declare the cluster you want — queues, daemon groups, instance types, counts — in a single lakeshore.yaml. One verb reconciles; another tears it down. The mental model is docker-compose for daemon fleets.

Prerequisites

RequirementWhy
CLI installed (npm i -g @dreamlake/lakeshore)up / down / ps / status / logs live here.
Authenticated (lakeshore auth login --server <url>)The CLI needs a control plane to reconcile against.
A provider registered (lakeshore providers add … --launcher EC2)Daemons need somewhere to launch.

Step 1 — Write a lakeshore.yaml

Create the file in the directory you will run lakeshore up from — the lookup is ./lakeshore.yaml with no upward walk.

lakeshore.yamlyaml
version: 1
project: my-first-cluster
provider: ec2-prod

queues:
  work:
    kind: fifo
    description: "General work queue"

daemons:
  worker:
    queues: [work]
    count: 2
    instance_type: t3.medium
    keep_alive_s: 600

Every field maps to an existing primitive:

  • project is written to composeProject on every Worker row this file spawns, so the compose verbs scope to just this cluster.
  • queues declares queues to create; creation is idempotent.
  • daemons declares groups. Each group gets labels <name>-0 through <name>-(count-1).

version: 1, project, and a daemons block with at least one group are all required. count is required per group and has no default.

Step 2 — Bring it up

bash
lakeshore up --wait

The CLI reads the file, creates any missing queues, then launches the delta of daemons. --wait blocks until each one reaches state=active.

Run it again and nothing happens — desired already matches actual. up is idempotent, and state is re-derived from the control plane on every run (there is no lock file).

Preview first with lakeshore up --plan, which prints the plan and takes no side effects.

Step 3 — Inspect with ps and status

bash
lakeshore ps          # every daemon owned by this compose
lakeshore status      # desired vs actual drift, read-only

Both take --json. Ad-hoc daemons launched with lakeshore daemon launch have a null composeProject and never show up here.

Step 4 — Scale

Edit count: 2 to count: 4 and re-run:

bash
lakeshore up --wait

It sees two existing daemons, wants four, and launches the delta of two. The originals are untouched. To reconcile only one group:

bash
lakeshore up worker

Step 5 — Tear it down

bash
lakeshore down                 # drain — instances stay alive
lakeshore down --terminate     # also terminate the cloud instances

Draining is the default: down removes the Worker rows this project owns without killing the machines. --terminate kills the EC2 / GCE / Kube instances too. Ad-hoc daemons are untouched either way.

Optionally prune the queues:

bash
lakeshore down --terminate --prune-queues

--prune-queues only deletes a declared queue when it has zero remaining members, so a queue another project also joined stays alive.

Multi-queue example

A CPU tier and a GPU tier:

lakeshore.yamlyaml
version: 1
project: training-cluster
provider: ec2-prod

queues:
  cpu-pool:
    kind: fifo
    description: "CPU preprocessing"
  gpu-t4:
    kind: priority
    description: "T4 GPUs for training"

daemons:
  cpu:
    queues: [cpu-pool]
    count: 5
    instance_type: t3.medium
    keep_alive_s: 600
  gpu:
    queues: [gpu-t4]
    count: 3
    instance_type: g4dn.xlarge
    image_id: ami-012ba162b9cd2729c
    keep_alive_s: 1800
bash
lakeshore up --wait          # 2 queues + 8 daemons
lakeshore up gpu             # re-reconcile only the gpu group
lakeshore up --plan          # dry-run
lakeshore down --terminate

Setup, startup, and Python version

setup is a list of bash snippets run on the host once after the daemon registers — they arrive as setup commands and drain FIFO, one per poll. startup is a Python-environment script; python names the target version.

lakeshore.yamlyaml
version: 1
project: ml-training
provider: ec2-prod

queues:
  training:
    kind: fifo

daemons:
  gpu:
    queues: [training]
    count: 2
    instance_type: g4dn.xlarge
    image_id: ami-012ba162b9cd2729c
    python: "3.12"
    setup:
      - apt-get update -y && apt-get install -y python3-pip
      - pip3 install --break-system-packages dreamlake-lakeshore torch torchvision --index-url https://download.pytorch.org/whl/cu121
    startup: "pip install -q transformers datasets"
setup must be a list, and unknown keys are fatal

A single YAML block scalar under setup: fails validation — split it into list entries. Any key inside a daemon group other than queues, count, provider, instance_type, image_id, runner, keep_alive_s, setup, startup, python, and mounts is a hard error naming the allowed set. In particular there is no mode: key and no tags: key on a daemon group.

Mounting storage into a group

lakeshore.yamlyaml
daemons:
  train:
    queues: [training]
    count: 4
    instance_type: g5.2xlarge
    mounts:
      - storage: checkpoints
        mount_path: /mnt/checkpoints
      - storage: datasets
        mount_path: /mnt/data
        readonly: true

storage names a Storage entry (optionally <namespace>/<name>); mount_path and readonly are the other two keys.

Declaring elasticity

Elasticity lives on the queue, not on the daemon group. In YAML the kinds are underscored:

lakeshore.yamlyaml
queues:
  training:
    kind: fifo
    elasticity:
      kind: pool_with_threshold
      min: 8
      max: 20
      threshold: 0.8

Only kind, min, max, and threshold survive the loader — pool, match, steal_delay_s, cost_weight, and any admission: block are silently dropped. Set those by PATCHing the queue over HTTP; see Elasticity. And remember the controller is off unless the control plane runs with ELASTICITY_ENABLED=true.

Verb reference

CommandWhat it does
lakeshore up [group]Reconcile to desired state (idempotent).
lakeshore up --planDry-run — print what would change.
lakeshore up --waitBlock until every spawned daemon is active.
lakeshore downRemove this project's workers; leave instances alive.
lakeshore down --terminateAlso terminate the cloud instances.
lakeshore down --prune-queuesDelete declared queues that have no members left.
lakeshore psList daemons owned by this compose.
lakeshore statusDesired vs actual drift (read-only).
lakeshore logs <group|label>Multiplexed cloud-console output. --follow, --poll-ms <n>.

All read ./lakeshore.yaml; override with -f <path>. up, down, ps, and status accept --json.

Read next

  • Compose — the concept page and full schema.
  • /dev/compose — the design doc.
  • Queues — the primitive compose wraps.
  • Fan out — put work on the queue you just created.