DreamLake

Compose — declarative cluster spec

Status: shipped. Phases 1–5 landed in CLI v0.1.16. See Rollout at the bottom for details.

A lakeshore.yaml at the project root declares the cluster you want to exist: which queues, which daemons, how many, in which mode. lakeshore up reconciles; lakeshore down tears down. The mental model is docker-compose for fleets — one file, one verb, idempotent.

Why

08-queues-cpu-vs-gpu/README.md is six imperative commands:

bash
lakeshore queues add cpu-queue …
lakeshore queues add gpu-t4 …
lakeshore daemon launch --queue cpu-queue --count 5 …
lakeshore daemon launch --queue gpu-t4   --count 5 …
…
lakeshore daemon kill --prefix cpu- --terminate --yes
lakeshore daemon kill --prefix gpu- --terminate --yes
lakeshore queues rm cpu-queue
lakeshore queues rm gpu-t4

That sequence is the literal serialization of a document we should have. With compose it collapses to:

bash
lakeshore up
… work …
lakeshore down --terminate

— and the document lives in the repo, version-controlled with the code that depends on it.

Schema (v1)

yaml
# lakeshore.yaml
version: 1
project: hello-cluster          # required — namespaces this compose's state on the CP
namespace: default              # optional — defaults to the namespace in ~/.config/lakeshore/auth.yml

# ─── Shared resources ──────────────────────────────────────────────
queues:
  cpu-queue:
    description: "CPU pool — batch + smoke"
    kind: fifo
  gpu-t4:
    description: "T4 GPUs (g4dn.xlarge)"
    kind: priority
    reserved_slots: 1

# ─── Daemon recipes (scoped to this file) ──────────────────────────
modes:
  cpu-batch:
    provider: ec2-prod
    instance_type: t3.medium
    runner: docker
  gpu-experiment:
    provider: ec2-prod
    instance_type: g4dn.xlarge
    image_id: ami-012ba162b9cd2729c       # DLAMI PyTorch 2.7
    runner: docker
    setup_scripts: [./host-setup/verify-gpu.sh]

# ─── The "services" ────────────────────────────────────────────────
daemons:
  cpu:                          # name prefix → cpu-0..cpu-4 daemons
    mode: cpu-batch
    queues: [cpu-queue]         # m:n — a daemon can be in multiple queues
    count: 5
    keep_alive_s: 600
  gpu:
    mode: gpu-experiment
    queues: [gpu-t4]
    count: 5
    keep_alive_s: 1800
    depends_on: [cpu]           # ordering hint for setup_scripts (optional)

Field reference

FieldRequiredNotes
versionyesinteger; v1 is the only one
projectyesslug; tags every Worker row spawned by this file
namespacenooverrides the auth.yml default
queues.<name>nosame shape as lakeshore queues add flags
modes.<name>noinline; merges over the same name from .dreamrc if both exist (inline wins)
daemons.<name>.modeyesreferences a key under modes: (inline) or .dreamrc
daemons.<name>.queuesnodefaults to [] (implicit default queue) if omitted
daemons.<name>.countnodefaults to 1
daemons.<name>.keep_alive_snoper-daemon override
daemons.<name>.setupnoshell command(s) run once after the daemon registers (e.g. pip install -r requirements.txt). Runs before any invocation is accepted.
daemons.<name>.startupnoshell command(s) run every time the daemon process starts (including restarts). Use for environment checks or service warm-up.
daemons.<name>.pythonnopath to the Python interpreter on the worker (e.g. /opt/conda/bin/python-sdk). Overrides the default python3.
daemons.<name>.depends_onnoordering hint for up; advisory only in v1

Naming rule

Daemon labels are <group>-<idx>. With daemons.cpu.count: 5, you get cpu-0, cpu-1, … cpu-4. Labels are global within a project; two composes with the same project name will collide — by design.

CLI surface

bash
lakeshore up                    # converge to desired state (idempotent)
lakeshore up gpu                # only the `gpu` group
lakeshore up --wait             # block until all daemons active
lakeshore up --plan             # dry-run — print what would change
lakeshore down                  # drain (graceful — accept no new work, finish in-flight)
lakeshore down --terminate      # actually terminate the EC2s
lakeshore down --force          # kill in-flight execs
lakeshore status                # current vs desired (per group)
lakeshore ps                    # like `docker compose ps`
lakeshore logs gpu-2 -f         # follow one daemon
lakeshore logs gpu -f           # follow all in a group (multiplexed)

Convention: lakeshore <verb> without flags reads ./lakeshore.yaml. -f <path> overrides.

State model

Three states matter:

┌──────────┐ lakeshore.yaml      ┌──────────┐
│  desired │ ──────read─────────▶│   CLI    │
└──────────┘                     └────┬─────┘
                                      │ reconcile
                                      ▼
┌──────────┐    /v1/.../workers   ┌──────────┐
│  actual  │ ◀───────────────────│controlplane│
└──────────┘   (Worker rows tagged└──────────┘
                with compose_project)
  • Desired is the YAML on disk.
  • Actual is the set of Worker rows on the CP filtered by compose_project == <project>.
  • up walks the diff: add what's missing, leave matching rows alone, optionally prompt to remove orphans (--prune).
  • No live controller. Reconciliation runs only when up is invoked. Drift is reported by status; the user decides when to reconcile.

Project isolation

A new column on Worker:

prisma
model Worker {
  …
  composeProject String?            // set when spawned via `lakeshore up`
  @@index([namespaceId, composeProject])
}

up filters and operates only on rows where composeProject == <yaml.project>. Daemons launched ad-hoc via lakeshore daemon launch have composeProject == null and are invisible to compose verbs. This is the same isolation pattern as docker-compose's com.docker.compose.project label.

Mode resolution

daemons.<name>.mode is resolved in this order:

  1. Inline under modes: in this lakeshore.yaml.
  2. .dreamrc at the project root.
  3. ~/.lakeshore/modes/<name>.yml (shared catalog — future).

Inline wins. Conflicting fields between sources are merged at the field level (inline instance_type: t3.large overrides a .dreamrc instance_type: t3.medium, but inherits the rest).

Queue resolution

daemons.<name>.queues is a hard reference to queue names. If a referenced queue doesn't exist on the CP:

  • If the queue is also declared under queues: in this file → up creates it.
  • If not → up errors and refuses to proceed.

Queues declared under queues: but not referenced by any daemon are created anyway (idempotent — queues add with same name+config is a no-op).

What about lakeshore exec from a compose?

Out of scope. Compose declares the cluster; exec dispatches to it. The two compose because queue references are stable: write lakeshore exec --queue cpu-queue … and it dispatches to whichever daemons compose put in cpu-queue.

Non-goals (v1)

  • Autoscaling. count: 5 is a hard number, not a target. No HPA controller. Phase 3 work — needs a metric source.
  • Live drift reconciliation. No daemon process that runs up continuously. Phase 4 — and only if pain demands it.
  • Cross-namespace composes. One file = one namespace.
  • Multi-file composes (compose overrides like docker-compose.override.yml). Phase 2 if requested.
  • Volume / network specs. We don't have those primitives.
  • Inline secrets. Reference them by env (${LAKESHORE_PROVIDER_KEY}) or rely on the worker's IAM role / provider creds.

Boundaries with existing concepts

ConceptLives inCompose role
Mode.dreamrc or inline modes:how a daemon is born
QueueCP (declared in queues:)where a daemon lives, who can dispatch to it
Projectproject: in this file"this cluster, owned by this file"
Daemon groupa key under daemons:the named recipe block
Daemona Worker rowone of count instances of a group

The compose document doesn't add primitives — it's a layout over the primitives we already have. Mode + queue + count is the existing data shape; this is just the document that says "these specific instances of those should exist."

Rollout

PhaseScopeStatus
0This design docdone
1lakeshore.yaml loader + schema validation (no side effects)shipped
2lakeshore up (creates queues + launches daemons; idempotent)shipped
3lakeshore down (drain + optional terminate)shipped
4lakeshore status / ps / logsshipped
5--plan, --wait, --prune flagsshipped
6Multi-file overrides, autoscaling, etc.not committed

Phases 1–5 shipped in CLI v0.1.16 (lakeshore/src/cli/get-started/compose/). 08-queues-cpu-vs-gpu now collapses to lakeshore up && lakeshore exec --queue cpu-queue ….

Open questions

  • Should up block on queue state == archived? Probably yes — an archived queue refuses new work, so daemons launched into it would be inert. Error and tell the operator to queues unarchive first.
  • What happens to in-flight execs on down? Default: graceful drain (refuse new work, finish in-flight, then terminate). --force kills immediately. Need to flesh out the "drain" wire — currently daemon kill --terminate is hard-stop only.
  • Do we surface compose_project on daemon list? Yes, as a column when present. Lets operators see which compose owns which daemons at a glance.
  • What about a lakeshore init to scaffold a lakeshore.yaml? Phase 5 — would template from an existing .dreamrc mode.