DreamLake

Architecture

Understand the three layers, the objects the control plane stores, and how an invocation travels from your laptop to a remote worker.

Lakeshore ships as three independently-deployed pieces, each swappable without touching the others:

  • lakeshore CLI — Node + TypeScript, distributed via npm (@dreamlake/lakeshore). The lakeshore binary registers providers, manages secrets / modes / queues / storages / mounts / tunnels, brings daemons up, and runs one-off commands against them.
  • Lakeshore control plane — Node + Fastify 5 + Prisma 5 + Mongo, deployed on Heroku. The single brain that owns canonical state and hands Invocations to Workers.
  • nymph daemon — Rust binary (the nymph submodule — dreamlake-ai/nymph), the host process that supervises Invocations under docker / gvisor / process runners. Closed-source for now; built from source.
┌─────────────────────────────────────────────────────────────────┐
│  Client surface                                                 │
│   ─ lakeshore CLI         (Node, on npm — operator verbs)       │
│   ─ Python SDK            (dreamlake-lakeshore — @udf, Futures) │
├─────────────────────────────────────────────────────────────────┤
│  Control plane  (Lakeshore, Node + Fastify + Prisma)            │
│                                                                 │
│   Durable store   Mongo    — every model below                  │
│   Hot lane        Redis    — result-chunk delivery, log cache   │
│   Blob store      S3       — spilled payloads, code archives,   │
│                              archived exec logs                 │
│                                                                 │
│   ─ Scheduling   Queue · Invocation · Function · ExecJob        │
│   ─ Compute      Worker · Provider · Mode · QueueServer         │
│   ─ Data         Storage · Mount · CodeRepo                     │
│   ─ Identity     Namespace · Token · EnrollToken · DaemonKey ·  │
│                  Secret · Tunnel                                │
│   ─ Audit        Event · ElasticityEvent · ResultChunk          │
├─────────────────────────────────────────────────────────────────┤
│  Runtime  (Rust daemon, on each compute host)                   │
│   ─ Daemon (nymph)        (long-polls; runs Invocations under   │
│                            process / docker / gvisor / slurm /  │
│                            kube runners)                        │
│   ─ Wire protocol         (msgpack long-poll; daemon-initiated) │
└─────────────────────────────────────────────────────────────────┘

prisma/schema.prisma defines exactly 21 models. All of them carry a namespaceId except three: Namespace itself (it is the boundary), QueueServer, and NymphRelease (a global binary registry, deduped by sha256).

Reading the layers

Client surface is what you talk to. The CLI is the operator surface: register providers, create queues, launch daemons, exec commands against them, inspect jobs. The Python SDK (dreamlake-lakeshore) is the developer surface — the @udf(queue=...) decorator, submit() returning a durable invocation id, and SyncQueue.result(id). See Python SDK.

Control plane is Lakeshore — a single Node.js + TypeScript service. Three stores sit behind it, each with a distinct job:

  • Mongo (via Prisma) is the durable record and the source of truth for every model on this page. It also carries the dispatch hot path: a raw findAndModify on the Invocation collection — { state: "queued", "runConfig.queue": … } sorted by { priority: -1, submittedAt: 1 } — atomically flips one row to running and stamps the workerId. That is why v0 needs no Lua scripts and no row locks.
  • Redis is a low-latency delivery channel, not the scheduler. It carries exec log streams (exec:<id>:stdout), per-invocation result journals (invocation:<id>:result), and the daemon-signature replay-nonce store. The ResultChunk rows in Mongo are the durable truth — trimming Redis never loses a frame, reads just degrade to Mongo. (A Redis hot lane for queue dispatch is described in the schema header as a designed-but-unbuilt v1+ optimization.)
  • S3 holds anything too big to inline: spilled payloads, CodeRepo archives, and archived exec logs. See Payload store for the concept and /dev/payload-store for the spec.

Runtime is what executes your code. The daemon — nymph, a Rust binary — runs on every compute host; it speaks the wire protocol to the control plane and supervises the Invocation under one of the registered runner kinds: process (alias subprocess), docker, gvisor, slurm, or kube. Daemons are outbound-only: they long-poll the control plane and never accept inbound connections, so a compute host needs egress and nothing else.

No relay exists in the shipped daemon

Older notes describe a nymph relay — a FIFO ↔ HTTPS forwarder for HPC clusters whose compute nodes have no egress. There is no relay subcommand, no FIFO code, and no LAKESHORE_RELAY_* handling in the nymph source. tui is the binary's only subcommand. Treat Relay config as a design note, not shipped behaviour.

The object model

Everything the control plane stores is namespaced. Start with the four you will meet first, then the supporting cast.

The core four

Queue

A named pipe carrying pending Invocations, and the first-class scheduling primitive. Workers subscribe by name — Worker.queues: string[] records which queues a daemon pulls from, and the empty array means the default queue (there is no sentinel string).

A Queue carries three policies:

PolicyFieldValues
Schedulingkindfifo · priority · boltzmann · filo
Elasticityelasticity.kindfixed · fully_elastic · pool_with_threshold · max_count
Admissionadmissionmax_depth, on_full (reject | block), rate_limit, deadline_cutoff_s

kindParams holds type-specific knobs — { temperature: 0.5 } for boltzmann, { priority_field: "p" } for priority.

Only fifo ordering is honoured today

All four kind values are accepted and stored, but the dispatcher's claim always sorts { priority: -1, submittedAt: 1 } regardless of kind. boltzmann and filo do not yet change dispatch order. Similarly, admission is validated and stored but never enforced at admit time.

Elasticity kinds use underscores

The stored values are fully_elastic, pool_with_threshold, and max_count — underscores, not hyphens. A hyphenated value is rejected by the API. (The lakeshore queues add --elasticity flag accepts the hyphenated spelling and converts it before the request.)

Lifecycle: the schema declares active, draining, paused, and archived, but only three are reachable through routes — /drain → draining, /archive → archived, /unarchive → active. Nothing writes paused.

See Queues for the concept page, Elasticity and Scaling rules for the autoscaling behaviour, and /dev/queues for the design note.

Worker

A daemon process registered with the control plane. Each Worker carries a stable machineId (reported on /hello and preserved across restarts), a queues array, a runners array, a capabilities object (runner list, max_invocations, workdir, plus the detected CPU/GPU inventory), and a tag set. Queue membership and elasticity match constraints decide which Invocations a Worker can claim.

Worker state runs pending → setting_up → joining → active → gone. pending is a row pre-created by daemons/launch before the host boots; setting_up means queued setup commands are still draining; joining is set by /hello; active is the first poll after that. draining is declared in the schema but no route writes it. A seventh label, stale, is derived at read time when lastSeenAt is older than 60 s — it is never persisted.

Two fields worth knowing: hostStatus tracks the cloud instance lifecycle independently of daemon poll state (reconciled against the provider's DescribeInstances), and composeProject records the project: slug when the Worker was spawned by lakeshore up — ad-hoc daemons leave it null and stay invisible to lakeshore ps / lakeshore down.

See Daemon runtime and Daemon lifecycle.

Invocation

One deferred execution of a Function. Keyed by a ULID generated by the producer — the same id that flows across the wire — with attempt, priority, submitted / started / finished timestamps, and a state:

queued → running → one of succeeded · failed · killed · timeout.

Arguments and results ride as inline msgpack in argsBlob / resultBlob. The model also declares argsRef / resultRef for spilling to S3, but no route writes them yet — the payload store is wired on the ExecJob path first. correlationId, parentInvocationId, and a denormalized rootInvocationId give you the pipeline tree — the root field equals the invocation's own id when it is the root.

There is no queueId column. Phase A.5 removed it; the queue name is denormalized onto runConfig.queue by POST /v1/producer/submit, and every read path (admin list, ?queue= filter, the dispatcher's claim) matches on that embedded path. Absent, it falls back to "default". idempotencyKey and codeRepoId exist on the model but no route writes either one yet.

Inspect via lakeshore jobs list / show / kill — see Jobs.

Worker host

The machine the daemon runs on. Brought online either by the control plane (via a Provider row — EC2 / GCE / SLURM / Kube) or by the CLI bootstrapping nymph onto an existing host over SSH. See Daemon launch and Bootstrap a daemon on a remote host.

Compute and placement

ModelWhat it is
ProviderWhere Worker hosts come from. The API field is launcher, and the only accepted values are SSH, SLURM, EC2, GCE, Kube — capitalized exactly like that. (launcher maps to the DB column Provider.kind; the API's kwargs maps to Provider.config.) Credentials ride as $secret references. Removal is a soft delete — the row is renamed with a [deleted <iso>] suffix and can be restored.
ModeA named compute config: backend, runner, image, resources, env, timeout, and an optional Provider link. The route does not validate backend or runner; it defaults them to "fabric" and "process" and otherwise stores any string. Managed via lakeshore modes ...; .dreamrc works as a local fallback.
QueueServerA dispatch tier — edge, region, or central — with a parent/child self-relation, and the one model that is not namespaced. Schema-only: no code reads it, and the only writes to Worker.serverNodeId are literal null. Do not plan around the tier hierarchy yet.
ExecJobAn ad-hoc command run on a Worker, separate from the Invocation path. Status runs pending → running → done | timeout | failed. Delivered through the same poll loop as a kind="exec" command. Exit code -1 maps to failed, -2 to timeout.
FunctionThe callable itself, content-addressed by (module, qualname, version) where version is the sha256 of the source. Captures the signature and, optionally, the full source for dashboard display.

Data and code movement

ModelWhat it is
StorageA registered S3-compatible bucket (s3) or scoped prefix (s3-prefix). The bucket / prefix / endpoint triple is immutable after creation — repointing would silently break every reference, so you create a new entry instead. Omitting endpoint means AWS; set it for Ceph, MinIO, R2, or B2. See Storages.
MountA shared filesystem — one of exactly ten kinds: nfs, samba, ftp, sftp, s3, s3fs, google_drive, dropbox, bind, configmap — that the runner attaches into the job's workdir. Metadata only today: the server stores and validates the config, but runner-side activation is a daemon follow-up. See Mounts.
CodeRepoA content-addressed snapshot of a git tree, deduped by (namespace, remote, commit). Dirty trees get a synthetic commit hash and a refs/snapshots/run-<id> ref. The archive lands in S3 or GitHub; the row maps (remote, commit) → archive location so later invocations skip the upload.

Identity and secrets

ModelWhat it is
NamespaceThe tenancy boundary — every user-owned collection scopes by namespaceId. Namespaces sharing an orgId can see each other's providers and queues. Also carries targetNymphVersion, the OTA version daemons converge to.
TokenA per-namespace bearer credential for the CLI and SDKs. Plaintext looks like dlk_<base64url-32-bytes>, is shown once, and is never stored; Mongo keeps sha256(plaintext) plus an 8-char display prefix. Revocation is a tombstone, so the listing can still show it.
EnrollTokenA one-shot, TTL-bounded credential consumed exactly once at /v1/daemon/hello to enroll a new daemon into a namespace. Plaintext prefix is dle_ — deliberately distinct from the CLI's dlk_. Default TTL 1 hour, hard cap 30 days.
DaemonKeyA per-host Ed25519 identity; the private key never leaves the host. Rotation runs active → retiring → revoked, with an overlap window so in-flight requests signed by the old key don't fail mid-rotation. See Key rotation.
SecretEncrypted credential material (SSH keys, AWS keypairs, GCP service-account JSON, opaque bytes), stored as AES-256-GCM ciphertext. Plaintext never touches disk or Mongo. Referenced elsewhere by $secret markers. See Auth and secrets.
TunnelA network tunnel (currently WireGuard only) the launcher brings up before reaching a Provider. Metadata-only today. See Tunnels.

Audit and observability

ModelWhat it is
EventThe tailable log, keyed by queue name, not id. The schema documents invocation.*, pool.*, and queue.state kinds, but the only kinds any code writes today are invocation.progress, invocation.log, and invocation.event — all from POST /v1/daemon/event.
ElasticityEventOne row per autoscaler decision (scale_up, scale_down, tick_noop) with the human-readable reason and before/after signal snapshots, so you can reconstruct what the controller saw. tick_noop is hidden unless you pass ?verbose=true.
ResultChunkThe durable journal of streamed result frames, one row per frame, ordered by a per-invocation seq cursor. Readers resume with ?since=<last seq>.
NymphReleaseAn operator-uploaded nymph binary (lakeshore nymph push), so debug fleets can OTA-update without the full release pipeline. See /dev/nymph-push.

How dispatch flows today

operator                control plane                  worker host
────────                ─────────────                  ───────────

lakeshore daemon launch
  --queue training-h100 ───▶ Worker row created
                             Provider boots host

                             /hello (EnrollToken ───▶ Daemon enrols,
                              consumed, DaemonKey       registers
                              registered)               queues=["training-h100"]
                                       │
                                       ▼
                                /poll ─┘  ◀── long-poll

lakeshore exec / run
  ───▶ Invocation enqueued ───▶ Queue("training-h100")
                                       │
                                       └─ poll response  ───▶ Daemon runs
                                                                command in
                                                                runner
                                                                (process /
                                                                 docker / gvisor)
                                  ◀── /result-chunk (streamed frames)
                                  ◀── /event         (progress, logs)
                                  ◀── /ack           (terminal result)

The daemon-facing surface is eight POST endpoints, all under /v1/daemon/ and all msgpack:

EndpointPurpose
/helloEnrol: consume an EnrollToken, register the DaemonKey, upsert the Worker row.
/pollThe long-poll. Returns commands + target_nymph_version.
/ackTerminal outcome of one invocation.
/eventProgress / log / event rows.
/keysKey rotation, signed by the outgoing key.
/result-chunkOne frame of a streamed result journal.
/exec/:exec_id/resultTerminal outcome of one ExecJob.
/exec/:exec_id/chunkOne buffered slice of exec stdout.

Three things to notice:

  1. The CLI never talks to a worker directly. Every operator action (exec, run, daemon update, daemon kill) goes through the control plane.
  2. The worker never accepts an inbound connection. It long-polls, and control commands ride back on the poll response. The complete set of command kinds is run, setup, exec, hibernate, ota_update, reset_backoff, and rotate_key — note the underscores on the last three. Only run comes from claiming an invocation; the others short-circuit the claim loop.
  3. Only /poll, /ack, and /event are covered by the signature pre-handler. /hello and /keys verify their own signatures in-route, and the two exec callbacks plus /result-chunk are authenticated by the unguessable id in the URL.

The Python SDK (/python-sdk) produces the same Invocation envelope and rides the same dispatch path.

Agents — not yet modeled

Design in progress — no Agent model exists today

An agent is a long-lived caller: it holds a session, issues many invocations over its life, and receives a streamed channel of messages rather than one terminal result. That shape does not fit the Invocation model, which is request/response and dies at the end.

Nothing in the schema models this yet. The design — an external | managed split mirroring Providers, explicit lifecycle states, and a persistent platform-owned workdir that composes with git mounts — is being worked out in Agent lifecycle and DreamLake ↔ Lakeshore federation. Treat both as forward-looking, not as documentation of shipped behaviour.

State machines at a glance

ObjectStatesNotes
Queueactive · draining · archivedpaused is in the enum but unreachable.
Workerpending · setting_up · joining · active · gonedraining is declared but never written; stale is derived at read time from lastSeenAt.
Invocationqueued · running · succeeded · failed · killed · timeoutThe last four are the terminal set.
ExecJobpending · running · done · timeout · failed
DaemonKeyactive · retiring · revokedactive and retiring both verify; revoked is a hard 403.
Worker host (cloud)starting · running · stopping · stopped · terminated · unknownNull for local/SSH daemons.

Glossary

TermDefinition
QueueThe unified scheduling primitive — a named pipe of pending Invocations. Workers subscribe by name; an empty queues array means the default queue.
WorkerA daemon process registered with the control plane. Carries queues: string[], runners, and a capabilities object.
InvocationOne deferred execution, keyed by a producer-generated ULID. Inspect via lakeshore jobs ....
FunctionThe callable, content-addressed by (module, qualname, sha256(source)).
ProviderA registered launch target the control plane uses to boot hosts. launcher is one of SSH, SLURM, EC2, GCE, Kube.
ModeA server-stored named compute config: backend + runner + image + resources + env, with an optional Provider link. Managed via lakeshore modes ...; .dreamrc is a local fallback.
QueueServerA dispatch tier (edge / region / central). Schema-only today — no code reads it.
Daemonnymph (Rust binary, repo dreamlake-ai/nymph). Long-running host process that runs and monitors Invocations. Its only subcommand is nymph tui; bare nymph runs the daemon.
RunnerThe execution sandbox the daemon uses for one Invocation: process (alias subprocess), docker, gvisor, slurm, or kube. An unknown kind fails the invocation rather than falling back.
StorageA registered S3-compatible bucket or prefix. Used for artifact upload, code shipping, and large-payload spill. Its S3 target is immutable after creation.
MountA registry entry for a shared filesystem the runner attaches into the workdir. Metadata-only today.
CodeRepoA versioned code snapshot pushed via lakeshore code push, deduped by (namespace, remote, commit). Daemons pull it at invocation time so your working tree arrives without manual rsync.
NamespaceThe tenancy boundary. Namespaces sharing an orgId can see each other's providers and queues.
LakeshoreThe control plane. Single Node + TS + Fastify + Prisma service over Mongo, with Redis for result streaming and S3 for blobs.

Read next

I want to…Page
Learn the queue modelQueues
Understand autoscalingElasticity
See all CLI commandsCLI
Try runnable examplesCLI examples
Launch daemons via providersDaemon launch
Understand the wire protocolDaemon protocol
Explore the control plane internalsLakeshore
Read the queue design noteQueue design
Read the payload store designPayload store