Architecture
Understand the three layers, the objects the control plane stores, and how an invocation travels from your laptop to a remote worker.
Lakeshore ships as three independently-deployed pieces, each swappable without touching the others:
lakeshoreCLI — Node + TypeScript, distributed via npm (@dreamlake/lakeshore). Thelakeshorebinary registers providers, manages secrets / modes / queues / storages / mounts / tunnels, brings daemons up, and runs one-off commands against them.- Lakeshore control plane — Node + Fastify 5 + Prisma 5 + Mongo, deployed on Heroku. The single brain that owns canonical state and hands Invocations to Workers.
nymphdaemon — Rust binary (thenymphsubmodule —dreamlake-ai/nymph), the host process that supervises Invocations under docker / gvisor / process runners. Closed-source for now; built from source.
prisma/schema.prisma defines exactly 21 models. All of them carry a
namespaceId except three: Namespace itself (it is the boundary),
QueueServer, and NymphRelease (a global binary registry, deduped by
sha256).
Reading the layers
Client surface is what you talk to. The CLI is the operator surface:
register providers, create queues, launch daemons, exec commands against
them, inspect jobs. The Python SDK (dreamlake-lakeshore) is the
developer surface — the @udf(queue=...) decorator, submit() returning
a durable invocation id, and SyncQueue.result(id). See
Python SDK.
Control plane is Lakeshore — a single Node.js + TypeScript service. Three stores sit behind it, each with a distinct job:
- Mongo (via Prisma) is the durable record and the source of truth
for every model on this page. It also carries the dispatch hot path:
a raw
findAndModifyon theInvocationcollection —{ state: "queued", "runConfig.queue": … }sorted by{ priority: -1, submittedAt: 1 }— atomically flips one row torunningand stamps theworkerId. That is why v0 needs no Lua scripts and no row locks. - Redis is a low-latency delivery channel, not the scheduler. It
carries exec log streams (
exec:<id>:stdout), per-invocation result journals (invocation:<id>:result), and the daemon-signature replay-nonce store. TheResultChunkrows in Mongo are the durable truth — trimming Redis never loses a frame, reads just degrade to Mongo. (A Redis hot lane for queue dispatch is described in the schema header as a designed-but-unbuilt v1+ optimization.) - S3 holds anything too big to inline: spilled payloads,
CodeRepoarchives, and archived exec logs. See Payload store for the concept and/dev/payload-storefor the spec.
Runtime is what executes your code. The
daemon — nymph, a Rust binary — runs on every
compute host; it speaks the wire protocol to the
control plane and supervises the Invocation under one of the registered
runner kinds: process (alias subprocess), docker, gvisor,
slurm, or kube. Daemons are outbound-only: they long-poll the
control plane and never accept inbound connections, so a compute host
needs egress and nothing else.
Older notes describe a nymph relay — a FIFO ↔ HTTPS forwarder for HPC
clusters whose compute nodes have no egress. There is no relay
subcommand, no FIFO code, and no LAKESHORE_RELAY_* handling in the
nymph source. tui is the binary's only subcommand. Treat
Relay config as a design note, not shipped
behaviour.
The object model
Everything the control plane stores is namespaced. Start with the four you will meet first, then the supporting cast.
The core four
Queue
A named pipe carrying pending Invocations, and the first-class
scheduling primitive. Workers subscribe by name —
Worker.queues: string[] records which queues a daemon pulls from, and
the empty array means the default queue (there is no sentinel string).
A Queue carries three policies:
| Policy | Field | Values |
|---|---|---|
| Scheduling | kind | fifo · priority · boltzmann · filo |
| Elasticity | elasticity.kind | fixed · fully_elastic · pool_with_threshold · max_count |
| Admission | admission | max_depth, on_full (reject | block), rate_limit, deadline_cutoff_s |
kindParams holds type-specific knobs — { temperature: 0.5 } for
boltzmann, { priority_field: "p" } for priority.
All four kind values are accepted and stored, but the dispatcher's
claim always sorts { priority: -1, submittedAt: 1 } regardless of
kind. boltzmann and filo do not yet change dispatch order.
Similarly, admission is validated and stored but never enforced at
admit time.
The stored values are fully_elastic, pool_with_threshold, and
max_count — underscores, not hyphens. A hyphenated value is rejected
by the API. (The lakeshore queues add --elasticity flag accepts the
hyphenated spelling and converts it before the request.)
Lifecycle: the schema declares active, draining, paused, and
archived, but only three are reachable through routes — /drain →
draining, /archive → archived, /unarchive → active. Nothing
writes paused.
See Queues for the concept page,
Elasticity and
Scaling rules for the autoscaling
behaviour, and /dev/queues for the design note.
Worker
A daemon process registered with the control plane. Each Worker carries
a stable machineId (reported on /hello and preserved across
restarts), a queues array, a runners array, a capabilities object
(runner list, max_invocations, workdir, plus the detected CPU/GPU
inventory), and a tag set. Queue membership and elasticity match
constraints decide which Invocations a Worker can claim.
Worker state runs pending → setting_up → joining → active →
gone. pending is a row pre-created by daemons/launch before the
host boots; setting_up means queued setup commands are still draining;
joining is set by /hello; active is the first poll after that.
draining is declared in the schema but no route writes it. A seventh
label, stale, is derived at read time when lastSeenAt is older
than 60 s — it is never persisted.
Two fields worth knowing: hostStatus tracks the cloud instance
lifecycle independently of daemon poll state (reconciled against the
provider's DescribeInstances), and composeProject records the
project: slug when the Worker was spawned by lakeshore up — ad-hoc
daemons leave it null and stay invisible to lakeshore ps /
lakeshore down.
See Daemon runtime and Daemon lifecycle.
Invocation
One deferred execution of a Function. Keyed by a ULID generated by the
producer — the same id that flows across the wire — with attempt,
priority, submitted / started / finished timestamps, and a state:
queued → running → one of succeeded · failed · killed ·
timeout.
Arguments and results ride as inline msgpack in argsBlob /
resultBlob. The model also declares argsRef / resultRef for
spilling to S3, but no route writes them yet — the
payload store is wired on the ExecJob path
first. correlationId, parentInvocationId, and a denormalized
rootInvocationId give you the pipeline tree — the root field equals
the invocation's own id when it is the root.
There is no queueId column. Phase A.5 removed it; the queue name
is denormalized onto runConfig.queue by POST /v1/producer/submit,
and every read path (admin list, ?queue= filter, the dispatcher's
claim) matches on that embedded path. Absent, it falls back to
"default". idempotencyKey and codeRepoId exist on the model but
no route writes either one yet.
Inspect via lakeshore jobs list / show / kill — see
Jobs.
Worker host
The machine the daemon runs on. Brought online either by the control
plane (via a Provider row — EC2 / GCE / SLURM / Kube) or by the CLI
bootstrapping nymph onto an existing host over SSH. See
Daemon launch and
Bootstrap a daemon on a remote host.
Compute and placement
| Model | What it is |
|---|---|
| Provider | Where Worker hosts come from. The API field is launcher, and the only accepted values are SSH, SLURM, EC2, GCE, Kube — capitalized exactly like that. (launcher maps to the DB column Provider.kind; the API's kwargs maps to Provider.config.) Credentials ride as $secret references. Removal is a soft delete — the row is renamed with a [deleted <iso>] suffix and can be restored. |
| Mode | A named compute config: backend, runner, image, resources, env, timeout, and an optional Provider link. The route does not validate backend or runner; it defaults them to "fabric" and "process" and otherwise stores any string. Managed via lakeshore modes ...; .dreamrc works as a local fallback. |
| QueueServer | A dispatch tier — edge, region, or central — with a parent/child self-relation, and the one model that is not namespaced. Schema-only: no code reads it, and the only writes to Worker.serverNodeId are literal null. Do not plan around the tier hierarchy yet. |
| ExecJob | An ad-hoc command run on a Worker, separate from the Invocation path. Status runs pending → running → done | timeout | failed. Delivered through the same poll loop as a kind="exec" command. Exit code -1 maps to failed, -2 to timeout. |
| Function | The callable itself, content-addressed by (module, qualname, version) where version is the sha256 of the source. Captures the signature and, optionally, the full source for dashboard display. |
Data and code movement
| Model | What it is |
|---|---|
| Storage | A registered S3-compatible bucket (s3) or scoped prefix (s3-prefix). The bucket / prefix / endpoint triple is immutable after creation — repointing would silently break every reference, so you create a new entry instead. Omitting endpoint means AWS; set it for Ceph, MinIO, R2, or B2. See Storages. |
| Mount | A shared filesystem — one of exactly ten kinds: nfs, samba, ftp, sftp, s3, s3fs, google_drive, dropbox, bind, configmap — that the runner attaches into the job's workdir. Metadata only today: the server stores and validates the config, but runner-side activation is a daemon follow-up. See Mounts. |
| CodeRepo | A content-addressed snapshot of a git tree, deduped by (namespace, remote, commit). Dirty trees get a synthetic commit hash and a refs/snapshots/run-<id> ref. The archive lands in S3 or GitHub; the row maps (remote, commit) → archive location so later invocations skip the upload. |
Identity and secrets
| Model | What it is |
|---|---|
| Namespace | The tenancy boundary — every user-owned collection scopes by namespaceId. Namespaces sharing an orgId can see each other's providers and queues. Also carries targetNymphVersion, the OTA version daemons converge to. |
| Token | A per-namespace bearer credential for the CLI and SDKs. Plaintext looks like dlk_<base64url-32-bytes>, is shown once, and is never stored; Mongo keeps sha256(plaintext) plus an 8-char display prefix. Revocation is a tombstone, so the listing can still show it. |
| EnrollToken | A one-shot, TTL-bounded credential consumed exactly once at /v1/daemon/hello to enroll a new daemon into a namespace. Plaintext prefix is dle_ — deliberately distinct from the CLI's dlk_. Default TTL 1 hour, hard cap 30 days. |
| DaemonKey | A per-host Ed25519 identity; the private key never leaves the host. Rotation runs active → retiring → revoked, with an overlap window so in-flight requests signed by the old key don't fail mid-rotation. See Key rotation. |
| Secret | Encrypted credential material (SSH keys, AWS keypairs, GCP service-account JSON, opaque bytes), stored as AES-256-GCM ciphertext. Plaintext never touches disk or Mongo. Referenced elsewhere by $secret markers. See Auth and secrets. |
| Tunnel | A network tunnel (currently WireGuard only) the launcher brings up before reaching a Provider. Metadata-only today. See Tunnels. |
Audit and observability
| Model | What it is |
|---|---|
| Event | The tailable log, keyed by queue name, not id. The schema documents invocation.*, pool.*, and queue.state kinds, but the only kinds any code writes today are invocation.progress, invocation.log, and invocation.event — all from POST /v1/daemon/event. |
| ElasticityEvent | One row per autoscaler decision (scale_up, scale_down, tick_noop) with the human-readable reason and before/after signal snapshots, so you can reconstruct what the controller saw. tick_noop is hidden unless you pass ?verbose=true. |
| ResultChunk | The durable journal of streamed result frames, one row per frame, ordered by a per-invocation seq cursor. Readers resume with ?since=<last seq>. |
| NymphRelease | An operator-uploaded nymph binary (lakeshore nymph push), so debug fleets can OTA-update without the full release pipeline. See /dev/nymph-push. |
How dispatch flows today
The daemon-facing surface is eight POST endpoints, all under
/v1/daemon/ and all msgpack:
| Endpoint | Purpose |
|---|---|
/hello | Enrol: consume an EnrollToken, register the DaemonKey, upsert the Worker row. |
/poll | The long-poll. Returns commands + target_nymph_version. |
/ack | Terminal outcome of one invocation. |
/event | Progress / log / event rows. |
/keys | Key rotation, signed by the outgoing key. |
/result-chunk | One frame of a streamed result journal. |
/exec/:exec_id/result | Terminal outcome of one ExecJob. |
/exec/:exec_id/chunk | One buffered slice of exec stdout. |
Three things to notice:
- The CLI never talks to a worker directly. Every operator action
(
exec,run,daemon update,daemon kill) goes through the control plane. - The worker never accepts an inbound connection. It long-polls, and
control commands ride back on the poll response. The complete set
of command kinds is
run,setup,exec,hibernate,ota_update,reset_backoff, androtate_key— note the underscores on the last three. Onlyruncomes from claiming an invocation; the others short-circuit the claim loop. - Only
/poll,/ack, and/eventare covered by the signature pre-handler./helloand/keysverify their own signatures in-route, and the two exec callbacks plus/result-chunkare authenticated by the unguessable id in the URL.
The Python SDK (/python-sdk) produces the same
Invocation envelope and rides the same dispatch path.
Agents — not yet modeled
An agent is a long-lived caller: it holds a session, issues many invocations over its life, and receives a streamed channel of messages rather than one terminal result. That shape does not fit the Invocation model, which is request/response and dies at the end.
Nothing in the schema models this yet. The design — an
external | managed split mirroring Providers, explicit lifecycle
states, and a persistent platform-owned workdir that composes with git
mounts — is being worked out in
Agent lifecycle and
DreamLake ↔ Lakeshore federation.
Treat both as forward-looking, not as documentation of shipped behaviour.
State machines at a glance
| Object | States | Notes |
|---|---|---|
| Queue | active · draining · archived | paused is in the enum but unreachable. |
| Worker | pending · setting_up · joining · active · gone | draining is declared but never written; stale is derived at read time from lastSeenAt. |
| Invocation | queued · running · succeeded · failed · killed · timeout | The last four are the terminal set. |
| ExecJob | pending · running · done · timeout · failed | |
| DaemonKey | active · retiring · revoked | active and retiring both verify; revoked is a hard 403. |
| Worker host (cloud) | starting · running · stopping · stopped · terminated · unknown | Null for local/SSH daemons. |
Glossary
| Term | Definition |
|---|---|
| Queue | The unified scheduling primitive — a named pipe of pending Invocations. Workers subscribe by name; an empty queues array means the default queue. |
| Worker | A daemon process registered with the control plane. Carries queues: string[], runners, and a capabilities object. |
| Invocation | One deferred execution, keyed by a producer-generated ULID. Inspect via lakeshore jobs .... |
| Function | The callable, content-addressed by (module, qualname, sha256(source)). |
| Provider | A registered launch target the control plane uses to boot hosts. launcher is one of SSH, SLURM, EC2, GCE, Kube. |
| Mode | A server-stored named compute config: backend + runner + image + resources + env, with an optional Provider link. Managed via lakeshore modes ...; .dreamrc is a local fallback. |
| QueueServer | A dispatch tier (edge / region / central). Schema-only today — no code reads it. |
| Daemon | nymph (Rust binary, repo dreamlake-ai/nymph). Long-running host process that runs and monitors Invocations. Its only subcommand is nymph tui; bare nymph runs the daemon. |
| Runner | The execution sandbox the daemon uses for one Invocation: process (alias subprocess), docker, gvisor, slurm, or kube. An unknown kind fails the invocation rather than falling back. |
| Storage | A registered S3-compatible bucket or prefix. Used for artifact upload, code shipping, and large-payload spill. Its S3 target is immutable after creation. |
| Mount | A registry entry for a shared filesystem the runner attaches into the workdir. Metadata-only today. |
| CodeRepo | A versioned code snapshot pushed via lakeshore code push, deduped by (namespace, remote, commit). Daemons pull it at invocation time so your working tree arrives without manual rsync. |
| Namespace | The tenancy boundary. Namespaces sharing an orgId can see each other's providers and queues. |
| Lakeshore | The control plane. Single Node + TS + Fastify + Prisma service over Mongo, with Redis for result streaming and S3 for blobs. |
Read next
| I want to… | Page |
|---|---|
| Learn the queue model | Queues |
| Understand autoscaling | Elasticity |
| See all CLI commands | CLI |
| Try runnable examples | CLI examples |
| Launch daemons via providers | Daemon launch |
| Understand the wire protocol | Daemon protocol |
| Explore the control plane internals | Lakeshore |
| Read the queue design note | Queue design |
| Read the payload store design | Payload store |