DreamLake

Living design doc. Tracks the plan to reduce server-side collections.

Current state

The control plane stores 17 Mongo collections. Several are pure config metadata that could live in the client's .dreamrc or lakeshore.yaml and be upserted (or sent inline) on each launch. Storing them server-side adds CRUD surface, migration burden, and forces users to run lakeshore <resource> add before they can do anything.

Classification

CollectionVerdictReason
NamespaceKeepMulti-tenancy root, auto-created
InvocationKeepJob state machine, server-authoritative
WorkerKeepDaemon registration, heartbeat, scheduling
QueueKeepScheduling state, worker membership, elasticity
FunctionKeepContent-addressed dedup, dashboard inspection
EventKeepAudit log, change feed for dashboard
ExecJobKeepAd-hoc command lifecycle
ElasticityEventKeepAutoscaler audit trail
QueueServerKeepDispatch tier hierarchy
TokenKeepAuth, per-user bearer tokens
NymphReleaseKeepOTA binary store
SecretKeepServer-side encryption, daemon access without shipping plaintext
ProviderKeepLaunch path needs it; holds $secret refs
StorageKeepServer-side presign + STS credential vending
ModeRemovePure config — inline in .dreamrc or lakeshore.yaml, sent with each invocation as runConfig
MountRemovePure config metadata — declare inline in compose spec, upsert on lakeshore up
TunnelInline into ProviderOnly used as a Provider attachment — store as a nested tunnel: block inside the Provider's config

Changes

1. Tunnel → inline in Provider config

Today:

Provider.tunnelId → Tunnel (separate collection)

After:

yaml
# Provider config (server-side JSON)
{
  "region": "us-east-1",
  "tunnel": {
    "kind": "wireguard",
    "interface": { ... },
    "peers": [ ... ]
  }
}

The tunnel config moves into Provider.config.tunnel. The Tunnel collection and tunnelId FK are removed. The launch path reads config.tunnel instead of joining.

CLI change: --tunnel <name> on providers add/update becomes --kwarg tunnel.kind=wireguard --kwarg tunnel.interface.address=... or --kwarg tunnel.$secret=wg-key. Same $secret resolution, no separate CRUD.

Migration: one-time script that reads each Provider's linked Tunnel row, copies { kind, config } into Provider.config.tunnel, and drops the FK.

2. Mode → client-side only

Modes are execution environment templates (backend, runner, image, resources, env). Today they're stored server-side and referenced by name in the Python SDK's @udf("mode_name") decorator.

The actual consumer is the invocation's runConfig — which is already a full inline copy of the mode's fields, merged with caller overrides. The Mode row is only used at submit time to resolve mode_name → runConfig fields.

Change: move mode resolution to the client. The Python SDK reads .dreamrc or lakeshore.yaml, resolves the mode locally, and sends the full runConfig inline with the invocation. The server never needs to know about modes.

CLI: lakeshore modes add/list/show/update/remove can be dropped. .dreamrc modes: block is the source of truth (already exists and works).

Dashboard: modes don't appear in the current dashboard. No change needed.

3. Mount → declare in compose, upsert on launch

Mounts are filesystem attach metadata (NFS, S3, bind, configmap). The daemon needs to know about them at launch time, but the server doesn't use them between launches.

Change: declare mounts inline in lakeshore.yaml:

yaml
daemons:
  gpu:
    queues: [training]
    count: 4
    mounts:
      - kind: s3
        bucket: ml-datasets
        prefix: imagenet/2012
        creds: { $secret: aws-creds }
        mount_path: /data/imagenet

The compose reconciler sends the mount specs as part of the launch body. The daemon receives them via the /hello response or setup commands. No server-side Mount collection needed.

Existing mount CRUD can stay as an optional server-side registry for operators who prefer centralized config, but the compose path doesn't require it.

What stays

After simplification, the server-side collections are:

  1. Namespace — identity
  2. Token — auth
  3. Secret — encrypted credentials
  4. Provider — compute backends (with tunnel inline)
  5. Storage — S3 bucket/prefix metadata + presign/credentials
  6. Queue — scheduling
  7. Worker — daemon fleet
  8. Function — content-addressed UDF registry
  9. Invocation — job state machine
  10. ExecJob — ad-hoc commands
  11. Event — audit log
  12. ElasticityEvent — scaler audit
  13. QueueServer — dispatch hierarchy
  14. NymphRelease — OTA binaries

Down from 17 to 14. Mode, Mount, and Tunnel (as a separate collection) are removed.

Migration order

  1. Tunnel → inline (smallest blast radius, Tunnel has fewer than 10 rows in prod)
  2. Mode → client-side (Python SDK change, no server migration)
  3. Mount → compose inline (optional — keep the CRUD as fallback)