DreamLake

Configuration

Four config files, owned by different processes and frequently confused with each other. Nothing merges across them.

FileRead byPurpose
.dreamrcCLI + Python SDK, producer-sideNamed compute modes (RunConfigs) and provider connections.
udf-daemon.tomlnymph, on the hostDaemon bootstrap — server URL, runners, workdir, poll tunables, identity.
.lakeshore / .lakeshore.localCLI, lakeshore daemon launch onlyPer-project launch defaults.
lakeshore.yamlCLI, up / down / ps / status / logsDeclarative queue + daemon-group topology.

.dreamrc

YAML. Declares named modes (RunConfigs) and providers — the account/cloud connections the control plane launches workers through.

Resolution order

The CLI resolver returns the first hit:

  1. --dreamrc <path> → source explicit
  2. $DREAMRC → env
  3. ./.dreamrc, walking up from cwd → workspace
  4. $HOME/.dreamrc → home
  5. Server pull (needs saved auth or LAKESHORE_URL) → server
  6. On-disk cache from a previous pull → cache
  7. Nothing → none (an empty DreamRc)

Inspect with lakeshore config show (the resolved config as YAML), lakeshore config refresh (force a server pull and update the cache), or lakeshore config source (one word: env, workspace, home, server, cache, none).

The server pull composes a DreamRc from GET /v1/namespaces/<ns>/providers and GET /v1/namespaces/<ns>/modes. A server mode literally named default is promoted to the root RunConfig. Providers whose launcher the CLI doesn't recognize are skipped silently rather than failing the pull. The cache lives at <cacheRoot>/<server-host-slug>/<namespace>/dreamrc.yml, where cacheRoot is $XDG_CACHE_HOME/lakeshore or ~/.cache/lakeshore (directories 0700, files 0600). A failed pull falls back to the cache and prints ⚠ server unreachable, using cached config from <ts>.

The Python SDK's own loader (dls.load_config()) is deliberately shorter — explicit path, $DREAMRC, ./.dreamrc walking up, ~/.dreamrc. It does not pull from the server.

Example

.dreamrcyaml
default_mode: local

providers:
  bos14: !providers.SLURM
    host: login.bos14.example.edu
    user: ge
    partition: gpu

  aws-prod: !providers.EC2
    region: us-east-1
    aws_credentials: { $secret: aws-prod-keys }

modes:
  local:
    backend: local

  bos14-gpu:
    backend: fabric
    provider: bos14
    runner: process
    resources: { cpu: 4, gpu: 1 }
    time_limit: "02:00:00" # unknown key → forwarded to the launcher

  aws-h100:
    backend: fabric
    provider: aws-prod
    tags: [gpu:h100]
    runner: docker
    image: ghcr.io/example/torch:2.4
    resources: { gpu: 1, mem: 64Gi }
    timeout_s: 7200

Provider entries

A provider entry must be a YAML-tagged mapping. The tag is the type discriminator — there is no type: field.

!providers.SSH   !providers.SLURM   !providers.EC2   !providers.GCE   !providers.Kube

Those five are the whole set (KNOWN_PROVIDER_TAGS). An unknown !providers.X tag is a hard parse error. The body is the launcher kwargs, deep-merged with the referencing mode's launcher-shaped fields at dispatch time.

Default dispatch per launcher, overridable with an explicit dispatch: key inside the kwargs:

LauncherDefault dispatch
SSHdirect
SLURMdirect
Kubedirect
EC2daemon
GCEdaemon

Mode fields

Top-level keys: default_mode (default "local"), modes (alias: mode), providers. Any other top-level key is parsed as a root RunConfig. A synthetic local mode is injected when the file declares none.

FieldTypeDefaultNotes
backendstring"local""local" runs in-process; "fabric" dispatches through the fabric.
serverstring | nullnoneControl-plane URL override for this mode.
tagslist[string][]Selector — matched against daemon tags.
runnerstring"docker""process", "docker", "gvisor", "slurm", "kube" ("subprocess" aliases "process").
resourcestable{}Reservation hints (cpu, mem, gpu, …).
imagestring | nullnoneOCI image. Required by the docker / gvisor runners.
envtable{}Extra env vars for the invocation.
timeout_sint | nullnoneWall-clock cap.
providerstring | nullnoneName of a provider entry.

Any key not in that list lands in extras and is forwarded to the launcher at dispatch time (instance_type, image_id, partition, time_limit, …).

Server-stored modes

Modes also live as rows on the control plane, managed with lakeshore modes add | list | show | update | remove. The override flag is --field, not --kwarg:

bash
lakeshore modes add aws-h100 \
  --field backend=fabric \
  --field provider=aws-prod \
  --field runner=docker \
  --field image=ghcr.io/example/torch:2.4 \
  --field resources.gpu=1 \
  --field timeout_s=7200

Dotted keys nest. lakeshore modes update <name> --field … deep-merges, so you can patch one field; --field foo=null deletes a key. --config-file <path> loads (on add) or replaces (on update) the whole RunConfig from YAML or JSON.

`backend` and `runner` are not validated server-side

The Mode create handler defaults backend to "fabric" and runner to "process" when unsupplied, and otherwise accepts any string. A typo is stored and echoed back rather than rejected. Unknown config keys round-trip through resources.__extras.

udf-daemon.toml

nymph's bootstrap config. Read once at process start and never re-read — everything mutable arrives over the long-poll channel.

Lookup order when --config <path> is absent:

  1. /etc/dreamlake/udf-daemon.toml
  2. $HOME/.config/dreamlake/udf-daemon.toml
  3. Built-in defaults

$LAKESHORED_CONFIG (note the trailing D) is the env form of --config.

The top-level Config is #[serde(deny_unknown_fields)], so an unknown section is a hard parse error. The sub-structs are not, so an unknown key inside a known section is silently ignored.

udf-daemon.tomltoml
[server]
url = "https://api.lakeshore.dreamlake.ai"
namespace = "default"
token_file = "/etc/dreamlake/token"
tls_verify = true

[daemon]
machine_id = "auto"
label = "gpu-box-01"
tags = ["gpu:h100", "region:us-east"]
lanes = []

[runtime]
workdir = "/var/lib/dreamlake/udf"
max_invocations = 100
runners = ["process", "docker"]
default_runner = "process"
keep_alive_s = 300
auto_update = false
auto_update_check_interval_s = 300

[poll]
wait_max_s = 20.0
retry_min_s = 1.0
retry_max_s = 60.0

[shutdown]
grace_s = 60

[log]
level = "info"
format = "human"
event_buffer = 1024

[introspect]
enabled = false
bind = "127.0.0.1"
port = 9876

[identity]
enabled = false
key_file = "/etc/dreamlake/daemon.key"
enroll_token_file = "/etc/dreamlake/enroll-token"

[mounts]
cache_dir = "/var/lib/dreamlake/cache"
cache_max_bytes = 21474836480

[server]

Exactly four keys.

FieldTypeDefaultNotes
urlstringhttp://localhost:8080Control-plane base URL.
token_filepath | nullnoneLegacy bearer token, sent as Authorization: Bearer on every request.
tls_verifybooltruefalse sets reqwest's danger_accept_invalid_certs. Local dev only.
namespacestring"default"Used only for namespace-scoped routes (the payload presign endpoint).
There is no `transport` or `socket_path` key

ServerConfig has no transport, no socket_path, and no Unix-socket mode. Because [server] is not deny_unknown_fields, a transport = line would be silently ignored rather than honored.

[daemon]

FieldTypeDefaultNotes
machine_idstring"auto""auto" resolves to the hostname on /hello.
labelstring""Display name; falls back to the worker id when empty.
tagslist[string][]Merged with auto-detected host tags; config tags win on a key= collision.
laneslist[string][]Queue memberships, sent as HelloRequest.lanes (the back-compat wire name for queues). Omitted from the wire entirely when empty, which the control plane reads as default-queue membership.
The launcher writes `queues`, this build reads `lanes`

The control plane's daemon-bootstrap renderer emits a queues = [...] key in [daemon]. The nymph source in this checkout has no queues field — only lanes — so that key parses without error and is then ignored. If you are hand-writing the TOML for this build, use lanes.

[runtime]

FieldTypeDefaultNotes
workdirpath/var/lib/dreamlake/udfRoot for per-invocation directories. See the workdir contract.
max_invocationsint100Concurrency cap, enforced by a semaphore. 0 = unbounded.
runnerslist[string]["process"]Advertised on /hello. Detected runners (slurm, kube) are appended if not already present.
default_runnerstring"process"Used when a run config names no runner.
keep_alive_sint300-1 = pool mode, never idle-exit. 0 = exit as soon as the idle clock starts. N = exit after N continuous idle seconds.
auto_updateboolfalseConverge to the namespace's target_nymph_version.
auto_update_check_interval_sint300Floor on how often the daemon reacts to a version advertisement.

Runner kinds accepted at dispatch: process, subprocess (alias for process), docker, gvisor, slurm, kube. An unknown kind fails the invocation with unknown runner kind: <k> — it does not fall back to the default runner.

There is no default_image key and no [runtime.runner] section. The image comes from run_config.image per invocation; the docker runner fails immediately with ImageRequired when it is missing.

[poll], [shutdown], [mounts]

FieldDefaultNotes
poll.wait_max_s20.0Requested long-poll window. The control plane hard-caps its own side at 20 s.
poll.retry_min_s1.0Backoff floor; a successful poll resets to this.
poll.retry_max_s60.0Backoff ceiling. Growth is min(backoff * 2, retry_max_s) with full jitter.
shutdown.grace_s60Seconds in-flight work gets after SIGTERM before the kill token fires.
mounts.cache_dir/var/lib/dreamlake/cacheMount cache root.
mounts.cache_max_bytes2147483648020 GiB.

[log], [introspect], [identity]

[log]: level (default "info"), format ("human" | "json"; an unrecognised value warns to stderr and falls back to human), event_buffer (default 1024; 0 disables the ring buffer). RUST_LOG overrides level; LAKESHORE_LOG_FORMAT overrides format.

[introspect]: enabled (default false), bind (default 127.0.0.1), port (default 9876; 0 = OS-assigned). The endpoint is unauthenticated, so a non-loopback bind is refused at startup — the failure is logged as a warning and introspection is skipped, not fatal.

[identity]: enabled (default false), key_file (default /etc/dreamlake/daemon.key), enroll_token, enroll_token_file. The inline token wins over the file. If enabled = true and the key cannot be loaded, the daemon fails closed rather than silently degrading to the bearer path.

Both are documented in full on Telemetry and introspection and Key rotation.

Daemon env vars

VarEquivalent
LAKESHORED_CONFIG--config
LAKESHORE_SERVER--server
LAKESHORE_TOKEN--token
DREAMLAKE_QUEUE--queue (default "default")
RUST_LOGoverrides log.level
LAKESHORE_LOG_FORMAToverrides log.format
`LAKESHORE_INTROSPECT` is not read

A comment in config.rs mentions LAKESHORE_INTROSPECT as a way to enable the introspection server. No code reads it. The only switch is [introspect] enabled = true.

.lakeshore project defaults

Two files at the project root, read only by lakeshore daemon launch: .lakeshore (checked in) and .lakeshore.local (gitignored override). Discovery walks up from cwd and stops at the filesystem root or at any directory containing .git. The first level where either file exists wins; .lakeshore.local layers on top with a shallow spread — arrays replace, they do not concatenate.

.lakeshoreyaml
daemon:
  provider: aws-prod
  name: trainer
  runners: [process, docker]
  tags: [gpu:h100]
  queues: [training]
  keep_alive_s: 600
  capacity: { instance_type: p5.48xlarge }
  bootstrap_script: ./scripts/cloud-init.sh
  setup_scripts:
    - ./scripts/install-cuda.sh

Every key is optional and unknown top-level keys are ignored. lanes is a legacy synonym for queues. bootstrap_script and setup_scripts paths resolve relative to the config file's directory, not cwd. Full walkthrough: Daemon lifecycle → Project defaults.

lakeshore.yaml compose

The declarative topology file for lakeshore up | down | ps | status | logs. Looked up as ./lakeshore.yaml from cwd only — there is no upward walk — or named explicitly with -f, --file.

Requires version: 1, a project: slug matching ^[a-z0-9][a-z0-9-]*$, and a non-empty daemons: mapping. Optional at top level: namespace, provider, queues.

lakeshore.yamlyaml
version: 1
project: research
namespace: default
provider: aws-prod

queues:
  training:
    kind: fifo
    elasticity: { kind: pool_with_threshold, min: 0, max: 8, threshold: 0.7 }

daemons:
  trainers:
    queues: [training]
    count: 2
    instance_type: p5.48xlarge
    runner: docker
Elasticity kinds are spelled differently on the two surfaces

lakeshore.yaml accepts only the underscored forms — fixed, fully_elastic, pool_with_threshold, max_count. The lakeshore queues add --elasticity flag documents the hyphenated forms — fixed, fully-elastic, pool-with-threshold, max-count. Neither spelling is universal.

Daemon-group keys: queues (required, non-empty), count (required positive integer, no default), plus optional provider, instance_type, image_id, runner, keep_alive_s, setup[], startup, python, mounts[]. Any other key is a hard error listing the allowed set. Full reference: Compose.

See also