DreamLake

Living dev note. Iterate freely.

Today nymph treats every command kind the same: one global semaphore of max_invocations slots. That's wrong for the actual workload — short interactive bash commands shouldn't queue behind a 30-minute training run, and a setup script shouldn't fight with a UDF dispatch for the same pool.

This doc proposes a work-class model: each kind of work has its own pool and its own lifecycle.

Principle: daemon is scheduler-free

The daemon does not decide who runs first, who gets capped, or who waits. It is a dumb compute substrate that runs whatever the controlplane hands it, up to OS resource limits.

  • Fairness (per user, per tenant, per group) → controlplane.
  • Admission control (queue-when-full, reject-when-overloaded) → controlplane.
  • Priority (interactive before batch, urgent before idle) → controlplane.
  • Quota enforcement → controlplane.

What the daemon does keep on its side:

  • Capacity caps per class — purely to prevent one class from exhausting OS resources before another. Not a fairness mechanism.
  • Setup serialization — 1-at-a-time for setup is a correctness invariant (host-setup scripts can't safely interleave), not a scheduling decision.

This split keeps the daemon stateless w.r.t. policy. Every scheduler knob lives one level up, where it can see the whole fleet.

The four classes

ClassLatency budgetTypical durationStateCancel signalStreaming
interactive< 100msminutes (until disconnect)persistent bash / PTYCtrl-C, disconnectalways
exec< 1ssecondsnonetimeout, CLI Ctrl-Cno
run< 1s startseconds → hoursper-invocationtimeout, killoptional
setup< 1s startseconds → minutesnone, but serializedtimeoutno

Each row is a different product surface:

  • interactive — lakeshore session <daemon> (planned). REPL-feel. Persistent bash subprocess + PTY on the daemon; bidirectional streaming. Today: doesn't exist. Punt to streaming-exec arc.
  • exec — lakeshore exec / daemon exec. One-shot bash, gets stdout/stderr/rc back. Shipped today.
  • run — lakeshore run <script> (shipped) and the canonical @dls.udf invocation path (the UDF dispatch). What the existing kind="run" already covers.
  • setup — lakeshore setup / cloud-init host-setup scripts. Serialized: only one setup runs at a time, blocks the worker from taking other work until done (this is intentional).

Daemon concurrency — single pool

One Semaphore(max_invocations), default 100. Every spawn (run / exec / interactive task) acquires from it. No per-class carve-out: there's no race condition that needs separation, and with 100 slots the priority-via-pools argument is moot (saturation is the user's problem, not the daemon's).

rust
let permits = if cfg.runtime.max_invocations == 0 {
    Semaphore::MAX_PERMITS  // unbounded
} else {
    cfg.runtime.max_invocations as usize
};
let sem = Arc::new(Semaphore::new(permits));

Setup is the one exception — kept as a separate, serialized queue (pendingSetup on the worker doc). One-at-a-time setup is a correctness invariant (host-setup scripts can't safely interleave), not a scheduling decision. It already lives outside the main spawn pool; no change.

Defaults are generous (100). The daemon is the compute substrate; it should run whatever it's handed up to OS resource limits. Fairness, admission control, per-user quotas → controlplane. The daemon stays scheduler-free on purpose: one less knob, one less place to introduce subtle ordering bugs. Operators who want unbounded set max_invocations = 0 (Tokio's Semaphore::MAX_PERMITS).

Priority delivery (controlplane side)

When the daemon polls, the controlplane pops the highest-priority pending command for which the daemon has capacity. Priority order:

  1. interactive (always preempts)
  2. setup (blocks the row state, ship ASAP)
  3. exec (cheap, latency-sensitive)
  4. run (everything else)

pendingExecs / pendingSetup / pendingCommand already exist as separate queues — adding pendingInteractive is the same pattern. The poll handler's existing "control commands short-circuit" branch already does roughly this for pendingCommand; generalize it.

What this enables

ScenarioTodayWith work-classes
daemon exec echo hi during a 30min training runQueues behind the run; up to 30min waitSub-second via the exec pool
10 users hitting the same daemon with lakeshore runSerialize on the 4-slot pool, fair sort ofStill serialize; add user-tag admission for fairness (separate work)
Setup script (cloud-init) lands while a UDF is runningSetup waits behind UDFSetup still waits — by design (serialized)
Interactive session <daemon> while UDFs runDoesn't existSession gets its own slot, no contention

Interactive sessions — the hard one

This is not just "exec with longer timeout". It's a fundamentally different protocol:

  • Bidirectional stdin/stdout (vs exec's "stdin captured at submit, stdout returned at end")
  • PTY allocation (vs exec's pipes — exec can't run vim or htop)
  • Session lifecycle: heartbeat to keep alive, cleanup on disconnect
  • Connection model: SSE / WebSocket from CLI to controlplane, proxied to a Unix socket on the daemon

It's a sibling effort to streaming exec — see Log storage (Tier 1 — Redis Streams). The chunked-result streaming there generalizes naturally to a session: stdin chunks flow in, stdout chunks flow out, never terminates until disconnect.

Implementation order

  1. Split the daemon's semaphore into per-class pools. Pure code refactor, no protocol change, no controlplane change. Configurable pool sizes in udf-daemon.toml. ~50 lines in drain.rs. Ship in the next nymph release.
  2. Priority delivery in the poll handler. Controlplane change, no daemon change. Pop pendingSetup before pendingExecs before pending queue invocations. Add an explicit priority list. ~30 lines in server.ts. Ship with the wake-up / batch PR.
  3. Per-user / per-tag admission control on the controlplane. Multi-tenant fairness. The controlplane decides which user's work gets dispatched to a daemon and in what order — daemon stays policy-free. Defer building this until we see contention in real use.
  4. Interactive sessions. Layer on the streaming-exec channel (Tier 1 — Redis Streams, see Log storage) once that lands. Separate verb, separate wire. Big effort; sequence after all of (1)–(3) plus streaming exec.

Why this ordering

(1) ships pure win with no risk — daemon-only, no protocol change. The smoke loop you're iterating on benefits immediately because exec stops competing with anything for slots.

(2) is the natural pair with the wake-up + batch design we just spec'd — once we're touching the poll handler, add the priority sort.

(3) waits for a real customer scenario; building it speculatively risks the wrong abstraction.

(4) is the big-product-bet — interactive sessions are the moment Lakeshore feels like a real platform, not just a job dispatcher. But it depends on streaming, so it's downstream of (1)–(3) and the Tier 1 wire in Log storage.