DreamLake

Scaling Rules

The elasticity controller runs a loop on the control plane. Every tick it reads each active queue's elasticity spec, gathers signals, and applies the rules below. This page is the reference for what the controller does and why; Elasticity is the narrative version.

Enable it explicitly

The loop starts only when the control plane is booted with ELASTICITY_ENABLED=true. Tick interval is ELASTICITY_TICK_S, default 5 seconds. Each tick iterates the queues of the default namespace only.

Tick lifecycle

partition active queues into (pooled by elasticity.pool) and (standalone)

for each standalone queue, then for each pool:
  1. Gather signals    →  daemon_count, utilization, pending
  2. Apply policy      →  scale_up | scale_down | tick_noop
  3. Execute decision  →  launch or kill exactly one daemon
  4. Write audit row   →  ElasticityEvent { action, reason, before, after }

A failed launch or a scale-down with no eligible victim is downgraded to tick_noop and the failure reason is appended to the audit row's reason — a broken provider never wedges the loop.

Signal gathering

SignalDefinition
daemon_countWorkers whose queues array contains the queue name. For the default queue, workers with an empty array count too. For pools, workers are de-duplicated by id across member queues.
utilizationsum(capabilities.current_invocations) / sum(capabilities.max_invocations) over those workers. max_invocations falls back to 1 per worker. If no worker reports current_invocations, the numerator falls back to counting state="running" invocations against those worker ids. 0 when there are no workers.
pendingstate="queued" invocations whose runConfig.queue equals the queue name. Summed across members for a pool.

Policy rules

Let in_flight be the number of launches recorded by the launch tracker for this queue.

fixed

Do nothing. Every tick is a tick_noop with reason fixed: no-op.

An unrecognised elasticity.kind also lands here, with the reason unknown elasticity.kind="…" — defaulting to no-op.

fully_elastic

target           = max( min(pending, max), min )      # max defaults to unbounded
scale_up   when  target - daemon_count - in_flight > 0
scale_down when  daemon_count > target AND daemon_count > min

pool_with_threshold

Hysteresis on utilization; the window is 2 consecutive ticks.

high = elasticity.threshold      or 0.8
low  = elasticity.threshold_low  or 0.3

utilization >= high  → consecutiveAbove += 1 ; consecutiveBelow = 0
utilization <= low   → consecutiveBelow += 1 ; consecutiveAbove = 0
otherwise            → both counters reset to 0        (dead band)

scale_up   when consecutiveAbove >= 2 AND daemon_count + in_flight < max
scale_down when consecutiveBelow >= 2 AND daemon_count > min

The acting counter resets to zero after a decision, so the next scale step needs two more qualifying ticks. Hitting max (or min) with the counter satisfied produces a tick_noop that names the bound, and does not reset the counter.

max_count

Same arithmetic as fully_elastic, with max defaulting to 0 instead of unbounded — an unset max therefore scales nothing up.

target           = max( min(pending, max), min )
scale_up   when  target - daemon_count - in_flight > 0
scale_down when  daemon_count > target AND daemon_count > min

Scale-up rules (all policies)

  1. Provider required. queue.providerRef wins, falling back to daemonTemplate.provider. Neither set → launch fails → tick_noop.
  2. Launch body comes from queue.daemonTemplate: runners (or a single runner, else ["process"]), capacity, cloudTags, bootstrap_script, and setup_scripts mapped to setup commands.
  3. Queue membership on the new worker is [queue.name], or [] for the default queue.
  4. One per tick. At most one daemon is launched per queue or pool per tick.

Scale-down rules (all policies)

  1. Idle-only. A worker is a candidate when capabilities.current_invocations == 0, or — when the field is absent — when it has zero state="running" invocations. Busy workers are never killed.
  2. Cost-first. Candidates sort by cost weight descending. capabilities.cost_weight wins if present; otherwise a worker with any gpu:* tag is 10 and everything else is 1.
  3. Then oldest. Ties break by lastSeenAt ascending.
  4. One per tick. At most one worker per queue or pool per tick.
  5. Kill is a row delete. The default kill path deletes the Worker row; a still-running daemon that hellos again re-creates it.

Shared pools

Queues with matching elasticity.pool values form a pool.

RuleBehavior
GroupingSame pool string → one group. No pool → ticked individually.
Signal aggregationpending summed across members; workers de-duplicated by id; utilization computed from the de-duplicated set.
Policy sourceThe first member queue's spec is canonical. Members should agree on kind / min / max.
Scale-upThe launched worker joins all member queue names (excluding default). Provider and template come from the first member.
Scale-downSame idle-only / cost-first / oldest rules, over the de-duplicated worker set.
Audit trailEvents are written with queueName = "pool:<name>".

Work-stealing rules

When a worker polls its own queues and finds no work, the dispatcher tries sibling queues in the same pool before sleeping. This is a poll-time decision in the request handler — the controller is not involved.

RuleBehavior
EligibilityOnly state="active" queues sharing a pool with one of the worker's own queues. The worker's own queues are excluded.
Match gateIf a sibling declares elasticity.match, the worker's tags must include every entry in match.tags and its runners every entry in match.runners.
Steal delayThe worker must have been idle for at least steal_delay_s seconds, taken as the maximum across its own queues.
Own-firstOwn queues are always tried first, every spin. Stealing is the fallback.

Dispatch latency

PathMechanism
Work already queued when the worker pollsImmediate findAndModify claim.
Work arrives while the worker is long-pollingPOST /v1/producer/submit calls bumpQueue(queueName), which interrupts the idle wait.
Work arrives between pollsPicked up on the next poll.

The poll's idle wait is min(5000 ms + 0–500 ms jitter, remaining), and the whole long-poll is capped at 20 seconds (min(body.timeout_s ?? 20, 20)).

Launch tracking

LaunchTracker prevents a double-spawn inside one launch window. When the controller launches a daemon it records an in-flight token keyed by queue name; later ticks subtract the in-flight count from the target. Tokens auto-expire after 15 seconds — long enough for the cloud launch plus the first /hello, short enough that a failed launch does not permanently consume a slot. The token is also resolved explicitly when the launch call returns, success or failure.

Configuration reference

FieldOnDefaultDescription
kindelasticity"fixed"Policy name. The only validated field.
minelasticity0Floor for daemon count.
maxelasticityunbounded (0 for max_count)Ceiling for daemon count.
thresholdelasticity0.8High watermark for pool_with_threshold.
threshold_lowelasticity0.3Low watermark for pool_with_threshold.
poolelasticity—Shared pool name.
match.tags / match.runnerselasticity—Worker requirements for work-stealing.
steal_delay_selasticity0Idle seconds before stealing from siblings.
cost_weightWorker.capabilities1 (10 inferred for gpu:* tags)Scale-down priority. Higher is killed first.
providerRefqueue row—Provider the controller launches under. Required for non-fixed.
daemonTemplatequeue row—{ instance_type, image_id, runner, … } for launches.
ELASTICITY_ENABLEDenv varfalseEnable the controller loop.
ELASTICITY_TICK_Senv var5Tick interval in seconds.

See also

  • Elasticity — concept overview and CLI examples.
  • Queues — the queue primitive and lifecycle.
  • Compose — declarative fleet sizing.
  • Source: src/elasticity/controller.ts in dreamlake-ai/lakeshore-controlplane.