Scaling Rules
The elasticity controller runs a loop on the control plane. Every tick
it reads each active queue's elasticity spec, gathers signals, and
applies the rules below. This page is the reference for what the
controller does and why; Elasticity is the
narrative version.
The loop starts only when the control plane is booted with
ELASTICITY_ENABLED=true. Tick interval is ELASTICITY_TICK_S, default
5 seconds. Each tick iterates the queues of the default namespace
only.
Tick lifecycle
A failed launch or a scale-down with no eligible victim is downgraded
to tick_noop and the failure reason is appended to the audit row's
reason — a broken provider never wedges the loop.
Signal gathering
| Signal | Definition |
|---|---|
daemon_count | Workers whose queues array contains the queue name. For the default queue, workers with an empty array count too. For pools, workers are de-duplicated by id across member queues. |
utilization | sum(capabilities.current_invocations) / sum(capabilities.max_invocations) over those workers. max_invocations falls back to 1 per worker. If no worker reports current_invocations, the numerator falls back to counting state="running" invocations against those worker ids. 0 when there are no workers. |
pending | state="queued" invocations whose runConfig.queue equals the queue name. Summed across members for a pool. |
Policy rules
Let in_flight be the number of launches recorded by the
launch tracker for this queue.
fixed
Do nothing. Every tick is a tick_noop with reason fixed: no-op.
An unrecognised elasticity.kind also lands here, with the reason
unknown elasticity.kind="…" — defaulting to no-op.
fully_elastic
pool_with_threshold
Hysteresis on utilization; the window is 2 consecutive ticks.
The acting counter resets to zero after a decision, so the next scale
step needs two more qualifying ticks. Hitting max (or min) with the
counter satisfied produces a tick_noop that names the bound, and does
not reset the counter.
max_count
Same arithmetic as fully_elastic, with max defaulting to 0 instead
of unbounded — an unset max therefore scales nothing up.
Scale-up rules (all policies)
- Provider required.
queue.providerRefwins, falling back todaemonTemplate.provider. Neither set → launch fails →tick_noop. - Launch body comes from
queue.daemonTemplate:runners(or a singlerunner, else["process"]),capacity,cloudTags,bootstrap_script, andsetup_scriptsmapped to setup commands. - Queue membership on the new worker is
[queue.name], or[]for thedefaultqueue. - One per tick. At most one daemon is launched per queue or pool per tick.
Scale-down rules (all policies)
- Idle-only. A worker is a candidate when
capabilities.current_invocations == 0, or — when the field is absent — when it has zerostate="running"invocations. Busy workers are never killed. - Cost-first. Candidates sort by cost weight descending.
capabilities.cost_weightwins if present; otherwise a worker with anygpu:*tag is10and everything else is1. - Then oldest. Ties break by
lastSeenAtascending. - One per tick. At most one worker per queue or pool per tick.
- Kill is a row delete. The default kill path deletes the
Workerrow; a still-running daemon that hellos again re-creates it.
Shared pools
Queues with matching elasticity.pool values form a pool.
| Rule | Behavior |
|---|---|
| Grouping | Same pool string → one group. No pool → ticked individually. |
| Signal aggregation | pending summed across members; workers de-duplicated by id; utilization computed from the de-duplicated set. |
| Policy source | The first member queue's spec is canonical. Members should agree on kind / min / max. |
| Scale-up | The launched worker joins all member queue names (excluding default). Provider and template come from the first member. |
| Scale-down | Same idle-only / cost-first / oldest rules, over the de-duplicated worker set. |
| Audit trail | Events are written with queueName = "pool:<name>". |
Work-stealing rules
When a worker polls its own queues and finds no work, the dispatcher tries sibling queues in the same pool before sleeping. This is a poll-time decision in the request handler — the controller is not involved.
| Rule | Behavior |
|---|---|
| Eligibility | Only state="active" queues sharing a pool with one of the worker's own queues. The worker's own queues are excluded. |
| Match gate | If a sibling declares elasticity.match, the worker's tags must include every entry in match.tags and its runners every entry in match.runners. |
| Steal delay | The worker must have been idle for at least steal_delay_s seconds, taken as the maximum across its own queues. |
| Own-first | Own queues are always tried first, every spin. Stealing is the fallback. |
Dispatch latency
| Path | Mechanism |
|---|---|
| Work already queued when the worker polls | Immediate findAndModify claim. |
| Work arrives while the worker is long-polling | POST /v1/producer/submit calls bumpQueue(queueName), which interrupts the idle wait. |
| Work arrives between polls | Picked up on the next poll. |
The poll's idle wait is min(5000 ms + 0–500 ms jitter, remaining), and
the whole long-poll is capped at 20 seconds
(min(body.timeout_s ?? 20, 20)).
Launch tracking
LaunchTracker prevents a double-spawn inside one launch window. When
the controller launches a daemon it records an in-flight token keyed by
queue name; later ticks subtract the in-flight count from the target.
Tokens auto-expire after 15 seconds — long enough for the cloud
launch plus the first /hello, short enough that a failed launch does
not permanently consume a slot. The token is also resolved explicitly
when the launch call returns, success or failure.
Configuration reference
| Field | On | Default | Description |
|---|---|---|---|
kind | elasticity | "fixed" | Policy name. The only validated field. |
min | elasticity | 0 | Floor for daemon count. |
max | elasticity | unbounded (0 for max_count) | Ceiling for daemon count. |
threshold | elasticity | 0.8 | High watermark for pool_with_threshold. |
threshold_low | elasticity | 0.3 | Low watermark for pool_with_threshold. |
pool | elasticity | — | Shared pool name. |
match.tags / match.runners | elasticity | — | Worker requirements for work-stealing. |
steal_delay_s | elasticity | 0 | Idle seconds before stealing from siblings. |
cost_weight | Worker.capabilities | 1 (10 inferred for gpu:* tags) | Scale-down priority. Higher is killed first. |
providerRef | queue row | — | Provider the controller launches under. Required for non-fixed. |
daemonTemplate | queue row | — | { instance_type, image_id, runner, … } for launches. |
ELASTICITY_ENABLED | env var | false | Enable the controller loop. |
ELASTICITY_TICK_S | env var | 5 | Tick interval in seconds. |
See also
- Elasticity — concept overview and CLI examples.
- Queues — the queue primitive and lifecycle.
- Compose — declarative fleet sizing.
- Source:
src/elasticity/controller.tsindreamlake-ai/lakeshore-controlplane.