Elasticity
A queue's elasticity decides whether the control plane is allowed to grow or shrink the fleet of daemons subscribed to it. Four named policies cover the shapes we have seen in real workloads.
The controller is a single async loop on the control plane. Every tick
it reads each state="active" queue's elasticity object, gathers three
signals — daemon count, utilization, pending invocations — and produces
exactly one decision: scale_up, scale_down, or tick_noop. Every
decision, noops included, lands a row in the ElasticityEvent
collection.
The controller starts only when the control plane is booted with
ELASTICITY_ENABLED=true. The tick interval comes from
ELASTICITY_TICK_S and defaults to 5 seconds.
Each tick resolves the default namespace and iterates its queues.
Queues in other namespaces are not scaled by the controller in this
build.
The four policies
| Policy | What the controller does | When to pick |
|---|---|---|
fixed | Nothing. Daemons are added and removed by hand. | You manage the fleet yourself or via Compose. |
fully_elastic | Targets min(pending, max) daemons, moving one step per tick. | Bursty jobs where spin-up cost is small next to runtime. |
pool_with_threshold | Hysteresis on utilization — scale up when it stays above threshold, down when it stays below threshold_low. | Steady traffic with bursts; you want warm capacity. |
max_count | Same target arithmetic as fully_elastic, framed as a hard cap. | Budget guardrails where the ceiling is the point. |
The stored values use underscores. The lakeshore queues --elasticity flag accepts the hyphenated spelling (fully-elastic,
pool-with-threshold, max-count) and converts before the request;
lakeshore.yaml and the HTTP API accept only the underscored form.
One decision per queue per tick. The controller never launches or kills more than one daemon per tick per queue, which caps the blast radius of a single bad signal reading.
Picking a policy
fixed
The default when no elasticity object is present, and the simplest.
Workers are added with daemon launch --queue … and removed with
daemon kill. Pair it with Compose to keep the
fleet shape in source control while still naming an exact count.
fully_elastic
Target is min(pending, max), floored at min. If the target exceeds
daemon_count + in_flight, the controller launches one daemon; if
daemon_count exceeds the target (and min), it kills one idle daemon.
pool_with_threshold
Keep min daemons warm. When utilization — sum(current_invocations) / sum(max_invocations) across members — stays at or above threshold
(default 0.8) for two consecutive ticks, launch one. When it stays at or
below threshold_low (default 0.3) for two consecutive ticks, kill one
idle daemon. In between, both counters reset and nothing happens.
max_count
Arithmetically the same as fully_elastic: target is
max(min(pending, max), min). The difference is intent — you set max
as a budget ceiling and let the pool drain to zero when work stops.
Scale-up requires a provider
scaleUp reads queue.providerRef, falling back to
daemonTemplate.provider. With neither set, the launch fails, the
decision is downgraded to tick_noop, and the reason recorded on the
audit row is queue "<name>" has no providerRef / daemonTemplate.provider — cannot launch. Any non-fixed policy therefore needs --provider.
Shared pools
Queues carrying the same elasticity.pool string are grouped and ticked
as one unit:
- Signals aggregate.
pendingis summed across member queues; workers are de-duplicated by id before counting and computing utilization. - Policy comes from the first member queue. All members should
declare the same
kind/min/max. - Scale-up joins every member queue. One launch, membership in all of them.
- The audit row uses
pool:<name>as itsqueueName, soqueues eventson an individual member will not show pool decisions.
Work-stealing
Work-stealing is a poll-time decision made by the dispatcher, not a controller decision. When a worker polls its own queues and finds nothing, the server resolves the sibling queues that share a pool with them and tries those before sleeping.
Match constraints
elasticity.match is a hard gate on stealing. A queue with
match: { tags: ["gpu:h100"] } only offers work to workers whose tags
contain every listed tag; match.runners works the same way against the
worker's runners. Queues with no match accept any worker in the pool.
Steal delay
steal_delay_s is how long a worker must be idle before it will steal
from a non-native sibling. The value used is the maximum across the
worker's own queues. It gives cheaper native workers a head start — an
expensive GPU worker with steal_delay_s: 5 will not grab a cheap CPU
job that a CPU worker might claim within five seconds. Default 0.
Cost-aware scale-down
When the controller scales down it only considers idle workers
(capabilities.current_invocations == 0, or zero running invocations
when the daemon does not report the field). Candidates are then sorted
by cost weight descending, ties broken by oldest lastSeenAt.
Cost weight is read from capabilities.cost_weight when the daemon
reports it; otherwise a worker carrying any gpu:* tag is inferred as
10 and everything else as 1.
ElasticitySpec declares a cost_weight field, and it round-trips
through the queue row, but the cull actually ranks workers using
Worker.capabilities.cost_weight and the gpu:* tag heuristic. Setting
cost_weight on a queue does not change which worker gets killed.
Setting the knobs
Only a subset of the spec has CLI flags:
| Field | Default | CLI flag |
|---|---|---|
kind | fixed | queues add --elasticity |
min | 0 | queues add/patch --min |
max | unbounded | queues add/patch --max |
threshold | 0.8 | queues add/patch --threshold |
threshold_low | 0.3 | — HTTP only |
pool | — | — HTTP only |
match.tags / match.runners | — | — HTTP only |
steal_delay_s | 0 | — HTTP only |
cost_weight | 1 | — HTTP only |
For the HTTP-only fields, PATCH the whole elasticity object:
kind is the only field the server validates; unknown keys inside
elasticity are stored and ignored.
The related admission knobs live on a separate object:
| Field | On | Meaning |
|---|---|---|
max_depth | admission | Depth at which on_full should kick in. |
on_full | admission | reject or block. Stored, not enforced today. |
Watching it work
tick_noop rows are hidden by default so a stable fleet does not bury
the audit log. If events --verbose shows nothing at all, the
controller is not running — check ELASTICITY_ENABLED on the server.
Patching a live queue
The controller picks up a new spec on its next sweep — no relaunch:
To take a queue out of service without touching its daemons:
Because the controller only iterates state="active" queues, draining
is also the way to freeze scaling for one queue.
See also
- Scaling rules — the complete rule book, signal by signal.
- Queues — the queue row the spec lives on.
- Compose — declarative fleet sizing when you
want an exact
count, not a policy. - Providers — the launchers the controller calls.
/dev/queues— the controller design note.