DreamLake

Payload store

When work crosses the wire, both the arguments going in and the return value coming out have to get there. Lakeshore keeps them in a two-tier cache:

TierWhereWhen
InlineOn the job row, as raw bytessize ≤ 256 KiB
S3Object store, reached by a presigned URLsize > 256 KiB

The threshold is a control-plane env var (INLINE_THRESHOLD_BYTES), not a per-queue knob — the same value across the fleet keeps the wire shape predictable.

Wired for exec jobs today

The payload store is live on the ExecJob path — POST /v1/namespaces/:ns/exec and the /jobs/:id/result publish/subscribe pair, which carry payloadInline / payloadRef and resultInline / resultRef. The Invocation model declares matching argsRef / resultRef columns, but no route writes them yet; invocation arguments and results still ride inline as argsBlob / resultBlob.

Everything S3-side is also gated on LAKESHORE_BLOBS_BUCKET. With that env var unset, bucketConfigured() returns false and the control plane stays entirely on the inline-bytes wire.

Why two tiers

The simple thing is "stick the bytes on the job row." That works for add(2, 3) and breaks the moment you pass a numpy array, an image, or a torch tensor.

The simple workaround — "spill big payloads to S3" — works, but proxying the bytes through the control plane just to write them to S3 puts the control plane on the bandwidth path. The fix: clients upload and download S3 directly via presigned URLs the control plane mints on demand. Bytes never touch the control plane; only the key does.

The presign route

POST /v1/namespaces/:ns/payloads/presign
  { "op": "put", "size": <n>, "contentType"?, "jobId"?, "kind"? }
  { "op": "get", "key": "payloads/…" }
→ 200 { url, bucket, key, expires_at }

On put the server mints the key — the caller cannot choose it:

payloads/<YYYY>/<MM>/<DD>/<job_id>/<kind>

<job_id> is the caller-supplied jobId (staged enqueue uploads before the job row exists, so the key must be stable) or a freshly minted ULID. <kind> defaults to args.

On get the server requires the key to start with the payloads/ prefix. That is the ownership gate today — it keeps callers out of the exec/ log-archive prefix, and the bearer-token gate in front of every /v1/namespaces/:ns/* route does the rest. Stricter per-namespace prefixing is a follow-up.

The four flows

Flow 1 — small payload in (≤ 256 KiB)

client                         control plane
  │ POST /v1/namespaces/:ns/exec     │
  │   { payloadInline: <base64> }    │
  │ ────────────────────────────────▶│
  │ ◀──── 202 { exec_id, worker_id } │

One HTTP call. The bytes ride inside the enqueue POST and are stored inline on the job row. The daemon polls, gets the bytes back, runs the command.

Flow 2 — large payload in (> 256 KiB)

client                         control plane                   S3
  │ POST payloads/presign            │                          │
  │   { op: "put", size: N }         │                          │
  │ ────────────────────────────────▶│                          │
  │ ◀──── { url, bucket, key, … }    │                          │
  │                                                             │
  │ PUT <url>  (binary body) ─────────────────────────────────▶ │
  │ ◀──── 200 ──────────────────────────────────────────────────│
  │                                                             │
  │ POST /v1/namespaces/:ns/exec     │                          │
  │   { payloadRef: { bucket, key,   │                          │
  │                   size, … } }    │                          │
  │ ────────────────────────────────▶│                          │
  │ ◀──── 202 { exec_id, worker_id } │                          │

Three hops. payloadInline and payloadRef are mutually exclusive, and the enqueue route rejects a payloadRef.key that does not start with payloads/ — so a client cannot point a job at an arbitrary S3 object.

Flow 3 — the daemon pulls a job

daemon                         control plane                   S3
  │ POST /v1/daemon/poll             │                          │
  │ ────────────────────────────────▶│                          │
  │ ◀── { exec_id,                   │                          │
  │       payload_inline|payload_ref}│                          │
  │                                                             │
  │ (if payload_ref:)                                           │
  │ POST payloads/presign            │                          │
  │   { op: "get", key }             │                          │
  │ ────────────────────────────────▶│                          │
  │ ◀──── { url }                    │                          │
  │ GET <url> ────────────────────────────────────────────────▶ │
  │ ◀──── 200 (binary) ─────────────────────────────────────────│

The daemon never sees long-lived AWS credentials. The control plane mints a short-lived presigned GET on demand; the daemon's fetch timeout is 60 s.

Flow 4 — the return value

Symmetric. The daemon presigns a PUT, uploads, and POSTs a resultRef to /v1/namespaces/:ns/jobs/:id/result. The client subscribes with GET /v1/namespaces/:ns/jobs/:id/result?wait=N and downloads via a presigned GET.

Typed exec results are not wired yet

The daemon's ExecResultPayload carries result_inline and result_ref fields, but the shell-exec path always sends both as None — a bash command has no typed return value. The bytes that flow back today are stdout/stderr and the exit code.

Encoding

Payload bytes are msgpack (with the SDK's ExtType codec for numpy / torch / PIL values). payloadInline is base64 on the JSON exec wire; the daemon's own msgpack wire carries raw bin.

Limits and safety

KnobEnv varDefault
Inline thresholdINLINE_THRESHOLD_BYTES256 KiB
Max single payloadMAX_PAYLOAD_BYTES100 MiB — larger requests get a 413
Presigned URL TTLPAYLOAD_PRESIGN_TTL_S900 s (15 min)
  • Size cap, twice. The route refuses to presign above MAX_PAYLOAD_BYTES, and the minted URL carries a Content-Length constraint so S3 itself rejects an oversized upload.
  • Auth. The presign route sits behind the same namespace bearer-token gate as the rest of /v1/namespaces/:ns/*.
  • Integrity. sha256 is optional on a payloadRef. When present, the daemon verifies it after download; a mismatch raises payload corrupt: sha256 mismatch (expected …, got …), short-circuits the exec with exit_code = -1, and never spawns the child process.

What this is not

The payload store is only for job arguments and return values. Streamed stdout/stderr is a different wire with a different bucket prefix (exec/<YYYY>/<MM>/<DD>/<exec_id>/out), a different retention policy, and a different access pattern. See /dev/log-storage.

See also

  • /dev/payload-store — the full design spec: row schema, failure modes, rollout phases, and what was deliberately left out.
  • Storages — user-registered buckets, which are a separate mechanism with their own credentials.
  • Daemon protocol — the daemon side of the poll/result wire.