# Payload store

When work crosses the wire, both the **arguments** going in and the
**return value** coming out have to get there. Lakeshore keeps them in a
two-tier cache:

| Tier | Where | When |
| --- | --- | --- |
| Inline | On the job row, as raw bytes | `size ≤ 256 KiB` |
| S3 | Object store, reached by a presigned URL | `size > 256 KiB` |

The threshold is a control-plane env var (`INLINE_THRESHOLD_BYTES`), not
a per-queue knob — the same value across the fleet keeps the wire shape
predictable.

> **Warning:** The payload store is live on the **ExecJob** path — `POST
> /v1/namespaces/:ns/exec` and the `/jobs/:id/result` publish/subscribe
> pair, which carry `payloadInline` / `payloadRef` and `resultInline` /
> `resultRef`. The `Invocation` model declares matching `argsRef` /
> `resultRef` columns, but no route writes them yet; invocation arguments
> and results still ride inline as `argsBlob` / `resultBlob`.
> 
> Everything S3-side is also gated on `LAKESHORE_BLOBS_BUCKET`. With that
> env var unset, `bucketConfigured()` returns false and the control plane
> stays entirely on the inline-bytes wire.

## Why two tiers

The simple thing is "stick the bytes on the job row." That works for
`add(2, 3)` and breaks the moment you pass a numpy array, an image, or a
torch tensor.

The simple workaround — "spill big payloads to S3" — works, but proxying
the bytes through the control plane just to write them to S3 puts the
control plane on the bandwidth path. The fix: **clients upload and
download S3 directly via presigned URLs** the control plane mints on
demand. Bytes never touch the control plane; only the key does.

## The presign route

```
POST /v1/namespaces/:ns/payloads/presign
  { "op": "put", "size": <n>, "contentType"?, "jobId"?, "kind"? }
  { "op": "get", "key": "payloads/…" }
→ 200 { url, bucket, key, expires_at }
```

On `put` the **server** mints the key — the caller cannot choose it:

```
payloads////<job_id>/<kind>
```

`<job_id>` is the caller-supplied `jobId` (staged enqueue uploads before
the job row exists, so the key must be stable) or a freshly minted ULID.
`<kind>` defaults to `args`.

On `get` the server requires the key to start with the `payloads/`
prefix. That is the ownership gate today — it keeps callers out of the
`exec/` log-archive prefix, and the bearer-token gate in front of every
`/v1/namespaces/:ns/*` route does the rest. Stricter per-namespace
prefixing is a follow-up.

## The four flows

### Flow 1 — small payload in (`≤ 256 KiB`)

```
client                         control plane
  │ POST /v1/namespaces/:ns/exec     │
  │   { payloadInline: <base64> }    │
  │ ────────────────────────────────▶│
  │ ◀──── 202 { exec_id, worker_id } │
```

One HTTP call. The bytes ride inside the enqueue POST and are stored
inline on the job row. The daemon polls, gets the bytes back, runs the
command.

### Flow 2 — large payload in (`> 256 KiB`)

```
client                         control plane                   S3
  │ POST payloads/presign            │                          │
  │   { op: "put", size: N }         │                          │
  │ ────────────────────────────────▶│                          │
  │ ◀──── { url, bucket, key, … }    │                          │
  │                                                             │
  │ PUT <url>  (binary body) ─────────────────────────────────▶ │
  │ ◀──── 200 ──────────────────────────────────────────────────│
  │                                                             │
  │ POST /v1/namespaces/:ns/exec     │                          │
  │   { payloadRef: { bucket, key,   │                          │
  │                   size, … } }    │                          │
  │ ────────────────────────────────▶│                          │
  │ ◀──── 202 { exec_id, worker_id } │                          │
```

Three hops. `payloadInline` and `payloadRef` are mutually exclusive, and
the enqueue route rejects a `payloadRef.key` that does not start with
`payloads/` — so a client cannot point a job at an arbitrary S3 object.

### Flow 3 — the daemon pulls a job

```
daemon                         control plane                   S3
  │ POST /v1/daemon/poll             │                          │
  │ ────────────────────────────────▶│                          │
  │ ◀── { exec_id,                   │                          │
  │       payload_inline|payload_ref}│                          │
  │                                                             │
  │ (if payload_ref:)                                           │
  │ POST payloads/presign            │                          │
  │   { op: "get", key }             │                          │
  │ ────────────────────────────────▶│                          │
  │ ◀──── { url }                    │                          │
  │ GET <url> ────────────────────────────────────────────────▶ │
  │ ◀──── 200 (binary) ─────────────────────────────────────────│
```

The daemon never sees long-lived AWS credentials. The control plane
mints a short-lived presigned GET on demand; the daemon's fetch timeout
is 60 s.

### Flow 4 — the return value

Symmetric. The daemon presigns a PUT, uploads, and POSTs a `resultRef`
to `/v1/namespaces/:ns/jobs/:id/result`. The client subscribes with
`GET /v1/namespaces/:ns/jobs/:id/result?wait=N` and downloads via a
presigned GET.

> **Note:** The daemon's `ExecResultPayload` carries `result_inline` and
> `result_ref` fields, but the shell-exec path always sends both as
> `None` — a bash command has no typed return value. The bytes that flow
> back today are stdout/stderr and the exit code.

## Encoding

Payload bytes are msgpack (with the SDK's ExtType codec for numpy /
torch / PIL values). `payloadInline` is base64 on the JSON exec wire;
the daemon's own msgpack wire carries raw `bin`.

## Limits and safety

| Knob | Env var | Default |
| --- | --- | --- |
| Inline threshold | `INLINE_THRESHOLD_BYTES` | 256 KiB |
| Max single payload | `MAX_PAYLOAD_BYTES` | 100 MiB — larger requests get a 413 |
| Presigned URL TTL | `PAYLOAD_PRESIGN_TTL_S` | 900 s (15 min) |

- **Size cap, twice.** The route refuses to presign above
  `MAX_PAYLOAD_BYTES`, *and* the minted URL carries a `Content-Length`
  constraint so S3 itself rejects an oversized upload.
- **Auth.** The presign route sits behind the same namespace
  bearer-token gate as the rest of `/v1/namespaces/:ns/*`.
- **Integrity.** `sha256` is optional on a `payloadRef`. When present,
  the daemon verifies it after download; a mismatch raises
  `payload corrupt: sha256 mismatch (expected …, got …)`, short-circuits
  the exec with `exit_code = -1`, and never spawns the child process.

## What this is not

The payload store is only for job arguments and return values. Streamed
stdout/stderr is a different wire with a different bucket prefix
(`exec////<exec_id>/out`), a different retention policy,
and a different access pattern. See
[`/dev/log-storage`](/dev/log-storage).

## See also

- [`/dev/payload-store`](/dev/payload-store) — the full design spec: row
  schema, failure modes, rollout phases, and what was deliberately left
  out.
- [Storages](/get-started/storages.md) — user-registered buckets, which are
  a separate mechanism with their own credentials.
- [Daemon protocol](/nymph/protocol.md) — the daemon side of the
  poll/result wire.
