# Storages

A **Storage** is a named reference to an S3-compatible bucket (or a
scoped prefix inside one) registered under a namespace. The control
plane keeps the long-lived credentials encrypted in its Secret store and
hands out **presigned URLs** or **short-lived STS credentials** on
request. Your scripts see a storage name; they never see an access key.

Two kinds, and only two:

| Kind | Config | What it is |
| --- | --- | --- |
| `s3` | `{ bucket, region?, endpoint?, creds }` | A whole bucket. |
| `s3-prefix` | `{ bucket, prefix, region?, endpoint?, creds }` | A scoped directory inside a bucket. |

Note the **hyphen** in `s3-prefix`. Omitting `endpoint` means AWS S3;
setting it targets Ceph, MinIO, R2, or Backblaze B2 (the client
switches to path-style addressing when an endpoint is present).

## Immutability

`bucket`, `prefix`, and `endpoint` are denormalized out of the config
onto the row and are **immutable after creation**. Repointing a storage
would silently break every reference to it, so you create a new entry
instead. Only `description`, `creds`, and `region` are mutable.

Two unique indexes back this: `(namespace, name)` and
`(namespace, bucket, prefix, endpoint)` — the second de-duplicates two
names pointing at the same target.

## Register a storage

```bash
lakeshore storage add my-data \
  --kind s3 \
  --bucket my-lab-datasets \
  --region us-east-1 \
  --creds aws-lab
```

`--creds <name>` is shorthand for `--kwarg 'creds.$secret=<name>'`. The
named secret must already exist in the namespace:

```bash
printf '%s:%s' "$AWS_ACCESS_KEY_ID" "$AWS_SECRET_ACCESS_KEY" \
  | lakeshore secrets add aws-lab --kind aws_keypair
```

The **storage** code path accepts an `aws_keypair` secret in either the
JSON form (`{"accessKeyId": …, "secretAccessKey": …}`) or the
colon-delimited `<access_key_id>:<secret_access_key>` form. Everywhere
else — the EC2/GCE/SSH `$secret` substitution — only the colon form
parses, so store the colon form and the same secret works for both.

The name must match `^[a-z0-9][a-z0-9-]*$`. You may prefix it with a
namespace — `lakeshore storage add team-robotics/training-data …`
targets that namespace instead of your default. The same
`<namespace>/<name>` form works on `show`, `update`, `remove`,
`presign`, and `credentials` (but not `list`).

### Provisioning a bucket

`--provision` has the control plane create the bucket for you. It is
only supported for `kind: s3`, requires a `creds.$secret` pointing at an
`aws_keypair`, and needs an explicit `--bucket` — the name is **not**
derived from the namespace.

```bash
lakeshore storage add scratch \
  --kind s3 \
  --bucket dreamlake-geyang-scratch \
  --region us-east-1 \
  --creds aws-lab \
  --provision \
  --s3-option ACL=private
```

The call is idempotent: a `HeadBucket` runs first, and an existing
bucket you own is adopted rather than recreated (the response reports
`bucketCreated: false`). `--s3-option key=value` entries are forwarded
verbatim as `CreateBucket` parameters.

## The verbs

`storage` is **singular**, unlike every sibling group. There are seven
subcommands:

| Command | What it does |
| --- | --- |
| `storage add <name>` | Register (and optionally provision). |
| `storage list [--json]` | List entries in the current namespace. |
| `storage show <name> [--json]` | Print one entry's config. |
| `storage update <name>` | `--kwarg k=v` (deep-merge, repeatable), `--config-file` (replace wholesale), `--description`. |
| `storage remove <name> [--purge]` | Delete the record. `--purge` also deletes the S3 bucket. |
| `storage presign <name> <key>` | Mint a presigned URL. |
| `storage credentials <name>` | Vend short-lived STS credentials. |

> **Warning:** The CLI does not ship `storage upload`, `download`, `duplicate`, `mv`,
> `ls`, or `sync`. Data movement goes through a presigned URL or STS
> credentials — use `curl`, the AWS CLI, `rclone`, or boto3 with the
> credentials the control plane vends. `add` takes bare `--<key> <value>`
> pairs as pass-through config fields, so an unknown flag will be swallowed
> as config rather than rejected.

## Presigned URLs

```bash
# Download URL (default), 1 hour
lakeshore storage presign my-data runs/exp-01/model.pt

# Upload URL
lakeshore storage presign my-data runs/exp-01/model.pt --put --expires-in 900
```

`--expires-in` defaults to 3600 seconds and is clamped to 86400. The key
you pass is relative to the storage's configured `prefix`; the server
joins them and returns the full key alongside the URL.

Then move the bytes yourself:

```bash
URL=$(lakeshore storage presign my-data runs/exp-01/model.pt --put)
curl -X PUT --upload-file ./model.pt "$URL"
```

## Temporary credentials

For anything that wants a real S3 client — boto3, the AWS CLI, `rclone`,
`torch.save` straight to S3 — ask for STS session credentials instead:

```bash
# Shell exports
eval "$(lakeshore storage credentials my-data --env)"
aws s3 sync ./checkpoints/ s3://my-lab-datasets/runs/exp-01/

# Or JSON, for scripting
lakeshore storage credentials my-data --json --duration 7200
```

`--duration` defaults to 3600 seconds and accepts 900–129600. The
control plane calls STS `GetSessionToken` with the storage's own
credentials.

## From Python

```python
from dreamlake.lakeshore import Storage

store = Storage("my-data")          # server/namespace/token resolved from env or auth.yml

url = store.presign("runs/exp-01/model.pt")                 # presigned GET
url = store.presign("runs/exp-01/model.pt", operation="put")

data = store.get("runs/exp-01/config.json")                 # -> bytes
store.put("runs/exp-01/out.json", b'{"loss": 0.01}')
text = store.get_text("runs/exp-01/notes.md")

creds = store.credentials(duration_seconds=7200)
session = boto3.Session(**creds.to_boto3_kwargs())
env = creds.to_env()      # AWS_ACCESS_KEY_ID / …_SECRET_ACCESS_KEY / …_SESSION_TOKEN / AWS_REGION
```

`Storage` resolves the control plane the same way the rest of the SDK
does: an explicit `base_url=`, then `LAKESHORE_URL`, then
`~/.config/lakeshore/auth.yml`, then `http://localhost:8080`. The token
comes from `LAKESHORE_CLIENT_TOKEN` (or `auth.yml`).

## Shared storage across namespaces

Register the storage under the team namespace, then reference it by
`<namespace>/<name>`:

```bash
# Admin registers it under team-robotics
lakeshore storage add team-robotics/training-data \
  --kind s3 --bucket dreamlake-team-robotics-data \
  --region us-east-1 --creds aws-team

# Any member of the org mints a URL against it
lakeshore storage presign team-robotics/training-data v2/dataset.tar
```

Cross-namespace access requires the two namespaces to share an `orgId` —
a token scoped to one namespace in an org authorises requests against
its siblings.

## HTTP

```
POST   /v1/namespaces/:ns/storages
GET    /v1/namespaces/:ns/storages
GET    /v1/namespaces/:ns/storages/:name
PATCH  /v1/namespaces/:ns/storages/:name
DELETE /v1/namespaces/:ns/storages/:name
POST   /v1/namespaces/:ns/storages/:name/presign      { key, operation?, expiresIn? }
POST   /v1/namespaces/:ns/storages/:name/credentials  { durationSeconds? }
```

Deletion is a soft delete: the row is renamed with a tombstone suffix
and `deletedAt` is stamped, so the unique index stays intact.

## See also

- [Payload store](/get-started/payloads.md) — the *other* S3 path, owned by
  the control plane and used for job arguments and results.
- [Mounts](/get-started/mounts.md) — the registry for filesystems the runner
  will attach into a job's workdir.
- [Auth and secrets](/api/auth-and-secrets.md) — how `$secret` references
  and `aws_keypair` secrets work.
- [Storage, code, and mounts examples](/cli/examples/storage-code-mounts.md)
  — runnable end-to-end commands.
