DreamLake

Storages

A Storage is a named reference to an S3-compatible bucket (or a scoped prefix inside one) registered under a namespace. The control plane keeps the long-lived credentials encrypted in its Secret store and hands out presigned URLs or short-lived STS credentials on request. Your scripts see a storage name; they never see an access key.

Two kinds, and only two:

KindConfigWhat it is
s3{ bucket, region?, endpoint?, creds }A whole bucket.
s3-prefix{ bucket, prefix, region?, endpoint?, creds }A scoped directory inside a bucket.

Note the hyphen in s3-prefix. Omitting endpoint means AWS S3; setting it targets Ceph, MinIO, R2, or Backblaze B2 (the client switches to path-style addressing when an endpoint is present).

Immutability

bucket, prefix, and endpoint are denormalized out of the config onto the row and are immutable after creation. Repointing a storage would silently break every reference to it, so you create a new entry instead. Only description, creds, and region are mutable.

Two unique indexes back this: (namespace, name) and (namespace, bucket, prefix, endpoint) — the second de-duplicates two names pointing at the same target.

Register a storage

bash
lakeshore storage add my-data \
  --kind s3 \
  --bucket my-lab-datasets \
  --region us-east-1 \
  --creds aws-lab

--creds <name> is shorthand for --kwarg 'creds.$secret=<name>'. The named secret must already exist in the namespace:

bash
printf '%s:%s' "$AWS_ACCESS_KEY_ID" "$AWS_SECRET_ACCESS_KEY" \
  | lakeshore secrets add aws-lab --kind aws_keypair

The storage code path accepts an aws_keypair secret in either the JSON form ({"accessKeyId": …, "secretAccessKey": …}) or the colon-delimited <access_key_id>:<secret_access_key> form. Everywhere else — the EC2/GCE/SSH $secret substitution — only the colon form parses, so store the colon form and the same secret works for both.

The name must match ^[a-z0-9][a-z0-9-]*$. You may prefix it with a namespace — lakeshore storage add team-robotics/training-data … targets that namespace instead of your default. The same <namespace>/<name> form works on show, update, remove, presign, and credentials (but not list).

Provisioning a bucket

--provision has the control plane create the bucket for you. It is only supported for kind: s3, requires a creds.$secret pointing at an aws_keypair, and needs an explicit --bucket — the name is not derived from the namespace.

bash
lakeshore storage add scratch \
  --kind s3 \
  --bucket dreamlake-geyang-scratch \
  --region us-east-1 \
  --creds aws-lab \
  --provision \
  --s3-option ACL=private

The call is idempotent: a HeadBucket runs first, and an existing bucket you own is adopted rather than recreated (the response reports bucketCreated: false). --s3-option key=value entries are forwarded verbatim as CreateBucket parameters.

The verbs

storage is singular, unlike every sibling group. There are seven subcommands:

CommandWhat it does
storage add <name>Register (and optionally provision).
storage list [--json]List entries in the current namespace.
storage show <name> [--json]Print one entry's config.
storage update <name>--kwarg k=v (deep-merge, repeatable), --config-file (replace wholesale), --description.
storage remove <name> [--purge]Delete the record. --purge also deletes the S3 bucket.
storage presign <name> <key>Mint a presigned URL.
storage credentials <name>Vend short-lived STS credentials.
There are no upload / download / sync verbs

The CLI does not ship storage upload, download, duplicate, mv, ls, or sync. Data movement goes through a presigned URL or STS credentials — use curl, the AWS CLI, rclone, or boto3 with the credentials the control plane vends. add takes bare --<key> <value> pairs as pass-through config fields, so an unknown flag will be swallowed as config rather than rejected.

Presigned URLs

bash
# Download URL (default), 1 hour
lakeshore storage presign my-data runs/exp-01/model.pt

# Upload URL
lakeshore storage presign my-data runs/exp-01/model.pt --put --expires-in 900

--expires-in defaults to 3600 seconds and is clamped to 86400. The key you pass is relative to the storage's configured prefix; the server joins them and returns the full key alongside the URL.

Then move the bytes yourself:

bash
URL=$(lakeshore storage presign my-data runs/exp-01/model.pt --put)
curl -X PUT --upload-file ./model.pt "$URL"

Temporary credentials

For anything that wants a real S3 client — boto3, the AWS CLI, rclone, torch.save straight to S3 — ask for STS session credentials instead:

bash
# Shell exports
eval "$(lakeshore storage credentials my-data --env)"
aws s3 sync ./checkpoints/ s3://my-lab-datasets/runs/exp-01/

# Or JSON, for scripting
lakeshore storage credentials my-data --json --duration 7200

--duration defaults to 3600 seconds and accepts 900–129600. The control plane calls STS GetSessionToken with the storage's own credentials.

From Python

python
from dreamlake.lakeshore import Storage

store = Storage("my-data")          # server/namespace/token resolved from env or auth.yml

url = store.presign("runs/exp-01/model.pt")                 # presigned GET
url = store.presign("runs/exp-01/model.pt", operation="put")

data = store.get("runs/exp-01/config.json")                 # -> bytes
store.put("runs/exp-01/out.json", b'{"loss": 0.01}')
text = store.get_text("runs/exp-01/notes.md")

creds = store.credentials(duration_seconds=7200)
session = boto3.Session(**creds.to_boto3_kwargs())
env = creds.to_env()      # AWS_ACCESS_KEY_ID / …_SECRET_ACCESS_KEY / …_SESSION_TOKEN / AWS_REGION

Storage resolves the control plane the same way the rest of the SDK does: an explicit base_url=, then LAKESHORE_URL, then ~/.config/lakeshore/auth.yml, then http://localhost:8080. The token comes from LAKESHORE_CLIENT_TOKEN (or auth.yml).

Shared storage across namespaces

Register the storage under the team namespace, then reference it by <namespace>/<name>:

bash
# Admin registers it under team-robotics
lakeshore storage add team-robotics/training-data \
  --kind s3 --bucket dreamlake-team-robotics-data \
  --region us-east-1 --creds aws-team

# Any member of the org mints a URL against it
lakeshore storage presign team-robotics/training-data v2/dataset.tar

Cross-namespace access requires the two namespaces to share an orgId — a token scoped to one namespace in an org authorises requests against its siblings.

HTTP

POST   /v1/namespaces/:ns/storages
GET    /v1/namespaces/:ns/storages
GET    /v1/namespaces/:ns/storages/:name
PATCH  /v1/namespaces/:ns/storages/:name
DELETE /v1/namespaces/:ns/storages/:name
POST   /v1/namespaces/:ns/storages/:name/presign      { key, operation?, expiresIn? }
POST   /v1/namespaces/:ns/storages/:name/credentials  { durationSeconds? }

Deletion is a soft delete: the row is renamed with a tombstone suffix and deletedAt is stamped, so the unique index stays intact.

See also