Living design doc. This is the starting point for implementing the DreamLake → Lakeshore integration. Edit freely as the design evolves.
Problem
DreamLake is the org-level platform (users, teams, billing). Lakeshore is the compute control plane (workers, queues, storage, jobs). Today they are separate systems. A user needs to know about both, configure credentials for both, and mentally map between them.
Goal: a DreamLake user should be able to run GPU workloads, access S3 storage, and monitor jobs without ever touching a Lakeshore API URL or bearer token directly. DreamLake owns the user identity; Lakeshore does the compute. Students in an org get a single login and see the clusters their advisor has set up.
Constraints
-
Lakeshore stays unaware of DreamLake. No Lakeshore code changes. DreamLake is just another API consumer with a token. This keeps the control plane deployable standalone (open-source, self-hosted, etc.).
-
Multiple Lakeshore CPs per org. A research lab might run one cluster on AWS us-east for GPU training and another on-prem for data preprocessing. DreamLake aggregates across them.
-
REST, not GraphQL. Lakeshore's data model is flat CRUD resources. GraphQL's value is reducing round-trips for deeply nested queries — that's not the shape here. Adding a GraphQL layer means maintaining two API surfaces. Federation (Apollo etc.) adds operational cost that doesn't pay off at this scale.
-
Token-per-user, not shared. Each DreamLake user gets their own Lakeshore bearer token. This gives per-user audit trails on the Lakeshore side (
lastUsedAttracking) and independent revocation.
Architecture
The proxy layer in DreamLake authenticates the user, resolves which cluster they're targeting, looks up (or mints) their Lakeshore token, and forwards the request. The response goes back to the client unmodified.
Data model (DreamLake side)
Cluster registry
One row per Lakeshore control plane connected to an org.
admin_token: the org admin creates a Lakeshore admin token via
lakeshore admin tokens create and pastes it here. DreamLake uses it
to mint per-user tokens. This is the only secret DreamLake stores for
the Lakeshore connection.
namespace: which Lakeshore namespace this org maps to. Multiple
DreamLake orgs could share a Lakeshore CP by using different
namespaces, but the default is one namespace per org.
User tokens
Per-user Lakeshore bearer tokens, minted lazily on first access.
Tokens are minted via POST /v1/namespaces/:ns/tokens using the
cluster's admin_token. The token_name is set to
dreamlake:<user_email> so the Lakeshore token list shows who each
token belongs to.
Cluster membership + RBAC
Roles
| Capability | admin | member | viewer |
|---|---|---|---|
| View workers, queues, jobs | yes | yes | yes |
Submit jobs, exec | yes | yes | no |
| Access storage (presign, credentials) | yes | yes | no |
| Create/delete providers | yes | no | no |
| Provision/delete storage buckets | yes | no | no |
| Manage secrets | yes | no | no |
| Invite/remove members | yes | no | no |
| Delete the cluster registration | yes | no | no |
Role enforcement happens in the DreamLake proxy layer, not in Lakeshore. Lakeshore sees a valid bearer token and serves the request. DreamLake checks the role before forwarding.
Proxy layer
Route pattern
Every request to this prefix is proxied to the corresponding Lakeshore CP. DreamLake adds auth, resolves the token, and forwards.
Pseudocode
Role permission map
Client configuration
Python SDK
The Python SDK already resolves its base URL from LAKESHORE_URL.
Point it at the DreamLake proxy:
The DreamLake frontend can inject these env vars (or write an
auth.yml) when a user clicks "Connect" on a cluster.
CLI
Same principle — the CLI reads from ~/.config/lakeshore/auth.yml:
DreamLake could provide a lakeshore dreamlake connect <cluster> command that
writes this file, or the user sets it up manually via
lakeshore auth login --server ... --token ....
Aggregation endpoints
For the DreamLake dashboard to show a unified view across clusters, add DreamLake-native endpoints that fan out to all of a user's clusters:
These are DreamLake endpoints, not Lakeshore proxies. They call each cluster's Lakeshore API in parallel and merge the results.
Token lifecycle
Minting
- Lazy: first time a user hits
/api/clusters/:slug/lakeshore/* - Token name:
dreamlake:<email>for audit trail - Stored in
lakeshore_user_tokens
Revocation
- When a user is removed from a cluster membership, DreamLake should
revoke their Lakeshore token:
DELETE /v1/namespaces/:ns/tokens/:nameusing the admin token - When a cluster is disconnected from the org, revoke all user tokens for that cluster
Rotation
- If a Lakeshore token stops working (401), DreamLake should
auto-mint a new one and update
lakeshore_user_tokens - The admin token itself must be rotated manually (create new on Lakeshore, update the cluster row in DreamLake)
Student onboarding flow
Namespace strategies
Option A: namespace-per-org (default)
All users in an org share one Lakeshore namespace. They see each other's jobs, share queues and storage. Simple, collaborative.
Option B: namespace-per-user
Each user gets their own namespace. Full isolation — they can't see each other's jobs or storage. More setup, less collaboration.
Option C: namespace-per-team
Teams within an org get namespaces. A student belongs to "yitong-lab" namespace. Their advisor can see all team namespaces.
Recommendation: start with Option A. Add Options B/C later by
making namespace a computed field on cluster_memberships instead
of a fixed field on lakeshore_clusters.
Open questions
-
Billing attribution: how does DreamLake track compute costs per user? Lakeshore doesn't have billing. Options: (a) tag EC2 instances with the DreamLake user ID via provider kwargs, (b) count invocation durations per user from the Lakeshore event log, (c) defer to the cloud provider's cost explorer with tags.
-
Quota enforcement: should DreamLake enforce per-user GPU-hour quotas? If so, it needs to track usage (option b above) and reject new submissions when the quota is exceeded — either in the proxy layer or via a pre-submission check.
-
SSO / identity federation: if DreamLake uses OAuth (Google, GitHub), can Lakeshore tokens carry the upstream identity? Today Lakeshore tokens are opaque bearer strings with no identity claim. Adding a
subfield to the token model would let the Lakeshore dashboard show "submitted by yitong@" without a DreamLake round-trip. -
WebSocket / streaming: the proxy layer above is request-response. Long-poll endpoints (admin change feeds) work because they're just HTTP requests with long timeouts. But if Lakeshore adds WebSocket endpoints later (log streaming, live dashboard), the proxy needs to handle upgrades.
-
Cross-cluster operations: can a user submit a job that reads data from cluster A's storage and runs on cluster B's workers? Today: no. Each cluster is independent. Future: DreamLake could orchestrate cross-cluster by presigning on cluster A and passing the URL to cluster B.
Lakeshore API surface
The complete set of endpoints DreamLake's proxy needs to handle. The proxy forwards all of these verbatim; role enforcement only needs to inspect the method and the first path segment.
Health
| Method | Path | Purpose |
|---|---|---|
| GET | /healthz | Liveness check |
| GET | /readyz | Readiness check |
Producer (msgpack wire, used by the Python SDK)
| Method | Path | Purpose |
|---|---|---|
| POST | /v1/producer/submit | Submit an invocation |
| POST | /v1/producer/await | Block until a result is ready |
| PUT | /v1/producer/get-started/queues | Upsert a queue (legacy) |
Daemon (msgpack wire, used by the nymph daemon)
| Method | Path | Purpose |
|---|---|---|
| POST | /v1/daemon/hello | Daemon registers itself |
| POST | /v1/daemon/poll | Long-poll for work |
| POST | /v1/daemon/ack | Acknowledge invocation completion |
| POST | /v1/daemon/event | Emit progress / log event |
Namespaced CRUD — Providers
| Method | Path | Purpose |
|---|---|---|
| POST | /v1/namespaces/:ns/admin/providers | Create provider |
| GET | /v1/namespaces/:ns/admin/providers | List providers |
| GET | /v1/namespaces/:ns/admin/providers/:name | Get provider |
| PATCH | /v1/namespaces/:ns/admin/providers/:name | Update provider |
| DELETE | /v1/namespaces/:ns/admin/providers/:name | Delete provider (soft) |
| POST | /v1/namespaces/:ns/admin/providers/:name/launch | Launch daemon via provider |
| POST | /v1/namespaces/:ns/admin/providers/:name/restore | Restore soft-deleted provider |
| GET | /v1/namespaces/:ns/admin/providers/:name/instances | List cloud instances |
| DELETE | /v1/namespaces/:ns/admin/providers/:name/instances/:id | Terminate cloud instance |
Namespaced CRUD — Workers
| Method | Path | Purpose |
|---|---|---|
| GET | /v1/namespaces/:ns/workers | List workers |
| GET | /v1/namespaces/:ns/workers/:id | Get worker detail |
| DELETE | /v1/namespaces/:ns/workers/:id | Remove worker row |
| POST | /v1/namespaces/:ns/workers/:id/get-started/queues | Update queue membership |
| POST | /v1/namespaces/:ns/workers/:id/hibernate | Hibernate worker |
| POST | /v1/namespaces/:ns/workers/:id/ota | Push OTA update |
| POST | /v1/namespaces/:ns/workers/:id/reset | Reset worker state |
| GET | /v1/namespaces/:ns/workers/:id/launch-log | Get cloud-init log |
| POST | /v1/namespaces/:ns/workers/cleanup | Garbage-collect stale workers |
| POST | /v1/namespaces/:ns/workers/reconcile | Reconcile host statuses |
Namespaced CRUD — Queues
| Method | Path | Purpose |
|---|---|---|
| POST | /v1/namespaces/:ns/get-started/queues | Create queue |
| GET | /v1/namespaces/:ns/get-started/queues | List queues |
| GET | /v1/namespaces/:ns/get-started/queues/:name | Get queue |
| PATCH | /v1/namespaces/:ns/get-started/queues/:name | Update queue |
| DELETE | /v1/namespaces/:ns/get-started/queues/:name | Delete queue |
| POST | /v1/namespaces/:ns/get-started/queues/:name/archive | Archive queue |
| POST | /v1/namespaces/:ns/get-started/queues/:name/unarchive | Unarchive queue |
| POST | /v1/namespaces/:ns/get-started/queues/:name/drain | Drain queue |
| GET | /v1/namespaces/:ns/get-started/queues/:name/stats | Queue statistics |
| GET | /v1/namespaces/:ns/get-started/queues/:name/events | Queue event log |
Namespaced CRUD — Storage
| Method | Path | Purpose |
|---|---|---|
| POST | /v1/namespaces/:ns/get-started/storages | Create storage (optional provision: true) |
| GET | /v1/namespaces/:ns/get-started/storages | List storages |
| GET | /v1/namespaces/:ns/get-started/storages/:name | Get storage |
| PATCH | /v1/namespaces/:ns/get-started/storages/:name | Update storage |
| DELETE | /v1/namespaces/:ns/get-started/storages/:name | Delete storage (soft, ?purge=true deletes bucket) |
| POST | /v1/namespaces/:ns/get-started/storages/:name/presign | Generate presigned S3 URL |
| POST | /v1/namespaces/:ns/get-started/storages/:name/credentials | Vend STS temporary credentials |
Namespaced CRUD — Secrets
| Method | Path | Purpose |
|---|---|---|
| POST | /v1/namespaces/:ns/secrets | Create secret |
| GET | /v1/namespaces/:ns/secrets | List secrets (metadata only) |
| GET | /v1/namespaces/:ns/secrets/:name | Get secret metadata |
| PATCH | /v1/namespaces/:ns/secrets/:name | Rotate secret |
| DELETE | /v1/namespaces/:ns/secrets/:name | Delete secret |
Namespaced CRUD — Modes
| Method | Path | Purpose |
|---|---|---|
| POST | /v1/namespaces/:ns/modes | Create mode |
| GET | /v1/namespaces/:ns/modes | List modes |
| GET | /v1/namespaces/:ns/modes/:name | Get mode |
| PATCH | /v1/namespaces/:ns/modes/:name | Update mode |
| DELETE | /v1/namespaces/:ns/modes/:name | Delete mode |
Namespaced CRUD — Mounts
| Method | Path | Purpose |
|---|---|---|
| POST | /v1/namespaces/:ns/get-started/mounts | Create mount |
| GET | /v1/namespaces/:ns/get-started/mounts | List mounts |
| GET | /v1/namespaces/:ns/get-started/mounts/:name | Get mount |
| PATCH | /v1/namespaces/:ns/get-started/mounts/:name | Update mount |
| DELETE | /v1/namespaces/:ns/get-started/mounts/:name | Delete mount |
Namespaced CRUD — Tunnels
| Method | Path | Purpose |
|---|---|---|
| POST | /v1/namespaces/:ns/get-started/tunnels | Create tunnel |
| GET | /v1/namespaces/:ns/get-started/tunnels | List tunnels |
| GET | /v1/namespaces/:ns/get-started/tunnels/:name | Get tunnel |
| PATCH | /v1/namespaces/:ns/get-started/tunnels/:name | Update tunnel |
| DELETE | /v1/namespaces/:ns/get-started/tunnels/:name | Delete tunnel |
Namespaced — Tokens
| Method | Path | Purpose |
|---|---|---|
| POST | /v1/namespaces/:ns/tokens | Create token (returns plaintext once) |
| GET | /v1/namespaces/:ns/tokens | List tokens |
| DELETE | /v1/namespaces/:ns/tokens/:id | Revoke token |
Namespaced — Invocations
| Method | Path | Purpose |
|---|---|---|
| GET | /v1/namespaces/:ns/invocations | List invocations |
| GET | /v1/namespaces/:ns/invocations/:id | Get invocation detail |
| DELETE | /v1/namespaces/:ns/invocations/:id | Cancel invocation |
| POST | /v1/namespaces/:ns/invocations/bulk-cancel | Bulk cancel |
Namespaced — Other
| Method | Path | Purpose |
|---|---|---|
| GET | /v1/namespaces/:ns/whoami | Token identity check |
| GET | /v1/namespaces/:ns/target-version | Get nymph target version |
| PUT | /v1/namespaces/:ns/target-version | Set nymph target version |
Admin (dashboard long-poll, uses cursor-based change feeds)
| Method | Path | Purpose |
|---|---|---|
| GET | /v1/admin/admin/providers | Provider list with daemon/invocation counts |
| GET | /v1/admin/admin/providers/:id | Provider detail with modes |
| POST | /v1/admin/admin/providers | Create provider (admin) |
| DELETE | /v1/admin/admin/providers/:id | Delete provider (admin) |
| GET | /v1/admin/invocations | Invocation list (long-poll) |
| GET | /v1/admin/invocations/:id | Invocation detail (long-poll) |
| GET | /v1/admin/pipelines | Pipeline list (long-poll) |
| GET | /v1/admin/pipelines/:id | Pipeline tree (long-poll) |
Implementation order
- Cluster registry + membership tables (DreamLake DB migration)
- Proxy route (
/api/clusters/:slug/lakeshore/*) - Token minting (lazy, on first proxy call)
- DreamLake dashboard — cluster list page, connect flow
lakeshore dreamlake connectCLI command — writes auth.yml- Aggregation endpoints — unified worker/queue/storage view
- Role enforcement in the proxy
- Token revocation on membership removal
- Billing attribution via tags or event log