DreamLake

Living design doc. This is the starting point for implementing the DreamLake → Lakeshore integration. Edit freely as the design evolves.

Problem

DreamLake is the org-level platform (users, teams, billing). Lakeshore is the compute control plane (workers, queues, storage, jobs). Today they are separate systems. A user needs to know about both, configure credentials for both, and mentally map between them.

Goal: a DreamLake user should be able to run GPU workloads, access S3 storage, and monitor jobs without ever touching a Lakeshore API URL or bearer token directly. DreamLake owns the user identity; Lakeshore does the compute. Students in an org get a single login and see the clusters their advisor has set up.

Constraints

  1. Lakeshore stays unaware of DreamLake. No Lakeshore code changes. DreamLake is just another API consumer with a token. This keeps the control plane deployable standalone (open-source, self-hosted, etc.).

  2. Multiple Lakeshore CPs per org. A research lab might run one cluster on AWS us-east for GPU training and another on-prem for data preprocessing. DreamLake aggregates across them.

  3. REST, not GraphQL. Lakeshore's data model is flat CRUD resources. GraphQL's value is reducing round-trips for deeply nested queries — that's not the shape here. Adding a GraphQL layer means maintaining two API surfaces. Federation (Apollo etc.) adds operational cost that doesn't pay off at this scale.

  4. Token-per-user, not shared. Each DreamLake user gets their own Lakeshore bearer token. This gives per-user audit trails on the Lakeshore side (lastUsedAt tracking) and independent revocation.

Architecture

┌─────────────────────────────────────────────────────────┐
│  DreamLake                                              │
│                                                         │
│  ┌──────────┐  ┌──────────────┐  ┌───────────────────┐  │
│  │ Users &  │  │ Cluster      │  │ Lakeshore Proxy   │  │
│  │ Orgs     │  │ Registry     │  │ (REST fan-out)    │  │
│  └──────────┘  └──────────────┘  └───────────────────┘  │
│       │              │                    │             │
└───────┼──────────────┼────────────────────┼─────────────┘
        │              │                    │
        │     ┌────────┴──────────┐         │
        │     │                   │         │
   ┌────▼─────▼──┐         ┌──────▼─────────▼─┐
   │ Lakeshore   │         │ Lakeshore        │
   │ CP "east"   │         │ CP "on-prem"     │
   │ (Heroku)    │         │ (internal)       │
   └─────────────┘         └──────────────────┘

The proxy layer in DreamLake authenticates the user, resolves which cluster they're targeting, looks up (or mints) their Lakeshore token, and forwards the request. The response goes back to the client unmodified.

Data model (DreamLake side)

Cluster registry

One row per Lakeshore control plane connected to an org.

sql
create table lakeshore_clusters (
    id            uuid primary key default gen_random_uuid(),
    org_id        uuid not null references orgs(id),
    slug          text not null,           -- "us-east", "on-prem"
    display_name  text not null,           -- "US East (GPU fleet)"
    api_url       text not null,           -- full URL to the CP
    admin_token   text not null,           -- Lakeshore admin token
    namespace     text not null default 'default',
    region        text,
    is_default    boolean not null default false,
    created_at    timestamptz not null default now(),

    unique(org_id, slug)
);

admin_token: the org admin creates a Lakeshore admin token via lakeshore admin tokens create and pastes it here. DreamLake uses it to mint per-user tokens. This is the only secret DreamLake stores for the Lakeshore connection.

namespace: which Lakeshore namespace this org maps to. Multiple DreamLake orgs could share a Lakeshore CP by using different namespaces, but the default is one namespace per org.

User tokens

Per-user Lakeshore bearer tokens, minted lazily on first access.

sql
create table lakeshore_user_tokens (
    id            uuid primary key default gen_random_uuid(),
    user_id       uuid not null references users(id),
    cluster_id    uuid not null references lakeshore_clusters(id),
    token         text not null,           -- Lakeshore bearer token
    token_name    text not null,           -- label on the LS side
    created_at    timestamptz not null default now(),

    unique(user_id, cluster_id)
);

Tokens are minted via POST /v1/namespaces/:ns/tokens using the cluster's admin_token. The token_name is set to dreamlake:<user_email> so the Lakeshore token list shows who each token belongs to.

Cluster membership + RBAC

sql
create table cluster_memberships (
    user_id       uuid not null references users(id),
    cluster_id    uuid not null references lakeshore_clusters(id),
    role          text not null default 'member',
    created_at    timestamptz not null default now(),

    primary key (user_id, cluster_id)
);

Roles

Capabilityadminmemberviewer
View workers, queues, jobsyesyesyes
Submit jobs, execyesyesno
Access storage (presign, credentials)yesyesno
Create/delete providersyesnono
Provision/delete storage bucketsyesnono
Manage secretsyesnono
Invite/remove membersyesnono
Delete the cluster registrationyesnono

Role enforcement happens in the DreamLake proxy layer, not in Lakeshore. Lakeshore sees a valid bearer token and serves the request. DreamLake checks the role before forwarding.

Proxy layer

Route pattern

/api/clusters/:slug/lakeshore/<...lakeshorePath>

Every request to this prefix is proxied to the corresponding Lakeshore CP. DreamLake adds auth, resolves the token, and forwards.

Pseudocode

typescript
app.all("/api/clusters/:slug/lakeshore/*", async (req, res) => {
  // 1. Authenticate the DreamLake user
  const user = requireAuth(req);

  // 2. Resolve the cluster
  const cluster = await db.lakeshore_clusters.findFirst({
    where: { org_id: user.orgId, slug: req.params.slug },
  });
  if (!cluster) return res.status(404).json({ error: "cluster not found" });

  // 3. Check membership + role
  const membership = await db.cluster_memberships.findFirst({
    where: { user_id: user.id, cluster_id: cluster.id },
  });
  if (!membership) return res.status(403).json({ error: "not a member" });
  if (!hasPermission(membership.role, req.method, lakeshorePath)) {
    return res.status(403).json({ error: "insufficient role" });
  }

  // 4. Resolve or mint a Lakeshore token
  let tokenRow = await db.lakeshore_user_tokens.findFirst({
    where: { user_id: user.id, cluster_id: cluster.id },
  });
  if (!tokenRow) {
    const minted = await lakeshoreAdmin(cluster).mintToken(
      `dreamlake:${user.email}`
    );
    tokenRow = await db.lakeshore_user_tokens.create({
      data: {
        user_id: user.id,
        cluster_id: cluster.id,
        token: minted.token,
        token_name: `dreamlake:${user.email}`,
      },
    });
  }

  // 5. Forward to Lakeshore
  const lakeshorePath = req.params["*"];
  const upstream = `${cluster.api_url}/v1/namespaces/${cluster.namespace}/${lakeshorePath}`;
  const resp = await fetch(upstream, {
    method: req.method,
    headers: {
      authorization: `Bearer ${tokenRow.token}`,
      "content-type": req.headers["content-type"] ?? "application/json",
      accept: "application/json",
    },
    body: req.method !== "GET" && req.method !== "HEAD"
      ? JSON.stringify(req.body)
      : undefined,
  });

  // 6. Relay the response
  res.status(resp.status);
  const body = await resp.text();
  res.type("application/json").send(body);
});

Role permission map

typescript
function hasPermission(
  role: string,
  method: string,
  path: string,
): boolean {
  // Viewers: read-only
  if (role === "viewer") return method === "GET";
  // Members: read + write, but not infrastructure
  if (role === "member") {
    const infraPaths = ["providers", "secrets", "tokens", "tunnels"];
    const resource = path.split("/")[0];
    if (infraPaths.includes(resource) && method !== "GET") return false;
    // Block storage provisioning (POST with provision:true) but
    // allow presign + credentials
    if (resource === "storages" && method === "POST" && !path.includes("/")) {
      return false; // block creating new storage entries
    }
    return true;
  }
  // Admins: everything
  return true;
}

Client configuration

Python SDK

The Python SDK already resolves its base URL from LAKESHORE_URL. Point it at the DreamLake proxy:

python
import dreamlake.lakeshore as dls
from dreamlake.lakeshore.storage import Storage

# Option A: env vars
# LAKESHORE_URL=https://app.dreamlake.ai/api/clusters/us-east/lakeshore
# LAKESHORE_CLIENT_TOKEN=<dreamlake-session-token>

# Option B: explicit
q = dls.Queue("training", kind="fifo",
              server="https://app.dreamlake.ai/api/clusters/us-east/lakeshore")

store = Storage("checkpoints",
                base_url="https://app.dreamlake.ai/api/clusters/us-east/lakeshore")
creds = store.credentials()

The DreamLake frontend can inject these env vars (or write an auth.yml) when a user clicks "Connect" on a cluster.

CLI

Same principle — the CLI reads from ~/.config/lakeshore/auth.yml:

yaml
server: https://app.dreamlake.ai/api/clusters/us-east/lakeshore
namespace: default
token: <dreamlake-session-token>

DreamLake could provide a lakeshore dreamlake connect <cluster> command that writes this file, or the user sets it up manually via lakeshore auth login --server ... --token ....

Aggregation endpoints

For the DreamLake dashboard to show a unified view across clusters, add DreamLake-native endpoints that fan out to all of a user's clusters:

GET /api/overview
  → fans out to all clusters the user belongs to
  → returns { clusters: [{ slug, workers, queues, storages, ... }] }

GET /api/overview/workers
  → aggregated worker list across all clusters
  → each row tagged with { cluster: "us-east", ... }

These are DreamLake endpoints, not Lakeshore proxies. They call each cluster's Lakeshore API in parallel and merge the results.

Token lifecycle

Minting

  • Lazy: first time a user hits /api/clusters/:slug/lakeshore/*
  • Token name: dreamlake:<email> for audit trail
  • Stored in lakeshore_user_tokens

Revocation

  • When a user is removed from a cluster membership, DreamLake should revoke their Lakeshore token: DELETE /v1/namespaces/:ns/tokens/:name using the admin token
  • When a cluster is disconnected from the org, revoke all user tokens for that cluster

Rotation

  • If a Lakeshore token stops working (401), DreamLake should auto-mint a new one and update lakeshore_user_tokens
  • The admin token itself must be rotated manually (create new on Lakeshore, update the cluster row in DreamLake)

Student onboarding flow

1. Advisor registers a cluster in DreamLake:
   POST /api/clusters
   { slug: "gpu-east", displayName: "US East GPUs",
     apiUrl: "https://cp-east.lakeshore.dreamlake.ai",
     adminToken: "dlk_..." }

2. Advisor invites students:
   POST /api/clusters/gpu-east/members
   { email: "yitong@...", role: "member" }

3. Student logs in, sees clusters:
   GET /api/clusters → [{ slug: "gpu-east", ... }]

4. Student clicks "Connect" → DreamLake writes auth.yml
   or shows the env var snippet:
     export LAKESHORE_URL=https://app.dreamlake.ai/api/clusters/gpu-east/lakeshore
     export LAKESHORE_CLIENT_TOKEN=<session>

5. Student runs code:
   uv run train.py --epochs 50

6. DreamLake proxies → Lakeshore handles the compute.

Namespace strategies

Option A: namespace-per-org (default)

All users in an org share one Lakeshore namespace. They see each other's jobs, share queues and storage. Simple, collaborative.

Option B: namespace-per-user

Each user gets their own namespace. Full isolation — they can't see each other's jobs or storage. More setup, less collaboration.

Option C: namespace-per-team

Teams within an org get namespaces. A student belongs to "yitong-lab" namespace. Their advisor can see all team namespaces.

Recommendation: start with Option A. Add Options B/C later by making namespace a computed field on cluster_memberships instead of a fixed field on lakeshore_clusters.

Open questions

  • Billing attribution: how does DreamLake track compute costs per user? Lakeshore doesn't have billing. Options: (a) tag EC2 instances with the DreamLake user ID via provider kwargs, (b) count invocation durations per user from the Lakeshore event log, (c) defer to the cloud provider's cost explorer with tags.

  • Quota enforcement: should DreamLake enforce per-user GPU-hour quotas? If so, it needs to track usage (option b above) and reject new submissions when the quota is exceeded — either in the proxy layer or via a pre-submission check.

  • SSO / identity federation: if DreamLake uses OAuth (Google, GitHub), can Lakeshore tokens carry the upstream identity? Today Lakeshore tokens are opaque bearer strings with no identity claim. Adding a sub field to the token model would let the Lakeshore dashboard show "submitted by yitong@" without a DreamLake round-trip.

  • WebSocket / streaming: the proxy layer above is request-response. Long-poll endpoints (admin change feeds) work because they're just HTTP requests with long timeouts. But if Lakeshore adds WebSocket endpoints later (log streaming, live dashboard), the proxy needs to handle upgrades.

  • Cross-cluster operations: can a user submit a job that reads data from cluster A's storage and runs on cluster B's workers? Today: no. Each cluster is independent. Future: DreamLake could orchestrate cross-cluster by presigning on cluster A and passing the URL to cluster B.

Lakeshore API surface

The complete set of endpoints DreamLake's proxy needs to handle. The proxy forwards all of these verbatim; role enforcement only needs to inspect the method and the first path segment.

Health

MethodPathPurpose
GET/healthzLiveness check
GET/readyzReadiness check

Producer (msgpack wire, used by the Python SDK)

MethodPathPurpose
POST/v1/producer/submitSubmit an invocation
POST/v1/producer/awaitBlock until a result is ready
PUT/v1/producer/get-started/queuesUpsert a queue (legacy)

Daemon (msgpack wire, used by the nymph daemon)

MethodPathPurpose
POST/v1/daemon/helloDaemon registers itself
POST/v1/daemon/pollLong-poll for work
POST/v1/daemon/ackAcknowledge invocation completion
POST/v1/daemon/eventEmit progress / log event

Namespaced CRUD — Providers

MethodPathPurpose
POST/v1/namespaces/:ns/admin/providersCreate provider
GET/v1/namespaces/:ns/admin/providersList providers
GET/v1/namespaces/:ns/admin/providers/:nameGet provider
PATCH/v1/namespaces/:ns/admin/providers/:nameUpdate provider
DELETE/v1/namespaces/:ns/admin/providers/:nameDelete provider (soft)
POST/v1/namespaces/:ns/admin/providers/:name/launchLaunch daemon via provider
POST/v1/namespaces/:ns/admin/providers/:name/restoreRestore soft-deleted provider
GET/v1/namespaces/:ns/admin/providers/:name/instancesList cloud instances
DELETE/v1/namespaces/:ns/admin/providers/:name/instances/:idTerminate cloud instance

Namespaced CRUD — Workers

MethodPathPurpose
GET/v1/namespaces/:ns/workersList workers
GET/v1/namespaces/:ns/workers/:idGet worker detail
DELETE/v1/namespaces/:ns/workers/:idRemove worker row
POST/v1/namespaces/:ns/workers/:id/get-started/queuesUpdate queue membership
POST/v1/namespaces/:ns/workers/:id/hibernateHibernate worker
POST/v1/namespaces/:ns/workers/:id/otaPush OTA update
POST/v1/namespaces/:ns/workers/:id/resetReset worker state
GET/v1/namespaces/:ns/workers/:id/launch-logGet cloud-init log
POST/v1/namespaces/:ns/workers/cleanupGarbage-collect stale workers
POST/v1/namespaces/:ns/workers/reconcileReconcile host statuses

Namespaced CRUD — Queues

MethodPathPurpose
POST/v1/namespaces/:ns/get-started/queuesCreate queue
GET/v1/namespaces/:ns/get-started/queuesList queues
GET/v1/namespaces/:ns/get-started/queues/:nameGet queue
PATCH/v1/namespaces/:ns/get-started/queues/:nameUpdate queue
DELETE/v1/namespaces/:ns/get-started/queues/:nameDelete queue
POST/v1/namespaces/:ns/get-started/queues/:name/archiveArchive queue
POST/v1/namespaces/:ns/get-started/queues/:name/unarchiveUnarchive queue
POST/v1/namespaces/:ns/get-started/queues/:name/drainDrain queue
GET/v1/namespaces/:ns/get-started/queues/:name/statsQueue statistics
GET/v1/namespaces/:ns/get-started/queues/:name/eventsQueue event log

Namespaced CRUD — Storage

MethodPathPurpose
POST/v1/namespaces/:ns/get-started/storagesCreate storage (optional provision: true)
GET/v1/namespaces/:ns/get-started/storagesList storages
GET/v1/namespaces/:ns/get-started/storages/:nameGet storage
PATCH/v1/namespaces/:ns/get-started/storages/:nameUpdate storage
DELETE/v1/namespaces/:ns/get-started/storages/:nameDelete storage (soft, ?purge=true deletes bucket)
POST/v1/namespaces/:ns/get-started/storages/:name/presignGenerate presigned S3 URL
POST/v1/namespaces/:ns/get-started/storages/:name/credentialsVend STS temporary credentials

Namespaced CRUD — Secrets

MethodPathPurpose
POST/v1/namespaces/:ns/secretsCreate secret
GET/v1/namespaces/:ns/secretsList secrets (metadata only)
GET/v1/namespaces/:ns/secrets/:nameGet secret metadata
PATCH/v1/namespaces/:ns/secrets/:nameRotate secret
DELETE/v1/namespaces/:ns/secrets/:nameDelete secret

Namespaced CRUD — Modes

MethodPathPurpose
POST/v1/namespaces/:ns/modesCreate mode
GET/v1/namespaces/:ns/modesList modes
GET/v1/namespaces/:ns/modes/:nameGet mode
PATCH/v1/namespaces/:ns/modes/:nameUpdate mode
DELETE/v1/namespaces/:ns/modes/:nameDelete mode

Namespaced CRUD — Mounts

MethodPathPurpose
POST/v1/namespaces/:ns/get-started/mountsCreate mount
GET/v1/namespaces/:ns/get-started/mountsList mounts
GET/v1/namespaces/:ns/get-started/mounts/:nameGet mount
PATCH/v1/namespaces/:ns/get-started/mounts/:nameUpdate mount
DELETE/v1/namespaces/:ns/get-started/mounts/:nameDelete mount

Namespaced CRUD — Tunnels

MethodPathPurpose
POST/v1/namespaces/:ns/get-started/tunnelsCreate tunnel
GET/v1/namespaces/:ns/get-started/tunnelsList tunnels
GET/v1/namespaces/:ns/get-started/tunnels/:nameGet tunnel
PATCH/v1/namespaces/:ns/get-started/tunnels/:nameUpdate tunnel
DELETE/v1/namespaces/:ns/get-started/tunnels/:nameDelete tunnel

Namespaced — Tokens

MethodPathPurpose
POST/v1/namespaces/:ns/tokensCreate token (returns plaintext once)
GET/v1/namespaces/:ns/tokensList tokens
DELETE/v1/namespaces/:ns/tokens/:idRevoke token

Namespaced — Invocations

MethodPathPurpose
GET/v1/namespaces/:ns/invocationsList invocations
GET/v1/namespaces/:ns/invocations/:idGet invocation detail
DELETE/v1/namespaces/:ns/invocations/:idCancel invocation
POST/v1/namespaces/:ns/invocations/bulk-cancelBulk cancel

Namespaced — Other

MethodPathPurpose
GET/v1/namespaces/:ns/whoamiToken identity check
GET/v1/namespaces/:ns/target-versionGet nymph target version
PUT/v1/namespaces/:ns/target-versionSet nymph target version

Admin (dashboard long-poll, uses cursor-based change feeds)

MethodPathPurpose
GET/v1/admin/admin/providersProvider list with daemon/invocation counts
GET/v1/admin/admin/providers/:idProvider detail with modes
POST/v1/admin/admin/providersCreate provider (admin)
DELETE/v1/admin/admin/providers/:idDelete provider (admin)
GET/v1/admin/invocationsInvocation list (long-poll)
GET/v1/admin/invocations/:idInvocation detail (long-poll)
GET/v1/admin/pipelinesPipeline list (long-poll)
GET/v1/admin/pipelines/:idPipeline tree (long-poll)

Implementation order

  1. Cluster registry + membership tables (DreamLake DB migration)
  2. Proxy route (/api/clusters/:slug/lakeshore/*)
  3. Token minting (lazy, on first proxy call)
  4. DreamLake dashboard — cluster list page, connect flow
  5. lakeshore dreamlake connect CLI command — writes auth.yml
  6. Aggregation endpoints — unified worker/queue/storage view
  7. Role enforcement in the proxy
  8. Token revocation on membership removal
  9. Billing attribution via tags or event log