DreamLake

13 · Cloud provider launch

The "spin a real cloud VM, install nymph, register" flow. The most complex path: cloud API call, secret decryption, bootstrap script, binary download, and daemon registration on a fresh image.

Exercises

  • The server-side launchers — EC2, GCE, and Kube (the only three the launch routes support; SSH and SLURM return 422).
  • Secret resolution at launch time: { "$secret": "…" } markers in provider kwargs decrypted in the handler and never serialized back.
  • POST /v1/namespaces/:ns/daemons/launch rendering the nymph bootstrap and handing it to the launcher as userdata / startup script.
  • Outbound-only registration from a fresh instance.

Requires

  • Cloud credentials stored as a Lakeshore secret — kind aws_keypair (plaintext formatted <access_key_id>:<secret_access_key>) or gcp_sa_json.
  • A provider registered with those credentials referenced by { "$secret": … }.
  • IAM / project permission to create, describe, and terminate instances.
  • Outbound HTTPS from the instance's network, so nymph can reach the control plane and R2.
  • LAKESHORE_PUBLIC_URL set on the control plane. The bootstrap script bakes in the URL from LAKESHORE_PUBLIC_URL, then LAKESHORE_SERVER_URL, then LAKESHORE_URL, then http://localhost:$PORT — and a daemon on EC2 pointed at localhost never registers.

Verifies

Cloud provisioning end to end. When green, you can stand up ephemeral workers on demand.

Run

bash
export LAKESHORE_URL=https://api.lakeshore.dreamlake.ai

# One EC2 daemon, named, waiting for it to come up:
lakeshore daemon launch --provider aws-us-east --name smoke --wait

# Four at once, into a queue, with a 30-minute idle timeout:
lakeshore daemon launch --provider aws-us-east --name gpu --count 4 \
  --queue training --runner docker --keep-alive 1800 --wait

With --count > 1 the --name becomes a prefix (gpu-00, gpu-01, …) and launches run at --concurrency (default 5, sized for the AWS RunInstances rate limit, with retry on throttle). Omitting --name falls back to <provider>-<short-ulid>.

daemon launch also reads project defaults from .lakeshore / .lakeshore.local, walking up from cwd and stopping at the git root. --no-config skips that discovery; --no-setup clears every setup script from both sources.

Then confirm:

bash
lakeshore daemon list
lakeshore providers instances aws-us-east
lakeshore daemon launch-log <worker-id>     # EC2 serial console

Expected output

A Worker row appears in daemon list — first pending or setting_up, then joining once /hello lands, then active on its first poll. providers instances shows the instance transitioning starting → running.

Tearing it down

bash
lakeshore daemon kill <id>              # terminates the cloud instance by default
lakeshore daemon kill --prefix gpu- --yes
lakeshore daemon kill <id> --no-terminate   # remove the row, leave the VM
The default is to terminate

daemon kill terminates the cloud instance unless you pass --no-terminate. Getting this backwards costs either money or data.

If it fails

SymptomLikely cause
502 with awsCode: UnauthorizedOperationIAM. The role needs ec2:RunInstances, ec2:Describe*, ec2:CreateTags, ec2:TerminateInstances, plus iam:PassRole when attaching an instance profile.
422 only EC2, GCE and Kube supportedThe provider's launcher is SSH or SLURM.
EC2 launch requires image_idNo image_id / ami in the merged provider + mode kwargs. There is no default AMI.
Instance created, never registersThe bootstrap script failed, or it points at the wrong control plane. Check /var/log/cloud-init-output.log on the host and confirm LAKESHORE_PUBLIC_URL.
secret '<name>' not found in namespaceSecrets resolve in the provider's own namespace, which is not necessarily the one you called from. Providers are org-shared; secrets are not.
Launches succeed until they suddenly do notA cloud account instance-count quota. Different problem.

Status

Manual. Costs real money — do not loop on failure.

Next

→ 14 · Dashboard observability