DreamLake

Living dev note. Iterate freely.

Where bytes live for every kind of daemon output (one-shot exec, long-running run, interactive session, daemon-itself logs).

The model

                  ┌──────────────────────────┐
                  │  CLI / Dashboard         │
                  │  stdin ↓     stdout ↑    │
                  └──────────┬───────────────┘
                             │ SSE / WS / HTTP
                  ┌──────────▼───────────────┐
                  │  Controlplane            │
                  │  ┌─────────────────────┐ │
                  │  │ TIER 1 — Redis      │ │  ← real-time
                  │  │   Streams + pub/sub │ │     hot, ms-latency
                  │  │   MAXLEN bounded    │ │     ephemeral
                  │  └──────────┬──────────┘ │
                  │             ↓ flush      │
                  │  ┌─────────────────────┐ │
                  │  │ TIER 2 — S3         │ │  ← durable archive
                  │  │   one key per exec  │ │     cold, after-end
                  │  │   lifecycle rules   │ │     queryable
                  │  └─────────────────────┘ │
                  └──────────┬───────────────┘
                             │ outbound HTTP only
                  ┌──────────▼───────────────┐
                  │  Daemon (nymph)          │
                  │  bash / docker / gvisor  │
                  └──────────────────────────┘

Tier 1 — Redis (real-time)

  • Producer (daemon) appends chunks via XADD.
  • Consumers (CLI tail, dashboard LogView, other daemons reading stdin) XREAD BLOCK for sub-millisecond latency.
  • MAXLEN ~256KB per stream caps memory; oldest bytes drop.
  • Pub/sub fan-out for multi-CLI attach (collaborative debug).
  • Stream IDs (<ms>-<seq>) enable reconnect-with-cursor — no lost bytes on flaky CLI connection.

Tier 2 — S3 (durable archive)

  • One key per exec/run/session, schema in S3 bucket provisioning.
  • Controlplane reads the full stream from Redis and writes the flattened buffer to S3 at exec/session end.
  • Lifecycle policies handle retention (exec=30d, session=7d, run=365d, daemon-logs=14d).
  • Consumers fetch via presigned GET — direct from S3, no controlplane in the byte path.

Stream layout

For each exec, run, or session:

<kind>:<id>:stdout      ← daemon writes
<kind>:<id>:stdin       ← daemon reads  (interactive only)
<kind>:<id>:control     ← signals: SIGWINCH, Ctrl-C, terminate

Three streams per active session keeps the wire orthogonal. One-shot exec only uses stdout + optionally stdin if --stdin was passed.

Why this enables TTY

A TTY session is:

  • Bidirectional bytes (keystrokes in, screen out)
  • A control channel (window resize, signals)
  • Persistent state across reconnects

Map directly onto the three Redis streams above:

rust
// Daemon's loop per interactive session
loop {
  tokio::select! {
    chunk    = redis.xread(stdin_stream)  => bash.stdin.write(chunk),
    bytes    = bash.stdout.read()         => redis.xadd(stdout_stream, bytes),
    bytes    = bash.stderr.read()         => redis.xadd(stdout_stream, bytes),
    signal   = redis.xread(control)       => apply(signal),
  }
}

(The daemon doesn't talk to Redis directly — it POSTs/long-polls to controlplane endpoints that wrap Redis. But the semantics are the same.)

What the daemon sees

ExecBody.log now points at the controlplane endpoints, not S3:

rust
struct LogSink {
    /// POST chunks here. Body: raw bytes.
    /// Content-Type: application/octet-stream.
    /// Example: /v1/daemon/exec/<exec_id>/stdout
    stdout_url: String,

    /// Only set for interactive sessions / exec --stdin. Long-poll
    /// GET — server holds until stdin bytes are available or timeout.
    /// Example: /v1/daemon/exec/<exec_id>/stdin
    stdin_url: Option<String>,

    /// Only set for interactive. Long-poll GET for control signals.
    /// Example: /v1/daemon/exec/<exec_id>/control
    control_url: Option<String>,

    /// Flush cadence on the daemon side.
    flush_ms: u32,        // 200
    flush_bytes: u32,     // 4096
}

Daemon doesn't know Redis exists. It POSTs to URLs. The controlplane puts them on Redis Streams.

What the consumer sees

The CLI / dashboard reads via SSE:

GET /v1/namespaces/:ns/exec/:id/log/stream
  → text/event-stream
  → events:
      id: 1716397800-0
      data: <base64 bytes>

      id: 1716397800-1
      data: <base64 bytes>

Last-Event-ID header on reconnect → server replays from cursor.

After exec ends, the stream returns the final chunks + a sentinel event, then closes. From that point on, consumers can fetch the durable archive at:

GET /v1/namespaces/:ns/exec/:id/log
  → 200 { url: <presigned S3 GET>, key: <s3 key>, expires_at }

Same endpoint as the v0 single-PUT design — only the source has changed (CP-driven flush from Redis instead of daemon-direct PUT).

Tier-1 provisioning

bash
heroku addons:create heroku-redis:mini -a lakeshore-controlplane-staging

Heroku injects REDIS_URL env var; controlplane reads it. Free tier covers 25MB which is plenty for the hot ring buffers (256KB cap × N concurrent execs = at most a few dozen KB to a few MB).

When we go multi-dyno, Redis is already the shared state — no migration needed.

Tier-2 provisioning

Already done — see S3 bucket provisioning. Schema + lifecycle rules unchanged.

Backward compatibility

  • No LogSink from the controlplane (old CP or REDIS_URL unset) → daemon falls back to inline-bytes-in-result-POST, same as today. Old behaviour intact.
  • Old daemon (no log field support) → CP doesn't get any stream chunks; falls back to reading stdout/stderr from the result POST and shipping THOSE to S3 at end. Slightly degraded (no real-time), but durable archive works.
  • No REDIS_URL → CP skips tier 1 entirely; nothing in LogSink.stdout_url. Daemon uses the legacy inline path.

This means real-time tier turns on/off with the env var. Useful for local dev (skip Redis, fall through to inline).

Implementation phases

PhaseScopeWhat ships
1. Redis cacheCP-side Map<exec_id, Buffer> (in-process, not Redis yet) + POST /v1/daemon/exec/:id/stdout endpoint + SSE consumer endpoint. Skip true Redis for v1 — keep state in-process.Real-time streaming on a single CP dyno
2. Daemon chunk writernymph flushes stdout/stderr to the CP endpoint every 200ms / 4KB instead of buffering in memoryDaemon producer
3. S3 archive on endCP flushes the cached buffer to S3 (existing presigning code from PR #3) on exec endDurable tier
4. CLI --streamlakeshore daemon exec --stream consumes the SSE; default still waits-then-prints for back-compatUser-visible real-time
5. Move tier 1 to RedisSwap the in-process Map for Redis Streams when we go multi-dyno or need cross-dyno fanoutMulti-dyno-ready
6. Stdin + control streamsAdd POST /stdin/chunk long-poll for daemon, POST /control/signal. Enables interactive exec --stdin and lays groundwork for sessions.Bidirectional
7. Sessionslakeshore session <daemon> — persistent bash + PTY on daemon, drives via streams from phase 6.TTY

Phases 1–4 are the v1 streaming win — phases 5–7 are the path to full TTY sessions.

What this replaces

This doc absorbed two earlier drafts (now deleted from the tree):

  • The daemon-direct-PUT design — daemon doesn't talk to S3 anymore; controlplane does. Simpler, no egress assumption, no S3 multipart-min-size constraint for real-time streaming.
  • The controlplane-buffered streaming draft — it was right in shape, but lived in a separate file as a "fallback". Promoted here as Tier 1 alongside the S3 archive (Tier 2).

The "where do bytes live" section of Sessions vs one-time commands now points at Tier 1 (Redis) for active sessions and Tier 2 (S3) for archive.

S3-direct from the daemon may come back as an optimization for huge-volume artifacts (training checkpoints, multi-GB log dumps) on daemons that do have egress — but it is not the universal path.

Open

  • Redis MAXLEN tuning — 256KB per stream is a guess. Could be per-class (exec=64KB, session=512KB).
  • Control-channel framing — text vs binary, JSON envelope vs msgpack. Defer until phase 6.
  • Stdin backpressure — if the daemon can't keep up with stdin chunks, what does the CLI see? Probably a soft cap on stdin stream MAXLEN with the CLI getting "queue full" feedback.
  • Tier-1 → Tier-2 atomicity — what if the CP crashes mid-flush? Could buffer the flush as a multipart upload that's committed only when the full stream is captured. Pre-Mature; revisit.