DreamLake

Bugs, fixes, and open issues

TL;DR. Four non-obvious failures hit during build-out — all resolved or worked around with acceptable production behaviour. Two longer-arc engineering items (zero-copy data path, async-yield for the xdist parent-blocks-child cycle) are tracked separately on the roadmap section below.

Status overview

#IssueSeverityStatus
1httpx ReadTimeout vs Future.resultHighFixed
2Function.source missing from PrismaMediumFixed
3pytest-xdist deadlockMediumWorkaround (-n 4 + 16 daemons) — accepted as the production posture
4Virtual-list polish (5 sub-issues)LowFixed

1. httpx ReadTimeout vs Future.result

FieldDetail
SymptomWorker requests httpx.ReadTimeout at 90 s while Future.result(timeout=300) was still waiting. The HTTP layer gave up before the future would have.
SurfaceSpurious timeout errors up the stack.
FixBumped ControlPlaneClient default timeout from 90.0 to 600.0.
EffectFuture.result(timeout=…) now bounds the wait, not the HTTP layer.

2. Function.source missing from Prisma schema

FieldDetail
SymptomClient sent source in function meta; the server's Prisma Function model lacked the column → upserts silently dropped it.
SurfaceUI couldn't show source for registered functions.
FixAdded source String? to the Prisma Function model. ensureFunction now upserts source. Dedup key is (module, qualname, version). Backfill on update so pre-existing rows pick up source the next time the function is seen.

3. pytest-xdist deadlock

FieldDetail
Symptompytest -n auto (8 workers on this box) deadlocks.
Root causeParent test invocation blocks on its children while still holding a daemon slot. The daemon pool is bounded, so children can't claim a slot to make progress. Classic resource-cycle deadlock.
Workaroundpytest -n 4 with 16 daemons. Eight free slots is plenty to break the cycle in practice.
Real fixDeferred (see O2). Structural async-yield in the parent invocation so it releases its slot while waiting on children.

4. Virtual-list polish

Five fixes from the pipeline-stream virtual-list review, all landed:

IssueFix
useLongPoll re-rendered on every pollTrack lastSeen cursor; skip setData when unchanged
getItemKey missing → scroll anchor broke on prependAdded ${pipeline.id}:${rowIndexInPipeline}
measureElement ref unnecessary (rows are fixed height)Removed
estimateSize: 28 didn't match actual row height (30)Bumped to 30
aggBreakdown unusedRemoved

Roadmap (not bugs)

These were earlier listed as "open" but they aren't bugs — current behaviour is acceptable in production. Tracking here so the design intent doesn't get lost.

Same-host zero-copy data path

Today: UDF-to-UDF data on the same host round-trips through cloudpickle even when the result never leaves the box. For most workloads this is invisible (the largest result the producer sees ends up shipped over msgpack anyway), but for chained UDFs that exchange large tensors / arrays this is wasteful.

Direction whenever we revisit: shared-memory channel (Python's multiprocessing.shared_memory or Arrow IPC) routed by the daemon when it detects co-locality. The daemon already owns the workdir contract, so a result_shm env var pointing at a shared block is the natural extension. Decision-deferring until the chained-UDF workload pattern is real in usage data.

Async-yield for parent → child invocations

pytest -n auto deadlocks today (§3) because the parent invocation holds a daemon slot while waiting on its children. Workaround (-n 4 with 16 daemons) is the production posture; the cycle breaks because 8 free slots are plenty. Documented for anyone who tries -n auto.

A "real" fix would have the parent release its slot during the wait — needs a new daemon control verb (yield / park) and a re-claim handshake when the child returns. Worth designing but not worth shipping until a workload actually saturates 16 daemons at depth ≥ 2. The current workaround documented in §3 is the recommended posture.