Bugs, fixes, and open issues
TL;DR. Four non-obvious failures hit during build-out — all resolved or worked around with acceptable production behaviour. Two longer-arc engineering items (zero-copy data path, async-yield for the xdist parent-blocks-child cycle) are tracked separately on the roadmap section below.
Status overview
| # | Issue | Severity | Status |
|---|---|---|---|
| 1 | httpx ReadTimeout vs Future.result | High | Fixed |
| 2 | Function.source missing from Prisma | Medium | Fixed |
| 3 | pytest-xdist deadlock | Medium | Workaround (-n 4 + 16 daemons) — accepted as the production posture |
| 4 | Virtual-list polish (5 sub-issues) | Low | Fixed |
1. httpx ReadTimeout vs Future.result
| Field | Detail |
|---|---|
| Symptom | Worker requests httpx.ReadTimeout at 90 s while Future.result(timeout=300) was still waiting. The HTTP layer gave up before the future would have. |
| Surface | Spurious timeout errors up the stack. |
| Fix | Bumped ControlPlaneClient default timeout from 90.0 to 600.0. |
| Effect | Future.result(timeout=…) now bounds the wait, not the HTTP layer. |
2. Function.source missing from Prisma schema
| Field | Detail |
|---|---|
| Symptom | Client sent source in function meta; the server's Prisma Function model lacked the column → upserts silently dropped it. |
| Surface | UI couldn't show source for registered functions. |
| Fix | Added source String? to the Prisma Function model. ensureFunction now upserts source. Dedup key is (module, qualname, version). Backfill on update so pre-existing rows pick up source the next time the function is seen. |
3. pytest-xdist deadlock
| Field | Detail |
|---|---|
| Symptom | pytest -n auto (8 workers on this box) deadlocks. |
| Root cause | Parent test invocation blocks on its children while still holding a daemon slot. The daemon pool is bounded, so children can't claim a slot to make progress. Classic resource-cycle deadlock. |
| Workaround | pytest -n 4 with 16 daemons. Eight free slots is plenty to break the cycle in practice. |
| Real fix | Deferred (see O2). Structural async-yield in the parent invocation so it releases its slot while waiting on children. |
4. Virtual-list polish
Five fixes from the pipeline-stream virtual-list review, all landed:
| Issue | Fix |
|---|---|
useLongPoll re-rendered on every poll | Track lastSeen cursor; skip setData when unchanged |
getItemKey missing → scroll anchor broke on prepend | Added ${pipeline.id}:${rowIndexInPipeline} |
measureElement ref unnecessary (rows are fixed height) | Removed |
estimateSize: 28 didn't match actual row height (30) | Bumped to 30 |
aggBreakdown unused | Removed |
Roadmap (not bugs)
These were earlier listed as "open" but they aren't bugs — current behaviour is acceptable in production. Tracking here so the design intent doesn't get lost.
Same-host zero-copy data path
Today: UDF-to-UDF data on the same host round-trips through cloudpickle even when the result never leaves the box. For most workloads this is invisible (the largest result the producer sees ends up shipped over msgpack anyway), but for chained UDFs that exchange large tensors / arrays this is wasteful.
Direction whenever we revisit: shared-memory channel (Python's
multiprocessing.shared_memory or Arrow IPC) routed by the daemon
when it detects co-locality. The daemon already owns the workdir
contract, so a result_shm env var pointing at a shared block is
the natural extension. Decision-deferring until the chained-UDF
workload pattern is real in usage data.
Async-yield for parent → child invocations
pytest -n auto deadlocks today (§3) because the parent invocation
holds a daemon slot while waiting on its children. Workaround
(-n 4 with 16 daemons) is the production posture; the cycle
breaks because 8 free slots are plenty. Documented for anyone who
tries -n auto.
A "real" fix would have the parent release its slot during the
wait — needs a new daemon control verb (yield / park) and a
re-claim handshake when the child returns. Worth designing but not
worth shipping until a workload actually saturates 16 daemons at
depth ≥ 2. The current workaround documented in §3 is the
recommended posture.