# F6 — scopes

Row F6 of the [execution matrix](/python-sdk/architecture/execution-matrix.md): scoped
contexts. The other rows are about invocation *shape* — this row is about
the **ambient context every call runs inside**. `dls.scope` sets a key
prefix; `dls.run` resolves prefix-relative string keys to local paths. A UDF
body never sees a bucket, a credential, or an absolute path — string keys
in, string keys out, file bytes never on the msgpack wire.

| Level | Shape | Spelling |
| --- | --- | --- |
| L0 | one call | `with dls.scope(p):` plus read, write, return keys |
| L1 | fan-out | each call gets its own key ledger |
| L2 | static pipeline | `with dls.scope(f"scenes/{i:04d}"):` per item |
| L3 | dynamic graph | a scope per data-discovered branch |

> **Warning:** Despite the name, `dls.run` is not an entry point. It is the author-facing
>   proxy onto the ambient run context, and its whole surface is `prefix`,
>   `root`, `read(key)`, `write(key, dir=False)`, `read_all(keys)`, and
>   `write_all(keys, dir=False)`. `read` and `write` return
>   `pathlib.Path` objects. The things that actually run work are
>   `f(...)`, `f.submit(...)`, `f.remote(...)`, `f.local(...)`, the queue
>   verbs, and `dls.run_worker(...)`.

Everything on this page is the local tier — bare `@dls.udf`, no queue, no
worker. Scopes work with zero infrastructure and touch no network.

## L0 — the contract

A data UDF does I/O through `dls.run.read(key)` and `dls.run.write(key)` and
**returns the string keys it wrote**. Keys are prefix-relative POSIX
strings; the ambient context maps them to real locations. In local mode a
key resolves to `root / prefix / key` — `write` hands out that path with
parents created, `read` resolves it and raises if it does not exist.

The root defaults to the current working directory; the outermost
`dls.scope(..., root=...)` of a run re-anchors it.

```python
import tempfile
from pathlib import Path

import dreamlake.lakeshore as dls

root = Path(tempfile.mkdtemp())

@dls.udf
def normalize(raw: str) -> str:
    text = dls.run.read(raw).read_text()
    out = dls.run.write("process/clean.txt")
    out.write_text(text.strip().lower())
    return "process/clean.txt"          # the keys you wrote ARE the result

with dls.scope("scenes/0007", root=root):
    dls.run.write("source/raw.txt").write_text("  Hello Lakeshore  ")
    key = normalize("source/raw.txt")

assert key == "process/clean.txt"
assert (root / "scenes/0007/process/clean.txt").read_text() == "hello lakeshore"
```

`normalize` names nothing outside its scope — the scene number lives in the
ambient prefix, the machine location lives in the root.

The return-your-keys contract is enforced after every call: each key handed
out by `dls.run.write` must appear in the returned manifest. The accepted
manifest shapes are `str`, `list[str]`, `tuple[str, ...]`, and
`dict[str, str]` (named outputs, where the values are the keys). Return the
wrong thing and the error states the fix:

```python
@dls.udf
def sloppy(raw: str) -> str:
    dls.run.write("process/mesh.ply").write_text("x")
    return "done"                       # a string, but not the key
```

```
KeyError_: UDF wrote key(s) it did not return: ['process/mesh.ply'];
include them in the returned manifest (or don't write them)
```

Return nothing at all and the message spells out the accepted shapes:

```
KeyError_: UDF wrote 1 file(s) via dls.run.write (process/mesh.ply) but
did not return their keys; return the key string(s) — str, list[str],
or dict[str, str] — so the outputs are wired into the graph
```

`KeyError_` is a subclass of `ValueError`, not of the builtin `KeyError`.

The check is **one-directional**: a returned string that was never written
is allowed, because it may be plain data or an upstream key passed through
by a selector, and the two are indistinguishable by inspection. A silent
write is not allowed — lineage and downstream wiring key off the returned
manifest, so an unreturned key is an invisible artifact. (A typo'd return
still gets caught, from the other side: the correctly written key goes
unreturned and fails the check.)

## L1 — a ledger per call

Every `@dls.udf` call runs in a **fresh child context**: same root, prefix,
storage, and ids as the ambient scope, but its own write ledger. The unit of
the return-your-keys check is one call, so sequential calls under one scope
never cross-contaminate.

```python
@dls.udf
def stamp(name: str) -> str:
    dls.run.write(f"out/{name}.txt").write_text(name)
    return f"out/{name}.txt"

with dls.scope("batch", root=root) as ctx:
    keys = [stamp(n) for n in ["a", "b", "c"]]

assert keys == ["out/a.txt", "out/b.txt", "out/c.txt"]
assert ctx.writes == {}                 # each call kept its own ledger
```

If the ledger were shared, the second call would fail its check for not
returning the first call's key. It is not — `stamp("b")` answers only for
`out/b.txt`. The scope's own context stays clean; it accumulates nothing
from the calls inside it.

Generator bodies follow the same contract stretched over time: each `yield`
is a partial manifest, and together the yields are THE manifest, checked
once when the body finishes.

```python
@dls.udf
def shards(n: int):
    for i in range(n):
        dls.run.write(f"shards/{i:02d}.txt").write_text(str(i))
        yield f"shards/{i:02d}.txt"     # partial manifests concatenate

with dls.scope("gen", root=root):
    keys = list(shards(3))

assert keys == ["shards/00.txt", "shards/01.txt", "shards/02.txt"]
```

Yields that are not key-shaped concatenate to nothing, so a generator that
streams plain data has no manifest. That is fine as long as it also wrote no
files — the check only fires when the ledger is non-empty. A body that both
writes files and yields plain data fails, with the same "did not return
their keys" message.

## L2 — scope per scene

Nesting joins prefixes. A pipeline scopes the run once at the top and each
item once inside the loop; the stage chain in the middle is written entirely
in relative keys.

```python
import json

@dls.udf
def simulate(params: str) -> str:
    seed = json.loads(dls.run.read(params).read_text())["seed"]
    frames = dls.run.write("process/frames.txt")
    frames.write_text("\n".join(f"frame-{seed}-{t}" for t in range(3)))
    return "process/frames.txt"

@dls.udf
def encode(frames: str) -> str:
    n = len(dls.run.read(frames).read_text().splitlines())
    clip = dls.run.write("outputs/clip.txt")
    clip.write_text(f"clip[{n} frames]")
    return "outputs/clip.txt"

with dls.scope("dataset-v2", root=root):
    for i in range(3):
        with dls.scope(f"scenes/{i:04d}"):
            dls.run.write("source/params.json").write_text(json.dumps({"seed": i}))
            frames = simulate("source/params.json")
            clip = encode(frames)

assert (root / "dataset-v2/scenes/0002/outputs/clip.txt").read_text() == "clip[3 frames]"
```

Inside the loop, `dls.run.prefix` is `dataset-v2/scenes/0002` — the inner
scope joined onto the outer. Neither UDF nor the stage chain mentions a
scene number or a directory. Move the run by changing `root=`, re-home it in
a dataset by changing the outer scope string, and every key still reads the
same.

For a stage that takes or produces several keys at once, `dls.run.read_all`
and `dls.run.write_all` mirror the shape you hand them: a `str` maps to a
`Path`, a list to a list, a dict to a dict.

## L3 — scope per branch

When the graph's width is discovered at run time, open a scope per branch.
Each branch gets its own namespace, and keys inside stay short and identical
across branches.

A branch often needs an input from its *parent* scope. Keys may climb with
`../` — resolution joins prefix and key, so `../` climbs the scope, and the
result is checked against the run root: a key may climb out of its prefix,
never out of the run. Absolute keys are rejected outright.

```python
import posixpath

@dls.udf
def detect(frame: str) -> list[str]:
    labels = dls.run.read(frame).read_text().split()
    keys = []
    for j, label in enumerate(labels):
        dls.run.write(f"crops/{j:02d}.txt").write_text(label)
        keys.append(f"crops/{j:02d}.txt")
    return keys

@dls.udf
def refine(crop: str) -> str:
    label = dls.run.read(crop).read_text()
    dls.run.write("refined.txt").write_text(label.upper())
    return "refined.txt"

manifest = []
with dls.scope("scan-17", root=root):
    dls.run.write("source/frame.txt").write_text("cat dog heron")
    crops = detect("source/frame.txt")      # width decided by the data
    for j, crop in enumerate(crops):
        with dls.scope(f"objects/{j:02d}"):
            key = refine(f"../../{crop}")   # climb to the parent scope
            manifest.append(posixpath.join(dls.run.prefix, key))

assert manifest == [
    "scan-17/objects/00/refined.txt",
    "scan-17/objects/01/refined.txt",
    "scan-17/objects/02/refined.txt",
]
assert (root / manifest[2]).read_text() == "HERON"
```

Three labels in the frame, three branch scopes — the structure of the key
space mirrors the structure the data dictated. `refine` is the same function
in every branch; only the ambient prefix differs. The manifest is
scope-qualified by joining `dls.run.prefix` onto each returned key — plain
strings, ready for a journal line.

Climbing past the root is refused:

```
KeyError_: key '../../outside' escapes the run root from prefix
"scan-17"; ../ may climb to a parent scope, not out of the run
```

## Why an ambient context

The alternative is threading a path or a context argument through every
call, at which point UDF signatures grow a parameter that is not data and
every caller becomes responsible for plumbing. With scopes, the *pattern*
rows (F1 through F5) stay pure: a sync call is a call, a generator is a
generator, and where the bytes land is decided entirely by the `with` blocks
around them. The same body runs under a temp directory in a test, under
`scenes/0042` in a pipeline, and under a bound storage on a worker —
unchanged.
