plumber · Discovery → Tracer handoff

Which components are ready to become tickets?

The plan breaks into 10 components. Six can go to tickets now. Four are blocked by two gaps in the decisions, and each gap needs one short grilling session.

✓ 6 ready ! 4 blocked by 2 gaps

Numbers like #5 point to tickets on the Discovery map.

How they depend on each other

An arrow means "the next one is built on this one." Dotted lines are support work.

flowchart LR
    P[1 · protocols] --> R[3 · resolver]
    P --> C[2 · config]
    R --> B[4 · builder / run]
    C --> B
    B --> LS[5a · LocalSequential]
    B --> MP[5b · LocalMultiprocess]
    MP --> CK[6 · checkpoints]
    LS --> CK
    R --> CH[7 · check]
    CK --> CLI[8 · CLI]
    CH --> CLI
    PK[0 · packaging + CI] -.-> P
    SH[9 · shim example + docs] -.-> R

Component by component

0Packaging + CI

Ready
  • src/plumber/, hatchling, published as ab-plumber
  • Nix + uv + direnv
  • ruff + pytest on Python 3.9 and 3.12

Copy maply's setup. Nothing to decide.

#8

1Protocols

Gap A
  • PhaseModule[Payload]
  • DatasetRef[Payload]
  • Result[Payload]

The shapes are locked. But what DatasetRef is used for is unclear. See Gap A.

#3 #5 #12

2Config loader

Gap A (small)
  • Load ./plumber.yaml, or --config PATH
  • Shallow-merge --params key=value over the file
  • Keys: pipeline, input, phases, output, execution, checkpoints

#4's YAML example has no output: block, because #14 added output later. Fix: give it the same shape as input (path + params).

#4 #14 #15

3Resolver

Ready
  • "pkg.mod:attr" → callable (:attr is required)
  • Adds the current folder to sys.path
  • name defaults to the path; duplicate names stop with an error
#2 #16

4Builder / run()

Gap A
  • Resolve paths
  • Pair each phase with its params
  • Call strategy.execute()

Depends on who reads the input. See Gap A.

#5

5aLocalSequential

Gap A
  • Runs the whole GeoDataFrame through every phase
  • No partitioning
#5

5bLocalMultiprocess

Ready after 5a
  • Splits rows by fields, by chunk size, or by worker count
  • Runs the pieces with ProcessPoolExecutor
  • partitionable: false → join the pieces, run once, split again
  • At the end: join and restore the original row order
#5 #13

6Checkpoints + --from/--to

Ready
  • checkpoint: true → <dir>/<name>.parquet (GeoParquet)
  • Under localmp, pieces are joined before saving
  • --from loads the previous phase's file; stops with an error if it's missing
  • output runs only if the run reaches the last phase
#15

7plumber check

Gap B
  • Three layers: structural → static (inspect.signature) → runtime probe
  • PASS/WARN/FAIL as text, plus report.json
  • run() calls it first, as a preflight

The runtime probe has no clear source of test data. See Gap B.

#6

8CLI

Ready
  • plumber run with --strategy --worker-count --chunk-size --params --from --to --config
  • plumber check with --config
#4 #8 #15

9Shim example + docs

Ready
  • plumber.examples.adapter_shim
  • Tested in plumber's own test suite so it can't drift
#7 #11

The two gaps

Gap A: who calls the user's input function, and what is DatasetRef for?

Two decisions on the map don't fit together:

#4 / #14:  input = the user's function → returns a GeoDataFrame
#5:        execute(phases, data: DatasetRef, ...)   # "LocalSequential loads the DatasetRef"

If the user's function does the read, execute never needs a DatasetRef to load. The same question applies to Result.ref and the user's output function. The fix touches components 1, 2, 4 and 5a.

Option 1: run() reads first

gdf = input_fn(**input_params)
execute(phases, data=gdf, ...)

Simple. DatasetRef would only name checkpoint files and Result.ref.

Option 2: the strategy reads

execute(phases,
        input=(input_fn, params), ...)

The strategy decides where the read runs. This keeps the Beam/Spark seam from #9: each worker reads its own share, so nothing has to be loaded in one place first.

Gap B: what data does check's runtime probe run on?

#6 says "a tiny synthetic GeoDataFrame." But a real phase like dissolve(by="zone_id") crashes on made-up data that has no zone_id column. That gives a false FAIL.

Sample the real input

Probe on input_fn(...).head(n). Uses real columns, but costs a real read (for example, a BigQuery query).

User-supplied sample

Let the user point to a small sample file in the config.

Skip without a sample

Skip the runtime probe when no sample is given. Structural and static checks still run.

Verdict

Ready for tickets now: 0 · 3 · 5b · 6 · 8 · 9

Waiting on two grilling tickets: 1 · 2 · 4 · 5a (Gap A) and 7 (Gap B)

Next step: open both gaps as grilling tickets on the map, and grill Gap A first, since it blocks the most components. Neither gap is in the plan.md draft yet, so both go into its Deferred list as open questions.