Sparse coverage & the cross-resolution read path¶
Status: draft. Tracks #198; temporal-partitioning amendments (D13–D15) ratified on #237; morton-only-storage amendment ratified on #262 (D16; O10 carved from it, resolved via #305); 2026-07 consolidation (D17–D24; O3/O10/O11 resolved, O8 constants blessed, O12–O14 opened, D23 tokens ratified, §8.3 test-obligations index added; normative grammars/constants relocated to the mortie spec page per the #305 acceptance criterion) recording decisions settled on the #251/#236/#209 and #296-family threads — per-entry provenance (thread-ratified vs in-session-recorded) is cited inline.
All design decisions (both made and open) are consolidated in the Decisions registry. Inline references use D# for decisions made and O# for open items needing input. Revisit before implementation issues are carved out.
1. Motivation¶
Three pressures converged in the mortie #48 v1.0 discussion, and one set of primitives answers all of them:
- S3 write contention at global scale. A CONUS-scale run is ~50,000 order-9 shards over ~2,000 concurrent workers, each shard issuing ≥8 PUTs. S3 throttles at ~3,500 PUT/s per partition, so a single-prefix output store is intractable. Morton indices decompose into a spatially-local, hive-like prefix hierarchy — that layout is Layer 0 below.
- Multi-dataset, multi-resolution reads. ICESat-2 (~12 m morton cells), GEDI (~25 m, one order coarser), Sentinel-2 (~10 m; possibly re-encoded finer) each get their own store at their own cell/shard orders — Lambdas are always per-dataset (D7). A reader wants cross-resolution joins: "for this fine ICESat-2 observation, what GEDI observation contains it?" Because a morton decimal string prefix is the spatial ancestor, that join is truncation — arithmetic, not I/O.
- xdggs assumes dense full-sphere coordinates. The legacy
fullspherelayout (deprecated, dense path removed in 0.x — D17) materialized a coordinate entry for every cell of the global grid. For sparse-coverage data (a continent, a flight campaign) this is waste at best and intractable at global orders. The fix is a domain declaration: a coverage MOC conservatively declaring where data exists, letting the xarray extension keep coordinates sparse and fabricate dense views lazily.
The stack, bottom to top: hive store layout (§2) → static manifest (§3) → coverage MOCs (§4) → reader architecture (§5) → xarray/xdggs extension (§6), with the pyramid sweep (§7) as the post-process phase that generates all derived artifacts.
2. Layer 0: the hive store layout¶
The layout convention is owned by the mortie spec (the frozen 1.x
contract: mortie #48 discussion →
mortie #62 →
docs/specification.md,
the normative page). Grammars and constants are normative there only;
this registry records decisions and rationale and cites the page (the #305
acceptance criterion — duplicated normative text drifts). The
array-level contracts inside a leaf — the ragged vlen layout, digest
payload bytes, composition word, overview attrs, O11 hashes — are normative
on zagg's own specification.md
(issue #340), under the same
division of labor. Summary of what zagg consumes:
{store_root}/ <- multi-product form (D19): a directory of
{name}/ <- NAMED product root — each product subtree
morton_hive.json is a COMPLETE morton-hive store: bare-named
coverage.moc manifest (§3) + optional root MOC (§4, O9)
aggregation.yaml <- canonical semantic core (D19); its hash is
a frozen manifest key, not a path name
<run records (parquet)> <- run-level telemetry, one row/shard;
timestamp-first names (D20)
{sign+base}/{d1}/{d2}/.../ <- one digit per level (D2; digit-chunking is
the manifest path_grouping param, D21)
{window}.zarr/ <- leaf, basename = time window (D23,
morton-hive/3); `all.zarr` for
schedule: none (reserved token, ratified;
excluded from the window grammar).
/1–/2 stores keep {full_id}.zarr and
{full_id}_{window}.zarr (D3/D13)
{window}.stats.json <- per-shard stats sidecar, sibling (D20,
D23-aligned for /3; all.stats.json for
schedule: none). /1–/2 stores keep
stats_{window}.json / stats.json
<sub-shardmap JSON> <- leaf sub-map for sweep rollups (D22)
A bare single-product store (today's layout — morton_hive.json at the
store root, no {name}/ level) remains fully valid: the D19 product root is
additive, and a product subtree is byte-identical to a bare store. A
reader distinguishes the two forms by what sits at the root (a manifest ⇒
bare store; only name-shaped prefixes ⇒ product directory). Product names
must not match the base-component grammar (-?[1-6]) so the walker's child
classification stays unambiguous; gridlook and other viewers enumerate
products by listing {store_root}/ and reading each
{name}/morton_hive.json directly — no name↔hash translation layer.
- Ids are morton decimal strings (D1): sign + base digit (constant width,
12 values
1..6/-1..-6), then one digit per order, digits1-4, never0. String prefix = spatial ancestor at every level. - One digit per path component (D2), because shards live at mixed orders —
across datasets and within a store (coarse shards in sparse regions) —
so every order must be a legal node. D21 makes the digit-chunking a
declared manifest parameter (
path_grouping, default1= this layout); readers chunk the digit string per the manifest, never by assumption. - Full morton id at the leaf (D3):
.../1/2/3/-31123.zarr/is self-describing without parsing its path, greppable in inventories, unambiguous if moved. (Unchanged by D19 — product identity lives at the product root, above the tree. Superseded by D23 formorton-hive/3stores: the basename becomes the time window; the full id stays recoverable from the path arithmetically and from the stamp/sidecarshard_key.) - Time-windowed leaves (D13, ratified on
#237): a store whose
manifest declares a temporal window schedule (§3) partitions each shard's
time series into one write-once leaf per window at the shard node,
rather than one growing leaf. The node invariant is unchanged (a node may
hold several
*.zarrobjects — mixed orders already require that; the walker classifies them as data as before). Every windowed leaf carries its own D4 commit stamp, so all append/retry semantics reduce to the existing ones: a torn window is debris, overwritable; backfill (extending the series to earlier data) is just a new leaf for an earlier window — noresize, no read-modify-write of committed objects, no time-axis reordering; concurrent runs on different windows share no object; the window is the unit of idempotent reprocessing. The rejected alternatives — a high-water time index, and per-run stamp entries with arrayresize— both reintroduce mutable shared state at the leaf and break the binary debris rule (rationale on the #237 thread). Cross-window reads open W leaves and concatenate along time; paths stay arithmetic because the schedule lives in the manifest (D10 preserved). The no-partitioning degenerate case (schedule: none) keeps the bare{full_id}.zarrname and is byte-identical to the pre-D13 layout — amorton-hive/1store is a/2store withschedule: none. (Leaf naming is revised by D23 formorton-hive/3stores; the windowing semantics here are unchanged.) - Node invariant: below a product root, a node contains only digit
children (
[1-4]/),*.zarrobjects, and the declared leaf-adjacent sidecars — the per-shard stats record (D20) and the sub-shardmap JSON (D22) — with nothing else, ever: the walker's child classification depends on the name set being closed. The product root alone also carries the manifest, MOC objects, the semantic core (aggregation.yaml, D19), and run-level telemetry records (D20); the store root of a multi-product directory carries only{name}/product roots (D19). - Termination condition: S3 has no empty directories (a prefix exists iff ≥1 object lies beneath it) and LIST is strongly consistent, so a delimiter-LIST returning no digit-shaped children is a definitive "nothing finer exists." Absence is trustworthy.
- Presence needs a commit stamp (D4): a worker that dies mid-shard has
already created the
.zarr/prefix. The shard's final write is a rootgroup.attrs.update(...)stamping completion (plus cheap payload: cell count, write timestamp, source granule count). A.zarr/prefix whose root metadata lacks the stamp is debris — incomplete, ignorable, safe to overwrite on retry. This is not consolidated metadata: one tiny PUT rewriting an object that must exist anyway, no store-wide aggregation. - The write path needs zero metadata above the leaf (D5). No zarr group objects at digit nodes, no shared mutable state, no create-group races across 2,000 workers.
- Concurrency contract (a derived property, recorded 2026-07-20 — assembled here so it exists in one place): concurrent runs against one product are supported whenever they are shard-disjoint — they share no objects (D4/D5), their run records are distinct by name (timestamp + run id, D20), and the manifest precheck admits both (frozen keys match by construction under D19). Same-(shard, window) concurrency is out of contract: sequential re-runs give clean D4 last-writer-wins replacement, but a leaf is several objects (per-array ShardingCodec objects D17, the ragged object D18, root attrs, coverage sidecar) and S3 PUTs are not transactional across them — two simultaneous different-content co-writers can interleave those objects and leave a stamped-but-torn leaf, which is neither replacement nor ignorable debris. Detection is content verification (the O11 hash) or a re-run; the supported contract stays shard-disjoint. Interior rollups may go stale during any of this, which D22's generation stamps make detectable, never silently wrong. Cross-product concurrency shares nothing at all.
Zarr-version note: implicit groups were a draft-v3 feature, dropped before
finalization; v2 requires explicit .zgroup objects too. Neither models the
digit tree for free — and we don't need either to. The hive tree is
effectively our own implicit-group layer (viable because names are constrained
and LIST is strongly consistent), sitting above completely vanilla zarr v3
leaf stores. No zarr-version coupling in either direction.
3. Layer 1: the static manifest (morton_hive.json)¶
Written asynchronously at init (issue #252 hybrid — the write comes off the synchronous pre-dispatch path: the Lambda leg posts a fire-and-forget setup invoke right after the fail-fast ping, so the manifest typically lands within seconds of init (best-effort: the Event invoke shares worker concurrency and runs retries-0, deferring to the finalize backstop under throttling or a dropped invoke) and readers can consume completed leaves mid-run; finalize keeps an idempotent backstop that self-heals a lost async write; a read-only frozen-key precheck keeps the up-front refusal on reruns — concurrent first writes now collide within seconds of init). O(1); otherwise never touched again during a run. Contents:
spec: convention version string (e.g."morton-hive/1") — the convention itself is versioned from day one (D6).- Dataset identity (short name, product/version).
cell_order,shard_order— each as a declared allowed set/range when the store permits region-dependent orders (D24: shard order is pure packaging; cell order is a resolution axis with per-leaf truth in the morton words + MOC). The allowed sets are the frozen keys; runs within the set pass the append precheck.semantic_hash(D19): sha256 of the canonical semantic core — a frozen key, so reusing a product name with different aggregation semantics refuses up front, exactly as an order mismatch does.- Split schedule (implicit under D2: one digit per level to
shard_order; recorded explicitly for forward compatibility). path_grouping(D21): how many morton digits each path component chunks. Existing stores are retroactively1; new stores default1; changing the default later is a parameter flip for new stores, never a schema break. When compared as a frozen key, an absent field normalizes to1on both sides, so appends to pre-D21 stores never refuse on it.- Rollup/pyramid declaration (D22 extends the original overview-only form):
which ancestor orders carry which derived artifact family — overview
zarrs, stats rollups, sub-shardmap rollups — declared per artifact
family (schedules may differ; display overviews as dense as rendering
warrants, metadata rollups sparser), populated/updated by the §7 sweep.
Tree shape (
path_grouping) and rollup schedules are deliberately decoupled. - Temporal block (D15, ratified on
#237): a store carrying this
block declares
spec: "morton-hive/2"(the version string of D6 covers these temporal/windowed-leaf extensions;/3= the same semantics under D23 window-only leaf naming). It records time encoding/units/epoch/calendar, the membership timestamp field, the window schedule (none|yearly|monthly|daily| explicit range list;quarterlygrammar-reserved), and the append policy. Label grammar and boundary semantics (UTC calendar terms, half-open[start, end), lexicographic = chronological) are frozen on the mortie spec page §6.3 (consolidated there from the original #62 thread record). Generative schedules keep the manifest write-once and static as data accrues: appending a new year to ayearlystore adds leaves the schedule already describes — no manifest touch, and no new manifests: each product tree has exactly onemorton_hive.json(under D19 a multi-product store is a directory of product trees, each with its own bare-named manifest), and each new windowed leaf brings only its own zarr metadata + D4 stamp. The explicit-range-list form is the noted exception: appending a window outside the declared list re-templates the manifest (a rare, single-writer, template-time operation, not a worker-race write) — append-heavy stores should prefer generative schedules. (This exception is an implication recorded here from the ratified schedule set, not a point separately ratified on the #237 thread.) Temporal extent is deliberately not manifest data: actual ranges live on the leaf stamps (truth) and in the root summary (cache), splitting static schema from accruing state exactly the way coverage splits under D9. The default schedule isnone(no temporal partitioning; a re-run replaces the leaf — the honest rename of the draftedmission, which made a completeness claim it couldn't keep for ongoing missions) — existing aggregation stores are unchanged. Ongoing missions with append intent declare a generative schedule (yearlyis the expected production default; t-digest mergeability makes mission-scale statistics a read-time merge over window leaves, the same approximation class as the worker's existing cross-buffer merge).
This file is the reader's bootstrap: with it, every shard path is computable arithmetically with zero requests.
4. Layer 2: coverage MOCs (hierarchical domain declaration)¶
Coverage is declared hierarchically — worker-owned at the leaves,
sweep-composed above — tiered the way cloud-geo formats split bbox /
geometry / data (GeoParquet's bbox covering column vs. the WKB column vs.
the data itself). Tracks #200.
The three tiers, per shard:
- Tier 0 — the morton box (fixed width): the minimal MOC with ≤ 4
members (mixed order allowed) covering the shard's occupied cells.
Existence is guaranteed: within one base cell, any coverage has a deepest
common ancestor whose ≤ 4 intersecting children form a valid cover — and a
shard's coverage is within one base cell by construction (a shard is a
single subtree; its id alone is the trivial 1-member cover, so the box is
what buys sub-shard resolution). Padded to exactly 4 slots for fixed width
(32 B raw; four decimal strings in attrs); the pad sentinel is null
(base-0 words / JSON
null), decided with the mortie-side spec (O8); repetition-padding would also have worked since repeats are idempotent under MOC algebra, but null is the frozen choice. Future store-level covers that cross base cells generalize to ≤ 12 members (the 12 base cells). Readers AOI-reject on the box without parsing anything larger. - Tier 1 — the exact shard bitmap (as O8 resolved it — the originally
drafted "budgeted, coarsen-to-fit MOC" was superseded by the
#202 item (6) measurement:
for linear-track occupancy, coarsen-to-fit ranges at KB budgets deliver
only box-level filtering while a compressed bitmap reaches exact at
~25 KB): the shard's cell-order occupancy as a zstd-compressed bit
field, one bit per subtree cell in ascending packed-word order, stored
as the in-leaf
coverage.mocsidecar. Raw size is deterministic (ceil(4^depth/8)) regardless of fragmentation; no coarsened variant is built (one code path, sidecar-only). The stamp attrs carry only the box + the bitmap's order + pointer + byte sizes (one extra GET, paid only by readers that pass the box test). The sidecar is the noted exception to §2's "vanilla zarr v3" leaf: one foreign key inside the leaf, ignored by zarr readers (data reads unaffected; member enumeration warns and skips it).
D14 amendment (ratified on
#237): the stamp envelope's
encoding discriminator gains a third value, "full", meaning
"coverage = the entire shard subtree" — no sidecar is written. Decided by
one popcount at stamp time. This is the fast path for dense-by-construction
workloads (pull-NN raster, #218/#237): interior shards skip the sidecar
and its GET entirely, while edge-of-scene/swath shards write the real
bitmap — the spatial union across the leaf window's acquisitions
(per-timestep validity stays in the data plane as nodata, D9 applied one
level down). One code path with a cheap branch, not a raster special
case; the tier-0 box is carried unconditionally either way.
- Tier 2 — exact: the morton coordinate array in the leaf is the
exact cell list. The MOC tiers are indexes, never truth (D9 discipline,
applied one level down).
Ownership and lifecycle:
- Leaf MOCs are worker-owned and ride the commit stamp (D4): the payload
lands on (or is finalized before) the shard's final root
attrs.update()PUT. Zero extra requests in the attrs case, and debris semantics are inherited automatically — a torn worker's MOC never becomes visible. - Ancestor and root MOCs are composed by the §7 sweep: union of children, re-coarsened to the same budget, in the same bottom-up level-by-level orchestration as overview zarrs. All regenerable caches.
- Optional end-of-run root
coverage.moc(O9): flag-gated. The dispatcher cannot write S3 (only the worker execution role can), so it posts its completion list to a fire-and-forget worker invoke (InvocationType="Event", ~10 ms of dispatcher wall clock, run-size independent) that writes a shard-order root MOC for the one-GET bootstrap. Under D15 this root summary also carries the time-range union alongside the MOC (cache, sweep-regenerable). Failure is harmless: readers degrade to the sweep MOC or the walk, never to wrong answers. Incremental runs: the leaves carry durable truth, so the sweep is always a correct rebuilder, and the end-of-run write may union with a prior root object. - Everything above the leaf is a cache, not truth (D9). Timestamped, regenerable (from leaf stamps or a tree walk). The strongly-consistent LIST walk (§2) remains ground truth; a run that crashes before any root MOC exists degrades to walking, never to wrong answers.
- This replaces consolidated metadata for extent/discovery. Consolidation measured +70 s per worker and is disabled by default; the coverage tiers cost effectively nothing and answer the actual question readers ask ("where is there data?") in one GET.
O8 and O9 are resolved (espg-ratified on the #200 thread, implemented on PR #208): the shard tier is an exact cell-order occupancy bitmap, zstd-compressed, as an in-leaf sidecar object — attrs carry only the tier-0 box + order + pointer + sizes, with a null pad sentinel — and the end-of-run root MOC defaults on for hive stores. The root object serializes per O1 as JSON ranges with decimal-string endpoints.
5. Layer 3: reader architecture¶
The hot path is arithmetic; LIST-walking is the fallback. This is substantial new wiring on the read side.
Single-dataset flow:
- GET
morton_hive.json(once per store, cacheable). - GET
coverage.moc; intersect with the query AOI's MOC (moc_and) → the populated shard set within the region. Zero LISTs. - For each shard id: compute the hive path by string arithmetic (digits → components), open the leaf zarr, check the commit stamp. The stamp's morton box / shard MOC (§4) lets the reader AOI-reject the leaf before touching chunk data.
- Fallback / discovery walk (no MOC, mixed orders, or verification):
from any node, delimiter-LIST; recurse on
[1-4]/children; a*.zarrentry is data at that node; no digit children ⇒ nothing finer. At each zarr encountered, theroleattribute (§7) says whether it's a summary you may stop at (display) or source you must not conflate (analysis).
Cross-resolution join (the ICESat-2 × GEDI case):
- Read the target dataset's manifest → its
cell_order/shard_order. - Truncate the fine observation's morton decimal string to the target's
cell order (equivalently
rust_mi_coarsenon the packed word) → the containing target cell id. Truncate further to the target's shard order → its shard's hive path. Zero requests to locate; one leaf open to read. - The nesting predicate is literally
fine_id.startswith(coarse_id). - Within the leaf zarr, the
mortoncoordinate locates the observation(s) in the containing cell.
(Under D24 a target's orders may be regionally heterogeneous — the declared sets bound them; the leaf's actual orders come from its morton words / MOC, and the truncation predicate is unchanged. Mixed-order reader support is gated on mortie#116 / moczarr#8.)
Per-observation LISTs are forbidden in the join loop: at 2,000 workers × millions of photons the walk is the robustness path, not the join path (D10).
6. Layer 4: the xarray/xdggs extension ("sparse DGGS")¶
This layer ships as zagg's own standalone extension — moczarr (O3
resolved: standalone-first, offered upstream; xdggs adoption of
MortonIndexDtype (#72) is a
follow-on, not a gate). Broad scope:
- Domain = MOC. A dataset declares its coverage as a MOC instead of materializing dense full-sphere coordinates. The accessor uses it to truncate the notional global grid to where data exists — top-level MOC coverage/polygons drive what the extension exposes.
- Coordinates stay sparse, and morton is the only stored coordinate
(D16). The
mortoncoordinate (packed u64 words,MortonIndexDtypein memory — upstream ask tracked in #72) is the sole cell labeling on disk. NESTEDcell_idsremains the interop encoding but is fabricated exactly, on demand (mort2healpixis vectorized arithmetic): moczarr fabricates it Python-side, the gridlook-jupyter hub proxy at serve time, a TS decode for browser-direct stores. Thecell_ids_encodingknob (#135) retires with the D16 writer flip. - Dense views are fabricated lazily, per-region, on demand — never stored.
- Multi-store alignment: opening several stores (datasets) over one AOI yields aligned sparse views whose join semantics are the §5 truncation rules. What this looks like as an xarray API (alignment? a join accessor?) is the biggest open design question (O4).
7. The pyramid / post-process sweep¶
Everything derived or stale-prone lives in a second pass, never at write time (D11) — overviews aggregate across worker-shard boundaries, so they can't be produced by shard workers anyway.
The sweep owns four derived artifact families (D22; the original scope was the first two):
- Overview zarrs at ancestor nodes, explicitly marked
(
role: overview+ source order + aggregation method in attrs). Never inferred from position: a shallow zarr may equally be coarse source in a sparse region. Full pyramid cost is a geometric ~⅓ extra storage (4 children per order). - MOC (re)generation — compose ancestor MOCs bottom-up from the leaf
stamps (union, re-coarsen to budget) and refresh the root
coverage.moc. - Sub-shardmap rollups (D22): each leaf prefix carries its shard's
sub-map as full ShardMap JSON; the sweep folds them up-tree via the
coarsen path (
ShardMap.reproject— exact pure regroup, granule union deduped by id; #294). - Stats/cost rollups (D22): the per-shard stats sidecars (D20) fold up-tree — the schema is associative by construction, so the rollup is the same fold shape as the pyramid.
- Optional interop materialization: if a use case ever demands a
zarr.open(store_root)-able hierarchy or a one-GET consolidated index, the sweep generates it as a derived artifact here. Round one ships without it (D12).
Every rollup is stamped with generation info (merged-leaf count + max leaf timestamp): after a leaf re-run, ancestors are detectably stale — staleness is detected, not prevented, and rollups are regenerated opportunistically (D9 semantics). The core test obligation is rollup == direct: the folded artifact at level N must equal direct computation at level N, for every family. Triggering is an end-of-run dispatcher hook plus a manual CLI; the sweep discovers work from the run record, not by listing.
The sweep is idempotent and can fail or lag without corrupting anything: the write path (§2) is load-bearing; this phase is optimization — deleting every rollup leaves all leaf reads intact.
Partitioning the sweep (issue #377)¶
One sweep invoke walks the whole store, so its wall clock and resident memory
both scale with the store — at state scale (the California o9 case in
#376: 2,721 leaves, entirely
within one HEALPix base cell) a single 900 s worker cannot finish. The pass
therefore decomposes into 2^n morton-subtree partitions, each swept by
an isolated worker (zagg.sweep_partition).
A D1 id is {sign+base}{digit}* with one base-4 digit — 2 bits — per order,
so a 2^n split lands on a digit boundary exactly when n is even: n = 2k
splits at order k, and partition i owns every id whose first k digits
rank i — one order-k subtree per base cell, twelve in all
(12 · 4^k / 2^n = 12). partitions is consequently a power of four; an
odd n would halve a digit and is rejected rather than rounded (the bit-level
alternative is an open fork on #377).
Ownership being a plain prefix test is what makes the workers need no
coordination: every node at order ≥ k and every (node, window) artifact
beneath it lies wholly inside one partition, so no two partitions can write
the same object, and the generation stamps above keep each partition
independently idempotent and resumable. Three things do span partitions and are
excluded from a partitioned pass by construction — nodes above the split order,
the store-root coverage.moc refresh, and the manifest pyramid.materialized
read-modify-write. They belong to a coarse-level finisher: one invoke after
the partitions land, folding the coarse levels from the partitions'
already-materialized overview slabs, which is bounded work only under #376's
cascade (hence the sequencing).
The root coverage.moc is on that list twice: it is a partitioned pass's
input as well as a deferred output. Overview discovery unions the dirty set
with it (LIST-free, above), so a partition reads whatever the last whole-tree
sweep or mode="coverage" leg left — never a refresh from its own pass. A
partition handed a run-scoped leaf set against a missing root MOC folds from
that run's leaves alone, and since a generation mismatch overwrites, it can
rewrite a complete node overview with fewer contributing leaves until the next
whole-tree pass repairs it (D9: regenerable). The degraded read is surfaced as
root_moc_stale in the partition's record, and ordering the refresh ahead of
the fan-out is part of the finisher's obligation.
Transport is unchanged (D8 — every store write stays worker-side): the client
fires one fire-and-forget mode="sweep" Event invoke per non-empty partition,
each carrying its disjoint slice of the work set plus a partition: {index, of}
block. An oversized slice degrades to discover: true as before, and the block
still rides, so worker-side discovery stays narrowed to that partition. On
windowed stores the window axis is a second, free parallelism dimension
(per-(partition, window) invokes); noted, not used yet.
Shipping status. The decomposition, the partition-scoped pass
(run_sweep(partition=…)), the fan-out event builder and the in-process
--partitions backstop are in. Two links of the Lambda leg are not yet
connected, and until both are, the fan-out must not be fired against a
deployed worker: nothing in a production dispatch path passes partitions
yet, and the worker's mode="sweep" handler does not forward the event's
partition block to run_sweep — so a partition-carrying event would today
be swept as a whole-tree pass over that partition's slice, walking up to the
base node and racing its siblings on exactly the shared coarse rollups this
section says cannot be shared. The single-process --partitions path has
neither gap and is safe now.
8. Decisions registry¶
8.1 Decisions made (rationale recorded)¶
- D1 — Ids are morton decimal strings; packed u64 is canonical storage.
Settled on mortie #48 (Option A): packed
uint64kernel as the compute substrate, the signed decimal string as the render-only repr and the external/path form. Type-stable: strings always, at every order; never data-dependent int-vs-string emits. - D2 — One digit per path component. Mixed shard orders are real (across
datasets and within a store), so every order must be a node boundary.
Grouped-digit and two-level schedules are dead ends. Note the schedule is
logical only: S3 partitioning is delimiter-blind, so slashes buy zero
throughput (#197 is the
throughput fix). Amended by D21: digit-chunking becomes the declared
path_groupingmanifest parameter (default1= this rule); the "grouped-digit schedules are dead ends" verdict is thereby softened to "grouping is a parameter, never a schema fork." - D3 — Full morton id at the leaf (
{full_id}.zarr), self-describing. (Unchanged by D19: product identity lives at the product root, above the tree. Superseded by D23 formorton-hive/3stores: the basename becomes the time window; the full id stays recoverable from the path and from the stamp/sidecarshard_key.) - D4 — Commit stamp via final root-attrs update. Absence (LIST) is trustworthy; presence requires the stamp. Torn shards are debris, overwritable on retry. One small PUT; not consolidation.
- D5 — Zero metadata above the leaf on the write path. No zarr groups at digit nodes, no shared mutable state during fan-out.
- D6 — The convention is versioned (
morton-hive/1) in the manifest. - D7 — One store per dataset, own orders. Workers never mix datasets; interop between any pair of stores is truncation against each manifest.
- D8 — Coverage MOCs are hierarchical and worker-owned at the leaves (amended per #200). Each shard's tier-0 morton box + tier-1 budgeted MOC rides the D4 commit stamp; ancestor/root MOCs are sweep-composed unions; an optional end-of-run root MOC is written by a fire-and-forget worker invoke from the dispatcher's completion list (the orchestrator has no S3 write access). Replaces consolidated metadata (disabled; measured +70 s/worker).
- D9 — MOC is a regenerable cache; the tree walk is ground truth.
- D10 — Arithmetic-first reads; no LISTs in join loops. Strengthened with D22 (espg-ratified in-session, recorded on #300): MOC-first is the reader contract — full recursive enumeration is out-of-contract for readers, reserved for prefix-sharded audit tooling. The §5 discovery walk remains the robustness/verification path, never the read path.
- D11 — Pyramids/overviews are a second-pass sweep,
role: overviewattrs, never inferred from tree position. Scope extended by D22 (four artifact families, generation stamps, per-family manifest schedules). - D12 — Plain manifest, not a zarr-native hierarchy, in round one.
Hierarchy metadata at nodes reintroduces the metadata-op storm
(#189,
#194) and couples the
layout to still-settling zarr v3 hierarchy semantics. One-way door avoided:
a root
zarr.jsoncan be added later by the §7 sweep without breaking anything. - D13 — Appendable time series = time-windowed, write-once leaves
(espg-ratified on #237).
One leaf per (shard, window) under a manifest-declared window schedule;
every leaf keeps full D4/D5 semantics (write-once, stamped, binary
debris, zero shared mutable state). Backfill is a new earlier-window
leaf; re-running a window is idempotent replacement. Rejected: a
high-water time index (can't extend backward) and per-run stamp entries
with array
resize(mutable shared state at the leaf — attrs RMW + metadata rewrite races, non-binary debris, unordered time axis on backfill). Not raster-specific:none(default; no partitioning, bare leaf names) reproduces today's aggregation stores byte-identically; per-year 88S runs, seasonal subsets, and append-as-acquired are all window leaves under the same convention. Leaf naming is frozen:{full_id}_{window}.zarr, underscore separator, split on the first_— grammar and boundary semantics normative on the mortie spec page §6.3 (originally recorded on the #62 thread) as part ofmorton-hive/2. (Unchanged by D19 — leaf naming survives the product-root design intact. Naming revised by D23 formorton-hive/3: the basename becomes the window alone; the/2grammar stays frozen for/2stores.) Reserved (lean, not decided): §7 overview/pyramid zarrs inherit window naming (per-window overviews, with an optional all-time overview as a derived artifact). - D14 — Coverage
encoding: "full"fast path (espg-ratified on #237). The stamp's coverage envelope discriminator becomes"ranges" | "bitmap" | "full";"full"= whole-subtree coverage, no sidecar written, chosen by a popcount at stamp time. Tier-0 box is always carried. Edge-of-scene/swath shards (the partial case) write the real bitmap as the spatial union across the leaf window's acquisitions; per-timestep validity remains data-plane nodata. - D15 — Temporal declaration splits like coverage
(espg-ratified on #237).
Manifest (
morton-hive/2, static): time encoding + window schedule + append policy. Leaf stamps (truth): each windowed leaf's actual time range + acquisition/granule count. Root summary (cache): the end-of-run root coverage object gains the time-range union alongside the MOC; sweep-regenerable, never truth. Extent never lives in the manifest, so the manifest stays write-once (§3; written async at init with a finalize backstop since issue #252); the noted exception is the explicit-range-list schedule, where appending outside the list re-templates the manifest — append-heavy stores should prefer generative schedules. (That exception is an implication recorded from the ratified schedule set, not separately ratified on the thread.) -
D16 — Morton-only storage: NESTED is fabricated, never stored (espg-ratified in-session 2026-07-17, filed on #262; both verification checks — artools clean, gridlook requirement satisfiable by fabrication — confirmed on the thread). zagg stops writing the
cell_ids(NESTED uint64) array to leaves;mortonis the only stored cell coordinate and the declared convention coordinate. Rationale: (1) morton words carry order intrinsically, so mixed-order arrays (D2 coarse shards, §7 pyramids) are first-class where NESTED u64 needs side-metadata or zuniq/nuniq; (2) kills the dual-encoding cost — one u64 array per leaf (8 B/cell, plus one array's objects per leaf in the #236/#240 object-count currency) and the NESTED↔morton double-encode maintenance (#72); (3) the reader stack is ready — the moczarr fabrication layer was the named gate and is merged (espg/moczarr PR #7), withopen_hive+ xdggsgrid_name: "morton"+ MOC-backed lazy index. Third instance of the derived-views principle (§6 dense views, D9 caches). Ratified refinements: the dggs attrs use a distinct gridname: "morton"— nevername: "healpix"+indexing_scheme: "morton", which scheme-blind readers silently misread as NESTED (garbage renders); a distinct name makes them hard-reject with a diagnostic, and matches moczarr's xdggs registration. Unknown-resolution point encodings at order 29 are clipped on the fly to order 24 for Number-safe browser paths (NESTED ids are float64-exact only through order 24; genuinely-finer-than-24 data takes other measures — hub-side fabrication, aggregation). The writer flip sequences as 0.x phases (emit knob default-on → default flip → removal). (An earlier plan bundled the flip with a #299 leaf-basename rename; D19's product-root revision made #299 additive, so this writer flip is the one remaining breaking store change in this family.) The order-29 discriminator metadata is O10 (final ruling: kind is encoding-carried, no metadata field; the clip is a viewer-side transient cast — see the O10 entry). -
D17 — Hive+sharded is the HEALPix default; flat/fullsphere deprecated; dense leaf arrays write through the ShardingCodec. (Phase-1 ratification recorded on #251; landed via PR #257 (default flip + dense-layout removal), #233 (
sharded: truedefault), #241 (hive dense parity); leaf object model ratified in-session, recorded on #236.) Every HEALPix config — point aggregation and raster — targets hive; the grid keys the default layout (hive for HEALPix; PR #257 / #253), and an explicitstore_layout: flatsurvives only as a deprecated escape hatch (emits aDeprecationWarning) until the O3-gated flat removal; rect keeps its bounded flat store. Hive leaf dense arrays are one object per array per leaf via the ShardingCodec (always-accumulate);sharded: falseremains a legitimate opt-out, and raster leaves skip the sharding codec by design — the only raster/aggregation difference at the leaf (espg on PR #257).grid.shard_order(sub-shard object granularity) is removed (#238); the manifest'sshard_order(tree/dispatch order) is unrelated and stays. Full flat-machinery removal remains gated on the O3 reader (the #251 phase-2 gate). - D18 — Ragged output = sharded vlen-bytes (measurement + recommendation
on
#209,
with espg's in-session ratification of the vlen-bytes adoption recorded in
the #210 body; tracking
#210). Ragged
t-digest arrays store as
bytesdtype + vlen-bytes codec under the ShardingCodec — one object per shard, 2-GET single-cell reads — with element interpretation as an attrs convention and a golden-bytes framing pin byte-compatible with numcodecsVLenArray. The CSR-per-inner-chunk fanout is deleted. A typedvlen-array<float32>ZDType is deferred behind three gates (upstream zarr-extensions convergence, at-scale proof, a second consumer); the framing pin guarantees any future migration is metadata-only. This is the second noted soft-exception to "vanilla zarr v3" leaves (after the §4 coverage sidecar): plain zarr readers see opaque bytes, not garbage. - D19 — Named product roots + semantic-core hash: a multi-product store
is a directory of stores; the name is the address, the hash is the
integrity check (concept settled across
#296
and #299 — espg
on-thread:
registry + per-product manifests;
revision history, all espg-ratified in-session 2026-07-20, trade
studies on the PR #306 thread: leaf-hash basenames (phase 2) →
hash-named product roots (phase 3) → named roots with the hash demoted
to metadata (this revision)). Each product lives under its own
human-readable root prefix
{name}/; a product subtree is a complete, unmodified morton-hive store — bare-named manifest and MOC,{full_id}.zarrleaves under/1–/2grammars ({window}.zarrunder/3— D23) — so existing single-product stores are already valid and the change is additive, not breaking. Readers distinguish the two root forms by content (manifest ⇒ bare store; name-shaped prefixes ⇒ product directory; names must not match the base-component grammar); viewers enumerate products by listing the root and reading each{name}/morton_hive.json— no name↔hash translation layer. Identity is split: the name addresses the product; thesemantic_hashverifies it — sha256 over the canonicalized output-defining subset only: theaggregationblock (functions + params + dtypes + fills + ragged kinds), thedata_sourcesemantics (dataset/product, groups, coordinates, variables, filters), the grid type + indexing scheme, and the pipeline type (spatial|temporal|event, absent ⇒spatial— espg-ruled on the PR #316 review, 2026-07-21: a temporal engine over the same aggregation block is a different product; the original list omitted it only because the temporal path wasn't in frame, and pre-merge was the free moment to add it — no store carries a hash yet). Excluded as packaging: cell order (a resolution axis — D24), parent/shard order,chunk_inner/sharded(shardedamended below — the D19 hash epoch), worker size, the wholeaggregation.streamingblock (mode AND its sizing knobs,buffer_granules/block_bytes— merge-vs-spill landsnp.iscloseand shares one store, with the actual mode recorded per-run; see the law-equivalence contract below), and read knobs — hashing the whole template would have made o8 and o9 runs different products and blocked mixed-order processing. Amended by the D19 hash epoch (espg-ruled on-thread 2026-08-07 as PR #397 questions (7)(b)/(8)©, recorded in #384 plan delta 2; implemented in #415): two corrections to the lists above, landed together because each one moves every existing digest. (a) Worker size was already named as packaging but its keys were never excluded, so the per-cellmin(K, n_granules)clamp made a small shard's worker-side hash differ from the run's —shard_workers/granule_workersnow exclude properly. (b) The leaf-shapingoutputknobs —aoi_mask,windowing,grid.sharded— are now included: they change what a leaf contains, so leaving them out made the D20 skip gate (#388) read a config-changed rerun asequal.shardedtherefore moves from the exclusion list above into the core; the orders it sits beside do not (D24 is untouched, and the object-layout hole is narrowed rather than closed —chunk_innerstill moves K without moving the digest, espg-ruled 2026-08-17 as the bounded state to ship). Thepyramidblock stays excluded: D11 keeps it out of the frozen manifest keys, and the leaf column it declares is verified by reading the artifact (§4.6) rather than by the digest — andgrid.emit_cell_idsstays excluded on the same argument (espg-ruled 2026-08-17): the #304 hatch is scheduled for removal, so a hashed hatch would strand every store built with it ON behind a digest no legal config can reproduce, in exchange for an array inventory that reading the leaf already verifies. © A third exclusion rode the same epoch (espg-ruled 2026-08-17, issue #449):credentials_provideris packaging — it selects how source bytes are fetched (the same class as the read knobs and asanonymous, already excluded), never what is computed, and a wrong credential fails the fetch loudly rather than silently. Ruled at the epoch so the first GEDI store's identity is born without an auth knob, and so a credential migration over unchanged data (the MERRA-2/gesdisc path) never rehashes. (d) The same ruling extended to the byte-movement knobsread_workers,write_bufferandsource_region(espg-ruled 2026-08-17 on the epoch PR): each selects how bytes are fetched or moved, never what is computed; each fails loudly rather than silently; andsource_regionsat in the same dict literal as the already-excludedanonymous, so the D19 line ran through the middle of one decision. The live demonstration is dated and covers the fan-out widths only: two GEDI flux builds of one shard on 2026-08-17 produced identicaltotal_obsandcells_with_datain the exact single-block spill regime and still hashed apart onread_workersand the two*_workersspellings.write_bufferandsource_regionare raster-only and a flux build is the point path, so those two rest on the arguments rather than on the measurement. (e) The epoch also canonicalized the OTHER half of the identity pair, the catalog hash below (espg-ruled 2026-08-17: "we want the granule to trigger the hash, not how that granule is fetched") — see that clause. Operator consequences — every pre-epoch hash invalidated, and the three migration paths — are indocs/hive_layout.md, "Migration: the D19 hash epoch".
Amended again — the streaming exclusion and its law-equivalence
contract (the PR #475 D19 question, resolved 2026-08-17 as option (b);
the decision record is on that PR's thread; issue #474): the whole
aggregation.streaming block — mode, buffer_granules,
block_bytes — is excluded from the core as packaging
(AGGREGATION_PACKAGING_KEYS), landing the code on the side this
record already stated twice while semantic_core still hashed the
block as spelled. The exclusion holds on condition of the
law-equivalence contract: every streaming mode MUST produce output
within the documented approximation law of the pooled path. Three law
classes are admitted, and a mode that cannot maintain the class its
channels fall in may not join the streaming block: (1) exact — the
single-block regime and the summation reducers; (2) the
kway/np.isclose class — the payload channel across merge flushes
and block closes; (3) the channel-specific documented bounds for
the companion channels, conservative rather than close — a located
companion word folds toward the contributors' common ancestor, so its
pin is an ancestor-or-equal hull bound (§9.1), and the packed
composition word keeps presence (lane > 0) exactly with counts within
one lane quantization (tol = 1 + n / 255.0) of pooled.
The contract, not the digest, is what makes one shared identity across
all streaming regimes honest; its enforcement is the single-block
exactness pins (tests/test_spill.py::TestSpillWorkerSingleBlock), the
cross-block suites (tests/test_spill_crossblock.py, being extended to
the temporal channel by issue #477), and tests/test_streaming.py's
pooled-parity pins.
Store compatibility was priced at decision time: no long-lived store
carries a streaming-declared digest (the deployed stores age out on a
30-day cycle), so the epoch moves nothing that outlives it — and the
fat-shard rescue that motivated the knob (issue #474) stays deployable
on an existing store, where the as-spelled hash would have tripped the
frozen-key refusal.
The hash is a
frozen manifest key (reusing a name with different aggregation
semantics refuses up front, like any frozen-key mismatch) and is
recorded in leaf attrs and D20 sidecars. The literal template is
deliberately not the product-level record (it carries run-varying
packaging): the product root holds the canonical semantic core as
aggregation.yaml (deterministic, valid YAML); each run archives its
literal template with its run record, and sidecars carry the run id to
join back. (This factoring formalizes a seam the code already has:
spatial_signature vs output_field_signature, the #89 split.)
Rationale for product-root prefixes (unchanged from the hash-first
study): S3's only cheap scoping primitive is the prefix, and
product-scoped operations (delete, lifecycle/expiration, access policy,
inventory) dominate — each is one prefix rule. Cost prediction for
appending new shards is likewise scoped (espg-noted in-session): the
product root holds its own telemetry history, so the pilot-first
estimator's priors (the #298 design) are exactly the product's own
records. What a shared spatial tree would have offered — one prefix
spanning all products — serves no planned workload (§5 reads are
arithmetic). Catalog identity lives in the sidecar, never the name:
granule count + sha256 of sorted granule ids + zagg version;
dedup/has_run consults the computed path, the semantic_hash, and
the sidecar catalog identity (a catalog-grown shard is "stale", not
"hit"). Amended at the D19 hash epoch (espg-ruled 2026-08-17, PR #420
question (1)(b)): the ids are canonicalized to the driver-stripped bare
granule id before hashing, and the recorded id list beside the hash is
written in that same space. One granule reaches these seams as an s3://
href, an https:// href or a bare catalog id depending only on
data_source.driver — which the semantic core has always treated as
packaging — so hashing the href form made a driver switch look like a
catalog change and rewrote whole stores over a fetch-mechanism edit. The
canonical form is the basename because it is the only component the s3 and
https spellings of one granule agree on; the accepted cost is that two
granules whose hrefs differ only in prefix collapse to one identity, which
every catalog zagg reads rules out by naming granules globally uniquely
(that is why the catalog's own id equals the basename). Where a collapse
does happen the cost is not only a collision: the contraction guard's
set diff loses per-granule resolution inside the collapsed group, so a
dropped member reads as id-multiset-drift and rewrites rather than
refusing — logged loudly per leaf, with the predicate question left
standing (PR #420 review finding (2)). Immutable-provenance naming (product root
{name}+{catalog-hash}/) stays an opt-in for frozen-catalog archival
runs. The output content hash that makes dedup verifiable is O11
(resolved — adopted; it complements the semantic hash — "intended identical" vs
"actually byte-identical"). A community registry maps names → semantic
cores + hashes; cross-deployment name collisions are disambiguated by
the hash in metadata rather than prevented by unreadable paths. D7
generalizes cleanly: "one store per dataset" becomes "one product tree
per semantic core" under a shared root, with cross-product joins
unchanged. Three refinements (espg-confirmed in-session, 2026-07-20):
the manifest semantic_hash is truth — aggregation.yaml is a
derived convenience (D9 cache class, sweep-regenerable), so divergence
resolves to the hash; the hash is stored full-length with git-style
ergonomics (the full digest is what is compared; a short fingerprint —
12 hex — is the display/CLI shorthand); and product names are
URL-safe by requirement (they appear in gridlook deep-links and web
paths — charset lowercase alphanumerics plus -/_, with the
base-component exclusion; the grammar is normative on the
mortie spec page §6.5).
- D20 — Per-shard telemetry sidecar + run records (espg on-thread:
sibling placement + envelope ride,
caller identity;
schema decisions recorded on
#296).
Each successful shard writes a versioned stats record as a sibling
object next to the leaf (not inside the .zarr/): timings
(read/index/aggregate/write/spill), counts, memory, cost (GB-s ×
price), catalog identity (D19), zagg version, and invoked_by (caller
identity resolved once per run by the dispatcher via STS and stamped
through the invoke payload — workers cannot see the caller). The schema
is mergeable by construction — only associative stats (counts, sums,
min/max, t-digests; never stored means) — so up-tree rollups are a pure
fold (D22). The record also rides the async result envelope; the
dispatcher writes a run-level parquet at the product root — a run maps
to one product ⇒ one product tree (D7/D19), so it lands under {name}/,
never the multi-product store root, whose node invariant admits only
{name}/ product roots — with one row per
shard, including failure rows sourced from the run report — sidecars
exist only on success; CloudWatch structured logs remain the failure
forensics channel). Run-record names are timestamp-first
(stats_{timestamp}_{run_id}.parquet) so lexicographic listing is
chronological and time-range queries prune on keys before reading;
per-user scoping stays the invoked_by column (names stay stable
identifiers). Sidecars carry the run_id (joining leaf → run record →
that run's archived literal template, D19) and the D19 semantic_hash;
they omit account-identifying fields (request ids, ARNs beyond the
caller identity).
- D21 — path_grouping is a manifest parameter, not a layout
(espg-ratified in-session, recorded on
#300).
The manifest declares how many morton digits each path component chunks;
existing stores are retroactively 1; new stores default 1; readers
chunk the digit string per the manifest. Rationale: hive paths are
computed (manifest + MOC → arithmetic paths → parallel GETs — zero
LISTs on every hot path), so grouping is walk ergonomics, not
performance; hard-adopting a grouped layout would have forked the path
dialect permanently for ~$1.60 and ~1.6 s per rare full walk. A future
default flip (e.g. to 3-order groups) is a parameter change, never a
schema break.
- D22 — One unified second-pass sweep owns all derived artifacts
(espg on-thread:
trigger + sub-map format + schedule discussion;
schedule decoupling ratified in-session, recorded on
#300).
Four families — overview zarrs, MOC regen, sub-shardmap rollups (full
ShardMap JSON at leaf prefixes, folded via the exact coarsen regroup,
#294), stats rollups
(D20 fold) — in one idempotent pass with per-family, manifest-declared
order schedules (decoupled from path_grouping). Generation stamps
(merged-leaf count + max leaf timestamp) make staleness detectable, not
prevented; rollup == direct is the standing test obligation per
family; trigger is end-of-run hook + manual CLI; the sweep discovers work
from the run record, never by listing. Nothing is load-bearing: deleting
every rollup leaves leaf reads intact (D9 semantics). (The refine
direction of ShardMap.reproject and its no-region semantics are still
under review on PR #295 — the sweep depends only on the exact coarsen
direction.) An optional fifth, audit-class family (espg-directed
in-session, 2026-07-20): debris collection — unstamped .zarr/
prefixes older than a declared horizon are deleted, prefix-sharded,
never load-bearing (D4 already makes debris ignorable; this stops it
accumulating as paid storage at fleet scale).
- D23 — Leaf basename = time window ({window}.zarr), morton-hive/3
(espg-proposed and ratified in-session, 2026-07-20; recorded here and on
the PR #306 thread). Completes the axis separation D19 began: product =
root prefix (D19), space = digit path (D1/D2), time = basename —
each identity axis in exactly one place, none encoded twice. This
removes the last path/basename redundancy (the #296 observation that
started the naming work, applied to its final instance): the morton id
currently appears in both the path and the leaf name. Listing a shard
node returns the temporal inventory directly (2019.zarr,
2020.zarr, …) — what append planning and time-series discovery
actually ask. Spec bump to morton-hive/3 under D6 versioning: /1
and /2 stores remain valid forever under their frozen grammars;
readers discriminate by the manifest spec string; the /3 grammar is
frozen on the
mortie spec page §6.4
(mortie#62, drafted via espg/mortie PR #118). Costs accepted
with the decision: the name==path self-check moves to the stamp attrs /
D20 sidecar shard_key (an fsck pays one GET per leaf); D3's
"unambiguous if moved" softens to "recoverable from attrs" — the moved
case reduces to a downloader that materializes the prefix tree into
folder names (espg-noted in-session). Both leans ratified
(espg, in-session 2026-07-20): the schedule: none reserved token is
all (all.zarr; reads as all-time; cannot collide with the
digit-shaped window grammar; excluded from the window grammar forever
on the mortie spec page), and the sidecar aligns to
{window}.stats.json / all.stats.json for /3 stores (the
spec-keyed seam shipped in PR #307 makes both a constant flip; /1–/2
stores keep the frozen stats_{window}.json / stats.json names).
Rejected alternative: time
as a path level ({name}/{window}/{morton…}) would make
window-scoped ops prefix-cheap but duplicates the morton tree per
window and shatters the dominant read — a time series at a location —
across W prefixes; reads dominate window expiry, so time stays at the
leaf.
- D24 — Resolution polymorphism: cell order is a query/packaging axis,
not product identity (espg-proposed and ratified in-session,
2026-07-20; rationale recorded on the PR #306 thread). Aggregation
composes across orders — finer cells fold to coarser under the same
merge law (exactly for count/sum/min/max; np.isclose for t-digest,
whose merge is order-dependent — the same epistemic class as
merge-vs-spill, already ruled one-store) — so a product's cell order is
excluded from the D19 semantic_hash, and one product tree may carry
regionally heterogeneous resolution (e.g. o19 cells in polar
shards, o17 mid-latitude). The design had already committed to the
pillars: D22's rollup==direct obligation is the composition claim;
D16 chose morton words because they carry order intrinsically; D11's
role attr anticipated coarse source; the #217 mergeable-reducer
machinery provides per-aggregator merge laws. Consequences frozen with
the decision: (1) composability class — exact | approximate |
none — is declared in the semantic core, derived from the product's
aggregator set; a none product pins its cell order (resolution is
identity there) and refuses mixed-order appends. The implementation
ships helper functions to derive the class from the product's
aggregator set (the existing mergeable-reducer merge-law flags) and to
enforce it — at template-validation time (the class is recorded into
the semantic core) and at append time (the mixed-order guard consults
it) (espg-directed in-session). (2) Per (shard,
window) there is one resolution at a time: heterogeneity is
regional, across shards. A same-cell-order rerun is D4 idempotent
replacement; writing a different cell order into an occupied leaf
refuses with a useful error (espg-directed) — intentional
re-resolution means rerunning at a parent order that isn't occupied, or
explicitly clearing the leaf. (3) Manifest cell_order and
shard_order become declared allowed sets/ranges (§3); per-leaf truth
is the morton words + MOC. (4) Coarse source and sweep-built overviews
unify — the same multi-order tree, distinguished only by role/
provenance attrs; the pyramid is the store's resolution axis, partially
materialized. Reader support is gated on mortie#116 mixed-order morton
(tracked as moczarr#8); the schema needs no bump — this is what the
coordinate system was built for. Sub-point ruled (espg-ratified
in-session, 2026-07-20, on the #201 thread): none (non-composable)
fields are allowed, with a loud warning at template validation,
and at overview orders the default is per-field exclusion —
composable fields roll up, none fields exist only at native
resolution (min-semantics product pinning rejected; "resolution of a
product" is per-field, which readers handle via the manifest's
per-family declarations). An explicitly declared derived summary
is available as the opt-in: e.g. an auto-digest of a roster field's
raw values under a different field name at overview orders, so
overview schema never silently differs from source; the declaration
lives in the pyramid block, never the semantic core (leaf truth is
unchanged — two products differing only in overview-summary
declarations are the same product), and derived fields carry their own
composability class and O11 hashes. Deterministic subsampling
(seeded reservoir per cell, seed = f(cell, window, field) for O11
reproducibility) is deferred until a concrete consumer (the #265
HHDC class) asks; concatenation ("roster pyramids") is rejected —
volume never shrinks, every level duplicates the raw data beneath it.
8.2 Open for review (input needed)¶
- O1 — MOC serialization format for
coverage.moc: JSON of nested-range pairs? Packed-word.npy? Needs to be frozen alongside the mortie spec (FITS/IVOA interop is an explicit non-goal per mortie #50). - O2 — MOC depth ceiling: resolved (mortie 0.9.0) (entry kept here
for O# id stability). The cap was a stale
MAX_DEPTH = 18constant, not a u64 limit; the coverage/MOC paths now reach the packed-u64 kernel ceiling (order 29), so cell-order MOCs at order 19 are representable today. - O3 — Upstream target: RESOLVED — standalone-first (moczarr)
(espg-ratified, recorded on
#251):
the sparse-DGGS reader ships as zagg's own standalone xarray extension
(espg/moczarr, published from our side and offered upstream); xdggs
adoption of MortonIndexDtype (#72) is a follow-on, not a 1.0 gate.
Near-drop-in
xr.open_zarr()-grade ergonomics is a hard acceptance criterion (it gates the #251 flat removal). - O4 — Multi-dataset join API: what does the cross-resolution join look like in xarray terms — alignment, an accessor method, a lazy index?
- O5 — Sentinel-2 encoding: native ~10 m order, one order finer (6 m, collision-free), or much finer + nearest-neighbor cell groups? Per-dataset choice; the tree doesn't care, but the join ergonomics might.
- O6 — Status channel layout: stays flat (
<store>.status/<run>/...) for now; revisit if poller LIST pagination becomes a bottleneck at global shard counts. - O7 — MOC staleness policy: stamp
generated_at+ source (dispatcher vs sweep); do readers warn, re-walk, or trust silently when stale? Current lean (#200 thread): trust silently on the hot path (false negatives only, per D9), detect lazily (warn on a stamped leaf the MOC doesn't list), regenerate explicitly (refresh=True/ the sweep); no wall-clock staleness horizon. The incremental-run half is settled by the §4 lifecycle (leaves are durable truth; the sweep rebuilds correctly). Implemented in this shape by PR #208's reader primitives (zagg.coverage:warn_if_stale— once per store, never auto-walk — andrefresh_root_coverage, the explicit walk); the concurrent-run GET-union-PUT race (last writer wins until re-union/sweep) is recorded there as accepted under this same lean. - O8 — shard-MOC budget, serialization, carrier, pad sentinel:
RESOLVED (espg-ratified,
from the #202 item (6) measurement):
no budgeted/coarsened tier — the leaf encoding is an exact cell-order
bitmap, zstd-compressed, as the in-leaf
coverage.mocsidecar (raw size deterministic atceil(4^depth/8), immune to the ragged worst case where exact ranges hit MB-scale on linear-track occupancy at ~1.1 cells/range). Stamp attrs carry only the tier-0 box + the bitmap's order + pointer + byte sizes; pad sentinel is JSON null; the envelope gains anencoding: "ranges" | "bitmap"discriminator (later extended by D14 to add a third value,"full"). Bit convention (normative on the mortie spec page §7, golden-vector-pinned on PR #208): bit i = the i-th shard-subtree cell in ascending packed-word order (base-4 D1 digit tail, digits 1..4 → 0..3), MSB-first per byte. Addendum (espg-blessed in-session, 2026-07-20 — closing the #276 item-5 flag): the PR #208 contract constants stand as-shipped — the bitmap bit convention and the root-MOC range ordering are contract (both golden-vector-pinned; spec page §7.2–§7.3); zstd level 3 is non-normative (any zstd level decodes identically; only "zstd stream" is contract). - O9 — end-of-run root MOC default: RESOLVED — on for hive stores
(espg:
coverage MOCs are the default for healpix templates; PR #208 implements
output.coverage_moc, default true understore_layout: hive, explicit true rejected elsewhere). The write is fail-open on both backends; the Lambda leg is one fire-and-forgetmode: "coverage"Event invoke with the pre-serialized ranges envelope. - O10 — order-29 kind: RESOLVED by encoding-carried convention
(espg-ratified 2026-07-21 on the mortie PR #118 review — the FINAL
ruling, superseding this entry's earlier declaration-based forms:
resolution: "exact" | "point", recorded on #305, and the interim three-valuemixedaddendum). Point-ness is carried by the encoding itself, never by store or array metadata: mortie's area words and point ids are distinct encodings — area words are exact cells at EVERY order including 29 (nothing unrepresentable); point ids are points, full stop. There is no metadata field; readers key on the packed word's suffix region. The 29→24 clip is not spec semantics — it is a viewer-side transient cast (a gridlook/JS float64-safety measure at the display layer; other viewers or future runtimes may not need it), and membership of a point in a coarser cell is likewise the ordinary transient truncation. The one remaining ambiguity — the order-29 decimal string, which denotes both kinds — is resolved by the normative parse tie-break pinned on the mortie spec page §4 (a parsed string always yields the AREA word; golden-pinned at the bit level in mortie'stest_spec_page.py). zagg emits no kind field and carries no kind constants. The companion identity ruling stands unchanged: morton-declared stores carry a self-declared convention UUID — minted once, permanent, normative value on the mortie spec page §5 (MORTON_CONVENTIONinzagg.grids.morton) — while the upstream dggs-registry ask (#72 ask 3) proceeds in parallel; the conventions attrs are a list, so both entries coexist. - O11 — logical content hash of outputs: RESOLVED — adopted as
proposed (espg-ratified in-session, 2026-07-20; carved from D19,
discussion trail on
#299
and
the expansion).
Per-array sha256 over decoded values — "per-array" means per named
zarr array (variable) in the leaf (each data field, the ragged bytes
array,
morton, every coordinate), hashed over that array's full decoded contents (raw C-order bytes at the declared dtype, after decompression); the ShardingCodec's inner chunks and the one-object-per-array packaging are invisible to it by construction — never stored object bytes, which churn on codec/library upgrades. Recorded in the D20 sidecar as{array_name: hash}plus one combined hash (hash of the sorted per-array hashes). Computed at write, in-worker (the data is already in memory — near-free); scope: all arrays including coordinates (cheap, deterministic). Exact bytes, no float tolerance: any value change — including flagged code changes that passnp.isclose(the PR #282 class) — flips the hash by design; interpretation pairs the hash with the sidecar's recorded zagg version. Three jobs: the verification half of D19'ssemantic_hash("intended identical" vs "actually byte-identical"); the mismatch localizer ("onlyh_li_tdigestdiffers in this leaf"); and the detection mechanism for stamped-but-torn leaves under the §2 concurrency contract's out-of-contract case. Writer-side addendum (espg-ratified 2026-07-31, #342 — the recipe itself was frozen in spec §5, pinned to the moczarr reference): (1) sweep-written overview leaves are in scope — the overview writer records the same §5content_hashessidecar, computed from the folded arrays it already holds; (2) the overview envelope's sweep-internal skip digest stays separate (different jobs: cheap pre-write idempotence check vs. verification record; convergence may be revisited as a simplification later); (3) O11 recording is hive-only —store_layout: flathas no leaf D20 sidecar to record into, so the writer no-ops there and flat stores stay verifiable by running the recipe manually; (4) the hash source is the staged arrays (no read-back GETs), guarded in CI by a write-read parity test (staged hashes == hashes recomputed from a full store read-back). - O12 — retention/expiration (proposal, espg-directed to record
2026-07-20): product-level retention is a prefix lifecycle rule
(one rule per
{name}/— delete or transition a whole product); window-level retention within a product is the tag mechanism — the expiration rule filters on an object tag (e.g.ephemeral=true) and keep-alive is a retag (PutObjectTagging), never an access-refresh or self-copy "touch" (plain S3 lifecycle cannot key on access; a self-copy is a full write per object). Pin/unpin is therefore explicit and cheap. Open: tag vocabulary, who tags at write time (worker vs dispatcher), and whether pinned windows are recorded in the manifest or only in tags. -
O13 — publication profile (shape espg-ratified in-session, 2026-07-20; details open): publishing a product to a public bucket is a prefix copy minus operator telemetry, with the telemetry replaced by a simple roll-up summary that preserves cost-estimation utility (per-product totals; no
invoked_byARNs, no request ids, no per-run operator detail). Defined once so publishing tooling never improvises. Open: the summary's exact schema, its placement in the published copy (product root), and whether sub-shardmap rollups publish as-is (they carry granule ids — likely fine) or coarsen. -
O14 — reader tree model: product and resolution axes as DataTree (direction espg-ratified in-session 2026-07-20; API details open with moczarr — espg/moczarr#1). The store's four hierarchical axes map to xarray
DataTreeunevenly, and the mapping is now declared: the product axis is a strong fit (open_store()→ one child node per{name}/, each node today'sopen_hiveDataset — heterogeneous schemas are exactly the non-alignable-groups case DataTree exists for, and D19's named roots make node names human-meaningful); the resolution axis is a good fit once #201/D24 materialize (one node per order, source + overviews distinguished byrole, riding the emerging multiscale-DataTree conventions) — with the honest caveat that a D24-heterogeneous product has no complete Dataset at any single order: the seamless order-k view is a computed compose (per the composability classes), never a tree node, so the DataTree represents what is stored and the single-resolution Dataset stays an accessor-fabricated view; the window axis stays a concatenated time dimension inside each node (windows-as-nodes would shatter the dominant time-series read); the spatial digit axis is never tree nodes — a node per morton digit recreates the D5/D12 metadata storm client-side. The Rust-backed fastopen_datatreepath applies to zarr-native hierarchies, which the hive tree deliberately is not (D12) — it plugs in via D12's escape hatch instead: the sweep's optional interop materialization can generate a consolidated,open_datatree-able hierarchy as a derived cache for non-moczarr consumers, while the MOC-first opener remains the truth path (the derived-views principle, fourth instance). Phasing: (1)open_hive→ Dataset stays the primitive, unchanged; (2)open_store()→ product-level DataTree after moczarr#11; (3) level nodes with #201/D24; (4) optional sweep-generated interop hierarchy.
8.3 Standing test obligations (what a conforming implementation proves)¶
Consolidated index of the normative test claims scattered through the registry — CI coverage should be auditable against this list:
- Rollup == direct, per derived-artifact family (D22): stats fold, shardmap coarsen (the #294 exactness property), MOC union, overview aggregation. Sweep idempotence (second run over an unchanged tree is a no-op) and nothing-load-bearing (deleting every rollup leaves leaf reads green).
- Merge-fold algebra (D20): associativity/commutativity of the stats
record fold; merge-of-children == direct — exactly for count/sum/min/max,
up to the approximate class (
np.isclose) for order-dependent t-digest members (cf. D24). - Semantic-hash canonicalization (D19): syntactic edits (whitespace,
key order, comments) never change the hash; packaging-knob edits
(orders, chunking, worker size, the
aggregation.streamingblock in every spelling — the law-equivalence contract above is the exclusion's condition) never change the hash; any semantic edit does; and, post-epoch (#415), a leaf-shapingoutputedit does while an explicit default hashes as absence. Name-grammar validation (base-component exclusion, URL-safe charset). - Manifest guard (§3): frozen-key match ⇒ idempotent accept, no
second PUT; mismatch ⇒ pre-dispatch refusal;
path_groupingabsent⇒1 normalization; allowed-set membership forcell_order/shard_order. - D24 guards: cross-cell-order write into an occupied (shard, window) leaf refuses with the useful error; same-order rerun replaces (D4); composability-class derivation from the aggregator merge-law flags.
- Naming dialects (D23):
/3round-trip ({window}.zarr↔ sidecar stem) with/1–/2grammars byte-unchanged; spec-string discrimination. - Golden byte pins: the D18 vlen-bytes framing (numcodecs
VLenArray-compatible) and the O8 coverage-bitmap bit convention (PR #208 vectors). - Stamp/debris semantics (D4): an unstamped leaf is invisible to readers and safely overwritable; a stamped leaf is complete.
9. References¶
Community precedents for the §4 tiered-coverage conventions (budgeted conservative summary in the metadata plane, exact geometry in the data plane):
- STAC best practices — footprint simplification discipline: github.com/radiantearth/stac-spec/blob/master/best-practices.md
- STAC item spec — bbox + geometry as the queryable summary: github.com/radiantearth/stac-spec/blob/master/item-spec/item-spec.md
- stactools raster-footprint — densify → reproject → simplify-to-tolerance: element84.com/geospatial/the-stactools-raster-footprint-utility/
- CMR ingest API — geometry-complexity constraints at ingest: cmr.earthdata.nasa.gov/ingest/site/docs/ingest/api.html
- GeoParquet —
bboxcovering column (fixed-size conservative cover in metadata, exact WKB in data; the direct analog of the tier-0 morton box): opengeospatial/geoparquet - PostgreSQL TOAST — the ~2 KB inline threshold (portability footnote for the tier-1 budget): www.postgresql.org/docs/current/storage-toast.html
- IVOA MOC recommendation — degraded-order MOC practice; NUNIQ int64 order-29 ceiling: www.ivoa.net/documents/MOC/
- H3
compactCells— mixed-resolution minimal covers: h3geo.org/docs/api/hierarchy