# Validation Report — Petrel Flow Accumulation Layer v0.1 Global (DEM-derived) ## Validation status: qualitative only **v0.1 has not been formally validated.** No quantitative cell-wise skill score against an independent drainage reference has been computed for this vintage. What follows is an honest record of the qualitative checks performed, the structural sanity anchors the product passes, and the quantitative validation scheduled for v1. Users who need demonstrated quantitative agreement should wait for v1 or validate v0.1 in their own region of interest first. ## Qualitative checks performed ### 1. Hydrological sanity check — Semeru (2026-07-27) The pysheds derivation was validated on a 2° (2400 × 2400 px) Copernicus-DEM window around Semeru volcano, Java: - **Runtime:** ~15 s for the window. - **Flow correctness:** the accumulation maximum (~953 000 cells) lands at a **0 m valley exit, not the summit** — i.e. water accumulates downhill and leaves the domain at the coast, which is the defining correctness property of a flow-accumulation grid. - **Channel density:** the derived channel network (`accumulation > 1 000 cells`) covers ~2.4% of the window — a plausible drainage density for mountainous tropical terrain. This anchors that the fill → flowdir → accumulation chain is behaving hydrologically; it is a structural check on one window, not a global skill score. ### 2. Drainage-structure agreement with the layer it replaces The derived channel network and distance-to-drainage were compared visually against the HydroRIVERS geography this Layer is built to replace. Around Semeru, the DEM-derived distance-to-drainage had a **local median on the order of ~1 km**, consistent with the HydroRIVERS-derived distances for the same area; the **global-land median is on the order of ~4 km**, reflecting the larger inter-channel spacing of arid and low-relief interiors. These are preliminary spot-check magnitudes, not a cell-wise benchmark (see *What v0.1 does not establish*). ### 3. Automated product integrity (promote audit) The Petrel promote audit machine-checks every release bundle: encoding tags (`PETREL_SCALE`, `PETREL_COAST_CLIP` provenance), overview pyramids, required documentation, value ranges (distance ≤ 50 km, TWI ≤ 35), and mask correctness (Antarctica + open ocean NoData, coastal land resolves to data). These are integrity checks on the shipped rasters, not accuracy checks on the underlying hydrology. ## What v0.1 does NOT establish - **No quantitative channel / distance benchmark vs HydroRIVERS** (or any independent hydrography). The planned benchmark — channel-overlap and distance-to-drainage correlation against HydroRIVERS, with per-continent breakdown — is the **flagged v1 work item**. - **No absolute-accumulation validation.** The dominant structural error is the halo-tiled trans-halo under-count; v0.1 does not quantify how far the absolute accumulation of a major river deviates from its true global value. The v1 Barnes-style two-pass accumulation is designed to remove it. - **No discharge / gauge validation** against river-flow observations. - **No DEM-artefact audit** — Copernicus DEM voids, canopy, and built-up biases propagate into flow paths and were not systematically assessed. ## ⚠ Quantified limitation — trans-tile accumulation shortfall (2026-09-09) Upslope contributing area is accumulated **per tile with a 1-degree halo**, so it under-counts wherever a river's catchment crosses tile boundaries. Measured against true basin areas: | Basin | captured | implied TWI bias | |---|---|---| | Amazon | 7.8 % | −2.55 | | Mississippi | 10.4 % | −2.26 | | Congo | 17.5 % | −1.74 | | Nile | 22.7 % | −1.48 | | Danube | 30.1 % | −1.20 | Because `TWI = ln(a / tan β)`, a shortfall in `a` propagates as a **logarithmic offset**: TWI reads roughly **1.2–2.6 units low** on those major basins, which on a 0–25 scale is a wetness class rather than a rounding error. Local maxima of accumulation sit exactly on tile boundaries (e.g. lat 30.00, lon 20.00), which is the halo showing through. **Not fixable by widening the halo** — the Amazon is ~3,000 km long. The fix is a global two-pass accumulation (per tile, then propagate cross-boundary inflow over the tile graph until it converges), which is the v1 item. **What this does NOT affect:** local and small-catchment terrain, where the contributing area is contained within a tile, and the *relative* drainage structure within a tile. The `distance_to_drainage` band is thresholded from the same accumulation, so it shares the dependency; its magnitude of error has not been measured and is not claimed here. ## Known limitations 1. **Halo-tiled derivation.** Each 10° tile is derived with a 1° (~110 km) halo. The drainage **network structure** (channels, distance-to-drainage, TWI) is correct; the **absolute accumulation** for a major trans-halo river is under-counted at tile interiors far from its headwaters (source beyond the halo). Barnes-style global two-pass accumulation is the v1 fix. 2. **Distance far-tail cap.** Distance-to-drainage is capped at **50 km**; deserts and endorheic interiors saturate at the cap and should be read as "far from any channel", not as exact distances. 3. **278 m native resolution.** Sub-pixel channels below the 1 000-cell threshold are unresolved; not for reach-scale hydraulic work. 4. **Channel-threshold dependence.** Network density is set by the fixed 1 000-cell accumulation threshold; a different threshold would thin or thicken the network and shift distance-to-drainage. 5. **Coverage:** global land 60° S – 84° N. Ocean is NoData (GSHHG-f coast clip, `all_touched`); Antarctica is excluded by Petrel standard. ## v1 validation plan 1. Quantitative HydroRIVERS benchmark: channel-network overlap + distance-to-drainage correlation, per continent. 2. Barnes-style global two-pass accumulation and a before/after comparison of absolute accumulation on major trans-halo rivers. 3. Discharge / gauge spot-checks where independent river-flow data exist. 4. DEM-artefact audit (voids, canopy, built-up) and its effect on derived flow paths. ## How to validate locally For your region of interest: 1. Overlay the derived channel network on satellite imagery or a national hydrography and check the channels track real river lines. 2. Compare distance-to-drainage against a local river-distance layer if you have one — expect close agreement on network structure. 3. Remember that absolute accumulation is under-counted for large trans-halo rivers in v0.1 — use the network structure, not the absolute count. 4. Report systematic artefacts to layers@petreldata.io — findings feed the v1 plan. ## References - Bartos, M. (2020). pysheds: simple and fast watershed delineation in Python. https://doi.org/10.5281/zenodo.3822494 - Barnes, R. (2016). Parallel non-divergent flow accumulation for trillion cell digital elevation models on desktops or clusters. *Environmental Modelling & Software*, 92, 202–212. https://doi.org/10.1016/j.envsoft.2017.02.022 - Wessel, P., & Smith, W. H. F. (1996). A global, self-consistent, hierarchical, high-resolution shoreline database. *JGR*, 101(B4), 8741–8743. - European Space Agency & Airbus (2022). Copernicus DEM (GLO-30). https://doi.org/10.5270/ESA-c5d3d65