Video - https://www.youtube.com/watch?v=GsYtVGVTPWM
GRIME is a multi-parameter optimization system that identifies optimal locations for deploying trash interception barriers ("nets") on urban waterways. It evaluates candidate sites across 27 geospatial parameters organized into 6 parameter families, producing 4 sub-scores that combine into a single composite ranking per candidate location.
Shipping today (July 2026): 449 pipeline-scored North Carolina regions serving 9,903 candidate sites, a 147-site Durham flagship with 25/27 parameters varying live, per-value provenance on every number the API serves, and a 94.1% Dirichlet-robustness result on the regenerated flagship — receipts in VALIDATION_LOG.md.
What makes it technically interesting:
- A two-level weighted scoring architecture (parameters → sub-scores → composite) that is both interpretable and tunable
- Hydrological pipeline built on real DEM data: pit-filling → depression-filling → flat resolution → D8 flow direction → flow accumulation → stream extraction
- Manning's equation applied to estimate flow velocity from DEM slope and channel geometry, with feasibility gates that eliminate sites where deployment is physically impossible
- Monte Carlo sensitivity analysis via Dirichlet-perturbed weight vectors to assess ranking robustness
- Three-phase candidate placement algorithm: spatial constraint satisfaction → full-parameter scoring → population-scaled risk-percentile filtering
- On-demand real waterway geometry from OpenStreetMap's Overpass API for 89,518 places across 239 countries
Scope: Site selection modeling and scoring. GRIME does not design the physical trap, predict trash composition, or model individual debris trajectories.
Non-goals: Real-time sensor integration, computer vision trash classification, trap mechanical design, economic cost optimization.
Conceptual visualizations generated in the Wolfram Language. The terrain base uses Wolfram's GeoElevationData for the Durham area (the pipeline itself uses USGS 3DEP); the parameter/score surfaces are illustrative reconstructions of GRIME's scoring functions over synthetic grids, not plots of pipeline output. See full documentation for the exact code behind each figure.
3D elevation surface of the Durham area (Wolfram
GeoElevationData; the scoring pipeline uses USGS 3DEP at 10 m). Stream valleys visible as low-elevation grooves are where GRIME extracts candidate net sites.
The final composite score as a function of Generation (trash input) and Impact (downstream consequence), with Flow and Feasibility held constant. The diagonal ridge shows that high scores require both trash presence and downstream consequence — neither alone is sufficient.
Deployment feasibility as a function of channel width and flow velocity. The green plateau marks the sweet spot: narrow, moderate-velocity channels where nets can be spanned and anchored. Red zones are eliminated by hard gates.
Flow velocity estimated from Manning's equation across slope and roughness parameter space. Steep, smooth channels (high slope, low roughness) produce dangerous velocities; flat, rough channels produce stagnant conditions.
Environmental justice burden across a synthetic metro region. Peaks identify overburdened communities where trash interception has the highest equity value — GRIME weights these areas higher in the Impact sub-score.
500 weight vectors sampled from a Dirichlet distribution (κ=10) projected onto a ternary diagram. The red dot is the baseline. GRIME recomputes rankings under each perturbation to test whether top-ranked sites are robust to weight assumptions.
Urban waterways accumulate trash from stormwater runoff, illegal dumping, combined sewer overflows, and bridge crossings. Deploying interception devices (nets, booms, trash traps) requires choosing locations that maximize debris captured per device while remaining physically feasible to install and maintain.
A naive approach — placing traps at the largest rivers — fails because:
- Large rivers are too wide (>30m) for stationary nets; debris passes around or damages the device
- High-traffic waterways (navigable canals, shipping channels) cannot be obstructed
- Upstream tributaries with high impervious surface and population density generate more trash per unit area than rural mainstems
- Downstream impact varies: a trap upstream of a drinking water intake has orders of magnitude more public health value than one upstream of an industrial canal
- Physical feasibility (road access, bank slope, flow velocity, land ownership) eliminates many otherwise-optimal locations
- All data sources must be free and low-friction (EPA/USGS/OSM are keyless; Census ACS uses a free, instant API key — see
.env.example) - The model assumes stationary barrier-style traps deployable in channels ≤30m wide (~100 ft)
- Scoring weights are set by informed heuristic and literature, not supervised learning (no ground-truth dataset of "correct" trap placements exists at scale)
- The system must produce results in <15 seconds per city for interactive demo use
graph TD
A[DEM Raster - USGS 3DEP] --> B[Hydrological Conditioning - pysheds]
B --> C[Flow Direction - D8 Algorithm]
C --> D[Flow Accumulation]
D --> E[Stream Network Extraction]
E --> F[Candidate Site Generation]
G[EPA APIs - TRI ECHO SDWIS FRS] --> H[Parameter Computation - 27 params x 6 families]
I[USGS APIs - NWIS StreamStats PAD-US] --> H
J[Census APIs - ACS TIGER incl. EJ index] --> H
K[OSM - Overpass API] --> H
F --> H
H --> L[Hard Gate Filtering]
L --> M[MinMax Normalization]
M --> N[Weighted Sub-scores x4]
N --> O[Composite Score]
O --> P[Risk-Percentile Filtering]
P --> Q[Ranked Deployment Sites]
Q --> R[FastAPI Backend]
R --> S[Mapbox GL JS Dashboard]
K --> S
| Subsystem | Location | Purpose |
|---|---|---|
| DEM Pipeline | core/pipeline.py |
Fetch elevation data, extract stream network, generate candidate points |
| Generation Params | core/generation.py |
Compute trash generation indicators (population, land use, industrial sources) |
| Flow Params | core/flow.py |
Compute hydraulic transport parameters (discharge, velocity, flood frequency) |
| Impact Params | core/impact.py |
Compute downstream consequence indicators (drinking water, EJ, protected areas) |
| Feasibility Params | core/feasibility.py |
Compute deployment constraint parameters (road access, channel width, slope) |
| Scoring Engine | core/scoring.py |
Normalize, weight, composite, sensitivity analysis |
| API Server | api/main.py |
REST + WebSocket endpoints serving scored GeoJSON |
| Landing page | dashboard/index.html |
Marketing/overview page (hero globe, dataset globe, illustrative UI mock) |
| Map explorer | dashboard/explore/index.html |
Interactive map: on-demand OSM waterway fetching, client-side demo scoring, real-pipeline overlay for the Durham pilot |
| Scored dataset | mock_data/candidates.geojson |
Frozen live-pipeline output (147 Durham sites) served by the API |
| Places Database | mock_data/places.json |
89,518 place records across 239 countries (~6MB compact JSON) |
About the places database. The 89,518 records come from
geonamescache(real cities/towns with population ≥ 100) plus procedurally generated nearby townships (scripts/generate_mock.py, seeded for reproducibility). They exist only to give the explorer a global set of clickable starting points — waterway geometry and all scoring are computed live from OpenStreetMap at click time, so the procedural names never enter a score. We say "places," not "cities," for this reason.
Two execution modes exist:
Mode 1 — Full Python pipeline (research/validation): DEM fetch → pysheds hydrology → stream extraction → candidate generation → API-based parameter computation → composite scoring → GeoJSON output
Mode 2 — Dashboard on-demand (demo/interactive): User clicks city → Overpass API returns real waterway geometry → client-side JS generates candidate positions with spatial constraints → client-side scoring using simplified parameter model → Mapbox GL renders results
Mode 2 exists because Mode 1 takes 3–5 minutes per watershed and requires installing pysheds (which has C dependencies that fail on some Windows machines). Mode 2 runs in <5 seconds anywhere with a browser.
The 27 raw parameters are not directly comparable (population density in persons/km² vs flow velocity in m/s vs a binary land ownership flag). The system handles this through two-level aggregation:
-
Parameter level: Each raw parameter is MinMax-normalized to [0, 1] independently within the candidate set, then multiplied by its within-family weight. The weighted sum produces a sub-score in [0, 100].
-
Sub-score level: The four sub-scores are combined via a second set of weights into the composite score in [0, 100].
This two-level structure has a specific advantage: it makes the model interpretable at the sub-score level. A judge or engineer can look at a candidate and immediately see "high generation, low feasibility" without needing to parse 27 individual numbers.
Some parameters act as binary disqualifiers rather than continuous scores. A channel wider than 50m cannot hold a net regardless of how much trash flows through it. These are implemented as hard gates that remove candidates before scoring, separate from the soft scoring that ranks survivors:
| Gate | Condition | Rationale |
|---|---|---|
| Velocity | V > 3.0 m/s | Trap will be damaged or torn loose |
| Channel width | W > 50m or W < 0.5m | Too wide to span or too narrow for meaningful accumulation |
| Land ownership | Confirmed private, no permission | Legal barrier to deployment |
Candidate placement is not random scatter. It is a three-phase algorithm:
- Constraint satisfaction: Generate all positions that pass spatial, width, and traffic constraints
- Full scoring: Evaluate every surviving position on the composite model
- Risk-percentile selection: Keep only the top N% by score, where N scales with city population
This separates "can we physically put a net here?" (phase 1) from "should we?" (phases 2–3).
A city of 20 million people needs more nets than a town of 10,000, but not linearly more. The risk percentile threshold scales in steps:
| Population | Percentile kept | Rationale |
|---|---|---|
| >10M | Top 35% | Mega-cities have extensive waterway networks; more sites are genuinely high-risk |
| >1M | Top 30% | Large cities still have substantial catchments |
| >100K | Top 25% | Mid-size cities, moderate network complexity |
| <100K | Top 20% | Small towns, fewer waterways, tighter selection |
A minimum floor of 5 deployed sites ensures the model always produces enough output to demonstrate ranking behavior.
Let x ∈ ℝ²⁷ be the raw parameter vector for a candidate site. The composite score S(x) is:
S(x) = Σ(k=1..4) ωk · Gk(x)
where ω = [0.30, 0.25, 0.30, 0.15] are the sub-score weights and each sub-score Gk is:
Gk(x) = 100 · Σ(j ∈ Fk) wj · x̂j
where Fk is the set of parameter indices belonging to family k, wj is the within-family weight for parameter j (renormalized to sum to 1 after filtering unavailable parameters), and x̂j is the MinMax-normalized value:
x̂j = (xj - min(xj)) / (max(xj) - min(xj))
For distance-based parameters where lower is better (estuary distance, beach distance), the normalization is inverted: x̂j = 1 − x̂j.
Important implementation detail: Normalization is computed across the candidate set, not against a global reference. This means scores are relative rankings, not absolute measures. A score of 80 means "top of this candidate pool," not "80% of some theoretical maximum."
Flow velocity at a candidate site is estimated via Manning's equation:
V = (1/n) · R^(2/3) · S^(1/2)
where:
- n is Manning's roughness coefficient (dimensionless), selected by channel type:
- Clean straight: 0.030
- Winding with pools (typical urban creek): 0.040
- Sluggish, weedy: 0.070
- Urban concrete-lined: 0.015
- (Source: Chow, V.T., 1959, Open-Channel Hydraulics)
- R is the hydraulic radius (m) = A_cross / P_wetted, approximated as rectangular channel: R = (W × D) / (W + 2D), where depth D ≈ 0.3W (bankfull approximation)
- S is the channel slope (dimensionless), computed from DEM as elevation difference over a 100m reach: S = (Z_here − Z_downstream) / 100, clamped to minimum 0.0001
A continuity estimate is blended in: V_continuity = Q_site / A_cross, where Q_site is the gauge discharge area-scaled to the candidate's own catchment (M4: site-specific, now that catchment area is real km²) and converted to m³/s. The final velocity is the geometric mean of the Manning and continuity estimates (a soft blend, not a fully independent measurement):
V_final = sqrt(V_Manning · V_continuity)
This hedges against errors in either the DEM slope (which can be noisy at 10m resolution) or the channel geometry assumption (rectangular approximation).
The velocity feasibility score is a piecewise function mapping velocity to a deployment viability multiplier:
f(V) =
0.3 if V < 0.05 m/s (stagnant — debris doesn't concentrate)
0.7 if 0.05 ≤ V < 0.30 (slow but workable)
1.0 if 0.30 ≤ V ≤ 1.50 (optimal interception range)
0.5 if 1.50 < V ≤ 2.50 (fast — heavy anchoring needed)
0.1 if V > 2.50 (too fast — trap damage likely)
Sites with V > 3.0 m/s are removed entirely by the hard gate before scoring.
The rational method runoff coefficient C is estimated from impervious surface percentage via a linear model:
C = 0.05 + 0.009 · I
where I is the NLCD impervious surface percentage [0, 100]. This yields C ∈ [0.05, 0.95], ranging from forest (≈5% runoff) to fully paved (≈95% runoff). This is the same linearization used in the WaterGate methodology.
Several parameters (CSO proximity, Superfund proximity) use an inverse-distance kernel to compute influence from point sources:
score = Σ(i=1..N) 1 / (1 + (di / h)²)
where di is the Euclidean distance (in UTM meters) from the candidate to source i, and h is the half-decay distance (500m default). This is a Cauchy kernel that gives full weight at distance 0 and half weight at distance h.
Proximity to downstream drinking water intakes uses an exponential decay:
score = Σ(i=1..N) exp(-di / 10)
where di is the distance in km. Intakes within 10km get weight ≈0.37, within 5km ≈0.61, within 1km ≈0.90. Intakes beyond 50km are ignored.
GRIME uses two related, deliberately separate robustness receipts:
- Runtime per-site
robustness_pct:core.scoring.sensitivity_analysis()samples N perturbed weight vectors from a Dirichlet distribution, recomputes scores, and records how often each candidate appears in the top 5. This is the strict per-site field stored in served GeoJSON. - Paper-validation top-25 stability:
scripts/validate_paper.py --dirichletmeasures whether the baseline top 10 candidates remain inside the top 25 across 10,000 seeded draws. This is the broader paper-style stability receipt.
Both use ω' ~ Dir(α), where α = 10 × [0.30, 0.25, 0.30, 0.15]. The α scaling factor controls perturbation magnitude: higher α concentrates samples closer to the baseline weights.
The two percentages should not be compared directly. A site can have low top-5 robustness_pct while the overall leading pool still has high top-25 stability.
The minimum-spacing constraint uses the haversine formula for geodesic distance:
a = sin²(Δφ/2) + cos(φ₁) · cos(φ₂) · sin²(Δλ/2)
d = 2R · atan2(√a, √(1−a))
where R = 6,371,000 m. This is used instead of Euclidean distance because the candidate set can span several kilometers, where flat-earth approximation introduces meaningful error at high latitudes.
Why this changed. EPA removed EJSCREEN — the tool, the data downloads, and the
ejscreenRESTbroker.aspxArcGIS server — on 2025-02-05, and the White House removed CEJST on 2025-01-22; neither has an official live replacement (the suit to restore EJSCREEN was dismissed on standing, 2026-03-13). Because EJSCREEN's demographic index is percentile-ranked ACS demographics, GRIME reconstructs it directly from the live Census ACS 5-year API (core.impact.get_ej_index), which we control, instead of calling a dead endpoint. Sources: EELP tracker, EDGI, CEJST removal, reconstruction mirror screening-tools.com.
We compute EJSCREEN's two-component core demographic index per block group:
demographic_index_bg = mean( pct_rank(% low-income) , pct_rank(% people of color) )
EJ = area-weighted mean of demographic_index_bg over the catchment ∈ [0, 1]
- % low-income — ACS table
C17002(income-to-poverty ratio < 2.0). - % people of color —
1 − (non-Hispanic white / total)fromB03002. - Percentile-ranked within the county, then area-weighted over the candidate's catchment — so the index varies across catchments (no longer a constant).
The six-component supplemental index (adding limited-English C16002,
< high-school B15003, under-5 and over-64 B01001) is a documented extension;
the two-component core index above is EPA's headline definition. The deprecated
get_ejscreen_index is retained only so old callers degrade to a neutral 0.5.
Purpose: Convert raw elevation data into a hydrologically consistent surface from which flow direction and stream networks can be extracted.
Input: DEM raster from USGS 3DEP (10m resolution)
Output: Flow direction grid, flow accumulation grid, stream network GeoJSON
Steps:
FUNCTION condition_dem(dem):
pit_filled ← fill_pits(dem) // remove single-cell sinks
flooded ← fill_depressions(pit_filled) // fill multi-cell depressions
inflated ← resolve_flats(flooded) // assign gradient to flat areas
flow_dir ← D8_flowdir(inflated) // each cell → 1 of 8 neighbors
accumulation ← flow_accumulation(flow_dir) // count upstream cells per cell
RETURN flow_dir, accumulation
Stream extraction: Cells where accumulation exceeds a threshold (default 500 cells = 500 × 10m × 10m = 0.05 km²) are classified as stream cells. Connected stream cells are vectorized into LineString geometries.
Complexity: O(n) for each step where n is the number of DEM cells. For the Ellerbe Creek bbox at 10m resolution: approximately 3000 × 1500 = 4.5M cells. Total conditioning time: ~30–60 seconds.
Why pysheds: It operates entirely in-memory on NumPy arrays without requiring ArcGIS or GRASS GIS. The D8 algorithm assigns each cell exactly one of 8 cardinal/diagonal flow directions based on steepest descent, which is the standard approach for stream extraction in computational hydrology.
Multi-source resolution: Channels narrower than the 10m DEM raster are captured through NHD and OSM waterway geometry, so sub-10m streams are still identified.
Purpose: Given a set of waterway geometries and city metadata, produce a set of spatially valid, risk-ranked candidate sites for trap deployment.
Input: Array of stream geometries (from Overpass API), city population, country code
Output: Ranked array of candidate objects with scores and parameters
The explorer is a fast, client-side scorer. The explorer scores client-side with clamped closed-form formulas (no MinMax), using OSM-estimated widths and deterministic per-candidate values seeded by each site's own coordinates (so rankings are stable across pan/zoom — no
Math.random()for any named quantity). It is a fast, client-side scorer; the Python pipeline (§5–6.1) is the full 27-parameter MinMax model.
FUNCTION generate_candidates(streams, pop, country):
// ── Phase 1: Constraint satisfaction ──
SAME_STREAM_SPACE ← 500m // wide gaps along one waterway (occlusion makes
CROSS_STREAM_SPACE ← 300m // a 2nd net 500m downstream largely redundant)
MAX_WIDTH ← 30m
placed ← []
valid_positions ← []
FOR EACH stream IN streams:
IF stream.width > MAX_WIDTH: CONTINUE
cap ← waterway_capacity(stream)
count_on_stream ← 0
dist_since_last ← SAME_STREAM_SPACE
FOR EACH point IN stream.coords:
dist_since_last += haversine(previous_point, point)
IF dist_since_last < SAME_STREAM_SPACE: CONTINUE
IF any p in placed where haversine(p, point) < CROSS_STREAM_SPACE: CONTINUE
IF count_on_stream >= cap: CONTINUE
placed.add(point); count_on_stream++; dist_since_last ← 0
valid_positions.add(point with metadata)
// ── Phase 2: Score every valid position (deterministic, coord-seeded) ──
FOR EACH pos IN valid_positions:
pr ← seeded_rng(hash(pos.lat, pos.lon) + seed) // STABLE per site
compute clamped sub-scores + composite using pr() (no Math.random)
// ── Phase 3: Greedy selection with multiplicative upstream occlusion ──
// NOT a static top-N%. Each time we place a net we discount every still-unplaced
// candidate that is DOWNSTREAM on the same river (chained across OSM ways by
// computeWayOrder, ordered by (wayOrder, coordIdx)) — a net already catches most
// debris, so a downstream net only sees the passthrough (1 − η)^k of trash.
η ← 0.65 // catch efficiency per net
target ← min(250, max(5, ceil(len * population_percentile)))
WHILE selected < target:
sort remaining by current composite; best ← pop highest; selected.add(best)
FOR EACH downstream c on best's river:
c.generation ← c.raw_generation × (1 − η)^(nets_upstream_on_river)
recompute c.composite
RETURN selected
Complexity: Phase 1 is O(n × m); Phase 2 is O(k); the greedy Phase 3 is O(target × k). In practice k < 250 and the whole function runs in <100 ms. Upstream occlusion (η = 0.65 compounding) is the original part of the system — it spreads nets across waterways instead of clustering them on the single highest-trash reach. See §6.3 for the robustness analysis.
FUNCTION sensitivity_analysis(candidates, n_perturbations=50):
baseline ← compute_composite_score(candidates)
top5_counts ← zeros(len(candidates))
REPEAT n_perturbations TIMES:
α ← [3.0, 2.5, 3.0, 1.5]
ω' ← sample_dirichlet(α)
composite' ← ω'[0]·gen + ω'[1]·flow + ω'[2]·impact + ω'[3]·feas
top5 ← indices of 5 highest composite'
top5_counts[top5] += 1
robustness ← top5_counts / n_perturbations × 100
RETURN baseline with robustness column
Real results (reproducible). python3 scripts/robustness_report.py runs this at
n = 500 on the offline demo candidate set (26 synthetic-parameter sites — not
the shipped live dataset) and writes an actual rank-stability histogram to
dashboard/docs/exports/robustness_hist.png (not a synthetic mock). On that demo
set, the top 4 sites retain a top-5 position in 81–99 % of perturbed-weight
runs. The shipped live 147-site dataset stores its own per-site robustness_pct
from the n = 200 run that produced it: the rank-1 site retains top-5 in 7.5 %
of draws, the maximum stored top-5 retention is 82 %, and only 23 of 147 sites
ever enter the top 5 (the rest report 0 %), so top-5 membership is confined to a
small stable pool while the exact order within
it is weight-sensitive at the margin.
Field-validation candidates (roadmap). The top sites the model flags — South Ellerbe (composite 68, 99 % robust) and the upper Sandy Creek reaches — are the natural targets for ground-truthing against photographed accumulation points; that validation is the next step before any deployment claim.
When ground-truth trap locations are available, weights can be optimized via scikit-optimize:
FUNCTION optimize_weights(candidates, known_good_sites):
FUNCTION objective(weights):
w ← normalize(weights)
scored ← recompute with w
penalty ← sum of ranks of known_good_sites in scored
RETURN penalty
search_space ← [Real(0.05, 0.60)] × 4
result ← gp_minimize(objective, search_space, n_calls=50)
RETURN normalize(result.x)
Status: Implemented in core/scoring.py as optimize_weights(candidates_df, known_good_indices) — a Gaussian-process gp_minimize over the four sub-score weights whose objective minimizes the mean rank of known-good sites. The shipped weights are GRIME's literature-informed heuristic defaults.
Decision: Aggregate 27 parameters into 4 sub-scores, then combine sub-scores into a composite.
Context: A flat 27-weight sum is opaque — changing one weight has a non-obvious effect.
Alternatives considered: (1) Flat weighted sum. (2) PCA dimensionality reduction. (3) Random forest classifier.
Chosen approach: Two-level weighted sum. Sub-scores map to real questions ("how much trash?", "how does it move?", "does it matter?", "can we deploy?").
Consequences: Interpretable and tunable, but assumes linear parameter contributions within each family. Non-linear interactions are not captured.
Decision: Use MinMax scaling to [0, 1] per parameter across the candidate set.
Alternatives considered: Z-score, percentile rank, log-transform + MinMax.
Chosen approach: MinMax. Simple, bounded, interpretable.
Consequences: Simple, bounded, and interpretable; extreme outliers on a parameter compress the others toward 0.
Decision: The dashboard computes scores in JavaScript, not by calling the Python backend.
Context: pysheds/rasterio have C dependencies that fail on Windows. Dashboard must work by opening one HTML file.
Chosen approach: A separate, simplified JS heuristic — not a parity port of the Python model. It uses clamped closed-form formulas (no MinMax normalization), OSM-estimated widths, and deterministic per-candidate values seeded by each site's own coordinates (stable across re-render). It shares the framing — four sub-scores and the greedy upstream-occlusion placement — and computes its numbers from OSM geometry. The Python pipeline is the full API-driven, MinMax-normalized implementation.
Consequences: The dashboard scores client-side for instant, keyless, global coverage; the Python pipeline is the full-fidelity implementation.
Decision: Fetch real waterway geometry from the Overpass API on each city click.
Alternatives considered: Pre-generated GeoJSON per city (storage), procedural random-walk rivers (alignment), NHD (US-only).
Chosen approach: Overpass API with 12s timeout and procedural fallback.
Consequences: Requires internet; waterway coverage follows OSM.
{
"type": "Feature",
"geometry": {"type": "Point", "coordinates": [-78.898, 35.994]},
"properties": {
"id": 0,
"city": "durham",
"city_name": "Durham, NC",
"stream_name": "Ellerbe Creek",
"composite_score": 47.08,
"generation_score": 42.5,
"flow_score": 38.1,
"impact_score": 35.2,
"feasibility_score": 82.0,
"population_density": 1424.3,
"impervious_pct": 51.8,
"usgs_mean_q_cfs": 37.0,
"flow_velocity_ms": 1.377,
"stream_order": 5,
"catchment_area_km2": 63.5,
"channel_width_m": 13.1,
"ej_index": 0.595,
"road_access_m": 366.4,
"bank_slope_deg": 19.6,
"robustness_pct": 72.6,
"rank": 1
}
}{"n":"Durham","c":"US","p":278993,"la":35.994,"lo":-78.8986}Fields: n=name, c=ISO country code, p=population, la=latitude, lo=longitude.
| # | Parameter | Unit | Family | Default Weight | Data Source |
|---|---|---|---|---|---|
| 1 | Population density | persons/km² | Generation | 0.18 | US Census ACS |
| 2 | Impervious surface % | % | Generation | 0.20 | NLCD 2021 |
| 3 | Road density | km/km² | Generation | 0.10 | Census TIGER / OSMnx |
| 4 | EPA TRI facility count | facilities/km² | Generation | 0.18 | EPA TRI API |
| 5 | NPDES discharge points | count | Generation | 0.12 | EPA ECHO API |
| 6 | CSO/storm outfall density | points/km² | Generation | 0.12 | EPA ECHO |
| 7 | Litter complaint density | reports/km² | Generation | 0.10 | Durham 311 / local GIS |
| 8 | USGS mean discharge Q | cfs | Flow | 0.22 | USGS NWIS |
| 9 | Flow velocity | m/s | Flow | 0.16 | Manning's eq from DEM |
| 10 | Stream order (confluence-degree heuristic) | ordinal | Flow | 0.14 | Network topology (H3: not true Strahler) |
| 11 | Catchment area A | km² | Flow | 0.18 | pysheds DEM analysis |
| 12 | Flood return period Q10 | cfs | Flow | 0.14 | USGS StreamStats |
| 13 | Seasonal flow variability | CV | Flow | 0.10 | USGS NWIS annual stats |
| 14 | Runoff coefficient C | dimensionless | Flow | 0.06 | Linear impervious→C (C=0.05+0.009·I) |
| 15 | Drinking water intake proximity | exp(-d/10) | Impact | 0.22 | EPA SDWIS / ECHO |
| 16 | Protected area proximity | score | Impact | 0.16 | USGS PAD-US |
| 17 | Environmental justice index | [0,1] | Impact | 0.18 | Census ACS (EJSCREEN demographic-index reconstruction) |
| 18 | Ocean/estuary proximity | km (inverted) | Impact | 0.14 | NHD terminus |
| 19 | Recreational beach proximity | km (inverted) | Impact | 0.12 | EPA BEACH Program |
| 20 | Tourism/recreation value | amenity count | Impact | 0.10 | OSM amenity density |
| 21 | Superfund site proximity | score | Impact | 0.08 | EPA FRS/CERCLIS |
| 22 | Road access distance | m | Feasibility | 0.25 | OSMnx routing |
| 23 | Channel width | m | Feasibility | 0.20 | NHD VAA + NBI span |
| 24 | Flow velocity (penalty) | m/s | Feasibility | 0.20 | Manning's eq |
| 25 | Land ownership | binary | Feasibility | 0.15 | USGS PAD-US |
| 26 | Bank slope stability | degrees | Feasibility | 0.10 | DEM gradient |
| 27 | Bridge/structure proximity | bonus | Feasibility | 0.10 | FHWA NBI |
The table above lists each parameter's designed source. In the committed
mock_data/candidates.geojson (the frozen June 2026 live run served by the API),
11 of the 27 parameters varied per-candidate — population density, road
density, flow velocity, stream order, catchment area, EJ index, estuary/beach
distance, tourism density, road access, velocity feasibility — and 16 were
constant fallbacks: all EPA-sourced parameters (TRI, NPDES, CSO, water intake,
Superfund = 0.0; ECHO/StreamStats endpoints were dead), single-gauge discharge
stats broadcast to every site (usgs_mean_q_cfs, flood_q10_cfs,
seasonal_cv), impervious_pct (whole-bbox NLCD mean, fallback 35.0) and its
derived runoff_coeff_C, litter_complaint_density (dead 311 endpoint),
land_ownership (PAD-US not wired), bridge_proximity_bonus (NBI not wired),
and channel_width_score/bank_slope_score (real inputs, uniform in this
area). Because compute_subscore drops constant columns and renormalizes, the
shipped ranking is driven by the 11 varying parameters — /api/weights reports
both the nominal and the effective weights, and the geojson's top-level
note + provenance block record the split for any regenerated output.
grime/
├── core/ # Python scoring pipeline (core deliverable)
│ ├── __init__.py # Constants, safe_call(), helpers
│ ├── pipeline.py # DEM → pysheds → stream extraction → candidates
│ ├── generation.py # 7 trash generation parameters + API integrations
│ ├── flow.py # 7 flow parameters, Manning's equation, USGS data
│ ├── impact.py # 7 downstream impact parameters, EJ scoring
│ ├── feasibility.py # 6 deployment feasibility parameters, hard gates
│ └── scoring.py # Normalization, weighting, composite, sensitivity
├── api/
│ └── main.py # FastAPI: REST + WebSocket + static serving
├── dashboard/
│ ├── index.html # Landing page (globes + illustrative UI mock)
│ ├── explore/index.html # Map explorer: Overpass fetch, client-side demo scoring
│ └── docs/ # Rendered documentation + figures
├── mock_data/
│ ├── candidates.geojson # Frozen live-pipeline output (147 sites) — API serves this
│ └── places.json # 89,518 places, 239 countries (6MB)
├── scripts/
│ ├── score_candidates.py # Offline/live scoring entry point (+ overwrite guard)
│ ├── check_model.py # model.json ↔ code drift guard
│ ├── robustness_report.py # Dirichlet sensitivity histogram
│ ├── healthcheck.py # External endpoint status
│ └── generate_mock.py # Builds places.json (skips if present)
├── notebooks/
│ └── validate_pipeline.ipynb # Pipeline validation notebook
├── requirements.txt
├── start.sh
└── README.md
Where critical logic lives:
| Logic | File | Function |
|---|---|---|
| Composite scoring formula | core/scoring.py |
compute_composite_score() |
| Manning's velocity | core/flow.py |
compute_flow_velocity() |
| Hard gate filtering | core/scoring.py |
apply_hard_gates() |
| Sensitivity analysis | core/scoring.py |
sensitivity_analysis() |
| Client-side placement | dashboard/explore/index.html |
generateCandidates() |
| Overpass waterway fetch | dashboard/explore/index.html |
fetchRealStreams() |
| Sub-score normalization | core/scoring.py |
compute_subscore() |
- User runs:
python -m core.pipeline --bbox "-79.05,35.90,-78.75,36.05" py3dep.get_map('DEM', bbox, resolution=10)fetches 3DEP raster- pysheds: fill_pits → fill_depressions → resolve_flats → flowdir → accumulation
extract_river_network(threshold=500)→ stream GeoJSONgenerate_candidates(spacing=200m)→ candidate points along streams- For each candidate: compute pixel coords, elevation, catchment area from DEM
- Output:
mock_data/candidates.geojson
- Browser opens
/explore(dashboard/explore/index.html, served by the API; the landing page at/is a separate marketing surface) - Fetches
places.json(6MB) → parses 89,518 places - Mapbox GL JS renders clustered city markers
- User clicks city →
openCity(idx)fires - POST to Overpass API → returns waterway geometry as JSON
fetchRealStreams()filters: tidal=no, width≤30m, top 15–25 by lengthgenerateCandidates()runs 3-phase placement + scoring- Mapbox renders: stream lines (cyan) + candidate dots (color-coded by score); where the city overlaps the committed live run's coverage (Durham NC), the real pipeline sites are drawn as dark-ringed dots alongside the demo estimates
- User clicks candidate → detail panel shows score breakdown + per-value provenance tags (EST / OSM / LIVE)
| Method | Path | Parameters | Response |
|---|---|---|---|
| GET | /api/candidates |
?min_score=N ?top_n=N ?subscore=field |
GeoJSON FeatureCollection (each feature carries its stable rank) |
| GET | /api/candidates/{id} |
{id} = the candidate's rank (1-based; 1 = top composite) |
Score breakdown with 4 sub-score parameter trees + rank/segment_id |
| GET | /api/weights |
— | Nominal design weights and effective weights for the served dataset (constant-fallback params show effective 0.0) |
| GET | /api/stats |
— | Count, score range, mean, top 5 |
| GET | /map |
— | Serves dashboard HTML |
| WS | /ws |
— | Scored-snapshot channel: sends the candidates GeoJSON on connect, answers ping, re-reads on refresh. No live pipeline producer pushes updates. |
| API | Auth | Rate Limit | Timeout | Fallback |
|---|---|---|---|---|
| Overpass API | None | Informal | 12s | Procedural generation |
| USGS NWIS | None | None published | 30s | Hardcoded Ellerbe Creek stats |
| EPA ECHO | None | None published | 30s | Empty GeoDataFrame |
| — | — | — | EJ index reconstructed from Census ACS instead (see §5.9) | |
| Census ACS | None | None published | 30s | Durham average (500/km²) |
| USGS 3DEP | None | None published | 60s | Fatal — no fallback |
Every external API call is wrapped in safe_call() with a default fallback value.
| Setting | Location | Default | Notes |
|---|---|---|---|
MAPBOX_TOKEN |
dashboard/config.js (gitignored) and .env |
Placeholder | Must replace — free at mapbox.com. Copy dashboard/config.example.js → dashboard/config.js, and .env.example → .env |
ELLERBE_BBOX |
core/__init__.py |
(-79.05, 35.90, -78.75, 36.05) |
Ellerbe Creek watershed |
ELLERBE_GAUGE |
core/__init__.py |
"02086849" |
USGS gauge site number |
UTM_CRS |
core/__init__.py |
"EPSG:32617" |
UTM zone 17N (Durham, NC) |
| DEM resolution | core/pipeline.py |
10m | Passed to py3dep |
| Accumulation threshold | core/pipeline.py |
500 cells | Stream extraction sensitivity |
| Candidate spacing | core/pipeline.py |
200m | Along-stream distance |
| Composite weights | core/scoring.py |
[0.30, 0.25, 0.30, 0.15] | Gen, Flow, Impact, Feas |
| Same-stream spacing (explorer) | dashboard/explore/index.html |
500m | Along one waterway (occlusion makes closer nets redundant) |
| Cross-stream spacing (explorer) | dashboard/explore/index.html |
300m | Haversine between nets on different waterways |
| Max width (explorer) | dashboard/explore/index.html |
30m | Skip channels wider |
| Risk percentile | dashboard/explore/index.html |
20–35% | Population-scaled |
| Min deploy floor | dashboard/explore/index.html |
5 | Always at least this many |
macOS only ships python3 (no bare python). The cleanest setup is a venv, which gives you a local python + pip that won't collide with system Python or other projects.
cd ~/Downloads/GRIME
# 1. Create + activate the venv (you'll see (.venv) in your prompt after activation)
python3 -m venv .venv
source .venv/bin/activate
# 2. Install web-server deps (FastAPI, uvicorn, websockets, etc.)
python -m pip install --upgrade pip
python -m pip install -r requirements.txt
# 3. Configure secrets
cp .env.example .env # then set MAPBOX_TOKEN=...
# Optional: only needed if opening dashboard HTML directly instead of through FastAPI
cp dashboard/config.example.js dashboard/config.js
# 4. Generate the places database (one-time)
python scripts/generate_mock.py
# 5. Start the API (this also serves the dashboard)
python -m uvicorn api.main:app --reload --port 8000Open http://localhost:8000/ — landing page. /explore for the map app. /api/swagger for API docs.
The repository includes a Procfile, .python-version, and app.json. Heroku
installs requirements.txt and starts the FastAPI app on its assigned $PORT.
Python 3.12 is selected because it is supported by Heroku and by the optional
scientific dependencies in requirements-full.txt.
# Creates the app and adds a "heroku" git remote.
heroku create --stack heroku-26
# Required for the map. Use a URL-restricted public Mapbox token.
heroku config:set MAPBOX_TOKEN=pk.your_public_token
# Optional: enables live Census-backed features.
heroku config:set CENSUS_API_KEY=your_key
# Deploy the current main branch.
git push heroku main
# Verify the dyno and open the app.
heroku ps
heroku open /healthz
heroku openIf deploying a branch other than main, use
git push heroku your-branch:main. Runtime logs are available with
heroku logs --tail.
Once setup is done, every subsequent run is just two commands:
cd ~/Downloads/GRIME
source .venv/bin/activate
python -m uvicorn api.main:app --reload --port 8000If you'd rather skip the venv, substitute python3 and python3 -m pip everywhere:
python3 -m pip install -r requirements.txt
python3 -m uvicorn api.main:app --reload --port 8000Use Python 3.9–3.12 for the full scientific/test stack; requirements-full.txt
currently pins NumPy below 2.0, whose wheels are not available for Python 3.13+.
For the full pipeline on Windows (rasterio, fiona, pysheds), use conda:
conda install -c conda-forge rasterio fiona geopandas pysheds
pip install -r requirements-full.txt| Error | Cause | Fix |
|---|---|---|
command not found: python |
macOS only has python3 |
Activate the venv, or use python3 |
No module named uvicorn |
Deps not installed in the active Python | python -m pip install -r requirements.txt (note python -m pip, not bare pip) |
| NumPy install fails on Python 3.13+ | Full stack pins numpy<2.0 |
Create the full-pipeline venv with Python 3.9–3.12 |
Mapbox token not configured (500 from /api/config) |
.env missing or MAPBOX_TOKEN= empty |
cp .env.example .env and fill in the token |
| Port 8000 in use | Old uvicorn still running | lsof -i :8000 to find PID, or pass --port 8001 |
Empty city list in /explore |
mock_data/places.json missing |
python scripts/generate_mock.py (skips when the committed file is present) |
/api/candidates returns 0 features |
mock_data/candidates.geojson missing |
Restore it from git — it is the frozen live-run dataset (regenerating needs the network-bound --live pipeline) |
- Open map → 89,518 places visible as clustered dots
- Search "Durham" → click → fly to Durham, NC
- Wait 3s → real Ellerbe Creek geometry appears with scored candidate dots
- Click top-ranked site → score breakdown panel shows 4 sub-scores
- Toggle "Waterways" → OSM overlay confirms alignment
- Toggle "Light theme" → clean beige presentation mode
- Explain the 27-parameter model and show weight sensitivity
python -m core.pipeline --bbox "-79.05,35.90,-78.75,36.05" --resolution 10 --threshold 500curl http://localhost:8000/api/candidates?top_n=10
curl http://localhost:8000/api/candidates/1 # {id} = the candidate's rank (1 = top site)
curl http://localhost:8000/api/weights # nominal + effective weights| Operation | Time | Bottleneck |
|---|---|---|
| Dashboard initial load | ~3s | Parsing ~6MB places.json |
| Overpass API query | 2–8s | Network + OSM server |
| Client-side placement + scoring | <100ms | Haversine collision checks |
| Python DEM pipeline (full) | 30–90s | py3dep HTTP fetch + pysheds |
| Python parameter computation | 3–5 min | Sequential EPA/USGS API calls |
No formal benchmarks have been run. Times above are observed during development.
| Failure | Impact | Mitigation |
|---|---|---|
| Overpass API down/slow | No real waterway data | Explorer distinguishes fetch failure (rate-limit/network → error panel with Retry + one automatic backoff retry) from a genuinely empty result ("no mapped rivers" panel); no procedural rivers |
| EPA/USGS API timeout | Missing parameter values | safe_call() returns fallback defaults; summarize_provenance flags the resulting constant columns each run |
| Census ACS needs a key | Population/EJ fall back | CENSUS_API_KEY (free) enables live values; otherwise neutral defaults |
| EPA EJSCREEN endpoint dead | EJ would be constant | EJ reconstructed from live Census ACS instead (C4, §5.9) |
| Mapbox token missing | Map blank | Fatal for dashboard; replace token |
| pysheds DEM fetch fails | No stream network | Fatal for Python pipeline; mock data works |
| MinMax on a constant column | sklearn maps it to 0 (not 0.5), silently deleting its weight | compute_subscore drops constant columns and renormalizes the surviving weights (logged); an all-constant family returns a neutral 50 |
- No authentication on the API — local development only
- Mapbox token is a publishable (
pk.*) client-side token. It lives indashboard/config.js(gitignored) and in.env(also gitignored). Usedashboard/config.example.jsand.env.exampleas templates. Restrict the token by URL in the Mapbox dashboard before deploying anywhere public. - Overpass queries use numeric interpolation only — no injection risk
- places.json contains only public geographic data, no PII
- External APIs are read-only; Census ACS uses the optional
CENSUS_API_KEY
Automated tests: pytest tests/ — 25 property and integration tests over the real scoring code:
composite ∈ [0, 100]; family weights sum to 1; compute_subscore drops a constant
column + renormalizes (and an all-constant family → neutral 50); hard gates remove
the right rows; Manning velocity is monotonic in slope; the velocity favorability
curve is peaked; occlusion (1−η)^k is non-increasing; estuary/beach distances are
decorrelated; the width-order fallback never trips the gate; the EJ area-weighting
varies across catchments; road density clips the cached OSM network to each
catchment; computed velocity reaches feasibility scoring; candidate detail lookups
are keyed on the stable rank (a re-sorted list can never remap them, verified on
the shipped dataset); /api/weights zeroes constant-fallback parameters and
renormalizes survivors; the offline scorer refuses to overwrite the live-provenance
dataset and writes to a separate path; the output provenance block reports the
actual varying/constant split; optimize_weights returns a valid simplex point.
Model drift guard: python3 scripts/check_model.py asserts the Python constants
match model.json: parameter + sub-score weights, the runoff formula, both velocity
curves (feasibility step and the Flow-side transport Gaussian), the channel-width
curve, hard-gate boundary semantics, pipeline/explorer spacing + occlusion constants,
and the width-order fallback.
CI: .github/workflows/ci.yml runs pytest + check_model.py on every push/PR.
Other validation:
notebooks/validate_pipeline.ipynb— pipeline verification + endpoint healthscripts/healthcheck.py— external endpoint connectivity (incl. the EJ source)scripts/score_candidates.py— offline (default) writes a deterministic synthetic-parameter dataset tomock_data/candidates_offline.geojson;--livere-runs the full network pipeline; overwriting the frozen live dataset requires--force- Dashboard visual QA — verify rivers align with satellite imagery
- Scores are relative, not absolute. MinMax normalization means scores across cities are not comparable.
- Client-side scoring is approximate. Dashboard uses heuristics; Python pipeline uses real API data.
- Weight values are heuristic. Not optimized against ground truth. Bayesian scaffold exists but hasn't been run.
- OSM coverage varies. Excellent in US/Europe, variable in developing nations.
- Channel width is estimated. <5% of OSM waterways have width tags. Estimated from type heuristic.
- No temporal modeling. Scores are static snapshots, not seasonal.
- US-centric data sources for the Python pipeline. Dashboard bypasses this with population-scaled heuristics globally.
- Spacing thresholds (500m same-stream / 300m cross-stream explorer, 200m pipeline) are heuristics — tuned by inspection, not engineering spec (no published interceptor-spacing guidance exists).
- The shipped Durham dataset is a frozen artifact. Regenerating it requires live federal APIs, a Census key, and DEM downloads (see §8 live-vs-fallback for what varied in each run).
- Ground-truth validation against actual trap deployment locations (Durham Stormwater Services)
- Bayesian weight optimization with real data
- Temporal scoring with USGS real-time discharge
- Computer vision trash detection from satellite/drone imagery
- Multi-city normalized scoring for global prioritization
- Cost modeling (deployment cost, maintenance frequency)
- Research publication: extend WaterGate paper, contact Wolfram Research authors
- University collaboration: Duke Environmental Engineering
- Chow, V.T. (1959). Open-Channel Hydraulics. McGraw-Hill.
- Leopold, L.B. (1964). Fluvial Processes in Geomorphology.
- Manning, R. (1891). On the flow of water in open channels and pipes.
Project story and build narrative: dashboard/docs/Overview.md. Full technical documentation with figures: dashboard/docs/documentation.md.
| Term | Definition |
|---|---|
| Candidate site | A point on a waterway evaluated for trap deployment |
| Composite score | Final weighted combination of 4 sub-scores (0–100) |
| CSO | Combined Sewer Overflow |
| D8 | Deterministic 8-direction flow algorithm |
| DEM | Digital Elevation Model |
| EJ | Environmental Justice |
| Hard gate | Binary feasibility check that eliminates a candidate |
| NPDES | National Pollutant Discharge Elimination System |
| NBI | National Bridge Inventory |
| Stream order | Confluence-degree heuristic (H3): order 1 = headwater/leaf, higher = more junctions. Not true Strahler. |
| Sub-score | Intermediate score (0–100) for one parameter family |
| TRI | Toxic Release Inventory |
MIT.
GRIME · 27 parameters · 6 families · 89,518 places · 239 countries





