Every forecast the system makes is recorded and later graded against what the river actually did — this page is that report card. It currently scores 120 physics crest predictions, 7440 empirical likelihood forecasts, and 770 recession countdowns (9 still maturing). Updated daily; treat it as experimental.
Predicts the crest a gauge will reach from rainfall + antecedent moisture. Scored on how close the predicted peak was to the actual peak.
| Predictor | Scored | Avg error (MAE) | Bias (mean / median) | Within ±20% |
|---|---|---|---|---|
| Cossatot River | 68 | 0.78 ft | 0.72 / 0.29 ft | 63.0% |
| Richland Creek | 32 | 0.87 ft | 0.87 / 0.49 ft | 16.0% |
| Hailstone (upper Buffalo) | 20 | 83.0 cfs | 83.0 / 62.0 cfs | 5.0% |
Bias is the mean / median signed error (predicted − actual). A large gap between them means a few outlier events — often a single flash-flood onset — are dragging the mean; the median is the typical miss.
Cossatot River
Richland Creek
Hailstone (upper Buffalo)
Predicts the likelihood of a rise to a given level from recent rainfall. Hit-rate = the river rose to at least the called level; no-rise-rate = the forecast rise never happened (lower is better).
| Basin | Forecasts | Hit-rate rose to ≥ called level | No-rise rate rise never came | Outcome detail |
|---|---|---|---|---|
| Big Piney Creek | 1062 | 5.6% | 94.1% | 52 exact · 8 higher · 3 short · 999 none |
| Buffalo at Boxley | 1055 | 0.0% | 100.0% | 0 exact · 0 higher · 0 short · 1055 none |
| Cossatot River | 1073 | 1.9% | 97.9% | 6 exact · 14 higher · 2 short · 1051 none |
| Hailstone (upper Buffalo) | 1055 | 0.0% | 100.0% | 0 exact · 0 higher · 0 short · 1055 none |
| Illinois River (Hwy 16) | 1057 | 10.6% | 89.1% | 112 exact · 0 higher · 3 short · 942 none |
| Mulberry River | 1067 | 9.7% | 89.5% | 104 exact · 0 higher · 8 short · 955 none |
| Richland Creek | 1071 | 0.0% | 100.0% | 0 exact · 0 higher · 0 short · 1071 none |
| Band | n | Hit-rate | No-rise rate |
|---|---|---|---|
| below_p25 | 6157 | 1.6% | 98.4% |
| p25_to_p50 | 277 | 18.4% | 81.6% |
| p50_to_p75 | 110 | 22.7% | 77.3% |
| above_p75 | 545 | 21.5% | 78.5% |
A well-calibrated engine would show hit-rate rising with the band. Where it does not, the engine is over-warning — a known selection-bias issue tracked for recalibration.
Predicts how long until a falling river drops to each threshold. HIT% = reached the level near the predicted time; MAE = average timing error (hours). Newly launched — most predictions are still maturing.
| Gauge | Target | Graded | HIT% | Never reached | Timing MAE | Bias |
|---|---|---|---|---|---|---|
| big_piney/below_longpool | low_floatable | 11 | 100.0% | 0.0% | 8.7h | 8.7h |
| big_piney/below_longpool | too_low | 75 | 100.0% | 0.0% | 10.9h | 7.1h |
| buffalo/harriet | low_floatable | 394 | 40.0% | 60.0% | 13.5h | -12.0h |
| buffalo/pruitt | too_low | 21 | 100.0% | 0.0% | 17.9h | -3.2h |
| buffalo/st_joe | low_floatable | 203 | 75.0% | 25.0% | 12.7h | -7.9h |
| cossatot/cossatot | low_floatable | 4 | 100.0% | 0.0% | 4.9h | 4.9h |
| cossatot/cossatot | too_low | 16 | 100.0% | 0.0% | 10.5h | 10.5h |
| mulberry/above_hwy_23 | too_low | 10 | 100.0% | 0.0% | 8.3h | -2.7h |
| mulberry/below_hwy_23 | too_low | 36 | 92.0% | 8.0% | 11.0h | 2.8h |
When rain arms it, predicts how big the Ponca gauge will get — a class (Fizzle/Moderate/High/Flood), a typical-peak band, and a flood-risk %. Graded against Ponca's actual crest over the 36 h after each call.
| Calls graded | Events | Class exact | Class within 1 | Peak in IQR | Peak error (MAE) |
|---|---|---|---|---|---|
| 623 | 19 | 89% | 99% | 83% | 446.0 cfs |
Flood-risk calibration (Brier score, lower is better; 585 calls since the raw analog was also logged): override floor 0.062 vs. raw k-NN 0.036 — the raw analog is currently the better-calibrated of the two — the deterministic override floor stays high through the post-crest recession. Small sample, directional only.
| Event (UTC) | Calls | Predicted class | Max flood-risk | Actual crest | Actual class |
|---|---|---|---|---|---|
| 2026-08-08T01:02 | 80 | Fizzle | 11% | 78 cfs | Fizzle |
| 2026-08-07T01:02 | 56 | Fizzle | 17% | 78 cfs | Fizzle |
| 2026-07-29T22:02 | 52 | Fizzle | 6% | 30 cfs | Fizzle |
| 2026-07-15T02:02 | 120 | Fizzle | 17% | 65 cfs | Fizzle |
| 2026-07-11T20:02 | 82 | Fizzle | 17% | 71 cfs | Fizzle |
| 2026-07-01T20:00 | 2 | Flood | 85% | 112 cfs | Fizzle |
For each Buffalo mainstem gauge, predicts a coming rise (slight / moderate / large) from local rain + upstream propagation, with a timing window. Graded on whether the gauge actually rose, and within the predicted window. Newly recording — predictions only fire during rain events, so this fills in over time.
| Gauge | Graded | Rise happened | On-time | No-rise (false alarm) |
|---|---|---|---|---|
| boxley | 0 | — | — | 0 |
| ponca | 76 | 25% | 89% | 57 |
| pruitt | 66 | 23% | 53% | 51 |
| st_joe | 123 | 33% | 2% | 83 |
| harriet | 95 | 23% | 9% | 73 |
Typical actual rise by predicted category: slight: ~347.0 cfs (n=45), moderate: ~373.6 cfs (n=50), large: ~31.4 cfs (n=1). (Categories come from rainfall, so this is how they map to real gauge rises — calibration that accrues over time.)
| When (UTC) | Gauge | Predicted | Window | Outcome | Actual rise |
|---|---|---|---|---|---|
| 2026-08-10T00:09 | harriet | slight | 0.0-0.8h | rose | +192.0 cfs @ 2.6h |
| 2026-08-09T23:24 | harriet | slight | 0.0-1.5h | rose | +340.0 cfs @ 3.3h |
| 2026-08-09T22:24 | harriet | slight | 0.0-2.5h | rose | +358.0 cfs @ 4.3h |
| 2026-08-09T21:24 | harriet | slight | 0.3-3.5h | rose | +358.0 cfs @ 5.3h |
| 2026-08-09T20:24 | harriet | slight | 1.3-4.5h | rose | +358.0 cfs @ 6.3h |
| 2026-08-09T19:24 | harriet | slight | 2.3-5.5h | rose | +354.0 cfs @ 7.3h |
When Ponca rises, predicts whether and how big a bump reaches Pruitt and St. Joe, and how many hours after Ponca peaks. Graded on the bump/no-bump call, the crest size, and the timing. Newly recording — only logs during a Ponca rise.
| Reach | Graded | Bump call right | Crest size hit | Timing hit |
|---|---|---|---|---|
| pruitt | 4 | 25% | 0% | 0% |
| st_joe | 4 | 25% | 0% | 0% |
When an upstream gauge crests, the wave tracker predicts the crest size and arrival window at each downstream gauge. Every confirmed wave is graded after its window closes: did a real rise arrive, was the crest inside the predicted band, and did it arrive inside the window. Waves are rare events — this fills in slowly.
Waves graded: 1 (significant: 1); crest inside the likely band 1/1; inside the wide band 1/1; arrived inside the window 0/1; median predicted/actual crest 1.0×. Alerts issued: 0; floods with no alert: 0.
| Wave | Upstream crest | Predicted band | Outcome | Actual peak | Lag (h) |
|---|---|---|---|---|---|
| st_joe → harriet | 445.0 cfs | 430–496 cfs (med 467) | significant | 467.0 cfs | 17.5 |
A neural network projects the level hour by hour from 48 h of 1 km radar rainfall (Ponca 6 h ahead, Boxley 4 h). Every hourly issuance is graded against what the gauge actually did: MAE = average error of the median; IN-BAND% = how often the likely range contained the truth (target ~50%); SKILL = error reduction vs assuming the level doesn't change, on hours where it moved.
| Ahead | Graded | MAE (cfs) | In-band | Skill vs no-change |
|---|---|---|---|---|
| +1 hr | 815 | 1.2 | 79.9% | 1.7% |
| +2 hr | 815 | 1.5 | 58.2% | 3.2% |
| +3 hr | 815 | 1.9 | 72.1% | 7.4% |
| +4 hr | 815 | 2.5 | 54.6% | 10.6% |
| +5 hr | 815 | 2.7 | 47.6% | 10.1% |
| +6 hr | 815 | 3.4 | 40.7% | 11.5% |
High-water crossings in scored windows: 0 of 815 (model mean probability 1.8%); flood: 0 of 815 (0.4%).
| Ahead | Graded | MAE (cfs) | In-band | Skill vs no-change |
|---|---|---|---|---|
| +1 hr | 811 | 0.4 | 59.6% | 0.1% |
| +2 hr | 810 | 0.5 | 61.9% | -2.5% |
| +3 hr | 809 | 0.7 | 58.5% | -1.7% |
| +4 hr | 808 | 0.9 | 70.7% | -2.1% |
High-water crossings in scored windows: 0 of 811 (model mean probability 0.4%); flood: 0 of 811 (0.2%).
The Cossatot pixel net projects the level hour by hour from 48 h of 1 km radar rainfall (6 h ahead); crossing probabilities come from its companion gradient-boosted ensemble. Graded the same way as the Buffalo net: MAE = average error of the median; IN-BAND% = how often the likely range contained the truth (target ~50%); SKILL = error reduction vs assuming the level doesn't change, on hours where it moved.
| Ahead | Graded | MAE (cfs) | In-band | Skill vs no-change |
|---|---|---|---|---|
| +1 hr | 755 | 1.1 | 67.2% | 4.0% |
| +2 hr | 755 | 2.0 | 64.0% | 10.2% |
| +3 hr | 755 | 3.0 | 67.9% | 11.1% |
| +4 hr | 755 | 3.8 | 74.4% | 11.8% |
| +5 hr | 755 | 5.0 | 66.4% | 7.5% |
| +6 hr | 755 | 5.6 | 75.5% | 10.0% |
Floatable (240 cfs) crossings in scored windows: 6 of 748 (model mean probability 1.0%); high water (1,270 cfs): 0 of 755 (0.2%).
What the scorecard grades today, and what is still being wired into the loop:
| Predictor family | Status | Notes |
|---|---|---|
| Physics rise predictors | graded | Cossatot, Richland, Hailstone — predicted crest height/flow vs. the actual peak. |
| Empirical forecast engine | graded | 6 basins — 'likelihood of rise to tier X' vs. the tier the gauge actually reached. |
| Recession countdowns | graded | 9 gauges — maturing; the longest horizons settle ~7-10 days after they're issued. |
| Ponca AI Rainfall Event Analysis | graded | Each armed call graded vs Ponca's actual crest over the next 36 h (ponca_analog_eval.py). |
| Downstream propagation forecast | graded | Ponca -> Pruitt -> St. Joe; predictions now logged + graded on bump/magnitude/timing (propagation_eval.py). |
| Buffalo per-gauge rise predictions | graded | Per-gauge rise nowcasts recorded + graded vs the gauge's actual rise (buffalo_predictions_archive.py). |
| Routed waves (wave_router) | graded | F4 tracked waves — each confirmed wave-target's frozen crest band + arrival window graded vs what the target gauge actually did (wave_eval.py). |
| Buffalo flood_risk labels | not yet logged | Lower-priority companion in buffalo_output; still ungraded. |
| Neural Net Predictions (pixel-v5) | graded | Ponca 6 h / Boxley 4 h hour-by-hour level tables — every hourly issuance graded against the observed level at each horizon (neural_eval.py). |
| Neural Net Predictions (Cossatot) | graded | Cossatot 6 h level tables + GBM-ensemble crossing probs — every hourly issuance graded at each horizon (neural_cossatot_eval.py). |