← Back to Home

Prediction Scorecard

How accurate Creek Intelligence's forecasts have been lately. Experimental — self-grading, updated daily.
Updated 2026-10-02 06:30 UTC · scoring window: last 45 days

Every forecast the system makes is recorded and later graded against what the river actually did — this page is that report card. It currently scores 294 physics crest predictions, 4370 empirical likelihood forecasts, and 0 recession countdowns (0 still maturing). Updated daily; treat it as experimental.

Where the truth comes from: USGS keeps revising its provisional gauge readings for months. The gauge archives these grades are checked against were last re-pulled from USGS on 2026-10-02 05:20 UTC (trailing 120 days, with approval status), and every settled grade whose readings are not yet USGS-approved is re-graded nightly; a grade that moves keeps its first verdict on the record. Last pass (2026-10-02 06:05 UTC): rise: 491 checked, 0 verdicts moved last night, 183 re-graded so far, 161 final; recession: 727 checked, 0 verdicts moved last night, 806 re-graded so far, 1145 final; physics: 1905 checked, 0 verdicts moved last night, 107 re-graded so far, 192 final; ponca: 968 checked, 0 verdicts moved last night, 678 re-graded so far, 0 final; empirical: 4827 checked, 0 verdicts moved last night, 2 re-graded so far, 0 final. USGS is retiring the data service these gauge feeds use; its replacement has been running alongside it since 2026-09-04: 663 hourly checks, 139204 readings compared between the two services with 0 value differences, and 117553 readings compared against what the live feeds stored with 10 differences. 569 readings were seen by the new service before the old one and 27 the other way round (timing, not disagreement). No gauge series is currently returning empty.

⚠️ Needs attention

Physics Rise Predictors

Predicts the crest a gauge will reach from rainfall + antecedent moisture. Since 2026-09-08 every hourly run is a claim — a rise, or no rise — and both are graded: detection (rises the engine called) and false alarms side by side, then how close the crest and its timing were when a called rise happened.

Sample: 294 records · 5 events · 42 on the new grading · last 45 days · truth as of 2026-10-02 05:20 UTC · events counted on v2 records only (fixed 18-h horizon, from 2026-09-04)

New grading from 2026-09-04: each prediction is checked over a fixed 18-hour window after it was issued (before, the window closed 2 h after the predicted peak, so a later crest was never seen). It is graded on three separate questions: did the gauge rise at all, how far off was the crest when it did, and how far off was the timing. Predictions whose crest was not above the current level are no longer archived as claims.

PredictorRise predictions
records / events
Gauge rose
records / events
Crest error given a riseTiming given a rise
median |err| / signed
Within ±20% (rose)
Cossatot River18 / 272.0% / 100.0%MAE 0.18 ft (median +0.11 ft)
in flow: typically ×1.53 (n=13)
6.7 h / 6.7 hn/a for stage
Richland Creek24 / 375.0% / 67.0%MAE 0.55 ft (median +0.08 ft)
in flow: typically ×2.29 (n=16)
4.6 h / 2.9 hn/a for stage

Claims (from 2026-09-08): every hourly run is archived as a claim — a rise, or no rise (quiet / rain too light) — and graded 18 h later against the gauge. Detection = rises the engine called, out of every rise that happened (a no-rise claim followed by a rise is a miss). False alarms = rise predictions the gauge did not follow. Events = episodes of consecutive hourly runs.

PredictorRise claimsNo-rise claims
graded / issued
Detection (POD)
records / events
False alarms (FAR)
records / events
Missed risesNo-rise verified
Cossatot River1888 / 55462.0% / 100.0%28.0% / 0.0%891.0%
Richland Creek2443 / 55086.0% / 100.0%25.0% / 33.0%393.0%

Rise calls by confidence (from 2026-09-09): three copies of the physics engine run every hour, from cautious to eager. The eager one is served, and each rise call is labelled by the most cautious copy that also calls a rise — HIGH (all three agree), MEDIUM (medium and eager), LOW (eager only). Quoted is the verification rate the card showed for that label (leave-one-year-out replay, 2014–2026, June–September figure in summer); Live is what the gauge has done since.

PredictorLabelRise calls
issued / graded
Verified / false alarmsLiveQuoted
Cossatot RiverHIGH16 / 1613 / 381.0%81%
Cossatot RiverMEDIUM1 / 10 / 10.0%46%
Cossatot RiverLOW1 / 10 / 10.0%35%
Richland CreekHIGH18 / 1818 / 0100.0%69%
Richland CreekMEDIUM3 / 30 / 30.0%34%
Richland CreekLOW1 / 10 / 10.0%22%
Legacy grading (window closed 2 h after the predicted peak; records issued before 2026-09-04)
PredictorScoredAvg error (MAE)Bias (mean / median)Within ±20%
Cossatot River1190.09 ft0.04 / 0.0 ft93.0% (n=119)
Richland Creek790.29 ft0.09 / 0.01 ft62.0% (n=79)
Hailstone (upper Buffalo)962.1 cfs1.15 / -0.4 cfs50.0% (n=96)

Bias is the mean / median signed error (predicted − actual). A large gap between them means a few outlier events — often a single flash-flood onset — are dragging the mean; the median is the typical miss.

Worst recent misses

Cossatot River

Richland Creek

Hailstone (upper Buffalo)

Empirical Forecast Engine

Each hourly forecast is a calibrated probability that the gauge rises to the next level within the promised window (12 h on the fast creeks, 30 h on the slow ones). Scored as probabilities: Brier = mean squared error (0 is perfect); skill (BSS) = improvement over always forecasting the long-run rate (>0 is skill); obs/forecast = how many rises happened per rise forecast (1.0 is calibrated); reliability = what actually happened when the engine said X%. Probabilistic scoring began 2026-09-02; earlier records are not comparable and are not shown.

Sample: 4370 records · 8 events · 77 claims / 4293 no-rise calls · last 45 days · truth as of 2026-10-02 05:20 UTC · events = claim episodes (issuances >= 8 % within 3 h of each other); non-claims do not form episodes

"No rise indicated" calls (forecast below 8%) are graded separately on their miss rate: 4293 such hourly calls in the window, 0.0% of which were followed by a rise to the target level. Real claims (8% and above): 77, of which 0 verified.

BasinForecasts
hourly / episodes
RisesMean forecastObservedBrier
lower is better
Skill (BSS)
>0 beats climatology
Obs / forecast
1.0 = calibrated
AUC
Big Piney Creek662 / 102.0%0.0%0.00316— (no rises in window)——
Buffalo at Boxley702 / 100.1%0.0%5e-05— (no rises in window)——
Cossatot River702 / 200.7%0.0%0.00185— (no rises in window)——
Hailstone (upper Buffalo)702 / 000.1%0.0%1e-05— (no rises in window)——
Illinois River (Hwy 16)219 / 324.6%0.9%0.01235— (no rises in window)——
Mulberry River685 / 000.4%0.0%0.00015— (no rises in window)——
Richland Creek698 / 100.2%0.0%0.00019— (no rises in window)——

All basins pooled: 4370 forecasts, 2 rises, Brier 0.00146, skill None, obs/forecast None, AUC None.

Reliability (when the engine said X%, how often did the rise come?)

Forecast binnMean forecastObserved
0-2%40760.2%0.0%
2-7%2175.0%0.9%
7-15%4512.4%0.0%
15-25%419.2%0.0%
25-40%2228.4%0.0%
40-60%645.8%0.0%
Per storm episode (highest probability issued)
Max forecast binEpisodesMean forecastRose
7-15%310.0%0.0%
15-25%318.1%0.0%
25-40%132.8%0.0%
40-60%145.8%0.0%
By target level
LevelnRisesBrierObs / forecastAUC
LOW415100.00088NoneNone
MEDIUM419520.0008NoneNone
HIGH437000.00011NoneNone

Recession Countdowns

Predicts how long until a falling river drops to each threshold. HIT% = reached the level near the predicted time; MAE = average timing error (hours). Newly launched — most predictions are still maturing.

Sample: 0 records · 0 events · last 45 days · truth as of 2026-10-02 05:20 UTC · events = recession episodes per gauge/target

All 0 recession predictions are still maturing — a countdown can only be graded once enough time has passed to see whether the river reached the level. First results expected early July 2026.

Ponca AI Rainfall Event Analysis

When rain arms it, predicts how big the Ponca gauge will get — a class (Fizzle/Moderate/High/Flood), a typical-peak band, and a flood-risk %. Scored per armed EVENT since 2026-09-02: the first call and the most-informed ('mature', most rain so far) call of each event are graded against the crest the event actually reached; the flood-risk % is scored as a probability (Brier, 0 is perfect; 'shown' is the number on the card, 'raw' the analog match alone). After the crest the card hides the % unless a second, higher crest is coming, so post-peak calls are graded only on whether a higher crest came. Events, not 15-minute re-issues, are the sample size. The class words are the analog library's own storm-outcome groups (an internal grouping the card never shows), not the Ponca levels on the gauge card.

Sample: 348 records · 6 events · last 45 days · truth as of 2026-10-02 05:20 UTC

Last 45 days — 6 armed event(s), 348 fifteen-minute calls, 0 flood(s).

EventsClass right (first call)Class right (mature call)Within one classCrest in likely bandFlood-risk Brier: shown / raw
6100%100%100%0%0.0 / 0.0 (n=6)

All time since 2026-06-08 — 16 event(s), 2 flood(s).

EventsClass right (first call)Class right (mature call)Within one classCrest in likely bandFlood-risk Brier: shown / raw
1688%94%100%20%0.0155 / 0.0155 (n=14)

Reliability of the shown flood-risk % on rising calls (348 fifteen-minute rows — rows within one event are near-duplicates, so this shows calibration shape, not sample size): 0-10%: said 1% → 0% flooded (n=348).

Recent events — first call, mature call, actual

Event (UTC)CallsFirst callMature callActual crestActual class
2026-09-27T11:0044Fizzle · 0% flood @ 0.26″, 2 cfsFizzle · 0% flood @ 0.33″, 2 cfs2 cfsFizzle
2026-09-21T06:0056Fizzle · 0% flood @ 0.33″, 2 cfsFizzle · 0% flood @ 0.55″, 2 cfs2 cfsFizzle
2026-09-12T17:0048Fizzle · 0% flood @ 0.33″, 2 cfsFizzle · 0% flood @ 0.33″, 2 cfs2 cfsFizzle
2026-09-10T19:0084Fizzle · 0% flood @ 0.28″, 2 cfsFizzle · 0% flood @ 1.44″, 2 cfs2 cfsFizzle
2026-08-26T22:0256Fizzle · 0% flood @ 0.45″, 3 cfsFizzle · 0% flood @ 0.9″, 10 cfs10 cfsFizzle
2026-08-24T07:0260Fizzle · 0% flood @ 0.3″, 0 cfsFizzle · 0% flood @ 0.6″, 3 cfs3 cfsFizzle
2026-08-08T01:0280Fizzle · 11% flood @ 0.63″, 42 cfsFizzle · 0% flood @ 1.66″, 78 cfs33 cfsFizzle
2026-08-07T01:0256Fizzle · 0% flood @ 1.19″, 9 cfsFizzle · 17% flood @ 2.34″, 45 cfs33 cfsFizzle

Radar Nowcast (Ponca card)

While rain is falling, the Ponca card reads the last 40 minutes of weather radar, moves the picture forward, and says roughly how much more rain lands on the creeks above Ponca in the next two hours (or that the rain is about done). Each radar frame's estimate is graded against the rain the next two hours actually delivered, using the card's own hourly rain accounting. Bias = estimate minus actual (negative means radar under-called); 'more coming' = the card said 0.1 in or more was on the way; 'about done' = it said little was left. The 'past storms like this went on to a real rise' perspective line is graded against whether the event rose at all.

Last 45 days — 328 radar frames graded across 7 rain event(s) (42 not gradable).

Amount claimFramesError (MAE)BiasActual ÷ estimate (median)
All frames3280.027″0.006″0.76
When it said ≥0.1″ was coming580.104″0.063″0.76
CallTimes madeVerifiedCaught
“More rain lined up” (≥0.1″ coming)5876%80%
“About done” (little left)25697%—

Verified = the call was right when made; caught = share of the real ≥0.1″ two-hour periods the card had called in advance.

Estimate saidFramesMean estimateMean actual≥0.1″ actually fell
0″2040.0″0.005″1%
0-0.1″660.035″0.045″12%
0.1-0.25″400.16″0.137″68%
0.25-0.5″140.317″0.255″93%
>=0.5″40.76″0.29″100%

By storm motion (estimates ≥0.1″): <8 mph (near-stationary): n=14, bias 0.056″, actual÷est 1.21; 8-20 mph: n=31, bias 0.05″, actual÷est 0.85; >=20 mph: n=13, bias 0.102″, actual÷est 0.57. Slow-moving storms grow over the watershed and deliver more than radar advection suggests; fast movers pass through and deliver less.

“Past storms like this went on to a real rise” line (2 lines in 2 events): “about half the time” (51%): said 1 times, the event rose 0% of them; “more often than not” (61%): said 1 times, the event rose 0% of them.

All time since 2026-07-11 — 706 radar frames graded across 12 rain event(s) (42 not gradable).

Amount claimFramesError (MAE)BiasActual ÷ estimate (median)
All frames7060.026″0.002″0.72
When it said ≥0.1″ was coming920.12″0.067″0.74
CallTimes madeVerifiedCaught
“More rain lined up” (≥0.1″ coming)9274%72%
“About done” (little left)58098%—

Verified = the call was right when made; caught = share of the real ≥0.1″ two-hour periods the card had called in advance.

Estimate saidFramesMean estimateMean actual≥0.1″ actually fell
0″4720.0″0.004″1%
0-0.1″1420.036″0.054″15%
0.1-0.25″580.159″0.144″67%
0.25-0.5″240.323″0.236″79%
>=0.5″100.71″0.384″100%

By storm motion (estimates ≥0.1″): <8 mph (near-stationary): n=20, bias 0.067″, actual÷est 0.95; 8-20 mph: n=50, bias 0.062″, actual÷est 0.85; >=20 mph: n=22, bias 0.081″, actual÷est 0.39. Slow-moving storms grow over the watershed and deliver more than radar advection suggests; fast movers pass through and deliver less.

“Past storms like this went on to a real rise” line (24 lines in 6 events): “more often than not” (63%): said 20 times, the event rose 0% of them; “about half the time” (53%): said 4 times, the event rose 0% of them.

Buffalo Rise Engine (per-gauge nowcast)

For each Buffalo mainstem gauge, predicts a coming rise (slight / moderate / large) from local rain + upstream propagation, with a timing window. Graded on whether the gauge actually rose, and within the predicted window. Predictions issued in the last 45 days, with the all-time totals beside them; predictions only fire during rain events, so this fills in over time.

Sample: 235 records · 11 events · last 45 days · truth as of 2026-10-02 05:20 UTC

Predictions issued in the last 45 days: 235 graded records, 11 events (4 rose). All-time: 652 graded records, 48 events (23 rose, 48%).

GaugeGraded recordsRise happened (records)On-timeNo-rise recordsEventsRise happened (events)Band held (rose)Chance skill (Brier)
boxley230%—233 (0 rose)0%—0.197 (n=23)
ponca240%—242 (0 rose)0%—0.404 (n=11)
pruitt270%—272 (0 rose)0%—0.061 (n=26)
st_joe919%0%832 (2 rose)100%0% (n=4)0.105 (n=55)
harriet7011%0%622 (2 rose)100%33% (n=6)0.085 (n=37)

Since 2026-09-03 the local-rain layer serves a rise probability and a typical rise band (25th–75th percentile, cfs) instead of a size word. "Band held" = the share of real rises that landed inside the band (target about 50%); Brier = mean squared error of the probability (0 is perfect, 0.25 is a coin flip).

A "rise" here means peak >= 1.25 x v0 and peak - v0 >= clamp(0.10 x v0, 10, 30) cfs — the same rule the engine uses to decide a predicted rise has arrived. Records graded before 2026-09-03 used a fixed +30 cfs floor, which read some real low-water rises as misses; the 45-day window mixes both vintages until mid-October. An event = a run of records less than 3 h apart.

Typical actual rise by predicted category: slight: ~28.1 cfs (n=13), moderate: ~28.0 cfs (n=3). (Categories come from rainfall, so this is how they map to real gauge rises — calibration that accrues over time.)

Recent rise predictions — predicted vs. actual

When (UTC)GaugePredictedWindowOutcomeActual rise
2026-09-21T20:54pruittslight4.0-18.3hno_rise—
2026-09-21T19:54pruittslight5.0-19.3hno_rise—
2026-09-21T18:54pruittslight6.0-20.3hno_rise—
2026-09-21T17:54pruittslight7.0-21.3hno_rise—
2026-09-21T16:54pruittslight8.0-22.3hno_rise—
2026-09-21T15:54pruittslight9.0-23.3hno_rise—

Downstream Propagation (Ponca → Pruitt → St. Joe)

Retired 2026-09-02. The Ponca card's own downstream model duplicated the wave tracker's reaches with weaker validation, so the card now narrates the wave tracker's bands and arrival windows for Pruitt, St. Joe and Harriet, and those are graded in the Routed Waves section below. The four events this section had graded included two synthetic host tests, which have been removed.

Retired — graded under Routed Waves below.

Routed Waves (wave tracker)

When an upstream gauge crests, the wave tracker predicts the crest size and arrival window at each downstream gauge. Every confirmed wave is graded after its window closes: did a real rise arrive, was the crest inside the predicted band, and did it arrive inside the window. Waves are rare events — this fills in slowly.

Sample: 1 records · 1 events · last 90 days · truth as of 2026-10-02 05:20 UTC · one grade per confirmed wave-target

Waves graded: 1 (significant: 1); crest inside the likely band 1/1; inside the wide band 1/1; arrived inside the window 0/1; median predicted/actual crest 1.0×. Alerts issued: 0; floods with no alert: 0.

WaveUpstream crestPredicted bandOutcomeActual peakLag (h)
st_joe → harriet445.0 cfs430–496 cfs (med 467)significant467.0 cfs17.5

Watershed Alerts (Q-bucket triggers)

The rain-triggered WATCH/WARNING words on the calibrated creeks (Upper Buffalo, Richland Main, Upper Cossatot, Upper Big Piney, Mulberry), graded against the creek's own gauge: an alert verifies when the creek reached its floatable floor within 24 h; a rise with rain and no alert is a miss. Same recipe as the leave-one-year-out numbers on the Watersheds page.

Sample: 31090 records · 6 events · last 90 days · truth as of 2026-10-02 05:20 UTC · records = 15-min Tier-1 decisions; events = alert episodes (+ rises with no alert)

Alert episodes: 6 (verified 1, false alarms 5, still open 0); rises with no alert: 1; median lead from first WARNING 3.2 h.

DrainageStartedWordStart bucketRain / triggerOutcomeCreek peakLead (h)
Upper Buffalo2026-09-11 14:17ZWATCHBONE-DRY1.63" / 2.10"false alarm2 / 500 cfs—
Mulberry (parent)2026-09-11 11:17ZWARNINGBONE-DRY1.39" / 1.25"false alarm4 / 387 cfs—
Upper Big Piney2026-09-11 09:32ZWARNINGBONE-DRY2.11" / 1.65"false alarm17 / 434 cfs—
Richland Main2026-09-11 08:17ZWARNINGBONE-DRY2.38" / 2.10"false alarm25 / 500 cfs—
Upper Cossatot2026-08-08 01:17ZWATCHBONE-DRY1.26" / 1.30"false alarm20 / 200 cfs—
Upper Cossatot2026-07-12 20:17ZWARNINGDRY1.52" / 1.20"verified253 / 200 cfs3.2

Rises with no alert: Upper Cossatot 2026-07-15 18:17Z (rain 0.77" of 1.20")

Neural Net Predictions (Ponca & Boxley)

A neural network projects the level hour by hour from 48 h of 1 km radar rainfall (Ponca 6 h ahead, Boxley 4 h). Every hourly issuance is graded against what the gauge actually did: MAE = average error of the median; IN-BAND% = how often the likely range contained the truth (target ~50%); SKILL = error reduction vs assuming the level doesn't change, on hours where it moved.

5 hour(s) in the window had no gauge data (USGS down): 5 re-served the previous projection with a note, the rest were offline.

558 issuances graded (model pixel-v5 / engine v3, 2026-09-08 → 2026-10-01).

Ponca (6 h)

AheadGradedMAE net / no-change (cfs)Skill, all hoursSkill, moving hoursIn-band calm / moving
+1 hr5580.01 / 0.0— (all 558 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+2 hr5580.02 / 0.0— (all 558 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+3 hr5580.04 / 0.0— (all 558 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+4 hr5580.08 / 0.01— (all 558 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+5 hr5580.1 / 0.01— (all 558 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+6 hr5580.13 / 0.01— (all 558 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —

800 cfs (Ponca: experienced only): 0 of 558 scored windows crossed (mean served probability 0.1%). Flood (1,600 cfs): 0 of 558 scored windows crossed (mean served probability 0.1%).

Boxley (4 h)

AheadGradedMAE net / no-change (cfs)Skill, all hoursSkill, moving hoursIn-band calm / moving
+1 hr5580.04 / 0.04— (all 558 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+2 hr5580.07 / 0.06— (all 558 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+3 hr5580.11 / 0.08— (all 558 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+4 hr5580.15 / 0.1— (all 558 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —

800 cfs (Ponca: experienced only): 0 of 558 scored windows crossed (mean served probability 0.8%). Flood (1,600 cfs): 0 of 558 scored windows crossed (mean served probability 0.2%).

Earlier model: 119 issuances graded (model pixel-v5 / engine v2, 2026-09-03 → 2026-09-08).

Ponca (6 h)

AheadGradedMAE net / no-change (cfs)Skill, all hoursSkill, moving hoursIn-band calm / moving
+1 hr1190.01 / 0.0— (all 119 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+2 hr1190.01 / 0.01— (all 119 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+3 hr1190.01 / 0.01— (all 119 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+4 hr1190.02 / 0.01— (all 119 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+5 hr1190.02 / 0.01— (all 119 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+6 hr1190.03 / 0.01— (all 119 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —

Boxley (4 h)

AheadGradedMAE net / no-change (cfs)Skill, all hoursSkill, moving hoursIn-band calm / moving
+1 hr1190.01 / 0.01— (all 119 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+2 hr1190.02 / 0.02— (all 119 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+3 hr1190.03 / 0.03— (all 119 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+4 hr1190.03 / 0.03— (all 119 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
Earlier model: 379 issuances graded (model pixel-v1 (pre-F5, untagged) / engine v1, 2026-08-18 → 2026-09-03).

Ponca (6 h)

AheadGradedMAE net / no-change (cfs)Skill, all hoursSkill, moving hoursIn-band calm / moving
+1 hr3790.21 / 0.21— (all 379 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+2 hr3790.24 / 0.26— (all 379 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+3 hr3790.26 / 0.3— (all 379 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+4 hr3790.29 / 0.34— (all 379 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+5 hr3790.31 / 0.39— (all 379 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+6 hr3790.35 / 0.43— (all 379 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —

Boxley (4 h)

AheadGradedMAE net / no-change (cfs)Skill, all hoursSkill, moving hoursIn-band calm / moving
+1 hr3790.02 / 0.02— (all 379 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+2 hr3790.04 / 0.03— (all 379 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+3 hr3790.07 / 0.03— (all 379 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+4 hr3790.08 / 0.04— (all 379 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —

All-time: 1923 issuances graded since 2026-07-10 (all models pooled).

Neural Net Predictions (Cossatot)

The Cossatot pixel net projects the level hour by hour from 48 h of 1 km radar rainfall (6 h ahead); crossing probabilities come from its companion gradient-boosted ensemble. Graded the same way as the Buffalo net: MAE = average error of the median; IN-BAND% = how often the likely range contained the truth (target ~50%); SKILL = error reduction vs assuming the level doesn't change, on hours where it moved.

8 hour(s) in the window had no gauge data (USGS down): 8 re-served the previous projection with a note, the rest were offline.

556 issuances graded (model cossatot-pixel-v6b / engine v3, 2026-09-08 → 2026-10-01; low-flow hold in 275).

Cossatot (6 h)

AheadGradedMAE net / no-change (cfs)Skill, all hoursSkill, moving hoursIn-band calm / moving
+1 hr5560.09 / 0.094.6% (540 below 10 cfs excluded)— (0 moving hrs) (n=0)18.8% / —
+2 hr5560.17 / 0.1615.7% (540 below 10 cfs excluded)-2.1% (n=1)26.7% / 0.0%
+3 hr5560.23 / 0.2417.5% (540 below 10 cfs excluded)11.5% (n=3)46.2% / 0.0%
+4 hr5560.28 / 0.328.8% (540 below 10 cfs excluded)15.6% (n=4)50.0% / 25.0%
+5 hr5560.32 / 0.3537.7% (540 below 10 cfs excluded)26.9% (n=3)38.5% / 0.0%
+6 hr5560.35 / 0.437.5% (540 below 10 cfs excluded)23.5% (n=3)61.5% / 0.0%

Floatable (240 cfs): 0 of 556 scored windows crossed (mean served probability 0.2%). High water (1,270 cfs): 0 of 556 scored windows crossed (mean served probability 0.3%).

Earlier model: 117 issuances graded (model cossatot-pixel-v6b / engine v2, 2026-09-03 → 2026-09-08).

Cossatot (6 h)

AheadGradedMAE net / no-change (cfs)Skill, all hoursSkill, moving hoursIn-band calm / moving
+1 hr1170.09 / 0.08— (all 117 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+2 hr1170.18 / 0.15— (all 117 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+3 hr1170.24 / 0.2— (all 117 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+4 hr1170.28 / 0.24— (all 117 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+5 hr1170.29 / 0.24— (all 117 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+6 hr1170.32 / 0.26— (all 117 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
Earlier model: 2 issuances graded (model cossatot-pixel-v1 / engine v2, 2026-09-03 → 2026-09-03; 2 issued with a missing input, graded but flagged).

Cossatot (6 h)

AheadGradedMAE net / no-change (cfs)Skill, all hoursSkill, moving hoursIn-band calm / moving
+1 hr20.0 / 0.0— (all 2 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+2 hr20.01 / 0.0— (all 2 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+3 hr20.03 / 0.0— (all 2 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+4 hr20.03 / 0.0— (all 2 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+5 hr20.0 / 0.0— (all 2 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+6 hr20.07 / 0.0— (all 2 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
Earlier model: 388 issuances graded (model pixel-v1 (pre-F5, untagged) / engine v1, 2026-08-18 → 2026-09-03).

Cossatot (6 h)

AheadGradedMAE net / no-change (cfs)Skill, all hoursSkill, moving hoursIn-band calm / moving
+1 hr3880.08 / 0.05— (all 388 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+2 hr3880.14 / 0.09— (all 388 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+3 hr3880.24 / 0.13— (all 388 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+4 hr3880.31 / 0.17— (all 388 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+5 hr3880.44 / 0.2— (all 388 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —
+6 hr3880.55 / 0.24— (all 388 issued below 10 cfs, not scorable)— (0 moving hrs) (n=0)— / —

All-time: 1874 issuances graded since 2026-07-14 (all models pooled).

Data health

Every grade above is checked against a USGS gauge feed. A single parameter can go silent while its site keeps reporting (USGS stops publishing discharge when the stage falls below its rating's lowest measured point); this table shows the age of the newest USGS reading per gauge and parameter, and whether a stage-derived stand-in is being served in its place.

21 series checked at 2026-10-02 06:30 UTC — 0 silent, 0 lagging.

All graded gauge series are current.

Level numbers: all 172 copies of the Boxley, Ponca and Hailstone levels (configs, countdown tables, forecast code and the last pages served) match the one shared list (registry 2026.9.28.1).

Coverage

What the scorecard grades today, and what is still being wired into the loop:

Predictor familyStatusNotes
Physics rise predictorsgradedCossatot, Richland, Hailstone — predicted crest height/flow vs. the actual peak; rise calls graded by confidence label (HIGH / MEDIUM / LOW, eager-plus-labels 2026-09-09) against the rate each label quoted.
Empirical forecast enginegraded6 basins — 'likelihood of rise to tier X' vs. the tier the gauge actually reached.
Recession countdownsgraded9 gauges — maturing; the longest horizons settle ~7-10 days after they're issued.
Ponca AI Rainfall Event AnalysisgradedScored per armed EVENT (first call + mature call vs the event crest), flood-risk % as a probability; post-peak rows only on 'did a higher crest come' (ponca_analog_eval.py v2, 2026-09-02).
Radar nowcast (Ponca card)gradedThe card's 'roughly another X in falls in the next 2 hours' radar estimate and its 'about done' / 'more coming' calls, graded per radar frame against the areal rain the next two MRMS hours actually delivered (radar_eval.py, 2026-09-03).
Downstream propagation forecastretiredRetired 2026-09-02: the Ponca card now narrates the wave router's own bands, which the Routed Waves section grades (wave_eval.py). propagation_eval.py kept read-only.
Buffalo per-gauge rise predictionsgradedPer-gauge rise nowcasts recorded + graded vs the gauge's actual rise (buffalo_predictions_archive.py).
Routed waves (wave_router)gradedF4 tracked waves — each confirmed wave-target's frozen crest band + arrival window graded vs what the target gauge actually did (wave_eval.py).
Buffalo flood_risk labelsretiredRetired from the pages 2026-09-28 (G11, Dave): replayed 2014-2026 it flagged only 8-19 % of flood onsets on either soil basis; the rise band, routed waves and the BIG FLOOD call cover it. Still computed in buffalo_output (unchanged) for anything that reads it; never graded.
Q-bucket / watershed alerts (/watersheds, Signal, Facebook mirror)gradedTier-1 alert episodes from creeks/logs/qbucket_shadow.log graded against the creek's own reading reaching its floatable floor within 24 h; rises with rain and no alert counted as misses (qbucket_eval.py, G4 2026-09-08). Tier-2 (reference-gauge) rows have no gauge of their own and stay ungraded.
Neural Net Predictions (pixel-v5)gradedPonca 6 h / Boxley 4 h hour-by-hour level tables — every hourly issuance graded against the observed level at each horizon (neural_eval.py).
Neural Net Predictions (Cossatot)gradedCossatot 6 h level tables + GBM-ensemble crossing probs — every hourly issuance graded at each horizon (neural_cossatot_eval.py).
Gauges Watersheds Ozark Rain Map
Cossatot Intelligence Richland Intelligence Mulberry Intelligence Big Piney Intelligence Illinois River Intel Branson Creeks St. Francis Intel Hailstone Intelligence Buffalo Intelligence
Guide Changelog Scorecard Page Suggestions and Corrections
Home
♥ Support this project
DISCLAIMER: This site provides creek condition estimates for informational purposes only. Gauge data, radar estimates, and forecasts may be delayed, inaccurate, or unavailable. Always exercise independent judgment. Whitewater kayaking is inherently dangerous — water conditions can change rapidly. This site and its maintainers assume no responsibility for decisions made based on information displayed here.