Model health &
deployment traceability.
The operator view tracks what computed today's served rating, the machine-learning challenger that trains beside it, source freshness, and current public coverage. Shorelife is a daily forecast product and does not detect intra-day spills.
Today's fit
logit(p) = logit(lookup) + b0 + b1·R + b2·R·F + b3·W + b4·S + b5·D + b6·R·D, refit every morning on four years of lab results. R is the latest result relative to its limit (log scale), F how fresh it is, W the 72 hours of rain before the update, S the season. D is 1 when the beach's most recent lab result was measured by ddPCR (San Diego's DNA method, judged at 1413 copies) and 0 for culture methods. If the fit fails a guard, the plain lookup serves instead and the pipeline raises an alert.
Live scoreboard
Within-beach AUROC asks: at the same beach, can the rating tell a dirty day from a clean day? 0.5 means it cannot. It is the deciding metric, so it comes first. Every row is scored on the same forecasts, and the last two columns are the change of each other row versus the served logistic, as a 95% interval (resampled by beach); an interval that includes 0 is not a real difference.
| Live, 1–3 days out | Within-beach AUROC | AUROC | AUCPR | Brier | Within-beach AUROC vs served (95% CI) | Brier vs served (95% CI) |
|---|---|---|---|---|---|---|
| Served logistic | 0.403 | 0.834 | 0.616 | 0.0914 | — | — |
| Per-beach lookup | 0.431 | 0.817 | 0.583 | 0.0964 | −0.250 to +0.260 | −0.0037 to +0.0142 |
| Persistence | 0.361 | 0.693 | 0.323 | 0.1530 | −0.266 to +0.163 | +0.0363 to +0.0890 |
| ML challenger | 0.292 | 0.822 | 0.501 | 0.0879 | −0.269 to +0.032 | −0.0139 to +0.0070 |
| Forward outcomes 1–3 days out | AUROC | AUCPR | Brier | Within-beach AUROC |
|---|---|---|---|---|
| Logistic (served) · Apr–Sep | 0.854 | 0.556 | 0.0531 | 0.512 |
| Lookup · Apr–Sep, same forecasts | 0.853 | 0.529 | 0.0569 | 0.382 |
| ML challenger · Apr–Sep, same forecasts | 0.790 | 0.307 | 0.0706 | 0.513 |
| Logistic · year-long backtest | 0.853 | 0.592 | 0.0729 | 0.646 |
| Lookup · year-long backtest | 0.828 | 0.542 | 0.0773 | 0.450 |
Performance on unobserved counties
These scores describe the machine-learning challenger, not the served estimate above.
Transparency documentation
Risk Labels
How exceedance probability maps to the four public bands, legal basis, and what the model can and cannot predict.
Data Sources
The six primary datasets feeding the pipeline — BeachWatch, NDBC, CDIP, Open-Meteo, USGS NWIS, and CEDEN — with freshness status.
Calibration
The ML challenger's metrics, spatial backtest results, and candidate model registry from the latest CI run.