Model performance
& validation.
All metrics on this page are read directly from the daily pipeline output. Spatial backtests use county-level GroupKFold holdout; the persistence baseline gives a lower bound on skill. The current public forecast covers 238 of 341 California beaches (70%).
Since 2026-09-22 the number users see is not this model: first the per-beach lookup, and since 2026-09-28 a logistic model on top of that lookup (latest result, 72-hour rain, season). The ML below still trains daily and is kept as a challenger; the metrics on this page describe it, not the served rating.
Discrimination & calibration scores.
These are the ML challenger's scores on days a lab sample was taken, with fresh lab history as input. They are not comparable with the served estimate's scores on the Research page, which grade every published day against the next lab result. AUC-PR in particular rises with the share of exceedances, which is higher on sample days. Scored the same way as the served estimate, on the same forecasts, the challenger is on the Research page.
Agreement with active advisories.
Each day the served forecast is checked against the beaches under an active health advisory or closure — the operational ground truth. A false negative is an advised beach the forecast did not flag; those are tracked with priority because they are the unsafe direction.
Performance on unseen geography.
County GroupKFold holdout: each fold withholds all beaches from one county. Tests generalisation to locations not seen during training.
Models evaluated each CI run.
All candidates must clear the block-bootstrap AUCPR gate to be eligible for production. The temporal validation winner becomes production if it also clears the spatial backtest promotion policy.