Skip to main content
Research · operator view · build 2026.10.07

Model health &
deployment traceability.

The operator view tracks what computed today's served rating, the machine-learning challenger that trains beside it, source freshness, and current public coverage. Shorelife is a daily forecast product and does not detect intra-day spills.

Served estimate
Logistic model on the per-beach lookup
logit-lookup-offset-v2 · computes the rating users see
Served estimate · AUROC / AUCPR / Brier
0.854 / 0.556 / 0.053
2026-04-23 – 2026-09-29, the same 22,741 forecasts as the challenger tile · year-long backtest 0.853 / 0.592 / 0.073 · live score since 2026-09-28: 393 forecasts with a lab result so far · scored on the first lab result 1–3 days after each forecast
Coverage
238 / 341
70% of California beaches
ML challenger
xgb-undersample-ensemble-curated-v0
registry winner; trains daily, does not set the served number
ML challenger · AUROC / AUCPR / Brier
0.790 / 0.307 / 0.071
what it served 2026-04-23 – 2026-09-29, scored on the same forecasts as the served estimate · its sample-day test score is on the Calibration page and is not comparable · first lab result 1–3 days after each forecast
Public release
Eligible
ML release gate
Served estimate · 2026-10-06

Today's fit

logit(p) = logit(lookup) + b0 + b1·R + b2·R·F + b3·W + b4·S + b5·D + b6·R·D, refit every morning on four years of lab results. R is the latest result relative to its limit (log scale), F how fresh it is, W the 72 hours of rain before the update, S the season. D is 1 when the beach's most recent lab result was measured by ddPCR (San Diego's DNA method, judged at 1413 copies) and 0 for culture methods. If the fit fails a guard, the plain lookup serves instead and the pipeline raises an alert.

b0
-0.227
b1 · R
0.315
b2 · R·F
0.519
b3 · W
0.719
b4 · S
0.338
b5 · D
0.103
b6 · R·D
0.655

Live scoreboard

Within-beach AUROC asks: at the same beach, can the rating tell a dirty day from a clean day? 0.5 means it cannot. It is the deciding metric, so it comes first. Every row is scored on the same forecasts, and the last two columns are the change of each other row versus the served logistic, as a 95% interval (resampled by beach); an interval that includes 0 is not a real difference.

Live, 1–3 days outWithin-beach AUROCAUROCAUCPRBrierWithin-beach AUROC vs served (95% CI)Brier vs served (95% CI)
Served logistic0.4030.8340.6160.0914——
Per-beach lookup0.4310.8170.5830.0964−0.250 to +0.260−0.0037 to +0.0142
Persistence0.3610.6930.3230.1530−0.266 to +0.163+0.0363 to +0.0890
ML challenger0.2920.8220.5010.0879−0.269 to +0.032−0.0139 to +0.0070
364 forecast–result pairs · 9 served days · reportable from 28 served days (numbers greyed until then) · versions pooled: logit-lookup-offset-v1, logit-lookup-offset-v2 · decision rule
Forward outcomes 1–3 days outAUROCAUCPRBrierWithin-beach AUROC
Logistic (served) · Apr–Sep0.8540.5560.05310.512
Lookup · Apr–Sep, same forecasts0.8530.5290.05690.382
ML challenger · Apr–Sep, same forecasts0.7900.3070.07060.513
Logistic · year-long backtest0.8530.5920.07290.646
Lookup · year-long backtest0.8280.5420.07730.450
Rows marked "same forecasts" are scored on identical forecast-days (Apr–Sep: 2026-04-23 – 2026-09-29, n = 22,741, the ML is what actually served then). Year-long backtest: 2025-09-01 – 2026-09-29, refit monthly, n = 76,889. Live rows are in the scoreboard above.
113,833 training rows · rain data for 100% of beaches · mean 0.075 vs lookup 0.114 · 49 of 478 beaches in a different band than the lookup
ML challenger · spatial backtests

Performance on unobserved counties

These scores describe the machine-learning challenger, not the served estimate above.

spatial beach persistence
AUCPR
0.793
Brier
0.166
15 folds · 12258 samples
spatial county persistence
AUCPR
0.483
Brier
0.136
6 folds · 67596 samples
spatial beach xgb undersample ensemble
AUCPR
0.949
Brier
0.099
15 folds · 12258 samples
spatial county xgb undersample ensemble
AUCPR
0.708
Brier
0.100
6 folds · 67596 samples
spatial beach hist gbm
AUCPR
0.946
Brier
0.103
15 folds · 12258 samples
spatial county hist gbm
AUCPR
0.637
Brier
0.118
6 folds · 67596 samples
spatial beach hist gbm positive persistence guard
AUCPR
0.829
Brier
0.142
15 folds · 12258 samples
spatial county hist gbm positive persistence guard
AUCPR
0.586
Brier
0.126
6 folds · 67596 samples
spatial beach hist gbm persistence blend
AUCPR
0.944
Brier
0.113
15 folds · 12258 samples
spatial county hist gbm persistence blend
AUCPR
0.673
Brier
0.099
6 folds · 67596 samples
spatial beach xgb undersample offset
AUCPR
0.908
Brier
0.193
15 folds · 12258 samples
spatial county xgb undersample offset
AUCPR
0.547
Brier
0.124
6 folds · 67596 samples
Deep dives

Transparency documentation

Labels · Thresholds

Risk Labels

How exceedance probability maps to the four public bands, legal basis, and what the model can and cannot predict.

READ →
Provenance · Cadence

Data Sources

The six primary datasets feeding the pipeline — BeachWatch, NDBC, CDIP, Open-Meteo, USGS NWIS, and CEDEN — with freshness status.

READ →
AUCPR · Brier · Spatial CV

Calibration

The ML challenger's metrics, spatial backtest results, and candidate model registry from the latest CI run.

READ →
Scientific artifacts

Open data & methodology

Technical Methodology

Model design, coverage limits, risk-band calibration, spatial holdout protocol, and known limitations.

Shorelife Team
READ →

Source Code & Data

Full pipeline, training scripts, and curated datasets available on GitHub under an open-source license.

Automated Pipeline
GITHUB →