How the forecast
is built.
Daily forecast across 9 of every 10 California surf spots — 237 of 260 ocean-facing beach groups, plus 68 of 81 bay and lake monitoring sites. Beaches without a lab test in the past 30 days show their latest official sample instead of a rating.
Numerically grounded by default.
Label policy
The forecast label is marine enterococcus exceedance, scored method-aware: culture samples (MPN/CFU) against the 104 single-sample STV, and ddPCR results (copies/100mL) against the molecular 1413-copy threshold — correcting a prior bug that compared ddPCR copies against the culture number and false-flagged most San Diego rapid-test samples. Freshwater E. coli and total/fecal coliform stay in the warehouse, outside the pooled label.
Daily forecast
The pipeline refreshes once each morning. It is a batch forecast that cannot react to an intra-day spill after publication.
What is served
Each beach’s rating starts from the share of its lab tests over the past 12 months that exceeded the safety limit, shrunk toward the statewide average (the per-beach lookup, p = (pos + 5·g)/(n + 5)). Since 2026-09-28 a logistic model adjusts that number for four things known by the morning update: how far the latest result was above or below its limit, how recent it is, the inches of rain in the past 72 hours, and the season. It is refit every morning on four years of lab results. A beach is never rated Low if its last test exceeded, and it is raised to High while an official county advisory is active. Bands keep the same cutpoints. If the model cannot be fit on a given day, the plain lookup is served and the pipeline raises an alert.
Calibration
Public-facing bands come from calibrated exceedance probabilities. Unsupported stations do not receive a colored risk claim or a fallback model badge on public surfaces.
Measured performance
In a year-long walk-forward test (Sep 2025 – Sep 2026, refit monthly, scored on the first lab result 1–3 days after each forecast, n = 74,727), the logistic model beat the plain lookup on AUROC (0.851 vs 0.827), AUC-PR (0.591 vs 0.548) and Brier score (0.077 vs 0.080), and its ability to tell a beach’s dirty days from its clean ones rose from chance (within-beach AUROC 0.45) to 0.64. On the window where the served log overlaps (Apr – Sep 2026) it scored AUROC 0.864, against 0.863 for the lookup and 0.804 for the ML that served then. Realized exceedance per band in that window: Low 4%, Moderate 19%, High 37%, Very High 87%. At the Low cutoff it misses about as many exceedances as the lookup, with roughly a quarter fewer false alarms.
Tested, shelved, and re-tested
We hold candidates to a verified-improvement bar and report what did not clear it. A per-beach layer on top of the fresh-sample forecast did not help on sample-days — the ensemble is already well-calibrated there, so it traded discrimination for nothing. But the same per-beach anchoring is exactly what the model needs between samples, where the ensemble otherwise decays toward “safe”; re-tested on that regime it earned its place, and served the stale majority in a two-tier system from July to September 2026. Cross-region (Texas) pooling and an LSTM were tested and rejected outright on held-out beaches. Honest negatives — and honest re-tests — keep the deployed model the one the evidence actually supports. On 2026-09-22 the two-tier ML itself was retired from serving after it lost to the per-beach lookup on forward outcomes, and on 2026-09-28 a four-term logistic layer on top of that lookup replaced it after beating it in a year-long walk-forward test.
Beach track record and official postings, side by side.
Shorelife runs two parallel feeds. The rating starts from each beach's own exceedance rate over its past 12 months of lab tests, adjusted for its latest result and recent rain. A second pipeline scrapes the live advisory page that each county health department publishes for itself — same-day visibility into Postings, Closures, and Rain Advisories, versus the statewide data.ca.gov export, which is frozen — its newest sample is dated 2026-03-05, so it publishes no status changes at all. Live results now arrive from the BeachWatch database (beachwatch.waterboards.ca.gov, added July 2026) and CEDEN/SafeToSwim.
12 counties live-scraped
San Diego, Orange, San Mateo, LA, Marin, Long Beach, East Bay Parks, Ventura, San Francisco, Humboldt, Sonoma, San Luis Obispo. Monterey and Santa Barbara still rely on the state feed (Monterey blocks datacenter traffic; SB does not publish current postings on the public web).
Direct or AI extraction
San Francisco ingests directly from its Socrata API; Humboldt and Sonoma are read by a structured-output LLM extraction; San Luis Obispo is rendered with a headless browser before extraction. All 12 fall back to the state feed if the county source flakes.
County-direct sample feed
A second parallel feed, county_direct_samples.parquet, captures the numeric enterococcus and coliform values each county publishes — independent of the BeachWatch sample archive. San Francisco’s results are now merged into the training labels from it, which cut that county’s serving anchor from ~30 days stale to ~2. Coverage is narrow — the other counties in the file (Sonoma, Humboldt) stopped publishing in May 2026.
AB411 single-sample triggers
For counties that publish raw values rather than a posted/not-posted flag (SF, Humboldt, Sonoma), Shorelife applies the same California AB411 thresholds the counties themselves apply: ENT ≥104, FECAL ≥400 MPN/100mL.
Advisory band, not model band
When the county marks a station Posted or Closed, the public surface shows a purple Advisory band with the cause and a deep-link back to the county source — never overridden by a model-derived risk band.
Stale records auto-demote
When a first-class county scraper succeeds, any state-feed record for that county that the county no longer affirms is moved to historical status — closes a real failure mode where year-old admin entries kept showing as active.
What this estimate can't do.
- 01It rates a beach's track record adjusted for its latest result and the 72 hours of rain before the morning update. It cannot see a spill, or rain that falls after the update; posted advisories can.
- 02Forecasts complement official monitoring; they do not replace county advisories.
- 03Only a subset of beaches currently has model coverage; the rest should be read as latest-official-sample views, not forecast gaps filled by guesswork.
- 04Heavy storm, sewage, or spill events can outrun any historical statistical model, especially after the morning publish cutoff.
- 05The numeric risk comes from each beach's lab history and measured rainfall, not from a weather forecast.
Sources and prior work.
- [1]Searcy, R. T. & Boehm, A. B. (2021). A day at the beach: enabling coastal water quality prediction with high-frequency sampling and machine learning. Water Research · DOI: 10.1016/j.watres.2021.117051 ↗
- [2]U.S. EPA. Virtual Beach 3 (VB3): User Guide. Enterococcus-based per-station MLR as production baseline. EPA VB3 Reference Manual ↗
- [3]California Department of Public Health. AB411 Annual Beach Report. County advisory thresholds and official culture-based sampling protocol. CDPH AB411 / Beach Monitoring Program ↗
- [4]NOAA National Data Buoy Center (NDBC). Real-time and historical wave, wind, and sea surface temperature observations used as model covariates. NDBC · noaa.gov/ndbc ↗
- [5]Scripps Institution of Oceanography. Coastal Data Information Program (CDIP). Nearshore wave model output and buoy telemetry. CDIP · cdip.ucsd.edu ↗
- [6]Open-Meteo. Open-source weather API — hourly UV index and solar radiation used for solar inactivation index feature. Open-Meteo API · open-meteo.com ↗