We began saving weather forecasts without being sure what we would find. Years of repeated forecast collection now let us follow how temperature predictions change from six days ahead to the near term, across stations around the United States and Puerto Rico. The clearest finding in this archive is geographic: middle America generally appears harder to predict than coastal America, with larger differences from the near-term forecast, especially several days ahead. We compare the same sampled months across stations and retain an audit of flagged readings. This pattern describes our archived forecasts; it does not establish accuracy against observed weather or explain what causes the regional differences.
i
Agreement with the near-term forecast. These comparisons do not establish observed-weather accuracy. An unusual station can reflect geography or forecasting difficulty; it does not establish bad data.
Mean absolute difference from the near-term forecast.
Loading state boundaries…
Mean absolute difference · fixed scale across years, leads and quality views
No comparable data
The map uses available pairs at the selected lead. Choose matched year to date to hold stations and calendar coverage fixed across years. A year-to-year change alone does not establish improving forecasts. Rankings and charts below retain their completed-year window and require all six leads at the same target hour.
Each contributing station receives equal weight within its state, after equal weighting of the selected calendar months. Thin coverage means fewer than 30 pairs or 10 sampled dates in a month. These are summaries of sampled stations, not statewide accuracy estimates. Alaska, Hawaii and Puerto Rico appear as insets. Boundaries: U.S. Census via us-atlas ↗
Potential outliers use |modified Z| > 3.5; the broader watchlist uses |ordinary Z| > 2.
See flags across all six leads and both quality views
Flags consider five correlated metrics at the selected lead. They are exploratory labels, not adjusted significance tests. A station stays in the analysis when flagged.
02 / COHORT RANKING
Mean absolute difference
Equal weights for the same sampled year-months. Select a bar to inspect a station.
03 / MAGNITUDE & DIRECTION
Different ways to disagree.
Horizontal: magnitude. Vertical: signed bias. Outlined points are flagged on any metric.
Positive bias = warmer than the zero-hour reference. Hover, focus the table below, or select a point for details.
View all station histories and comparison eligibilitySelection, weighting & interpretation
Archive and comparison eligibility
Include all station records with forecast history, except the explicitly identified duplicate Salt Lake City #80. Keep Salt Lake City #1. Other same-city records remain distinct and display their IDs. Short or incomplete histories remain downloadable and explorable.
A fairer calendar comparison
Use the latest three completed calendar years. Each retained site/month/lead must have at least 30 screened pairs on 10 local dates. Use the same supported year-months at every station and lead; average each month's metric with equal weight. Each station's target hours must have valid pairs at all six leads. Exact hours still differ across stations.
Reading-level and station-level flags
The existing two-sigma temperature screen is unchanged. Flagged target-hour share means any of the six leads or its reference was flagged. Station-level scores compare these summary metrics across the cohort; no station is deleted.
Uncertainty and limits
Pointwise 95% intervals use 1,000 paired year-month block resamples within calendar-month strata. The seasonal mix stays fixed. Peer-gap intervals compare a station with the other comparable stations' median. Neighboring months may still be dependent, and intervals do not correct sampling bias, regional differences or multiple comparisons.