← Back to the articleSnapshot through 2026-09-19 · All 70 station summaries · No individual forecast recordsAbout this edition
FORECAST RESEARCH / WEATHER

Forecasts, in hindsight.

How the outlook changes between six days away and the hour itself.

All-station analysis ↗Get the dataset ↓Read the methodology
Reading the archive…

Forecast reference, not measured weather. Differences use the stored zero-hour forecast. Gaps and changing seasonal coverage limit claims about improvement.

Loading station summaries…
01 / THROUGH TIME

Does the outlook get closer?

Mean absolute difference from the zero-hour forecast.

Nearest forecast within ±4 hours
Whiskers / shading: exploratory 95% intervals. Empty years stay empty.
View chart data and sample counts
02 / DISTANCE TO TARGET

What changes as the day approaches?

Selected metric across all six nominal leads.

Inspect the selected lead offsets

Offsets are rounded to the nearest hour for this display; matching uses exact seconds.

03 / DIRECTION

Warmer, or cooler?

Share of paired forecasts at the selected lead.

04 / SEASONAL SIGNATURE

A different bias in summer?

Mean signed difference by target month and lead. Select a cell to focus the study.

CoolerWarmer

Positive values mean warmer than the near-term forecast. All selected years combines the page's year range, weighted by paired-hour count. This year selector affects only this chart; all twelve months stay visible. Hover or focus a cell for its sample count and dates.

05 / YEAR BY YEAR

How does the seasonal bias change?

Mean signed difference by year and target month · 6 days ahead, ±4 hours · °F

CoolerWarmer

Rows follow the page's year range; columns always show January–December. This grid has its own lead selector and follows the station, quality and sample settings. Each cell is a paired-hour-weighted mean; missing or future months stay empty. Select a cell to focus the study on that year, month and lead. Both bias grids use the same color scale.

06 / WHAT THE ARCHIVE CAN SUPPORT

Coverage is part of the result.

How many tracked hours have a stored near-term forecast? Select a month to explore it.

No data More coverage

Download station summaries ↓

07 / INSPECT THE EVIDENCE

Compare the leads for a day.

Daily mean differences and sample counts across all six leads. Quality and sample settings apply.

Methodology & interpretation Reproducible choices, visible limitations

A consistent comparison

Each target hour contributes once per lead. Select the nearest nominal lead within ±4 hours; exact matches win and equal-distance ties use the older forecast. A missing zero-hour reference stays missing. Temperatures are normalized to °F.

What “closer” means

The reference is another forecast from the same system. Shared errors can remain invisible. These results measure forecast agreement, not accuracy against a sensor.

Uncertainty, without false precision

Pointwise 95% percentile intervals use 1,000 resamples of fixed calendar blocks. At least 30 sampled days and eight occupied blocks are required. Try 7, 14, and 28 days. Intervals condition on retained samples; gaps and unequal seasons can bias comparisons. They are exploratory and unadjusted for multiple comparisons.

Quality screening

Temperatures more than two sample standard deviations from the site/month/local-hour baseline are flagged. Baselines pool retained zero-hour forecasts across years within ±2 local hours and require 30 samples and nonzero variation. Sparse groups skip this screen. Broad temperature bounds and adjacent reference jumps are also flagged. Flagged inputs are excluded by default, preserved in the audit log, and can be restored with the quality filter. These rules can remove real extremes and affect apparent performance.

Time & coverage

Dates follow each site's timezone; repeated daylight-saving hours stay distinct. Nominal lead uses the first hour in the saved bundle, not a verified issuance time. Winter in the year chart means January, February, and December of that calendar year.

Sources & provenance

Recorded NWS hourly forecasts from Trackit. NWS product documentation ↗ · Verification metrics ↗

READING THE COVERAGE GRID

How much data do we have?

Coverage measures data availability. Each cell tells you what share of tracked hours has a stored near-term temperature forecast to use as a reference.

25%
For example

180 reference hours ÷ 720 tracked hours = 25%. We have a near-term forecast for one in four tracked hours.

Read the grid

  • Rows are years; columns are months. Each station has its own grid.
  • Brighter teal means more coverage. The color reaches its brightest at 35%, so higher percentages can share the same color. Read the number to compare them.
  • A dash means no reference hours are shown. It can mark a collection gap, a month outside that station’s tracking interval, or a future month.

What counts as a tracked hour?

We count hours within the station’s tracking interval, including gaps. The first and last months can cover only part of a month. Dates follow the station’s local timezone.

Why coverage matters

Thin coverage means the results describe fewer sampled hours. Different collection gaps can affect comparisons between months or years.

A reference alone does not guarantee a usable historical comparison: a forecast must also exist at the chosen lead and pass the selected quality rules. The paired-hour counts elsewhere can therefore be lower than the reference count here.

Try it: hover or focus a cell to see its reference and tracked-hour counts. Select a month to explore it. Cell percentages are rounded.

This grid shows the full tracked history for the selected station or stations. Year, season, lead and quality filters do not change these reference-availability counts.