Field Notes

Forecasting

Why Temperature Forecasts Go Wrong in the Mountains

Cold pools and smoothed terrain can make mountain temperatures hard to forecast. Here is what station records and elevation data can—and cannot—show, and what we do about it.

A mountain valley at dawn filled with a pool of cold cloud lying below the surrounding peaks, with a forecast grid drawn over the terrain and stretched flat across the sky above it

Drive north from Gunnison, Colorado to Crested Butte and you climb 1,245 feet. Yet across the eleven Decembers from 2015 through 2025, Crested Butte reported the higher daily minimum on 145 of the 338 dates with data at both stations. Across the same eleven Julys, that happened twice.

Those two stations are about 39 kilometers apart, so this is not a measurement of one vertical column of air. It is still a vivid example of a well-established winter pattern: cold air can collect on mountain-valley floors while nearby higher terrain is milder.

Two sources of mountain forecast error are especially tractable. A model’s surface height can differ from the ground at a forecast location, and any height adjustment can be misleading when the near-surface temperature profile is weak or inverted. We will take them in order, and then describe what we do about each. Precipitation is a separate and harder problem, one that we will tackle in a subsequent post.

Problem one: the model is standing at the wrong elevation

A weather model divides the atmosphere into a horizontal grid and vertical layers. Each surface grid point has a representative surface elevation, so surface variables such as 2-meter temperature refer to the model’s represented terrain rather than every slope, bench and valley floor inside the cell. Models can also carry sub-grid land or orographic information, but that does not turn one grid-point temperature into a separate forecast for every elevation inside the cell.

The cells are large beside narrow valleys. NOAA’s Global Forecast System runs at about 13 kilometers. The European Centre’s IFS medium-range control forecast, formerly called HRES, runs at 9 kilometers. NOAA’s High-Resolution Rapid Refresh, among the finest widely used operational numerical guidance over the contiguous United States, runs at 3 kilometers.

Representing real terrain on those grids necessarily smooths features smaller than the grid spacing. The following calculation illustrates that loss of detail; it is not an extraction of any model’s orography.

7,000 8,500 10,000 11,500 13,000 14,500 ft 0 10 20 30 40 50 60 70 80 90 kilometers along 39.19°N (west → east) Aspen · 7,890 ft a 13 km model stands 1,200 ft higher 9,095 ft, where the town is 7,890 ft real ground 3 km 9 km 13 km
Terrain along 39.1911°N through Aspen, Colorado, from the Elk Mountains to the Sawatch crest. White is the ground profile sampled every 250 meters from the Mapzen Terrain Tiles service in one-arc-second HGT format. Mapzen combines several source datasets, including USGS 3DEP and SRTM; “one arc second” describes the sampling format, not a claim that every US value comes from SRTM. The colored lines are moving averages of that one-dimensional profile over 3, 9 and 13 kilometer windows. They illustrate smoothing at those scales; they are not HRRR, ECMWF or GFS orography, and a one-dimensional average is not equivalent to a model's two-dimensional terrain processing.

In this one-dimensional illustration, much of the range-scale shape remains after a 3-kilometer moving average, while the 9- and 13-kilometer curves smooth narrower terrain much more strongly. East of Aspen the sampled ground climbs 2,000 feet in a little over two kilometers. At Aspen, the 13-kilometer moving average is 9,095 feet while the sampled town coordinate is 7,890 feet. Across the 90-kilometer section, the range of elevations in the illustrated profile falls from 7,129 feet to 4,044 feet after the 13-kilometer averaging operation. Those are properties of this illustration, not the surface heights of a named forecast model.

Another way to see the same thing is to sample the ground inside a coarse-grid-sized square.

8,000 9,000 10,000 11,000 12,000 ft 0% 3% 6% 9% 12% share of the ground inside one 13 km grid box the one number a 13 km model gets · 9,199 ft 4,405 ft of real ground 7,385 to 11,781 ft
Ground elevation at 625 regularly spaced points inside an illustrative 13-kilometer square centered on Aspen, sampled from the same Mapzen terrain service. The square is not an actual GFS cell, and its arithmetic mean is not a model orography value.

The illustration puts 4,396 feet of sampled relief, several thousand people and part of a ski mountain inside one coarse-grid-sized square.

The consequence is a representativeness mismatch. If a model’s surface is 1,200 feet above the requested ground, its 2-meter temperature describes a different elevation. The resulting temperature error is not automatic or fixed: its size and even its sign depend on the simulated temperature profile, surface state and subsequent downscaling.

Finer grids help. The National Weather Service publishes forecasts on office grids of roughly 2.5-kilometer spacing, and its API reports the elevation associated with each forecast point. At 39.1911°N, 106.8175°W in Aspen, the API reports 7,897 feet, seven feet above the 7,890-foot terrain value used in our illustration.

Large differences can remain at narrow-valley coordinates. At 40.5883°N, 111.6383°W near the base of Alta, Utah, the API reports 9,800 feet while the terrain value used by our service is 8,543 feet. That comparison is coordinate-specific; it does not imply that every point in Alta, or every NWS grid cell in mountain terrain, has the same error.

Problem two: temperature does not always fall with height

One way to reduce an elevation mismatch is to transfer a grid-point temperature to the requested ground height. That requires an assumed relationship between temperature and elevation. The 6.5 °C-per-kilometer environmental lapse rate is the tropospheric value in the US Standard Atmosphere and is used in some fixed height-correction schemes. It is not a universal rule inside weather forecasts: numerical models diagnose 2-meter temperature from their evolving atmosphere and land surface, and post-processing systems use different, sometimes stability-dependent, adjustments. For example, ECMWF describes its forecast 2-meter temperature as a nonlinear interpolation that depends on near-surface stability.

A fixed 6.5 can be a poor description of air near mountain ground. Minder, Mote and Lundquist estimated regional lapse rates across the Cascades and found annual means of 3.9 to 5.2 °C per kilometer on the windward side of the range, falling to 2.5 to 3.5 for late-summer overnight temperatures, with large differences between the windward and lee sides. They also cautioned that individual station pairs can reflect local conditions and may not represent a regional lapse rate. For that reason we call the station-pair figures below apparent lapse rates: each is the temperature difference between two separated sites divided by the difference in their elevations, not a measurement of the air column above either one.

On many clear, calm nights—especially long winter nights—the near-surface relationship can flatten or reverse. The surface loses energy through net longwave radiation, the air near it cools, and dense air can drain downslope and collect on a valley floor. The result is a cold-air pool, with warmer air above colder valley air. Sunlight and convective mixing after sunrise, stronger wind, cloud changes or a passing weather system can weaken or remove it. Cold pools are not present on every winter night, and they can occur outside winter.

7,500 8,000 8,500 9,000 9,500 10,000 10,500 11,000 ft -20° -10° 10° 20° 30° December 2024 mean temperature (°F) what 6.5 °C/km predicts daily high overnight low Gunnison Crested Butte Butte Schofield Pass 32 °F
Mean reported daily maximum and minimum for December 2024 at four stations spanning 3,078 feet of elevation in the Gunnison region. These are geographically separated sites, not a vertical sounding. Gunnison 3SW and Crested Butte are NOAA cooperative stations; Butte and Schofield Pass are SNOTEL sites. The Gunnison daily maximum has 30 observations; the other plotted monthly means have 31. Daily records are from the Applied Climate Information System.

Read the blue line from the bottom up. The mean reported daily minimum at Butte is 23 °F higher than at Gunnison, whose listed elevation in the Applied Climate Information System (ACIS) is 2,538 feet lower. A hypothetical transfer using 6.5 °C per kilometer would instead make Butte about 9 °F colder, a 32 °F difference between that fixed-elevation calculation and the observed monthly contrast. The contrast is a mean of the available daily minima, not one unusual event. Because the sites are about 41 kilometers apart and use different observing networks and daily windows, it should not be interpreted as a measured vertical temperature profile.

To see how often that paired-site contrast occurs, we ran the same calculation across eleven and a half years of daily records.

inverted — warmer above 0 50 100 150 200 250 300 350 400 450 nights -16 -12 -8 -4 0 +4 +8 +12 observed overnight lapse rate, Gunnison → Schofield Pass (°C per km) 6.5 °C/km the value most gridded forecasts assume
Each nominal archive date from 1 January 2015 through 30 June 2026 with a reported daily minimum at both Gunnison 3SW and Schofield Pass: 4,189 paired dates. The calculation uses the ACIS-listed elevations of 7,622 and 10,700 feet; current NRCS metadata lists Schofield Pass at 10,640 feet. Positive means the higher site reported a lower minimum. The stations are about 55 kilometers apart, and their daily reporting windows differ, so “date” does not guarantee the same overnight interval. SNOTEL temperatures also have documented measurement limitations; NRCS is applying a temperature-dependent sensor-conversion correction, which does not itself correct radiation, shielding or siting effects, to affected historical sensors. The distribution is descriptive paired-site evidence, not a direct count of atmospheric inversions.

The median apparent lapse rate for the paired daily minima is exactly zero. On 48.3 percent of paired dates, Schofield Pass reports the higher minimum. Only 5.3 percent fall within 1 °C per kilometer of 6.5. This shows that a fixed standard-atmosphere transfer does a poor job reproducing the relationship between these two sites’ daily minima. It does not establish the frequency of inversions over Gunnison, because that would require observations aligned in time and space through the same air column.

The paired daily maxima are much closer to the reference: their mean apparent lapse rate is 5.5 °C per kilometer. The larger and more operationally important departures occur in daily minima and are especially common during winter cold-pool regimes. Valley-floor towns are exposed to those regimes, but the all-season histogram above should not be described as a winter-only result.

The difficulty also appears in national guidance. A 2026 study by National Weather Service and NCEP scientists examined four 2023 cold-air-pool cases in Canaan Valley, West Virginia using National Blend of Models version 4.1; three of the four featured a pronounced cold pool. The blend failed to distinguish the valley floor from the rim, and its largest warm bias across those cases reached nearly 15 °C. This is case-study evidence about NBM v4.1, not proof that every cold pool is missed. NBM v5 became operational in May 2026, and that paper did not evaluate it.

The authors identified low weights assigned by the analysis quality-control system to unusually cold valley observations as a likely contributor. Such a reading can resemble a bad sensor when compared with warmer surrounding observations, so a defensible general-purpose quality-control system may give it too little influence. The paper did not prove that those observations were always discarded or that this was the only cause.

What we do about it

We are not claiming to have solved this. Cold pools are an open research problem and mountain forecasts will remain harder than plains forecasts.

The narrower claim we make is that part of mountain temperature error is a correctable representativeness and calibration problem: model and ground elevations can differ, a fixed height transfer can be inappropriate, and local forecast biases can vary by time and weather regime. The observations above do not quantify what share of total forecast error comes from each cause; that requires direct forecast verification and ablation tests.

Two things follow from that. First, forecast for real ground rather than for a grid cell. Second, keep a record of your own errors at each place, and use it.

For the first, our public locations are represented by forecast regions built from a 30-meter terrain model, so a valley floor and the ridge above it can be separate places with separate answers. The region containing our Aspen coordinate runs from 7,825 feet to 8,009 feet: 184 feet of relief, compared with 4,396 feet in the illustrative 13-kilometer square above. That comparison is between our region and the illustration, not between our region and an actual GFS cell.

For the second, here is what happens to a raw model field on its way to becoming a temperature we will publish.

RAW MODEL FIELDS HRRR 3 km NBM 2.5 km ECMWF 9 km GFS 13 km 1 Move it to the right ground Each model's temperature is carried to the region's real elevation using that model's own lapse rate for that hour, measured from its own neighbouring cells. 2 Subtract the bias we already know about Held separately for each model, variable, lead time and three-hour block of the local day, because a valley's error flips sign between night and afternoon. 3 Blend by how each model has actually been doing here Weights come from the last 60 days of verification at this location, capped so a thin run of observations cannot let one model take over the forecast. 4 Anchor the near term to what is being observed The first hours are pulled toward the surface analysis, and then toward the reports of nearby weather stations. 5 Correct what is left over A per-location regression on cloud, wind, hour, season, model spread and lead, with an explicit clear-and-calm-at-night term: the cold-pool signature. The forecast we serve Archived, then scored against nearby station observations what the record teaches, fed back in
What happens to a raw model field on its way to becoming a temperature we publish, for one forecast region.

Four of those steps are doing most of the work, and they map directly onto the two problems above.

Step one addresses the elevation-transfer problem. For each location and hour, we regress a model’s 2-meter temperatures in the surrounding grid cells against the corresponding surface heights. HRRR and GFS supply their own surface-height field; for NBM and ECMWF, whose files in our pipeline do not, we estimate cell heights from our terrain service. The resulting horizontal neighborhood slope is a proxy for the model’s local temperature–terrain relationship, not a direct vertical sounding. A regression is omitted when fewer than five usable cells remain or their height range is under 100 meters. Otherwise the computed slope is clamped to the configured range of −12 to +10 °C per kilometer rather than discarded. If no guarded slope is available, the pipeline uses its fixed-lapse fallback.

Step one also carries a rule that exists purely because of the physics above. When the forecast has to move downhill on a clear, calm night, the warming it applies is deliberately reduced. Lower terrain is not automatically warmer under those conditions, and a naive lapse-rate transfer would walk the forecast straight into the cold pool and get the sign wrong.

Step two is why we keep bias in three-hour blocks of the local day rather than as a single number per place. A valley’s forecast error can change sign between an overnight inversion and afternoon mixing. One all-hours correction can average opposite errors into a value that is useful at neither end of the day.

Step four anchors the near term. The first hours are nudged toward the latest RTMA surface analysis and nearby station observations, subject to age, distance, elevation and magnitude guards. This occurs before the learned residual temperature correction.

Step five is where our own record comes back in. After those temperature layers have been applied, a per-location ridge regression removes part of the residual using cloud cover, wind speed, local hour, season, model spread and forecast lead. One input is an explicit clear-and-calm-at-night term. It is trained on the preceding 45 days of selected archived forecasts scored against station observations, requires at least 120 samples, is shrunk toward zero when samples are thin and is capped at ±3 °C. It is skipped inside two hours, where fresh observations own the near-term adjustment.

That correction is deliberately linear and inspectable rather than a neural network. Rasp and Lerch’s study of 2-meter temperature at surface stations shows that auxiliary predictors and station-specific information are valuable. Their nonlinear network also improved materially over a linear network with similar inputs, so the paper does not establish that architecture is unimportant in our setting. Our choice of a ridge regression is an engineering choice for inspectability, sample efficiency and bounded behavior, not a conclusion proved by that paper.

We archive selected issued forecasts at seed and active verification locations, rather than every value served to every user. Hourly temperature snapshots use lead times of 0, 1, 2, 3, 6, 12 and 24 hours; daily cards and minute-precipitation products have separate sampling schedules. Eligible snapshots are scored against nearby station observations, and a rolling summary is on the site by location, with the mean error, bias and number of comparisons. That selected record supplies the training data for step five and contributes to the weights in step three.

In our next post, we will tackle mountain precipitation, which is even harder.