Forecasting
Why Temperature Forecasts Go Wrong in the Mountains
Cold pools and smoothed terrain can make mountain temperatures hard to forecast. Here is what station records and elevation data can—and cannot—show, and what we do about it.
Drive north from Gunnison, Colorado to Crested Butte and you climb 1,245 feet. Yet across the eleven Decembers from 2015 through 2025, Crested Butte reported the higher daily minimum on 145 of the 338 dates with data at both stations. Across the same eleven Julys, that happened twice.
Those two stations are about 39 kilometers apart, so this is not a measurement of one vertical column of air. It is still a vivid example of a well-established winter pattern: cold air can collect on mountain-valley floors while nearby higher terrain is milder.
Two sources of mountain forecast error are especially tractable. A model’s surface height can differ from the ground at a forecast location, and any height adjustment can be misleading when the near-surface temperature profile is weak or inverted. We will take them in order, and then describe what we do about each. Precipitation is a separate and harder problem, one that we will tackle in a subsequent post.
Problem one: the model is standing at the wrong elevation
A weather model divides the atmosphere into a horizontal grid and vertical layers. Each surface grid point has a representative surface elevation, so surface variables such as 2-meter temperature refer to the model’s represented terrain rather than every slope, bench and valley floor inside the cell. Models can also carry sub-grid land or orographic information, but that does not turn one grid-point temperature into a separate forecast for every elevation inside the cell.
The cells are large beside narrow valleys. NOAA’s Global Forecast System runs at about 13 kilometers. The European Centre’s IFS medium-range control forecast, formerly called HRES, runs at 9 kilometers. NOAA’s High-Resolution Rapid Refresh, among the finest widely used operational numerical guidance over the contiguous United States, runs at 3 kilometers.
Representing real terrain on those grids necessarily smooths features smaller than the grid spacing. The following calculation illustrates that loss of detail; it is not an extraction of any model’s orography.
In this one-dimensional illustration, much of the range-scale shape remains after a 3-kilometer moving average, while the 9- and 13-kilometer curves smooth narrower terrain much more strongly. East of Aspen the sampled ground climbs 2,000 feet in a little over two kilometers. At Aspen, the 13-kilometer moving average is 9,095 feet while the sampled town coordinate is 7,890 feet. Across the 90-kilometer section, the range of elevations in the illustrated profile falls from 7,129 feet to 4,044 feet after the 13-kilometer averaging operation. Those are properties of this illustration, not the surface heights of a named forecast model.
Another way to see the same thing is to sample the ground inside a coarse-grid-sized square.
The illustration puts 4,396 feet of sampled relief, several thousand people and part of a ski mountain inside one coarse-grid-sized square.
The consequence is a representativeness mismatch. If a model’s surface is 1,200 feet above the requested ground, its 2-meter temperature describes a different elevation. The resulting temperature error is not automatic or fixed: its size and even its sign depend on the simulated temperature profile, surface state and subsequent downscaling.
Finer grids help. The National Weather Service publishes forecasts on office grids of roughly 2.5-kilometer spacing, and its API reports the elevation associated with each forecast point. At 39.1911°N, 106.8175°W in Aspen, the API reports 7,897 feet, seven feet above the 7,890-foot terrain value used in our illustration.
Large differences can remain at narrow-valley coordinates. At 40.5883°N, 111.6383°W near the base of Alta, Utah, the API reports 9,800 feet while the terrain value used by our service is 8,543 feet. That comparison is coordinate-specific; it does not imply that every point in Alta, or every NWS grid cell in mountain terrain, has the same error.
Problem two: temperature does not always fall with height
One way to reduce an elevation mismatch is to transfer a grid-point temperature to the requested ground height. That requires an assumed relationship between temperature and elevation. The 6.5 °C-per-kilometer environmental lapse rate is the tropospheric value in the US Standard Atmosphere and is used in some fixed height-correction schemes. It is not a universal rule inside weather forecasts: numerical models diagnose 2-meter temperature from their evolving atmosphere and land surface, and post-processing systems use different, sometimes stability-dependent, adjustments. For example, ECMWF describes its forecast 2-meter temperature as a nonlinear interpolation that depends on near-surface stability.
A fixed 6.5 can be a poor description of air near mountain ground. Minder, Mote and Lundquist estimated regional lapse rates across the Cascades and found annual means of 3.9 to 5.2 °C per kilometer on the windward side of the range, falling to 2.5 to 3.5 for late-summer overnight temperatures, with large differences between the windward and lee sides. They also cautioned that individual station pairs can reflect local conditions and may not represent a regional lapse rate. For that reason we call the station-pair figures below apparent lapse rates: each is the temperature difference between two separated sites divided by the difference in their elevations, not a measurement of the air column above either one.
On many clear, calm nights—especially long winter nights—the near-surface relationship can flatten or reverse. The surface loses energy through net longwave radiation, the air near it cools, and dense air can drain downslope and collect on a valley floor. The result is a cold-air pool, with warmer air above colder valley air. Sunlight and convective mixing after sunrise, stronger wind, cloud changes or a passing weather system can weaken or remove it. Cold pools are not present on every winter night, and they can occur outside winter.
Read the blue line from the bottom up. The mean reported daily minimum at Butte is 23 °F higher than at Gunnison, whose listed elevation in the Applied Climate Information System (ACIS) is 2,538 feet lower. A hypothetical transfer using 6.5 °C per kilometer would instead make Butte about 9 °F colder, a 32 °F difference between that fixed-elevation calculation and the observed monthly contrast. The contrast is a mean of the available daily minima, not one unusual event. Because the sites are about 41 kilometers apart and use different observing networks and daily windows, it should not be interpreted as a measured vertical temperature profile.
To see how often that paired-site contrast occurs, we ran the same calculation across eleven and a half years of daily records.
The median apparent lapse rate for the paired daily minima is exactly zero. On 48.3 percent of paired dates, Schofield Pass reports the higher minimum. Only 5.3 percent fall within 1 °C per kilometer of 6.5. This shows that a fixed standard-atmosphere transfer does a poor job reproducing the relationship between these two sites’ daily minima. It does not establish the frequency of inversions over Gunnison, because that would require observations aligned in time and space through the same air column.
The paired daily maxima are much closer to the reference: their mean apparent lapse rate is 5.5 °C per kilometer. The larger and more operationally important departures occur in daily minima and are especially common during winter cold-pool regimes. Valley-floor towns are exposed to those regimes, but the all-season histogram above should not be described as a winter-only result.
The difficulty also appears in national guidance. A 2026 study by National Weather Service and NCEP scientists examined four 2023 cold-air-pool cases in Canaan Valley, West Virginia using National Blend of Models version 4.1; three of the four featured a pronounced cold pool. The blend failed to distinguish the valley floor from the rim, and its largest warm bias across those cases reached nearly 15 °C. This is case-study evidence about NBM v4.1, not proof that every cold pool is missed. NBM v5 became operational in May 2026, and that paper did not evaluate it.
The authors identified low weights assigned by the analysis quality-control system to unusually cold valley observations as a likely contributor. Such a reading can resemble a bad sensor when compared with warmer surrounding observations, so a defensible general-purpose quality-control system may give it too little influence. The paper did not prove that those observations were always discarded or that this was the only cause.
What we do about it
We are not claiming to have solved this. Cold pools are an open research problem and mountain forecasts will remain harder than plains forecasts.
The narrower claim we make is that part of mountain temperature error is a correctable representativeness and calibration problem: model and ground elevations can differ, a fixed height transfer can be inappropriate, and local forecast biases can vary by time and weather regime. The observations above do not quantify what share of total forecast error comes from each cause; that requires direct forecast verification and ablation tests.
Two things follow from that. First, forecast for real ground rather than for a grid cell. Second, keep a record of your own errors at each place, and use it.
For the first, our public locations are represented by forecast regions built from a 30-meter terrain model, so a valley floor and the ridge above it can be separate places with separate answers. The region containing our Aspen coordinate runs from 7,825 feet to 8,009 feet: 184 feet of relief, compared with 4,396 feet in the illustrative 13-kilometer square above. That comparison is between our region and the illustration, not between our region and an actual GFS cell.
For the second, here is what happens to a raw model field on its way to becoming a temperature we will publish.
Four of those steps are doing most of the work, and they map directly onto the two problems above.
Step one addresses the elevation-transfer problem. For each location and hour, we regress a model’s 2-meter temperatures in the surrounding grid cells against the corresponding surface heights. HRRR and GFS supply their own surface-height field; for NBM and ECMWF, whose files in our pipeline do not, we estimate cell heights from our terrain service. The resulting horizontal neighborhood slope is a proxy for the model’s local temperature–terrain relationship, not a direct vertical sounding. A regression is omitted when fewer than five usable cells remain or their height range is under 100 meters. Otherwise the computed slope is clamped to the configured range of −12 to +10 °C per kilometer rather than discarded. If no guarded slope is available, the pipeline uses its fixed-lapse fallback.
Step one also carries a rule that exists purely because of the physics above. When the forecast has to move downhill on a clear, calm night, the warming it applies is deliberately reduced. Lower terrain is not automatically warmer under those conditions, and a naive lapse-rate transfer would walk the forecast straight into the cold pool and get the sign wrong.
Step two is why we keep bias in three-hour blocks of the local day rather than as a single number per place. A valley’s forecast error can change sign between an overnight inversion and afternoon mixing. One all-hours correction can average opposite errors into a value that is useful at neither end of the day.
Step four anchors the near term. The first hours are nudged toward the latest RTMA surface analysis and nearby station observations, subject to age, distance, elevation and magnitude guards. This occurs before the learned residual temperature correction.
Step five is where our own record comes back in. After those temperature layers have been applied, a per-location ridge regression removes part of the residual using cloud cover, wind speed, local hour, season, model spread and forecast lead. One input is an explicit clear-and-calm-at-night term. It is trained on the preceding 45 days of selected archived forecasts scored against station observations, requires at least 120 samples, is shrunk toward zero when samples are thin and is capped at ±3 °C. It is skipped inside two hours, where fresh observations own the near-term adjustment.
That correction is deliberately linear and inspectable rather than a neural network. Rasp and Lerch’s study of 2-meter temperature at surface stations shows that auxiliary predictors and station-specific information are valuable. Their nonlinear network also improved materially over a linear network with similar inputs, so the paper does not establish that architecture is unimportant in our setting. Our choice of a ridge regression is an engineering choice for inspectability, sample efficiency and bounded behavior, not a conclusion proved by that paper.
We archive selected issued forecasts at seed and active verification locations, rather than every value served to every user. Hourly temperature snapshots use lead times of 0, 1, 2, 3, 6, 12 and 24 hours; daily cards and minute-precipitation products have separate sampling schedules. Eligible snapshots are scored against nearby station observations, and a rolling summary is on the site by location, with the mean error, bias and number of comparisons. That selected record supplies the training data for step five and contributes to the weights in step three.
In our next post, we will tackle mountain precipitation, which is even harder.