Benchmarking Generative Models for Weather Data Assimilation on Real Station Observations

MIT S.M. Thesis · 2025

Ruizhe Huang1  ·   Qidong Yang1  ·   Jonathan Giezendanner1  ·   Sherrie Wang1

1Massachusetts Institute of Technology

Side-by-side assimilation outputs across six methods, for 10 m wind ($u$, $v$), 2 m temperature, and 2 m dewpoint. Flow Guidance and Diffusion‑SDA sharpen local structure around stations without artifacts; D‑Flow and stop‑gradient collapse under sparsity; 3D‑Var under‑corrects far from stations.

Abstract

In weather forecasting, the WeatherBench benchmark catalyzed progress by providing a common evaluation protocol that enabled fair comparison across deep learning architectures. No equivalent exists for generative data assimilation (DA) — the task of integrating sparse, real‑world observations into atmospheric fields using learned priors. Existing generative approaches have been evaluated under disparate, often idealized conditions (typically synthetic pseudo‑observations sampled from the grid), making it hard to compare design choices.

This thesis establishes the first unified benchmark for generative weather DA on real, noisy station observations. We evaluate flow‑matching methods (Flow Guidance, D‑Flow, FlowDPS), diffusion (Diffusion‑SDA), and classical 3D‑Var, along with latent‑space variants of Flow Guidance and 3D‑Var, under a single experimental protocol: 11,849 ground weather stations from NCEP MADIS (Meteorological Assimilation Data Ingest System) across CONUS (contiguous United States), ERA5 reanalysis fields for four surface variables (10 m u- and v-wind, 2 m temperature, 2 m dewpoint), 360 evaluation snapshots in 2023, and consistent train/val/test station splits with 1,778 held‑out test stations.

The central finding is that guidance with a pretrained generative prior is an effective recipe for weather DA on real observations. Starting from pure Gaussian noise and with no access to ERA5 at inference, Flow Guidance outperforms 3D‑Var (35.7 % vs. 33.3 % root-mean-square error (RMSE) reduction relative to ERA5) despite 3D‑Var being given the ERA5 field twice, both as background regularizer and as initial condition. The learned prior carries more information about atmospheric state structure than the combination of background covariance and background field in classical DA.

Key Findings

  1. A learned prior beats the classical background.

    Flow Guidance (35.7 %) exceeds 3D‑Var (33.3 %) in RMSE reduction over ERA5, without ERA5 at inference. The generative prior encodes atmospheric structure more richly than $\mathbf{B} + \bar{\mathbf{x}}_b$.

  2. Flow Matching $\approx$ Diffusion via complementary tradeoffs.

    Both are instances of the same guidance template $\nabla\log p(\mathbf{x}\mid\mathbf{y}) \approx \nabla\log p(\mathbf{x}) + \nabla\log p(\mathbf{y}\mid\mathbf{x})$ with different network parameterizations. Flow matching integrates accurately but has a weaker, indirectly‑derived score; diffusion has a strong Langevin corrector but a less accurate predictor. Each framework compensates for its weaker component through its stronger one, leaving them at empirical parity.

  3. The Langevin corrector is thermodynamic‑specific.

    Removing it costs $-9.7$ to $-10.6$ pp on $T_{2\mathrm{m}}/D_{2\mathrm{m}}$ (dense) and $-16.6$ to $-20.7$ pp under sparsity, but barely touches wind. A $\sim\!2.4\times$ normalized‑space residual imbalance explains why off‑manifold drift disproportionately hurts thermodynamic variables.

  4. Variable‑mixing in the latent autoencoder is unnecessary.

    A spatial‑mixing‑only autoencoder matches a variable+spatial mixer on both reconstruction and assimilation, with 4 × fewer AE parameters.

  5. Latent DA: memory and speed win.

    Under sparsity, latent Flow Guidance matches pixel Flow Guidance at $\sim 40\%$ lower memory and $\sim 19\%$ lower wall‑clock time.

  6. Generative corrections propagate farther.

    At the median 5‑NN station distance (48 km), Flow Guidance and Diffusion‑SDA achieve $\sim 28\%$ RMSE reduction vs. $\sim 22\%$ for 3D‑Var.

Method

All guidance methods sample from the posterior $p(\mathbf{x}\mid\mathbf{y}) \propto p(\mathbf{x})\,p(\mathbf{y}\mid\mathbf{x})$ by combining a pretrained unconditional generator (flow or diffusion) with an observation‑likelihood gradient along the sampling trajectory. Methods differ in how they approximate the score of the prior and how the measurement term is injected.

Flow matching vs. diffusion
Figure 1. Flow matching learns a deterministic velocity field; diffusion learns a score and denoises via the reverse SDE. Both define a path from noise to data; guidance steers that path with $\nabla\log p(\mathbf{y}\mid\mathbf{x})$.
Unified schematic of Flow Guidance, D‑Flow, FlowDPS, Diffusion‑SDA, and 3D‑Var showing their sampling trajectories and where the observation term is injected
Figure 2. Unified view of the five generative methods plus 3D‑Var. Flow Guidance and Diffusion‑SDA share the same guidance template; D‑Flow replaces inner guidance with an outer noise‑optimization loop; stop‑gradient drops the Jacobian of the network.

Benchmark

DataERA5 (0.25°), surface variables $u_{10}$, $v_{10}$, $T_{2\mathrm{m}}$, $D_{2\mathrm{m}}$.
Observations11,849 MADIS ground stations (CONUS), real noisy measurements.
SplitsStation-level train / val / test; 1,778 held-out test stations.
Evaluation360 snapshots from 2023 — dense & sparse protocols.
MetricRMSE reduction over the ERA5 background, averaged across variables.
MADIS station coverage across CONUS for dense (11,849) and sparse (1,000) benchmarks with train/val/test splits
Figure 3. MADIS station coverage across CONUS.

Results

Distance Decay

How far do corrections propagate from observation locations?

Dense benchmark: RMSE improvement over ERA5 as a function of distance to the nearest observation, all methods
Dense benchmark. Generative corrections stay positive out to ${\sim}120$ km, well beyond the median 5-NN distance.
Sparse benchmark: RMSE improvement over ERA5 as a function of distance to the nearest observation, all methods
Sparse benchmark. The gap widens; D‑Flow and stop-gradient collapse.

Latent DA Pareto

Accuracy vs. compute Pareto frontier comparing pixel and latent Flow Guidance across compression depths
Latent Flow Guidance matches pixel‑space quality at ~40 % lower memory and ~19 % lower wall-clock.

Regional Case Studies

Florida peninsula u10 case study: ERA5 background, observations, and assimilation outputs across methods
Florida ($u_{10}$). Coastal winds over ungauged water: generative methods $\sim\!72\%$ RMSE reduction vs. $66\%$ for 3D‑Var.
Great Lakes u10 case study: ERA5 background, observations, and assimilation outputs across methods
Great Lakes ($u_{10}$). ERA5 overestimates wind by $\sim\!4.5$ m/s; generative methods $53\%$ RMSE reduction vs. $45\%$ for 3D‑Var.
LA basin t2m case study: ERA5 background, observations, and assimilation outputs across methods
LA Basin ($T_{2\mathrm{m}}$). Dense stations, complex terrain: all methods converge at $\sim\!36\%$ RMSE reduction (parity regime).
Rocky Mountains t2m case study: ERA5 background, observations, and assimilation outputs across methods
Rocky Mountains ($T_{2\mathrm{m}}$). Near-uniform $\sim\!6.5$ K cold bias; generative RMSE 3.3–3.5 K edges 3D‑Var (4.0 K).

Perturbation Analysis

Perturbation amplification vs. injection time for flow and diffusion scores under matched noise schedules
Perturbation amplification ($\log_{10}$ scale) vs. injection time $t$. Top: score-quality isolation (flow vs. diffusion score on the same flow ODE). Bottom: noise-schedule isolation (flow ODE vs. VP‑SDE with diffusion score). Left: guidance-direction perturbations; right: random-direction.

BibTeX

Cite the paper:

@article{huang2025benchmark,
  title   = {Benchmarking Generative Models for Weather Data Assimilation
             on Real Station Observations},
  author  = {Huang, Ruizhe and Yang, Qidong and Giezendanner, Jonathan and Wang, Sherrie},
  year    = {2025}
}

Cite the thesis:

@mastersthesis{huang2025thesis,
  title   = {Benchmarking Generative Models for Weather Data Assimilation
             on Real Station Observations},
  author  = {Huang, Ruizhe},
  school  = {Massachusetts Institute of Technology},
  year    = {2025},
  type    = {{S.M.} thesis},
  address = {Cambridge, MA, USA}
}