Gas Optics Benchmark
This section documents the experimental results comparing different neural network architectures for emulating gas optics computations in the RTnn framework, using the CAMS RRTMGP Gas Optics dataset. The goal is to predict shortwave Rayleigh scattering optical depth (tau_sw_ray) from seven input atmospheric state variables.
Note
This benchmark uses a different configuration than the other benchmarks:
Output channels: 224 (g-points) vs 4 or 2 for other benchmarks
Hidden size: 32
Layers: 3
Learning rate: 0.001
Input features: 7 (tlay, play, h2o, o3, co2, n2o, ch4)
Target: tau_sw_ray (Rayleigh scattering optical depth)
Model Performance Comparison
This presents a comprehensive comparison of three different neural network architectures trained on the CAMS RRTMGP Gas Optics dataset. All models were trained with identical hyperparameters where applicable to ensure fair comparison.
Experiment Configuration
All models were trained with the following common configuration:
Dataset: CAMS RRTMGP Gas Optics data
Training years: 2009-2018
Testing year: 2014
Input features: 7 channels (tlay, play, h2o, o3, co2, n2o, ch4)
Output channels: 224 (g-points for tau_sw_ray)
Sequence length: 60 vertical levels
Normalization: minmax scaling (per-feature)
Loss function: Huber loss (β=0.5, δ=1.0)
Learning rate: 0.001
Batch size: 4
Epochs: 500
Hidden size: 32
Number of layers: 3
Dropout: 0.1
Model Architectures
Three different architectures were evaluated:
LSTM (Long Short-Term Memory) - Traditional recurrent architecture - Model identifier: lstm_h32_l3_d0d1 - Parameters: 75,232
GRU (Gated Recurrent Unit) - Simplified recurrent architecture - Model identifier: gru_h32_l3_d0d1 - Parameters: 60,064
FCN (Fully Connected Network) - Baseline dense architecture - Model identifier: fcn_h32_l3_d0d1 - Parameters: 460,416
Note
The Transformer architecture was not evaluated for this benchmark due to the high-dimensional output space (224 g-points).
Performance Metrics
The following metrics were used for evaluation (validation set, final epoch):
Loss: Huber loss value
NMAE: Normalized Mean Absolute Error (normalized by target range)
NMSE: Normalized Mean Squared Error (normalized by target variance)
MAE: Mean Absolute Error (in physical units)
MSE: Mean Squared Error (in physical units)
R²: Coefficient of determination
Runtime: Training time per epoch (in seconds)
Quantitative Comparison
The table below shows the performance comparison for the tau_sw_ray prediction.
Model |
Loss ↓ |
NMAE ↓ |
NMSE ↓ |
R² ↑ |
MAE ↓ |
MSE ↓ |
Runtime (s/epoch) |
|---|---|---|---|---|---|---|---|
LSTM |
8.96e-05 |
5.47e-03 |
1.33e-02 |
0.999822 |
2.52e-03 |
1.34e-02 |
~4.3 |
GRU |
1.18e-04 |
6.39e-03 |
1.54e-02 |
0.999763 |
2.93e-03 |
1.54e-02 |
~4.3 |
FCN |
1.18e-03 |
2.22e-02 |
5.01e-02 |
0.997493 |
1.02e-02 |
5.02e-02 |
~4.0 |
Note: ↓ indicates lower is better, ↑ indicates higher is better. MAE and MSE are reported in physical units.
Quantitative Comparison - Training vs Validation
Model |
Train Loss |
Valid Loss |
|---|---|---|
LSTM |
9.58e-05 |
8.96e-05 |
GRU |
1.25e-04 |
1.18e-04 |
FCN |
2.23e-03 |
1.18e-03 |
Key Findings
Best Overall Performance: The LSTM model achieves the highest R² score (0.999822), lowest loss (8.96e-05), lowest MAE (2.52e-03), and lowest MSE (1.34e-02) for the tau_sw_ray prediction, demonstrating excellent capability in capturing gas optics processes.
Parameter Efficiency: The GRU model uses only 60,064 parameters (20% fewer than LSTM’s 75,232) while still achieving excellent performance (R² = 0.999763), making it the most parameter-efficient choice.
Runtime Efficiency: All models are extremely fast (~4 seconds per epoch) due to the small model size and efficient data loading.
Generalization Gap: LSTM shows excellent generalization with minimal gap between training and validation. GRU also shows good generalization. The FCN shows a larger gap and significantly lower accuracy.
Model Complexity: The FCN has the largest parameter count (460,416) but performs the worst, suggesting that the sequential nature of the atmospheric profiles is important for this task.
Recommendations
Based on the comparison results:
For maximum accuracy: Use LSTM.
For parameter efficiency: Use GRU (20% fewer parameters, only slightly lower accuracy).
For real-time applications: All models are suitable (~4 s/epoch), with the FCN being slightly faster but with significantly lower accuracy.
For memory-constrained environments: Use GRU which has the smallest parameter count (60,064).
Diagnostic Visualizations
This section presents diagnostic plots generated at the final epoch for the LSTM model, showing prediction quality across all validation samples.
Aggregated Results (All Samples)
The following figure shows the density scatter plots (hexbin) for all validation samples, comparing predicted vs observed values for tau_sw_ray. The diagonal red dashed line represents perfect prediction (y=x), and the R² score is displayed in each panel.
Figure 1: Aggregated validation results for LSTM model at epoch 590. The figure shows predicted vs observed tau_sw_ray values. The color scale represents the logarithm of point density.
The aggregated results demonstrate excellent agreement between predictions and observations, with R² exceeding 0.9998.