Gas Optics Benchmark ==================== This section documents the experimental results comparing different neural network architectures for emulating gas optics computations in the RTnn framework, using the CAMS RRTMGP Gas Optics dataset. The goal is to predict shortwave Rayleigh scattering optical depth (``tau_sw_ray``) from seven input atmospheric state variables. .. note:: This benchmark uses a different configuration than the other benchmarks: - **Output channels**: 224 (g-points) vs 4 or 2 for other benchmarks - **Hidden size**: 32 - **Layers**: 3 - **Learning rate**: 0.001 - **Input features**: 7 (tlay, play, h2o, o3, co2, n2o, ch4) - **Target**: tau_sw_ray (Rayleigh scattering optical depth) Model Performance Comparison ---------------------------- This presents a comprehensive comparison of three different neural network architectures trained on the CAMS RRTMGP Gas Optics dataset. All models were trained with identical hyperparameters where applicable to ensure fair comparison. Experiment Configuration ~~~~~~~~~~~~~~~~~~~~~~~~ All models were trained with the following common configuration: - **Dataset**: CAMS RRTMGP Gas Optics data - **Training years**: 2009-2018 - **Testing year**: 2014 - **Input features**: 7 channels (tlay, play, h2o, o3, co2, n2o, ch4) - **Output channels**: 224 (g-points for tau_sw_ray) - **Sequence length**: 60 vertical levels - **Normalization**: minmax scaling (per-feature) - **Loss function**: Huber loss (β=0.5, δ=1.0) - **Learning rate**: 0.001 - **Batch size**: 4 - **Epochs**: 500 - **Hidden size**: 32 - **Number of layers**: 3 - **Dropout**: 0.1 Model Architectures ~~~~~~~~~~~~~~~~~~~ Three different architectures were evaluated: 1. **LSTM** (Long Short-Term Memory) - Traditional recurrent architecture - Model identifier: `lstm_h32_l3_d0d1` - Parameters: 75,232 2. **GRU** (Gated Recurrent Unit) - Simplified recurrent architecture - Model identifier: `gru_h32_l3_d0d1` - Parameters: 60,064 3. **FCN** (Fully Connected Network) - Baseline dense architecture - Model identifier: `fcn_h32_l3_d0d1` - Parameters: 460,416 .. note:: The Transformer architecture was not evaluated for this benchmark due to the high-dimensional output space (224 g-points). Performance Metrics ~~~~~~~~~~~~~~~~~~~ The following metrics were used for evaluation (validation set, final epoch): - **Loss**: Huber loss value - **NMAE**: Normalized Mean Absolute Error (normalized by target range) - **NMSE**: Normalized Mean Squared Error (normalized by target variance) - **MAE**: Mean Absolute Error (in physical units) - **MSE**: Mean Squared Error (in physical units) - **R²**: Coefficient of determination - **Runtime**: Training time per epoch (in seconds) Quantitative Comparison ~~~~~~~~~~~~~~~~~~~~~~~ The table below shows the performance comparison for the ``tau_sw_ray`` prediction. .. list-table:: Performance comparison for tau_sw_ray predictions (validation set) :header-rows: 1 :widths: 15, 12, 12, 12, 12, 15, 15, 12 :align: center * - Model - Loss ↓ - NMAE ↓ - NMSE ↓ - R² ↑ - MAE ↓ - MSE ↓ - Runtime (s/epoch) * - LSTM - 8.96e-05 - 5.47e-03 - 1.33e-02 - 0.999822 - 2.52e-03 - 1.34e-02 - ~4.3 * - GRU - 1.18e-04 - 6.39e-03 - 1.54e-02 - 0.999763 - 2.93e-03 - 1.54e-02 - ~4.3 * - FCN - 1.18e-03 - 2.22e-02 - 5.01e-02 - 0.997493 - 1.02e-02 - 5.02e-02 - ~4.0 *Note: ↓ indicates lower is better, ↑ indicates higher is better. MAE and MSE are reported in physical units.* Quantitative Comparison - Training vs Validation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ .. list-table:: Training vs Validation metrics (final epoch) :header-rows: 1 :widths: 20, 20, 20 :align: center * - Model - Train Loss - Valid Loss * - LSTM - 9.58e-05 - 8.96e-05 * - GRU - 1.25e-04 - 1.18e-04 * - FCN - 2.23e-03 - 1.18e-03 Key Findings ~~~~~~~~~~~~ **Best Overall Performance**: The **LSTM** model achieves the highest R² score (0.999822), lowest loss (8.96e-05), lowest MAE (2.52e-03), and lowest MSE (1.34e-02) for the tau_sw_ray prediction, demonstrating excellent capability in capturing gas optics processes. **Parameter Efficiency**: The **GRU** model uses only 60,064 parameters (20% fewer than LSTM's 75,232) while still achieving excellent performance (R² = 0.999763), making it the most parameter-efficient choice. **Runtime Efficiency**: All models are extremely fast (~4 seconds per epoch) due to the small model size and efficient data loading. **Generalization Gap**: LSTM shows excellent generalization with minimal gap between training and validation. GRU also shows good generalization. The FCN shows a larger gap and significantly lower accuracy. **Model Complexity**: The FCN has the largest parameter count (460,416) but performs the worst, suggesting that the sequential nature of the atmospheric profiles is important for this task. Recommendations ~~~~~~~~~~~~~~~ Based on the comparison results: 1. **For maximum accuracy**: Use **LSTM**. 2. **For parameter efficiency**: Use **GRU** (20% fewer parameters, only slightly lower accuracy). 3. **For real-time applications**: All models are suitable (~4 s/epoch), with the FCN being slightly faster but with significantly lower accuracy. 4. **For memory-constrained environments**: Use **GRU** which has the smallest parameter count (60,064). Diagnostic Visualizations ------------------------- This section presents diagnostic plots generated at the final epoch for the LSTM model, showing prediction quality across all validation samples. Aggregated Results (All Samples) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The following figure shows the density scatter plots (hexbin) for all validation samples, comparing predicted vs observed values for tau_sw_ray. The diagonal red dashed line represents perfect prediction (y=x), and the R² score is displayed in each panel. .. figure:: ../../images/validation_epoch590_hexbin.png :align: center :figwidth: 90% **Figure 1:** Aggregated validation results for LSTM model at epoch 590. The figure shows predicted vs observed tau_sw_ray values. The color scale represents the logarithm of point density. The aggregated results demonstrate excellent agreement between predictions and observations, with R² exceeding 0.9998.