ClimateOrbitOrbital forcing · seasonal climate · attribution

Model comparison

The best predictor is not the best explanation

Eleven models in three families, from 2-parameter physics to Gaussian processes, on the same targets, folds and metrics. Prediction ranks them one way; transfer to forcing they have never seen ranks them another.
Regime A dataset

Monthly temperature, NH extratropics | land

Skill vs climatology, blocked CV · Berkeley Earth

  • physical (insolation-driven)
  • statistical (descriptive)
  • flexible (GP, boosting)
Skill is the MSE improvement over a monthly climatology refitted on the same training years: 0 means no better than climatology. R² differences of 0.003 in the extratropics are 15–75% differences on this anomaly scale. The flexible models' CV gains come from interpolating the warming trend inside held-out blocks, and they do not survive extrapolation.

Transfer to latitudes the model never saw

Seasonal-cycle R², Berkeley Earth

  • one-pole, shared τ and k
  • zero-lag insolation
  • mean cycle of training rows
One τ and k shared across the training rows, driven by each held-out row's own insolation. Whole 10° bands are held out because neighbouring rows are dependent. The forcing-free comparator copies the mean training cycle.

…and to a different surface or hemisphere

Seasonal-cycle R², Berkeley Earth

  • one-pole, shared τ and k
  • zero-lag insolation
  • mean cycle of training rows
No model transfers across domains: τ and k are properties of a surface and hemisphere, not universal constants. NH land learns τ = 32 d; the ocean would choose 106 d. SH land wants k = 0.032 against NH land's 0.088 K per W m⁻².

Orbital time scales

R² of each model on the proxy records

  • LR04
  • EPICA Dome C
  • Cheng 2016
Sinusoids at the La2004 periods with free amplitude and phase beat the physical insolation response for LR04 and EPICA by fitting the 100-kyr band, which no linear insolation pathway can produce: predictive association outperforming a physical account. The GP on age cannot extrapolate at all. Only the monsoon record is well explained by a physical linear response.

What the comparison can and cannot show

  • Every seasonal model hits the same ceiling, the monthly climatology, because orbital inputs in the modern era are a calendar: they carry no year-to-year information. Predictive skill cannot tell a physical model from a description.
  • Transfer can. Only the insolation-driven model predicts a latitude it has not seen, and only inside one physical regime.
  • Neural forecasting was deliberately not run. With a 3-parameter model at the ceiling and no interannual signal in the inputs, a network could only rediscover the climatology or the trend.
  • Nothing here is an intervention. Attribution rests on the physics of the forcing and on counterfactual forcings (see sensitivity), never on skill.