Model comparison
The best predictor is not the best explanation
Eleven models in three families, from 2-parameter physics to Gaussian processes, on the same targets, folds and metrics. Prediction ranks them one way; transfer to forcing they have never seen ranks them another.
Regime A dataset
Monthly temperature, NH extratropics | land
Skill vs climatology, blocked CV · Berkeley Earth
- physical (insolation-driven)
- statistical (descriptive)
- flexible (GP, boosting)
Transfer to latitudes the model never saw
Seasonal-cycle R², Berkeley Earth
- one-pole, shared τ and k
- zero-lag insolation
- mean cycle of training rows
…and to a different surface or hemisphere
Seasonal-cycle R², Berkeley Earth
- one-pole, shared τ and k
- zero-lag insolation
- mean cycle of training rows
Orbital time scales
R² of each model on the proxy records
- LR04
- EPICA Dome C
- Cheng 2016
What the comparison can and cannot show
- Every seasonal model hits the same ceiling, the monthly climatology, because orbital inputs in the modern era are a calendar: they carry no year-to-year information. Predictive skill cannot tell a physical model from a description.
- Transfer can. Only the insolation-driven model predicts a latitude it has not seen, and only inside one physical regime.
- Neural forecasting was deliberately not run. With a 3-parameter model at the ceiling and no interannual signal in the inputs, a network could only rediscover the climatology or the trend.
- Nothing here is an intervention. Attribution rests on the physics of the forcing and on counterfactual forcings (see sensitivity), never on skill.