Cross-validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure

David R. Roberts, Volker Bahn, Simone Ciuti, Mark S. Boyce, Jane Elith, Gurutzeta Guillera-Arroita, Severin Hauenstein, José J. Lahoz-Monfort, Boris Schröder, Wilfried Thuiller, David I. Warton, Brendan A. Wintle, Florian Hartig, Carsten F. Dormann

Ecography, 40(8), 913–929 (2017)
Cite this
@article{roberts2017cross,
  author = {David R. Roberts and Volker Bahn and Simone Ciuti and Mark S. Boyce and Jane Elith and Gurutzeta Guillera-Arroita and Severin Hauenstein and José J. Lahoz-Monfort and Boris Schröder and Wilfried Thuiller and David I. Warton and Brendan A. Wintle and Florian Hartig and Carsten F. Dormann},
  title = {Cross-validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure},
  journal = {Ecography},
  volume = {40},
  number = {8},
  pages = {913–929},
  year = {2017},
  doi = {10.1111/ecog.02881},
}

DOI: 10.1111/ecog.02881
Cited by 3,547 (Google Scholar) · 2,966 (OpenAlex), as of 07 September 2026

View article (DOI)

Abstract

Ecological data often show temporal, spatial, hierarchical (random effects), or phylogenetic structure. Modern statistical approaches are increasingly accounting for such dependencies, but when performing cross-validation, these structures are regularly ignored, resulting in serious underestimation of predictive error. One cause for the poor performance of uncorrected (random) cross-validation is dependence structures in the data that persist as dependence structures in model residuals, violating the assumption of independence. Even more concerning, because often overlooked, is that structured data also provides ample opportunity for overfitting with non-causal predictors. Block cross-validation, where data are split strategically rather than randomly, can address these issues, but the blocking strategy must be carefully considered: blocking may unwittingly induce extrapolations by restricting the ranges or combinations of predictor variables available for model training, thus overestimating interpolation errors, while deliberate blocking in predictor space may improve error estimates when extrapolation is the modelling goal. Here we review the ecological literature on non-random and blocked cross-validation approaches, and provide a series of simulations and case studies showing that, for all instances tested, block cross-validation is nearly universally more appropriate than random cross-validation if the goal is predicting to new data or predictor space, or for selecting causal predictors. We recommend that block cross-validation be used wherever dependence structures exist in a dataset, even if no correlation structure is visible in the fitted model residuals.

What the paper shows and why it matters (AI-generated)

Standard cross-validation assumes data points are independent, an assumption ecological data — spatial, temporal, hierarchical, phylogenetic — routinely violates, and violating it silently makes a model look far better than it really is. This paper works through why, and shows across simulations and real case studies that block cross-validation, splitting data to respect the dependence structure rather than at random, corrects the problem almost universally when the goal is predicting to new data. With well over 2,000 citing papers spanning remote sensing, landslide susceptibility mapping and species distribution modelling, it has become one of the standard references for honest model evaluation across quantitative ecology and beyond.