Model averaging in ecology: a review of Bayesian, information-theoretic, and tactical approaches for predictive inference
Abstract
In ecology, the true causal structure for a given problem is often not known, and several plausible models and thus model predictions exist. It has been claimed that using weighted averages of these models can reduce prediction error, as well as better reflect model selection uncertainty, but these claims are often demonstrated by isolated examples. Here we review the mathematical foundations of model averaging along with the diversity of approaches available. We explain that the error in model-averaged predictions depends on each model’s predictive bias and variance, as well as the covariance in predictions between models, and uncertainty about model weights. We show that model averaging is particularly useful if the predictive error of contributing model predictions is dominated by variance, and if the covariance between models is low – conditions that will often be met for noisy data, which predominate in ecology. A general recommendation on which of the many averaging methods to use is difficult, because performance is often context dependent, and estimating weights creates some additional uncertainty, so estimated model weights may not always outperform arbitrary fixed weights such as equal weights. When averaging a set of models with many inadequate models, however, estimating model weights will typically be superior to equal weights. We also investigate the quality of the confidence intervals calculated for model-averaged predictions, showing that they differ greatly in behaviour and seldom manage to achieve nominal coverage. Our overall recommendations stress the importance of non-parametric methods such as cross-validation for a reliable uncertainty quantification of model-averaged predictions.
What the paper shows and why it matters (AI-generated)
Model averaging had been recommended in ecology on the strength of isolated success stories rather than a settled understanding of when it actually helps. This review works out the conditions: averaging pays off when a model's error is dominated by variance rather than bias and when the models being combined are only weakly correlated, both common for noisy ecological data — but its warning about confidence intervals is arguably the more consequential finding, since model-averaged intervals rarely reach their nominal coverage regardless of method. With over 400 citing papers spanning ecotoxicology, species distribution modelling and disease emergence, it has become one of the field's default references for the practice.