Statistical inference for stochastic simulation models — theory and application
Abstract
Statistical models are the traditional choice to test scientific theories when observations, processes or boundary conditions are subject to stochasticity. Many important systems in ecology and biology, however, are difficult to capture with statistical models. Stochastic simulation models offer an alternative, but they were hitherto associated with a major disadvantage: their likelihood functions can usually not be calculated explicitly, and thus it is difficult to couple them to well-established statistical theory such as maximum likelihood and Bayesian statistics. A number of new methods, among them Approximate Bayesian Computing and Pattern-Oriented Modelling, bypass this limitation. These methods share three main principles: aggregation of simulated and observed data via summary statistics, likelihood approximation based on the summary statistics, and efficient sampling. We discuss these principles as well as advantages and caveats of these methods, and demonstrate their potential for integrating stochastic simulation models into a unified framework for statistical modelling.
What the paper shows and why it matters (AI-generated)
Stochastic simulation models had a well-known catch: their likelihood usually can't be written down, cutting them off from standard maximum-likelihood and Bayesian machinery. This paper lays out how newer methods — Approximate Bayesian Computation, Pattern-Oriented Modelling — sidestep that catch via summary statistics and efficient sampling, giving ecologists a principled way to test simulation models against data instead of treating them as unfalsifiable. Nearly 400 citing papers later, the framework is still being applied well outside its original ecological context, from epidemiological microsimulation to deep-learning-based diversification inference.