Using synthetic data to evaluate the benefits of large field plots for forest biomass estimation with LiDAR

Fabian Fassnacht, Hooman Latifi, Florian Hartig

Remote Sensing of Environment, 213, 115–128 (2018)
Cite this
@article{fassnacht2018using,
  author = {Fabian Fassnacht and Hooman Latifi and Florian Hartig},
  title = {Using synthetic data to evaluate the benefits of large field plots for forest biomass estimation with LiDAR},
  journal = {Remote Sensing of Environment},
  volume = {213},
  pages = {115–128},
  year = {2018},
  doi = {10.1016/j.rse.2018.05.007},
}

DOI: 10.1016/j.rse.2018.05.007
Cited by 60 (Google Scholar) · 43 (OpenAlex), as of 07 September 2026

View article (DOI)

Abstract

With the maturation of methods for estimating aboveground forest biomass by remote sensing, researchers increasingly need test data, particularly ground reference data, that are large enough to fine-tune existing approaches and test their robustness under diverse conditions. In this context, realistic synthetic datasets present an interesting alternative to costly and limited field data. Here, we present a new approach to simulate realistic canopy height and cover type data by combining an individual-tree forest simulator with real LiDAR point clouds of individual trees, and demonstrate its utility by re-examining the influence of field plot size on the predictive power of remote-sensing models for biomass estimation. Our approach with a complete (wall-to-wall) field reference dataset and matching synthetic remote sensing data allowed us to not only perform internal cross-validations with field plots that were used to fit the model, as in studies with real data, but to also consider the quality of model predictions on a standardized spatial grid or across the entire region. Our results confirm earlier reports of smaller predictive errors with increased field plot sizes under internal model validation, but we show that this is mainly an artifact of comparing the models with the same data they were fit on. When validating on a grid with standardized scale, smaller field plots performed almost equally well as larger field plots, and even outperformed them once we accounted for the fact that increasing the plot size means fewer field plots can be obtained for the same effort. We conclude that synthetic remote sensing datasets are a useful tool for method testing, and could be instrumental in improving remote sensing methodology more broadly.

What the paper shows and why it matters (AI-generated)

The received wisdom in forest remote sensing was that bigger field plots make for better biomass models — evidence largely drawn from validating those models on the same data used to fit them. By simulating realistic LiDAR data with a matching wall-to-wall ground truth, the authors could validate independently for the first time, and found smaller plots perform almost as well once that circularity is removed, and better once effort per plot is accounted for. The synthetic-data approach itself has aged well: it now underpins deep-learning LiDAR pipelines the original paper never anticipated.