Hierarchical Nearest-Neighbor Gaussian Process Models for Large Geostatistical Datasets

Abhirup Datta, Sudipto Banerjee, Andrew O. Finley, Alan E. Gelfand

Research output: Contribution to journalArticle

Abstract

Spatial process models for analyzing geostatistical data entail computations that become prohibitive as the number of spatial locations become large. This article develops a class of highly scalable nearest-neighbor Gaussian process (NNGP) models to provide fully model-based inference for large geostatistical datasets. We establish that the NNGP is a well-defined spatial process providing legitimate finite-dimensional Gaussian densities with sparse precision matrices. We embed the NNGP as a sparsity-inducing prior within a rich hierarchical modeling framework and outline how computationally efficient Markov chain Monte Carlo (MCMC) algorithms can be executed without storing or decomposing large matrices. The floating point operations (flops) per iteration of this algorithm is linear in the number of spatial locations, thereby rendering substantial scalability. We illustrate the computational and inferential benefits of the NNGP over competing methods using simulation studies and also analyze forest biomass from a massive U.S. Forest Inventory dataset at a scale that precludes alternative dimension-reducing methods. Supplementary materials for this article are available online.

Original languageEnglish (US)
Pages (from-to)800-812
Number of pages13
JournalJournal of the American Statistical Association
Volume111
Issue number514
DOIs
Publication statusPublished - Apr 2 2016
Externally publishedYes

    Fingerprint

Keywords

  • Bayesian modeling
  • Gaussian process
  • Hierarchical models
  • Markov chain Monte Carlo
  • Nearest neighbors
  • Predictive process
  • Reduced-rank models
  • Sparse precision matrices
  • Spatial cross-covariance functions

ASJC Scopus subject areas

  • Statistics and Probability
  • Statistics, Probability and Uncertainty

Cite this