Gaussian Predictive Process Models for Large Spatial Data Sets
Top Cited Papers
- 9 July 2008
- journal article
- Published by Oxford University Press (OUP) in Journal of the Royal Statistical Society Series B: Statistical Methodology
- Vol. 70 (4), 825-848
- https://doi.org/10.1111/j.1467-9868.2008.00663.x
Abstract
Summary. With scientific data available at geocoded locations, investigators are increasingly turning to spatial process models for carrying out statistical inference. Over the last decade, hierarchical models implemented through Markov chain Monte Carlo methods have become especially popular for spatial modelling, given their flexibility and power to fit models that would be infeasible with classical methods as well as their avoidance of possibly inappropriate asymptotics. However, fitting hierarchical spatial models often involves expensive matrix decompositions whose computational complexity increases in cubic order with the number of spatial locations, rendering such models infeasible for large spatial data sets. This computational burden is exacerbated in multivariate settings with several spatially dependent response variables. It is also aggravated when data are collected at frequent time points and spatiotemporal process models are used. With regard to this challenge, our contribution is to work with what we call predictive process models for spatial and spatiotemporal data. Every spatial (or spatiotemporal) process induces a predictive process model (in fact, arbitrarily many of them). The latter models project process realizations of the former to a lower dimensional subspace, thereby reducing the computational burden. Hence, we achieve the flexibility to accommodate non-stationary, non-Gaussian, possibly multivariate, possibly spatiotemporal processes in the context of large data sets. We discuss attractive theoretical properties of these predictive processes. We also provide a computational template encompassing these diverse settings. Finally, we illustrate the approach with simulated and real data sets.Funding Information
- National Science Foundation (DMS-0706870, DEB05-16198)
- National Institutes of Health (1-R01-CA95995, 2-R01-ES07750)
This publication has 33 references indexed in Scilit:
- Computational techniques for spatial logistic regression with large data setsComputational Statistics & Data Analysis, 2007
- Approximate Likelihood for Large Irregularly Spaced Spatial DataJournal of the American Statistical Association, 2007
- Approximately optimal spatial design approaches for environmental health dataEnvironmetrics, 2006
- Spatial modelling using a new class of nonstationary covariance functionsEnvironmetrics, 2006
- Functional Data AnalysisPublished by Wiley ,2005
- Sequential, Bayesian Geostatistics: A Principled Method for Large Data SetsGeographical Analysis, 2005
- Space–Time Covariance FunctionsJournal of the American Statistical Association, 2005
- Flexible Spatial Models for Kriging and Cokriging Using Moving Averages and the Fast Fourier Transform (FFT)Journal of Computational and Graphical Statistics, 2004
- Spatially Balanced Sampling of Natural ResourcesJournal of the American Statistical Association, 2004
- Nonconjugate Bayesian Estimation of Covariance Matrices and its Use in Hierarchical ModelsJournal of the American Statistical Association, 1999