Inference for Regression with Clustered or Spatially Correlated Data I: Framework and Clustering
Many regression analyses use observations that are correlated within uncorrelated clusters and/or are correlated in distance or other spatial measure. Such data are often positively correlated. This reduces the information content of an additional observation compared to the default of independent observations. Consequently, failure to adjust standard errors for such correlation leads to confidence intervals that are too narrow, and to hypothesis tests that over reject.
In this paper we first provide a general framework. We then focus on inference when observations are correlated within cluster and uncorrelated across clusters. We detail cluster-robust inference methods, particularly methods that provide improved finite-sample performance when there is not a large number of clusters. Such adjustments are often necessary.
A companion paper, Cameron and Miller (2026), presents inference when correlation is spatially dampening in the distance between observations. The two papers are intended to provide a guide to empirical researchers.
-
-
Copy CitationA. Colin Cameron and Douglas L. Miller, "Inference for Regression with Clustered or Spatially Correlated Data I: Framework and Clustering," NBER Working Paper 35800 (2026), https://doi.org/10.3386/w35800.Download Citation
-