Linear Regressions with Combined Data
We study linear regressions in a context where the outcome of interest and some of the covariates are observed in two different datasets that cannot be matched. Traditional approaches obtain point identification by relying, often implicitly, on exclusion restrictions. We show that without such restrictions, coefficients of interest can still be partially identified, with the sharp bounds taking a simple form. We obtain tighter bounds when variables observed in both datasets, but not included in the regression of interest, are available, even if these variables are not subject to specific restrictions. We develop computationally simple and asymptotically normal estimators of the bounds. Finally, we apply our methodology to estimate racial disparities in patent approval rates and to evaluate the effect of patience and risk-taking on educational performance.
-
-
Copy CitationXavier D'Haultfoeuille, Christophe Gaillac, and Arnaud Maurel, "Linear Regressions with Combined Data," NBER Working Paper 34507 (2025), https://doi.org/10.3386/w34507.Download Citation
-