Harmonizing and Combining Large Datasets - An Application to Firm-Level Patent and Accounting Data

Grid Thoma; Salvatore Torrisi; Alfonso Gambardella; Dominique Guellec; Bronwyn H. Hall; Dietmar Harhoff

doi:10.3386/w15851

Harmonizing and Combining Large Datasets - An Application to Firm-Level Patent and Accounting Data

Grid Thoma, Salvatore Torrisi, Alfonso Gambardella, Dominique Guellec, Bronwyn H. Hall & Dietmar Harhoff

Working Paper 15851

DOI 10.3386/w15851

Issue Date March 2010

This paper discusses methods for the harmonization and combination of large-scale patent and trademark datasets with each other and other sources of data. Dictionary- and rule-based approaches to the consolidation of applicant names in patent data are presented and shown to have both benefits and drawbacks in isolation. We combine the two methods and develop a set of rules and dictionaries to consolidate European, Patent Cooperation Treaty (PCT) and US patent data with firm accounting data. The resulting data encompass about 131,000 patent applicant names from 46 countries, covering 58.8 percent of EPO applications and 50.6 percent of PCT applications by business organizations during the time period from 1979 to 2008. For US data, the resulting dataset includes around 54,000 assignee names and 51.3 percent of US granted patents during approximately the same time period.

We thank Jim Bessen, Hélène Dernis, Megan MacGarvie, Paola Giuri, Stine Grodal, Myriam Mariani, Kazu Motohashi, Teruo Okazaki, James Rollinson, Philipp Sander, Georg von Graevenitz, Stefan Wagner, Norihiko Yamano, Maria Pluvia Zuniga, and the participants at the PATSTAT Users' Meeting in Paris in March 2008, seminars at the Ludwig-Maximilians-Universität München and Università L. Bocconi in Milan for very fruitful discussions. We also thank Armando Benincasa and Luisa Quarta from Bureau Van Dijk for clarifications about the structure of the Amadeus database and its changes over time. Thoma and Hall are grateful to the Kauffman Foundation for support of some of this work. The views expressed herein are those of the authors and do not necessarily reflect the views of the National Bureau of Economic Research.
Copy Citation

Grid Thoma, Salvatore Torrisi, Alfonso Gambardella, Dominique Guellec, Bronwyn H. Hall, and Dietmar Harhoff, "Harmonizing and Combining Large Datasets - An Application to Firm-Level Patent and Accounting Data," NBER Working Paper 15851 (2010), https://doi.org/10.3386/w15851.

Download Citation

MARC RIS BibTeΧ

Harmonizing and Combining Large Datasets - An Application to Firm-Level Patent and Accounting Data

Related

Topics

Programs

Working Groups

More from the NBER