Tuesday, September 29, 2026

Priem, J., Piwowar, H. and Orr, R. (2022) ‘OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts’, STI Conference 2022, Granada.



Priem, Piwowar and Orr provide the clearest demonstration that corpus scale becomes analytically useful only through explicit data modelling. OpenAlex is defined as a heterogeneous directed graph composed of scholarly entities and their relations, rather than as a gigantic list of publications. At the time of the paper it indexed approximately 209 million works, with around 50,000 being added daily, alongside authors, venues, institutions and concepts. Every entity receives a persistent OpenAlex identifier, while external canonical identifiers—DOI for works, ORCID for authors, ISSN-L for venues, ROR for institutions and Wikidata identifiers for concepts—support interoperability beyond the platform. The conceptual importance lies in the graph schema: scale is governed by typed nodes, edges, disambiguation and normalisation. Openness extends to data, API and source code, making the infrastructure inspectable and reusable. OpenAlex thus bridges bibliometrics and open scholarly infrastructure, showing how numerical magnitude becomes knowledge architecture only when entities remain addressable, relational and computationally traversable.