OpenCitations Meta RDF dataset of all bibliographic metadata and its provenance information
This dataset contains all the bibliographic metadata and its provernance information (in JSON-LD format) included in OpenCitations Meta.
The data and the provenance are organized through a complex structure of folders and subfolders, that allows you to quickly find any entity from its URI. The first level consists of the following folders, that are provided zipped and separately:
- [folder "ar"]: contains the data and provenance of the responsible agent type entities (http://purl.org/spar/pro/RoleInTime);
- [folder "br"]: contains the data and provenance of the entities of type bibliographic resource (http:///purl.org/spar/fabio/Expression);
- [folder "id"]: contains the data and provenance of the identifier entities (http://purl.org/spar/datacite/Identifier);
- [folder "ra"]: contains the data and provenance of the responsible agent type entities (http://xmlns.com/foaf/0.1/Agent);
- [folder "ra"]: contains the data and provenance of resource embodiment entities (http://purl.org/spar/fabio/Manifestation).
The inner folders are named through the supplier prefix of the contained entities. It is a prefix that allows you to recognize the entity membership index (e.g., OpenCitations Meta corresponds to 06*0).
After that, the folders have numeric names, which refer to the range of contained entities. For example, the 10000 folder contains entities from 1 to 10000. Inside, you can find the zipped RDF data.
At the same level, additional folders containing the provenance are named with the same criteria already seen. Then, the 1000 folder includes the provenance of the entities from 1 to 1000. The provenance is located inside a folder called prov, also in zipped JSON-LD format.
For example, data related to the entity is located in the folder /br/06250/10000/1000/1000.zip, while information about provenance in /br/06250/10000/1000/prov/1000.zip
This version of the dataset contains:
- 105,953,699 bibliographic entities
- 338,173,282 authors and 2,523,200 editors (counted by their roles, without disambiguating individuals)
- 691,262 publication venues
- 36,679 publishers
The weight of the sum of the archives is 38.1 GB on an NTFS filesystem, which does not vary once extracted because they contain zipped JSON files. We recommend processing such files as zipped without extracting them.
Additional information about OpenCitations Meta at the official webpage.