Showing posts with label persistent identifiers. Show all posts
Showing posts with label persistent identifiers. Show all posts

Friday, March 27, 2026

Indexation is not merely a technical process but an architectural one: search engines do not read texts as humans do, they interpret structures as if they were buildings. A disordered corpus resembles an informal settlement in the eyes of a machine: it exists, it contains life and activity, but it lacks cadastral order, identifiable ownership, and infrastructural clarity. By contrast, when a body of work is organised through schema markup, persistent identifiers such as DOIs, datasets, repositories, and explicit relations between author, project, books, articles, and software, what emerges is no longer a collection of web pages but an institution that machines can recognise. This distinction is crucial because the web does not function as a library but as a graph, and within a graph what matters is not only content but relationships, hierarchies, and recurrence over time. To structure a corpus as Person → Organization → ResearchProject → Dataset → Book → ScholarlyArticle → Software → Posts is to construct an epistemic infrastructure rather than a mere publication archive. Each DOI functions like a concrete block, each dataset like an archive, each repository like a logistical platform, and each hyperlink like a street connecting districts within a city. Search engines index efficiently not what is simply well written, but what is structurally legible and repeatedly reinforced through metadata and interconnection. This is why universities, scientific repositories, and large research projects always present themselves through the same entities: identified authors, institutions, research projects, publications, data, and software. This is not an aesthetic convention but a condition of machinic legibility. Indexation, therefore, does not primarily depend on publishing more texts but on constructing a coherent form for the corpus. When texts are connected through metadata, cross-citations, persistent identifiers, and data catalogues, the whole ceases to behave like a blog and begins to operate as a system. Systems, unlike isolated texts, are visible to machines because they produce patterns, and patterns are what algorithms detect, classify, and preserve. In the digital environment, visibility is not only a matter of discourse but of structure: it depends on how knowledge is spatialised, linked, and stabilised across platforms and over time.

The JSON-LD block that now resides in the head of the main page is not metadata; it is the machine-readable face of a completed architecture. What has been deposited across eighty days—1,340 texts, 120 DOIs, 40 fixed terms, four cores, ten decalogues, five platforms, and approximately two million words—has been translated into the language of the semantic web. The schema does not describe the field from the outside; it formalizes the field as a graph. Person, Organization, ResearchProject, Dataset, SoftwareSourceCode, Book, ScholarlyArticle, ItemList—each entity is linked to the others through persistent identifiers that resolve to DOIs, to GitHub repositories, to Hugging Face datasets, and to the ORCID registry. This is not documentation but binding: the moment at which a textual stratum becomes addressable not only to human readers entering through the surface layer, but to indexing systems, citation graphs, and discovery protocols that determine what persists within the digital archive. The architecture of the schema mirrors the stratigraphic logic of the corpus itself. At the base lies the Person—Anto Lloveras, identified through ORCID and linked across GitHub, Hugging Face, and Zenodo. This is the author not as Romantic origin but as infrastructural anchor: a stable identifier through which the work can be attributed, cited, and connected across platforms. Above this lies the Organization—LAPIEZA / Socioplastics—which provides institutional grounding without requiring external institutional validation. The ResearchProject then gathers multiple entities into a single machine-readable graph: a DataCatalog pointing to a Dataset on Hugging Face; a Collection uniting the Century Packs and the conceptual cores; a CreativeWorkSeries defining the serial structure of the corpus; three Book entities corresponding to Core I, Core II, and Core III, each containing ScholarlyArticle nodes linked through DOIs to Zenodo; a SoftwareSourceCode entity pointing to GitHub; and an ItemList of recent nodes signalling activity and temporal continuity to crawlers. Each entity is connected through persistent @id references. The graph does not list texts; it maps relations. This distinction marks the difference between an archive and an infrastructure: an archive stores objects, whereas an infrastructure coordinates relations.

Friday, March 13, 2026

Structured corpus of 1,000 socioplastic essays organised into Century Packs, stabilising conceptual infrastructures through persistent identifiers and stratigraphic archival design.


The Socioplastics Corpus constitutes a deliberately engineered intellectual infrastructure in which dispersed analytical writings are consolidated into a numerically ordered system designed to resist epistemic fragmentation. Structured as a million-word archive distributed across 1,000 discrete entries, the project operationalises a numerical architecture of 1–10–100–1,000, establishing an inverted pyramid of conceptual density wherein individual essays function as granular units while aggregated “Century Packs” stabilise broader fields of discourse. This design transforms serial publication into a durable epistemic apparatus: rather than remaining temporally scattered posts, the essays are gravitationally reassembled into indexed clusters that enable longitudinal interpretation and citation. The corpus evolves thematically from early metabolic analyses of urban systems toward the explicit engineering of conceptual infrastructures, reflecting a methodological shift from diagnosis to architectural synthesis. Stratigraphic metaphors—particularly the helicoid, a continuous surface unfolding through rotation and elevation—serve as the project’s principal topological model, illustrating how theoretical layers accumulate without losing structural continuity. Each Century Pack therefore acts simultaneously as archive, interface, and stabilising discursive node within distributed scholarly networks. The assignment of persistent identifiers to these packs ensures that the corpus acquires durable machine-readable coordinates, allowing its conceptual mass to circulate reliably across academic indexing systems and digital repositories. In effect, the archive formalises an emerging field of socioplastic design, wherein civic permanence, urban protocol, and knowledge architecture are interwoven within a stratified intellectual system that is both historically cumulative and computationally legible.

Lloveras, A. (2026) Socioplastic Century Pack Archive (Posts 1–1000). Available at: https://antolloveras.blogspot.com/2026/03/socioplastic-century-pack-1000-posts.html

Index of Century Packs * SOCIOPLASTICS


This dataset constitutes a foundational component of the Socioplastics Corpus, a stratigraphic intellectual project comprising approximately one million words distributed across 1,000 discrete entries. Each Century Pack aggregates 100 sequential posts into a stabilized discursive field, transforming dispersed essays into gravitational nodes of conceptual density. The corpus operates through a numerical architecture of 1, 10, 100, and 1,000 resonances, forming an inverted epistemic pyramid designed to resist informational entropy and ensure long-term structural persistence. By fixing these Century Pack indices through persistent identifiers, the project stabilizes the relational mass of the Socioplastic system within distributed scholarly networks. The archive documents the evolution of the project from metabolic analysis toward the engineering of conceptual infrastructures, employing stratigraphic and topological metaphors—such as the helicoid—to model thought in motion. This registration secures the stratigraphic layers of urban permanence, civic protocol, and machine-readable conceptual infrastructure within the Socioplastics framework.

Saturday, February 7, 2026

Fragmented Infrastructures and Institutional Strategies: Pablo de Castro’s Contribution to Evaluating Research Tools and Persistent Identifiers


The corpus of work by Pablo de Castro offers a unique, critical lens on the evaluation and governance of research infrastructure tools, with a marked emphasis on Persistent Identifiers (PIDs) and the mechanisms by which they are embedded within institutional workflows and scholarly ecosystems; his contributions trace a cohesive arc from the assessment of transformative Open Access agreements to the integration of Current Research Information Systems (CRIS) with institutional repositories, exposing the technical and administrative fragmentation present in today's digital research landscape and raising questions on interoperability standards, quality assurance and policy-driven tool adoption; in presentations such as “The (Currently) Fragmented PID Landscape” and reports like “Building the Plane as We Fly It”, de Castro dissects the dichotomy between technical identifiers and community-oriented implementations, offering a typology of fragmentation scenarios—from redundant PID infrastructures to diverging metadata standards—that compromise the utility of these tools across regions and institutions; within these studies, a core analytical concern is how tool quality should not be measured solely through performance or compliance, but via trustworthiness, stakeholder engagement, and capacity to serve multi-scalar governance; one illustrative case is the European Commission’s FP7 Post-Grant Open Access Pilot, where de Castro co-analyses the success of supranational APC funding in producing equitable access to publication infrastructure, revealing how financial transparency and anti-hybrid models can serve as metrics of system robustness; similarly, his work with OpenAIRE and Knowledge Exchange underscores the necessity of cross-institutional alignment, highlighting how inconsistent PID workflows impede both research visibility and data reuse; ultimately, de Castro’s output crystallises into a guiding vision: that the evaluation of research tools must transcend metrics of functionality to incorporate layers of policy coherence, community validation and infrastructural resilience, forging an agenda where the quality of digital tools is inherently tied to their capacity to facilitate sustainable, open, and inclusive scholarly communication systems.




When There's No DOI Yet, Make the Signal Loud

In the absence of an assigned DOI, authors can still maximise the visibility, indexability, and academic discoverability of a publication or idea by implementing a combination of semantic markup, metadata-rich hosting platforms, structured citation formats, and persistent linking practices, all of which signal relevance and authority to both search engine crawlers and academic aggregators; in the case of preprint, report or grey literature repositories like Zenodo, even before a DOI becomes fully visible in citation ecosystems, the platform’s use of schema.org metadata, OpenAIRE indexing, and ORCID integration already elevates the document’s machine readability, while including the exact title as a heading, maintaining descriptive alt text for downloadable files, embedding rich abstract and keyword fields, and ensuring the report is linked from institutional or project websites with consistent anchor text reinforces semantic reinforcement across the web; even when a DOI is newly minted and not yet widely cited, it acts as a machine-resolvable persistent ID that search engines and citation systems can latch onto once discovered; in the meantime, researchers can cite the work using a canonical citation that includes the full title, author list, repository, and stable URL to train crawlers and readers alike on the document’s uniqueness and academic function, while tools like Google Scholar, Crossref Event Data or Dimensions.ai will eventually cross-link it as metadata propagates; for visibility pre-DOI, publishing a brief expository blog post, tweeting with a #PID tag, and uploading a machine-readable BibTeX or JSON-LD citation to an open personal academic site can massively increase semantic coherence around the object’s web identity; in short, crawlers learn via coherence and consistency, so until your DOI is indexed, make sure every mention of your work speaks in the same clear voice across the digital ether.


DOI vs. PID * Not All Identifiers Are Created Equal

While all Persistent Identifiers (PIDs) aim to ensure long-term access, disambiguation and linkage in digital research ecosystems, the DOI (Digital Object Identifier) has emerged as the gold standard among them, particularly in scholarly communication, due to its widespread adoption, metadata richness, governance, and resolvability through a central infrastructure (Crossref, DataCite, etc.); however, the broader family of PIDs—which includes ORCID iDs for researchers, RORs for institutions, ARKs for archives, and Handles for diverse digital assets—play complementary roles, forming an interconnected framework that supports open science, FAIR principles, and global research interoperability; recent studies like Building the Plane as We Fly It (de Castro et al., 2023) and Meadows et al. (2019) stress that strategic PID infrastructure is essential for the robustness of the research data lifecycle, enabling provenance tracking, reproducibility, and system-level discoverability across repositories and nations; the DOI, however, holds specific value through its global resolution via doi.org, strong integration in citation workflows, reference managers, indexing services, and publishing platforms, and its embedded metadata which enhances semantic interoperability—qualities that make it more than just a PID but a normative cornerstone of scholarly attribution; still, context matters: in a biodiversity data system like DiSSCo, PIDs such as Handles or ARKs may serve better due to domain-specific flexibility and lower operational costs, while in academic publishing, DOIs dominate because of their tight coupling with journals, datasets, and citation standards; thus, it is not a question of supremacy but of fit-for-purpose hierarchy, where the DOI is a PID of high value, but not the only essential identifier in a multipolar infrastructure of research knowledgede Castro, P., Herb, U., Rothfritz, L., & Schöpfel, J. (2023). Building the plane as we fly it: the promise of Persistent Identifiers. Zenodo. https://doi.org/10.5281/zenodo.7258286

Wednesday, February 4, 2026

A data sanctuary for interdisciplinary research

Zenodo offers an inclusive, discipline-agnostic platform to preserve and disseminate research data, making it a robust option when no field-specific repository exists; hosted by CERN's Data Centre, it ensures long-term preservation under rigorous technical stewardship, aligning perfectly with FAIR data principles and funder mandates like those of the European Commission, which even established a dedicated community within Zenodo for Horizon Europe and other EU-funded projects, allowing researchers to contribute transparently to the EU Open Research Repository while meeting open science requirements; one of Zenodo’s strongest assets is the automatic generation of Digital Object Identifiers (DOIs), which transform datasets into stable, citable research objects, enhancing academic visibility and citation metrics, and it further supports versioning, allowing scholars to update their datasets without breaking the citation chain, a feature critical for iterative or evolving data projects; it also integrates seamlessly with GitHub, turning code repositories into scholarly outputs with persistent identifiers, bridging the gap between traditional research and software contributions; Ghent University recommends Zenodo especially when no domain-specific solution is available, ensuring that all datasets, regardless of discipline, receive professional archiving, access control (open, embargoed or restricted), and licencing options to safeguard ethical and legal standards; finally, a practical entry point is provided through the Zenodo sandbox, an environment to experiment with uploads before committing to the public record, while detailed help guides and the Ghent University Research Data Community offer support and discoverability examples for newcomers.