Showing posts with label DOIInfrastructure. Show all posts
Showing posts with label DOIInfrastructure. Show all posts

Friday, March 27, 2026

Indexation is not merely a technical process but an architectural one: search engines do not read texts as humans do, they interpret structures as if they were buildings. A disordered corpus resembles an informal settlement in the eyes of a machine: it exists, it contains life and activity, but it lacks cadastral order, identifiable ownership, and infrastructural clarity. By contrast, when a body of work is organised through schema markup, persistent identifiers such as DOIs, datasets, repositories, and explicit relations between author, project, books, articles, and software, what emerges is no longer a collection of web pages but an institution that machines can recognise. This distinction is crucial because the web does not function as a library but as a graph, and within a graph what matters is not only content but relationships, hierarchies, and recurrence over time. To structure a corpus as Person → Organization → ResearchProject → Dataset → Book → ScholarlyArticle → Software → Posts is to construct an epistemic infrastructure rather than a mere publication archive. Each DOI functions like a concrete block, each dataset like an archive, each repository like a logistical platform, and each hyperlink like a street connecting districts within a city. Search engines index efficiently not what is simply well written, but what is structurally legible and repeatedly reinforced through metadata and interconnection. This is why universities, scientific repositories, and large research projects always present themselves through the same entities: identified authors, institutions, research projects, publications, data, and software. This is not an aesthetic convention but a condition of machinic legibility. Indexation, therefore, does not primarily depend on publishing more texts but on constructing a coherent form for the corpus. When texts are connected through metadata, cross-citations, persistent identifiers, and data catalogues, the whole ceases to behave like a blog and begins to operate as a system. Systems, unlike isolated texts, are visible to machines because they produce patterns, and patterns are what algorithms detect, classify, and preserve. In the digital environment, visibility is not only a matter of discourse but of structure: it depends on how knowledge is spatialised, linked, and stabilised across platforms and over time.

The JSON-LD block that now resides in the head of the main page is not metadata; it is the machine-readable face of a completed architecture. What has been deposited across eighty days—1,340 texts, 120 DOIs, 40 fixed terms, four cores, ten decalogues, five platforms, and approximately two million words—has been translated into the language of the semantic web. The schema does not describe the field from the outside; it formalizes the field as a graph. Person, Organization, ResearchProject, Dataset, SoftwareSourceCode, Book, ScholarlyArticle, ItemList—each entity is linked to the others through persistent identifiers that resolve to DOIs, to GitHub repositories, to Hugging Face datasets, and to the ORCID registry. This is not documentation but binding: the moment at which a textual stratum becomes addressable not only to human readers entering through the surface layer, but to indexing systems, citation graphs, and discovery protocols that determine what persists within the digital archive. The architecture of the schema mirrors the stratigraphic logic of the corpus itself. At the base lies the Person—Anto Lloveras, identified through ORCID and linked across GitHub, Hugging Face, and Zenodo. This is the author not as Romantic origin but as infrastructural anchor: a stable identifier through which the work can be attributed, cited, and connected across platforms. Above this lies the Organization—LAPIEZA / Socioplastics—which provides institutional grounding without requiring external institutional validation. The ResearchProject then gathers multiple entities into a single machine-readable graph: a DataCatalog pointing to a Dataset on Hugging Face; a Collection uniting the Century Packs and the conceptual cores; a CreativeWorkSeries defining the serial structure of the corpus; three Book entities corresponding to Core I, Core II, and Core III, each containing ScholarlyArticle nodes linked through DOIs to Zenodo; a SoftwareSourceCode entity pointing to GitHub; and an ItemList of recent nodes signalling activity and temporal continuity to crawlers. Each entity is connected through persistent @id references. The graph does not list texts; it maps relations. This distinction marks the difference between an archive and an infrastructure: an archive stores objects, whereas an infrastructure coordinates relations.