The open web was built upon a productive ambiguity: publishing a page made it available to strangers, search engines and unknown future readers without requiring the author to specify every legitimate encounter in advance. Crawling initially reinforced this settlement by creating indexes that returned attention to original locations, but generative systems have altered the purpose and scale of retrieval. Texts, images, recordings and code may now be collected not merely to help users locate them, but to construct models capable of summarising, imitating and partially substituting for the materials ingested. The difference is functional rather than mystical. Reading, indexing, copying, statistical analysis and the production of competing outputs are not identical operations simply because they occur inside one technical pipeline. The United States Copyright Office separates data acquisition, training, retrieval and outputs because each stage creates distinct questions of consent, market effect and lawful use; Stober and Dornis argue that generative training cannot simply inherit every legal assumption attached to text-and-data mining; Shanklin and colleagues propose contextual copyleft as an attempt to make licensing obligations travel from training data toward resulting models. These approaches disagree on important points, but they share a recognition that accessibility is not a universal licence. CitationalCommitment becomes the practical threshold at which computational ingestion leaves enough evidence to identify major sources, dataset composition, rights assumptions, exclusions and subsequent transformations. This would not solve compensation by itself, but it would make negotiation possible. Creators could distinguish between discovery, non-commercial research, model training, stylistic reproduction and direct commercial substitution; collective organisations could negotiate where individual licensing is impossible; repositories could expose machine-readable permissions; and model providers could demonstrate how opt-outs were handled without pretending that one technical signal resolves every jurisdiction. The future of crawling will depend less upon a universal prohibition or universal licence than upon layers of declared use. Search engines, cultural archives, scientific corpora and commercial generative models serve different purposes and should not inherit identical permissions merely because they employ similar collection technologies. The web remains valuable because it permits unexpected circulation, quotation and recombination, but openness survives only when contributors can recognise the terms under which their work travels. Once extraction becomes invisible, participation begins to resemble surrender. The objective is neither to close the web nor to freeze culture into proprietary fragments, but to make the passage from publication to dataset sufficiently inspectable that public availability and industrial appropriation no longer collapse into the same category.
U.S. Copyright Office (2025). Copyright and Artificial Intelligence, Part 3: Generative AI Training.
Stober, S. and Dornis, T. W. (2026). Generative AI Training and Copyright Law.
Shanklin, G., Hine, E., Novelli, C., Schroder, T. and Floridi, L. (2025). The Case for Contextual Copyleft.
Birhane, A., Prabhu, V. U. and Kahembwe, E. (2021). Multimodal Datasets: Misogyny, Pornography, and Malignant Stereotypes.