Sunday, August 9, 2026

A dataset is an argument about what counts — Media Theory / Digital Humanities · Dataset Logic · datasets, classification, data epistemology · AssumptionPresence — Anto Lloveras · Socioplastics · LAPIEZA-LAB, Madrid · 2026




'Raw data' is one of the more successful fictions in contemporary research culture, and it survives mainly because almost nobody stops to notice that raw is doing all the work. By the time a dataset exists as a dataset — as rows, columns, a schema, an inclusion criterion — someone has already decided what counts as a data point and what counts as noise unworthy of capture. That decision is an argument, even when it is never written down as one, because every exclusion silently asserts what does not matter enough to record. AssumptionPresence names the moment an unstated methodological choice becomes structurally present in a dataset without ever appearing in its documentation — invisible in the columns, load-bearing in the conclusions those columns make possible. Treating a dataset as pre-theoretical, rather than as a compressed argument about categorisation, is how bad inferences get laundered into apparently neutral numbers. The cleanest-looking spreadsheet is often the one hiding the most contestable decision.


Halpern, O. (2014) Beautiful Data: A History of Vision and Reason since 1945. Durham, NC: Duke University Press.

Bender, E.M., Gebru, T., McMillan-Major, A. and Shmitchell, S. (2021) 'On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?' In: Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency. New York: ACM, pp. 610-623.

boyd, d. and Crawford, K. (2012) 'Critical Questions for Big Data', Information, Communication & Society, 15(5), pp. 662-679.

Devlin, J. et al. (2018) 'BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding'. arXiv:1810.04805.

Burrell, J. (2016) 'How the Machine Thinks: Understanding Opacity in Machine Learning Algorithms'. Big Data & Society, 3(1), pp. 1-12.