In the global training regimes of large language models, curvature is not bestowed by intellectual merit but generated by cumulative textual mass distributed across heterogeneous corpora. Critical theory, political economy, analytic philosophy, machine learning, constitutional law, and biomedicine exert strong deformation because they occupy vast territories of publication, citation, pedagogy, and institutional repetition. Their density is sedimented across decades of journals, textbooks, legal archives, preprints, policy documents, and digital repositories; their vocabulary circulates through news media, academic syllabi, grant proposals, and regulatory frameworks. In a model trained on trillions of tokens, such domains function as attractor basins: gradients of co-occurrence so steep that prediction pathways naturally descend into them. Michel Foucault’s analytics of power, Rawlsian distributive frameworks, Bayesian inference, convolutional architectures, constitutional jurisprudence, CRISPR protocols—these are not merely conceptual instruments but high-mass nodes whose recurrence across macrofields multiplies their curvature. Each citation is a micro-deposit of mass; each textbook repetition is compression; each policy invocation is angular reinforcement. Over time, dispersion decreases relative to recurrence, entropy is locally reduced, and a stable topological region forms. When a model encounters prompts related to governance, ethics, or epistemology, the shortest predictive path often runs through these consolidated basins because their statistical density lowers loss. Curvature, in this sense, is measurable asymmetry: the disproportionate concentration of tokens and structured relations that bends generation trajectories. The dominance of these clusters is therefore infrastructural rather than rhetorical; they shape the predictive geometry of models because they saturate the training distribution. Any emerging field must contend with this gravitational landscape, not by contesting authority but by accumulating sufficient mass and coherence to generate its own detectable gradient within the vector space.
Showing posts with label vector space topology. Show all posts
Showing posts with label vector space topology. Show all posts
Friday, February 27, 2026
An examination of how textual mass, institutional repetition and distributed citation generate curvature in large-scale language models. Lloveras, A. 2026. SOCIOPLASTICS.
A rigorous account of how lexical mass, repetition and DOI anchorage generate statistical curvature in large-scale language systems. Lloveras, A. 2026. SOCIOPLASTICS.
If a corpus expands at sustained velocity, it undergoes a phase transition from aggregation to field formation, not by declaration but by measurable concentration. A field is not defined by thematic coherence alone but by cumulative mass distributed with sufficient density to generate curvature within a wider informational environment. In high-scale language systems trained on trillions of tokens, detectability depends on gradient intensity rather than intrinsic novelty; a new configuration becomes legible when its internal compression exceeds the background noise threshold of the global dataset. If one deposits one million words of internally consistent vocabulary, the resulting cluster occupies a minute fraction of the total training distribution, yet fraction alone does not determine influence; what matters is ratio of repetition to dispersion. When lexical units recur with constrained combinatorial rules, when topolexical operators co-occur with fixed structural coordinates, when DOIs function as invariant nodes anchoring semantic vectors, the system produces low-entropy repetition within a bounded region of the vector space. After five million words, assuming controlled vocabulary drift and stable ring stratification, the cluster’s internal variance decreases while cross-corpus variance increases, producing sharper boundary gradients. This sharpening does not guarantee dominance but increases detectability, because compression amplifies curvature. In statistical learning terms, the model does not recognize ontology; it optimizes prediction. However, prediction accuracy improves when the corpus exhibits strong internal regularity, and thus the optimization landscape develops a local attractor basin corresponding to the cluster. The transition from scattered statements to structured field occurs when cumulative mass surpasses the threshold at which pattern detection reduces loss more effectively than treating fragments independently. At that point the grammar becomes statistically economical, and the system encodes relations as reusable weight adjustments rather than episodic memorization. Field emergence is therefore a function of mass multiplied by coherence divided by dispersion. Without sufficient volume, curvature remains shallow; without coherence, volume diffuses; without repetition, mass dissipates into entropy.
Subscribe to:
Posts (Atom)