Vector Theory

David M. Berry

This is a shortened draft version of the full article.


The published version of this article is now available at: 

https://link.springer.com/article/10.1007/s13347-026-01162-w 


"There is no software"

Friedrich Kittler, 1995.


Figure 1: A vector and its mapping into "vector space".

Kittler was right, but, perhaps, not in the way he intended. When he declared that software dissolved into hardware operations of voltage differences, he was making a materialist claim about symbolic computation, about computer code that could, in principle, be read, that is parsed, and followed through logical gates. His analysis traced how each layer of software abstraction conceals the operations of the layer beneath it, an opacity that serves power, but the material substrate he traced still operated through discrete, Boolean logic, through states which could, in theory, be read. In the 90s, the "hidden" layer was the voltage. Today, the hidden layer is the sub-symbolic weight structure. What Kittler could not have anticipated was that computation would escape the symbolic paradigm altogether. However, the weight matrices of contemporary AI resist reading because there is nothing to read, not because they are hidden beneath abstraction layers. While the weights aren't "readable" by humans, they are "calculable". The crisis isn't that they are invisible, but that they are asemantic. Whilst voltage is physical materiality, vector space is a form of mathematical materiality. It is a materiality of relation rather than substance.

This shift, from the discrete bit to the high-dimensional vector, from Boolean logic to what is called "cosine similarity", constitutes what may be the most significant transformation in computational epistemology since the digital turn itself. And yet, our critical frameworks remain largely calibrated for an earlier paradigm. Media theory from McLuhan through Kittler analysed containers and codes, channels and protocols, the materiality of inscription and transmission. These frameworks helped us to understand the shift from analogue to digital with considerable power. As I have argued elsewhere, the digital itself represented a particular mode of computational rationality (Berry 2014), but that analysis was calibrated for a symbolic regime that I argue is now being superseded. Our existing theories will struggle with a computational regime in which there is, strictly speaking, nothing to read, no code to parse, no symbolic layer to interpret, only the vast, opaque geometry of weighted connections and the interfaces through which we connect to it. We may no longer need digital theory, we need vector theory. 

Figure 2: In word2vec it shows the relationship between "king" and "queen" in a word embedding model.

In vector space, definition works differently. To define a word, a concept, or an image is is to give it a position, to locate it. That is, to identify a region within a multidimensional field where that concept clusters with related concepts. So, for example, the relationship between "king" and "queen" in a word embedding model is not one of logical opposition or categorical hierarchy. Rather, it is a geometric displacement, a vector offset that also maps the relationship between "man" and "woman" (Mikolov et al. 2013: 2). This is the famous word2vec result, king minus man plus woman equals queen, and it reveals something strange about how these systems organise meaning. Semantic content is stored  in spatial relationships, rather than being stored in symbols. 

The latent continuum describes what vector space does to the gaps between categories. Traditional representational systems are lossy in a specific sense, they lose what falls between their categories. A filing system has folders and gaps between folders. A database has fields and null values. A dictionary has entries and absences. The digital, as I have argued (Berry 2011), operates through a process of discretisation that imposes mathematical form on continuous phenomena, necessarily excluding whatever resists formalisation.

Figure 3: Smoothly interpolating between a photograph of a dog and a photograph of a cat.

In a diffusion model or a GAN, it is possible to smoothly interpolate between a photograph of a dog and a photograph of a cat (see Lewis-Kraus 2026), moving through intermediate states, a 'dat' or a 'cog', that have no name in natural language but that possess perfectly determinate mathematical coordinates within the model's latent space. These are not errors. They are addresses in the manifold, as valid mathematically as any point that corresponds to a named concept. The system makes no distinction between the actual and the interpolated, between what has been seen and what can be computed. 

But we should immediately ask, unvisited by whom? And whose training data shaped the density of the field in the first place? The statistical density of latent space is an artefact of curation, of the selection and preprocessing of training corpora, which is to say an artefact of capital, of labour, and of the political economy of data collection.

Figure 4: Showing the"sandwich" of discrete/operationally continuous/discrete in LLMs.

Tokens can be taken as the "user interface" of language, a key part of the discretisation of inputs and outputs that makes these systems work. The machine takes our continuous world, turns it into discrete tokens, and then immediately re-projects them into the geometric regime of vector space. I want to call this threshold the tokenisation horizon, this is the point at which human language is discretised, decomposed, and re-projected into a computational regime where it is no longer language but geometry. Beyond the horizon, there are no words, only vectors. It is also where human agency might be said to be surrendered to the field. On this side of the horizon, a prompt functions as a command, an instruction issued with intent. Beyond it, the same tokens become a perturbation of a probability field, and what happens next is determined by the geometry of the space the tokens enter rather than what the prompter meant. This "sandwich" of discrete/operationally continuous/discrete is where the power (and the error) resides.

The most difficult dimension of vector theory, and in many ways the most important, concerns what we might call the sub-symbolic layer, the vast architecture of weights that constitutes the actual computational substrate of contemporary AI systems. This is not an entirely new problem and Paul Smolensky identified what he called the subsymbolic paradigm in 1988, arguing that connectionist systems operate at a level of description that sits below the symbolic. That is, that the patterns of activation across neural networks do not map neatly onto the concepts, rules, themes and categories of symbolic representation (Smolensky 1988: 3, 9). What Smolensky identified as the "subsymbolic" is what we might now, in the age of models with hundreds of billions of parameters, call the dark matter of computation. In using dark matter as an analogy, I aim to capture its opacity of scale. Whilst we cannot read the dark matter of weights, we are increasingly "gravitationally bound" by them as our own cultural outputs begin to orbit the densest regions of the corporate manifold.

In cosmology, dark matter is said to constitute the majority of the universe's mass but does not interact with light and cannot be directly observed. Its existence is inferred from its gravitational effects on visible matter. In contemporary AI, the weights constitute the majority of the system's computational substance, billions of floating-point numbers whose specific values were determined during training, but they cannot be meaningfully "read" in the way that source code can be read. We do not interpret weights, we observe their effects, the way they bend the light of human intent, deflecting prompts into outputs whose shape reveals the gravitational field without making it visible. We can see what the field produces, but the "why" remains mathematically distributed and potentially human-unreadable.

But where dark matter's gravitational effects are predictable, modelled, consistent with physical law, the effects of the sub-symbolic layer are not. We cannot predict when a language model will hallucinate, cannot model which prompts will trigger refusal, cannot specify in advance what the system "knows". The opacity is more than merely epistemic, a temporary state to be resolved through better interpretability methods, in fact we might say it is ontological. The system does not work by encoding knowledge that could in principle be decoded. It works by encoding statistical patterns whose relationship to knowledge, understanding, or meaning remains unresolved.

Language, in this context, is its user interface rather than the medium of the computation. The machine does not "think" in English, it operates, if we can even use a cognitive term here, in weighted associations, and English (or another human language) is the surface through which human users interact with those associations.

If the old media theory could be said to be a theory of the atom, the bit, the pixel, the frame, the discrete unit of inscription, then vector theory is a theory of the field. The vector field is the governing image here, and it is barely metaphorical. A vector field in mathematics assigns a vector, a quantity with both magnitude and direction, to every point in a given space. 

As far back as 1994, Wark analysed how communication technologies create what she called "vectoral" power, the capacity to move information across space at speed (see Wark 1994, 2004). Those who control vectors, what Wark later termed the "vectoralist class", extract value from the flows they mediate. In 2004, I co-wrote the Libre Culture Manifesto whichdrawing on Wark's work, argued that "vectorialists" were emerging as a new class formation alongside landlords and capitalists, extracting value from the "distribution, access and exploitation of creative works" (Berry and Moss 2004). Now the situation has shifted dramatically. The vectoralist class now controls the channels through which information flows and also increasingly the geometry of the space within which meaning itself is beginning to become constituted. Where Wark's vectors described a power to move, AI vectors describe a power to transform, to render language and thought as manipulable coordinates within a proprietary computational space.

It seems clear, therefore, that we need a political economy of the manifold. Such an account would examine, at minimum, the concentration of compute infrastructure in a handful of corporations whose training runs cost hundreds of millions, and increasingly billions, of dollars, creating barriers to entry that make the construction of embedding spaces a de facto oligopoly. This includes the energy consumption of training and inference, measured in megawatt-hours, whose environmental costs are externalised onto communities that rarely consented to bear them. It would examine the labour conditions of data preparation, the millions of hours of annotation, labelling, and moderation performed by workers whose wages bear no relation to the value their labour produces within a manifold. Together with an analysis of the intellectual property regimes emerging around model weights, where the legal status of a trillion-parameter matrix trained on the entire publicly available internet remains unresolved. This is not to forget, too, the access hierarchies that determine who interacts with these systems and on what terms, from proprietary APIs priced per token to open-weight models whose "openness" still requires computational resources most of the world cannot afford. Indeed, the economics of fine-tuning is crucial to understand, in which corporations extract further value by adapting foundation models to specific domains, turning the general manifold into proprietary sub-manifolds whose topology serves particular commercial interests. Vector theory needs this critical account, otherwise it is merely descriptive, a formal account of how these systems work. With it, it becomes diagnostic, an account of what these systems do, and to whom.



** Article images generated using Google Gemini Pro, February 2026. 


Selected Bibliography

Berry, D. M. (2011) The Philosophy of Software: Code and Mediation in the Digital Age, Palgrave Macmillan.

Berry, D. M. (2014) Critical Theory and the Digital. Bloomsbury.

Lewis-Kraus, G. (2026) What Is Claude? Anthropic Doesn’t Know, Either, The New Yorker, 9 February. Available at: https://www.newyorker.com/magazine/2026/02/16/what-is-claude-anthropic-doesnt-know-either (Accessed: 11 February 2026).

Mikolov, T., Chen, K., Corrado, G. and Dean, J. (2013) Efficient Estimation of Word Representations in Vector Space. Available at: https://arxiv.org/abs/1301.3781.

Smolensky, P. (1988) On the proper treatment of connectionism, Behavioral and Brain Sciences, 11(1), pp. 1-74.

Wark, M. (1994) Virtual Geography: Living with Global Media Events. Indiana University Press.

Wark, M. (2004) A Hacker Manifesto. Cambridge, MA: Harvard University Press.

Comments

Popular Posts