Every time a recommendation system surfaces a video, a post, or a piece of news, it is making a small bet about what you will do next. If it is right, it learns. If it is wrong, it also learns. Across billions of such transactions per day, the system is continuously updated by the behaviour it is simultaneously shaping. This is not a feature or a bug. It is the definition of a feedback loop.
Cybernetics has been describing systems like this since the 1940s. What it offers is not a critique of recommendation systems, but a precise vocabulary for understanding why they behave the way they do, and what would need to change for them to behave differently.
The ecosystem around the algorithm
Before examining the machine itself, it is worth mapping the ecology in which it operates. Recommendation algorithms do not exist in isolation. They are embedded in a network of actors with partially overlapping, partially competing interests.
The network makes a structural argument visible that prose alone struggles to convey. Regulators and civil society have real influence, but their paths to the algorithm run through the platform, and the platform’s primary feedback channel is the advertiser relationship, not the public interest one. Ashby’s Law of Requisite Variety applies here directly: the regulatory system cannot control what it cannot model, and the algorithm changes faster than any regulatory model of it.
The feedback architecture
Inside the ecosystem, two amplifying loops drive system behaviour. Neither contains a stabilising counterpart with comparable strength.
The first loop is the engagement loop: user behaviour generates signal, signal trains the model, the model refines recommendations, recommendations shape user behaviour. The second is the supply loop: recommendations drive demand for a type of content, creators produce more of it, users encounter it, behaviour shifts further in that direction.
What cybernetics would call a stabilising loop (a feedback path that damps oscillation and pulls the system back toward equilibrium) is largely absent. Policy intervention (content moderation, demotion) exists, but it operates on outputs after the fact, not on the optimisation target itself. The system is not broken; it is working exactly as designed. The question is whether the design objective is adequate.
Stafford Beer observed that the purpose of a system is what it does, not what it says it does. These systems say they connect people to content they love. What they do is maximise time-on-surface as measured by engagement signals. Those are not the same thing.
The technical control architecture
Understanding why intervention is so difficult requires looking at how the system is actually built. The recommendation pipeline is a deeply layered technical structure, with the critical control path running through components that operate at machine speed, far below the layer where human editorial judgement could intervene.
The architecture diagram makes explicit something that platform transparency reports routinely obscure: safety filtering sits between the ranker and the user interface, not between the training signal and the model. It catches some outputs after scoring. It does not shape what the model learns to value. The label store (the implicit signal of what counts as a positive outcome) sits at the base of the system, upstream of everything. Whoever controls what goes into the label store controls the system’s value function. That is not a product decision; it is a governance decision. It is rarely discussed as one.
The adversarial surface
A system with this architecture (one that learns continuously from behavioural signal and optimises for engagement) presents a predictable adversarial surface. Any actor who can generate coordinated, inauthentic engagement signal at scale can, in principle, distort what the system learns.
The diagram shows a structural problem: the appeals process is used as an exhaustion tactic, flooding the moderation queue to delay removal while reach accumulates. The content classifier is compromised not by defeating it technically, but by submitting enough volume that borderline content gains traction before the classifier can act. The ranker is compromised not by hacking it, but by inflating the signal it trusts.
This is a cybernetic attack: it does not break the system’s components, it corrupts the information environment those components depend on. Wiener called this kind of vulnerability the corruption of the signal channel. It is not a new problem. It is as old as the concept of feedback itself.
What would a well-regulated system look like?
A cybernetic reading does not lead automatically to a prescriptive answer. But it does clarify what the question actually is. The problem is not that these systems use feedback: all adaptive systems do. The problem is that the feedback loop is closed around a single variable (engagement) that is a poor proxy for the values we actually care about (wellbeing, informed public discourse, epistemic diversity).
A system regulated at the level of its optimisation target, rather than its outputs, would look structurally different. It would have multiple, competing feedback signals (some representing engagement, some representing other values) with a governance layer that determines their relative weights. That governance layer would need to be legible, contested, and accountable in ways that a private product roadmap is not.
This is not a utopian proposal. It is an engineering description of what requisite variety actually requires. The system being regulated is extraordinarily complex. The regulatory apparatus matching that complexity does not yet exist.
Building it is one of the defining systems challenges of this generation.