IIT as a Lens on AI Systems
Could a metric of integration signal when a system is too entangled to predict or control?
Sometimes the most consequential properties of a system are invisible. Large language models speak fluently and answer questions with convincing confidence, yet what might make them powerful is not only what they output, but how their internal parts interact to generate it.
Integrated Information Theory, or IIT, was developed to describe this feature of biological systems. It’s goal is audacious: to measure consciousness not through behaviour but through the causal structure of the system itself. The key quantity, Φ, represents the amount of irreducible information in a system. Conceptually, it is a measure of how much of the whole is more than the sum of its parts. In biological terms, a high Φ indicates a system whose parts are interdependent in such a way that no subpart could generate the same causal effect on its own. It is not intelligence or behavior that matters. It is entanglement of cause.
IIT begins by considering all possible ways to partition a system into independent subsystems. For each partition, it compares the probability distribution over system states in the intact system with the distribution assuming the partition. Φ is then defined as the minimum difference across all partitions, representing the amount of information lost when the system is split. In simple terms, the higher the Φ, the more a system’s causal structure is integrated, and therefore irreducible. We can write this as:
\[ \Phi = \min_{partitions}D_{KL}\left(P_{whole}||P_{partitioned}\right) \]Where \(D_{KL}\) is the Kullback-Leibler divergence measuring how distinguishable the intact system’s states are from those of a partitioned system.
When we consider AI systems, IIT is both inspiring and confounding. LLMs are designed to optimise predictive accuracy, not to mirror biological causality. So, while the outputs might look the same, what’s going on under the hood might be far from comparable in terms of integrated information. Computing Φ exactly in a network on millions or billions of parameters is infeasible, yet, the concept remains suggestive. Integration may signal where a system is internally interdependent enough that simple interventions or interpretability techniques fail. It may reveal point of structural opacity where emergent behaviours arise not from explicit rules but from irreducible interactions.
This opens an ethical and strategic question: should we care about integration even if we are not concerned with consciousness? A system can operate reliably, generate coherent outputs, and appear intelligent, yet the underlying causal structure may hide fragility or amplification. Monitoring only outputs risks missing this entanglement. IIT encourages us to think in terms of structure as well as behaviour, of the unseen architecture of influence within a system. Perhaps we do not need to assert that AI is conscious to recognise that internal interdependence matters, and structural awareness can be a practical lens for governance and risk. Just as in biological systems where high integration correlates with emergent properties, AI systems may reach thresholds of internal entanglement beyond which they are difficult to predict or manage. Understanding the limits of decomposition, even approximately, may help us anticipate consequences before they become visible in outputs.
In this sense, IIT might not be a tool for measuring artificial minds, but as a conceptual guide for thinking about power, opacity, and responsibility in design. It reminds us that some risks are hidden not in what a system says, but in the patterns that bind its internal world together.