Synthetic Data and the Chain of Consent

How consent is inherited in synthetic descendants of data.

Harriet Job

Synthetic data feels smooth, like plastic. It’s man-made, it’s modern, and it resists imperfections unlike its organic predecessor.

Plastic surfaces don’t hold much memory. Fingerprints fade, and marks don’t accumulate into anything that changes how the object is treated later. When plastic breaks, it doesn’t really carry the sense that something specific has been lost - it’s disposable. Consequently, how we treat these objects remains largely unchanged with time. There is less hesitation, and less pause before contact. Synthetic data, through virtue of more than just its name, is at risk of being considered much of the same. Once reproducibility is taken as a low-friction given, the conditions under which something was produced no longer sit close to how it is later used. A dataset begins to read less like a record of a specific moment of human response, and more like something that could be produced again under similar constraints. The fact of human participation remains true, though it starts to matter less in the handling of the output. If something can be generated again, it loses some of its singular weight while still remaining materially the same. It can be easy to see this disentanglement of the data from its original generator as inconsequential. But, as always, just because something is easy, does not mean it is necessarily right, and in the process of this decoupling we break the epistemic spine of trust and authorship: the chain of consent.

The artificial nature of synthetic data dominates talking points, but with this we risk forgetting the important nuance that synthetic data is fuelled by real-world participants with real-world considerations. The bridge between human respondent and final synthetic model might feel byzantine, however the path is still complete, and the latter could never exist without the former in some capacity.

We know that in market research there are clear and non-negotiable consent frameworks - rules that cover things like data storage and use, to the end of preservation of the principles of fundamental human rights. Yet, somewhere along the labyrinthine path of synthetic data generation, through chains of transformations, we reach a consent dilution problem. One in which the further the system moves from the original act of participation, the less clear it becomes whose permission is still ethically relevant. While synthetic data might not contain identifiable individuals, the model itself is shaped by the distribution of real human behaviour, and these models can be repurposed and scaled far beyond the original context in which consent was given. Thus raising the question: did participants consent only to being analysed, or also to their behaviour becoming part of a generative system that produces future populations?

We have already encountered a related version of this problem elsewhere in generative AI. Image generation systems have prompted widespread discomfort around training data drawn from photographs and artworks uploaded long before large-scale generative models entered public consciousness. In many cases, the original act of sharing took place under an entirely different technical landscape, one in which participation did not carry an obvious implication that the material might later contribute to systems capable of generating new synthetic outputs at scale. This issue was never purely about copyright or ownership, it was also about contextual continuity: the conditions under which people originally contributed material no longer matched the conditions under which that material was later being operationalised. A participant might consent to their responses being analysed within the bound context of a study, while never anticipating that those same responses could contribute to the production of synthetic populations or simulated respondents.

A common argument in favour of synthetic data is that it has the potential to eliminate privacy risk because no real person can be identified in the output. However, this conflates two different ideas: of identifiability and of derivation. While synthetic data might solve the first of these, it can leave the second intact - data is still based on patterns learned from individuals. In this way, synthetic data might reduce individual data exposure risk significantly, but it does not eliminate inference risk, such that the risk that human behaviour, preferences, and vulnerabilities are still being extracted, encoded, and operationalised, even when no individual data is visible. This is key, as many ethical intuitions around consent are not only about identifiability, they are also about control over how one’s participation is used to shape future systems. Particularly in cases where a dataset is used to train a model that later generates synthetic populations used in domains like healthcare or public policy, the original participants have indirectly contributed to systems that may extend far beyond the context they agreed to.

We have to answer: is agency over consequences of one’s data a fundamental tenet of the existing consent frameworks? Because if it is, it might be time to reconsider how we are working with synthetic data.

So, what could this look like in practice? The point is not that synthetic data must be treated as equivalent to identifiable personal data, rather, it suggests that ethical responsibility does not end at anonymisation and transformation. At a minimum, it likely needs to expand to include consent transparency and bias/representational audits.

Synthetic data does not remove real people from the equation, it removes their visibility. What becomes harder to see is not just where the data came from, but where permission was ever supposed to stop mattering.