This piece first appeared in German. The English version is a rewrite rather than a line-by-line translation, and the German original stays online in the archive: Warum der Zugang zu guten Daten über gesellschaftlichen Einfluss entscheidet.


A mid-sized machine builder invests in AI-supported maintenance models, and technically that is no longer difficult. The obstacle is the history: years of service reports, consistent fault classification, structured machine telemetry. A company holding that record can forecast the probability of a failure, and one without it stays in reaction mode.

The model is interchangeable and the dataset is not, which decides more about economic and social influence than any argument about parameter counts. Debates about AI are usually conducted in technical terms, and that misses where the advantage actually forms: who has access to good data, and who does not. Good here means relevant, structured, dependable and above all proprietary.


Many companies are discussing tools they have no foundation for

A great deal of AI conversation happens inside organizations whose operational reality is documented only in fragments. The data base underneath is neither consistent nor strategically curated, and the model becomes the hoped-for answer to a structural omission that took years to accumulate.

In parallel, large platform companies negotiate data use with hospitals, insurers and logistics groups. Formally the subject is efficiency, while what actually happens is a transfer: the model learns from the partner's data, and the aggregated intelligence stays with the platform operator. Value does not distribute symmetrically. Dependencies rarely arrive spectacularly; they arrive through architecture.


Capital, then technology, now curated reality

For a long time capital was the decisive bottleneck, later technology. Today it is curated reality. AI systems produce no knowledge in a vacuum; they compress patterns already present in the data. Their quality reflects the quality and the exclusivity of that base, along with the conditions under which it can be reached.

That changes competitive dynamics. Companies with exclusive datasets build entry barriers without being monopolies in any classical sense. States with access to population-wide behavioral and movement data can train models that remain out of reach for everyone else. Industries that hold their data back out of uncertainty lose innovation speed, and end up dependent on those that did not.

Data access has stopped being a technical side question, and has become strategic infrastructure.


The counter-argument deserves a hearing

There is a serious position on the other side. It holds that models are becoming commoditized, that open-source architectures spread quickly and that compute keeps getting cheaper, so the real competitive advantage lies in adapting models fast and integrating them into processes.

That argument carries in part, because in some application areas, language models among them, quality differences genuinely level out, and many applications can be trained on generic datasets.

Even there, though, the decisive difference forms at the interface with real use. A financial investor who has collected proprietary transaction data, restructuring histories and governance information over decades can build risk and valuation systems that reach far past anything publicly available. Nothing about the model is extraordinary, and the reality it represents is documented more densely than anyone else's.

Look only at the model and you underestimate that preparatory work, and overestimate how interchangeable the result is.


Definitional power is the part that should worry people

The social version of the question is sharper. When few actors hold aggregated health, mobility or education data, economic power is not the only thing that moves. Definitional power moves with it: what counts as normal, what counts as risk, what counts as deviation. Models reproduce the structures present in their data, which is a logical consequence and not a moral accusation.

Regulation tries to limit the effect through data protection, data trustee models and sectoral access rules. Access that is too restrictive slows innovation, and access that is too free concentrates power. So the real decision is not between open and protected. What matters is the design of access architectures, and who defines and controls that design.


Historical discipline beats analytical brilliance

There is a parallel in restructuring work. Companies that have understood their own key figures structurally over years, and not only reported them, behave differently in a crisis than those who start ordering their data under pressure. What separates them is seldom analytical brilliance. It is historical discipline.

AI amplifies that effect. It rewards long preparation more than spontaneous creativity, and no model makes an organization strategically superior overnight if the record underneath is thin.

The decisive question is not who develops the technically best model. It is who has documented reality over years in a way that makes it modelable, and under what conditions that access gets shared or withheld.

Data is not a raw material in the classical sense. It arises in relationships: between companies and customers, between the state and citizens, between a platform and its partners. Whoever controls those relationships structurally controls the data trail, and with it a growing share of economic and social power.

Whether that produces open ecosystems or new dependencies will not be decided in code. It gets decided in governance, and in the willingness to shape the distribution of power deliberately.


Sources

Grouped by the section they support.

Definitional power is the part that should worry people

  • Regulation (EU) 2016/679 (General Data Protection Regulation), eur-lex.europa.eu. Regulation (EU) 2022/868 of 30 May 2022 (Data Governance Act), which sets a framework for data intermediation services, eur-lex.europa.eu. Regulation (EU) 2025/327 of 11 February 2025 on the European Health Data Space, eur-lex.europa.eu. Examples of the data protection, data trustee and sectoral access rules the text refers to.

All other sections

  • No external source. The essay draws on the author's professional experience in restructuring and corporate finance.
The link has been copied!