Who decides what good AI means
A committee approves a model because it leads a leaderboard. Nobody in the room asked who decided what that leaderboard measures.
A committee approves a model because it leads a leaderboard. Nobody in the room asked who decided what that leaderboard measures.
This piece first appeared in German. The English version is a rewrite rather than a line-by-line translation, and the German original stays online in the archive: Gute KI ist keine technische Frage – sondern eine Machtfrage.
A committee signs off on a model because it sits near the top of the relevant leaderboard. The procurement is justified in one line, the governance box is ticked, and nobody in the room has asked who decided what that leaderboard measures.
Discussion of artificial intelligence usually follows a familiar pattern. Models get compared, performance values get debated and progress gets quantified, so that whoever scores higher counts as superior and whoever scores lower counts as behind. The logic is understandable, it connects to everything else, and it is deceptive.
The decisive question is who establishes what counts as capability in the first place, and at that point the debate leaves the technical level and enters a political one. Good AI is no objective state; it is the result of evaluation decisions.
Benchmarks, scores and leaderboards suggest objectivity, and they supply numbers, rankings, curves of progress, and the impression that quality can be determined unambiguously.
What gets overlooked is that any measure is a decision. It settles which capabilities are relevant, which errors are tolerated, and which contexts stay out of view. A benchmark measures what it has declared measurable, and everything else disappears. That is a consequence of standardization, and no methodological error.
The trouble starts where that limitation stays invisible, because a measuring instrument then turns into a frame of interpretation, and a comparison turns into a quiet standard for good and bad.
Yardsticks do more than describe; they steer. Development orients itself toward what gets measured, models get optimized to score better on the tests that matter, and whatever goes untested loses priority.
Implicit priorities form out of that: consistency ahead of context, scalability ahead of appropriateness, measurability ahead of judgment. Those priorities are not wrong. They are also not neutral, because they favor some use cases and displace others without anybody deciding explicitly.
The power does not sit in the model; it sits in the yardstick the model is measured against.
That shift is no accident. Evaluation and benchmark systems structurally favor actors who have large training resources, serve standardized use cases, and need global comparability.
Platforms, large providers and capital-rich organizations benefit when quality is defined through uniform metrics, because standardization is an advantage for them and optimizing for benchmarks pays.
Organizations with specific contexts, regulatory particularities or high demands on case-by-case appropriateness fall behind. Their requirements are no less legitimate, only harder to measure, and normalization takes the place of differentiation.
Clear rankings are attractive inside an organization. They simplify decisions, justify procurement, convince committees and cover responsibility. A model counts as state of the art because it leads the relevant benchmarks, and for most committees that is enough.
What gets lost is the question of fit. The useful question is whether a model suits this context, which differs from whether it is capable.
Fit cannot be standardized. It demands judgment, and standardized evaluation systems increasingly replace that judgment, quietly and never openly. That is convenient, and it is risky.
A common argument holds that benchmarks are necessary to secure quality and avoid arbitrariness. That is true as far as it goes, because comparability is an advance on mere assertion.
The problem begins where measurability replaces judgment in place of supporting it, and where numbers stop serving as a basis for decisions and start serving as a substitute. Responsibility shifts at that moment.
The question stops being whether this solution is appropriate and becomes why anybody should deviate from a well-rated standard, which means deviation now requires justification and conformity does not.
Yardsticks often exert more steering effect than formal regulation does. Laws set limits and benchmarks set incentives, and it is the incentives that define what development orients itself toward.
A model that scores badly in the relevant rankings counts as second-class, even where it would be superior in a particular context. Development then follows measurability, never need.
A quiet standardization forms out of that, through technical comparison systems and not political decision, through ranking tables and not debate. It works precisely because it is rarely recognized as standard-setting at all.
The problem is not that standards get set. Standards are necessary, and the problem is that they get set without a clear mandate.
Who decides which criteria are relevant? Which risks count as tolerable, and who determines that? And who carries responsibility for the blind spots of an evaluation framework?
Those questions are rarely asked, out of habit more than out of bad faith. Technology gets treated as a neutral space where norms are allowed to form implicitly. What is actually happening is decision-making with considerable organizational, economic and social effect.
Without benchmarks, the objection runs, no systematic progress would be possible, because comparability is a precondition for development. That is also true.
Comparability is a means and not an end, and it never removes the need to reflect on the measures themselves. Benchmarks are useful for as long as they are understood as tools. They become a problem once they become authorities. Good evaluation supports judgment, and replaces none of it.
In the end, good AI is not a property of a model. It is the result of a deliberate decision about which criteria count and which do not. That decision cannot be delegated, to developers or to rankings.
Organizations that adopt evaluation standards without examining them hand over a piece of responsibility. Not formally, and in practice. They follow externally set norms without making those norms their own, which may be efficient and is no substitute for governance.
So the question is worth asking differently. Which model is better matters less than who defines what better means in this context. While that question goes unasked, the evaluation of AI stays technical and its effect stays political.
This piece is part of a workshop series.
No external sources. The essay draws on the author's professional experience in restructuring, turnaround work and consulting.
Your link has expired. Please request a new one.
Your link has expired. Please request a new one.
Your link has expired. Please request a new one.
Great! You've successfully signed up.
Great! You've successfully signed up.
Welcome back! You've successfully signed in.
Success! You now have access to additional content.