Making data AI-ready: why clean data still guesses

Data quality alone is not enough. Without domain context an AI system translates metrics into the most probable meaning – not the right one. Four layers make data truly AI-ready.

← Back to overview
Adrian Bourcevet

Adrian Bourcevet

9 Min. Lesezeit

askbeyond chaotic analytics
·Teilen

SAP Sapphire in America, a few years ago. The dry run with the German slides had gone well; the English translation came back only the evening before the demo, polished and cleanly set. A quick look, relief – this looks professional. In the morning, the final pass through the whole presentation, then the shock: from Lift, the metric for the discriminative power of a classification model, had become Elevator. The elevator, it said roughly, had risen to 2.7. Everyone burst out laughing; the slide was fixed in minutes. The question afterwards was no longer funny: Which errors are still in there?

The translator knew the word but not the domain context. Exactly this error happens today in every organisation that wants to make its data AI-ready before terms are clarified: a system without domain context translates your metrics into the most probable meaning, not the right one. On a slide translation such an error is visible and fixable. In an AI system that answers your numbers, it stays invisible. It becomes expensive when a decision rests on it and your signature sits underneath.

The question behind this translation faux pas is on almost every board agenda today, only larger: How do we make our data AI-ready? The common answers, from consultancies to data-platform vendors, come as a list: improve data quality, introduce governance, clarify access rights. Correct, but incomplete. The elevator slide would have passed every one of those checks – the numbers were right, the source was clean, nothing was wrong with the process. What was missing is something that appears on none of those lists: domain context. This article shows why that hurdle sits before the technology, which four layers actually make data AI-ready, and who in the house decides where to start.

What “AI-ready” really means for data

AI-ready means: a system does not only answer the question about a metric; it can also justify the answer with exactly the meaning that metric has in your house. Not data availability is the measure, but the traceability of the answer. A database can be complete, current and technically flawless and still deliver answers nobody in the house can defend, because the system guesses where a definition is missing.

This confusion leads many organisations to invest in technology before they have clarified what their most important terms mean. The effect first shows not as an error but as uncertainty: nobody fully trusts the AI answer, so someone checks again, and the time gain of automation is spent again.

Why data quality alone is not enough

Data governance and data quality are necessary but not sufficient. Databricks CEO Ali Ghodsi put it sharply at the Data and AI Summit keynote in June 2026: AI does not have an intelligence problem but a context problem. Experience from 124 projects in regulated and data-intensive industries matches that. Systems rarely fail on compute power or missing fields. They fail because nobody wrote down beforehand what a term means in-house, which boundary applies and which example belongs to it.

How often this happens shows in projects we check before an AI introduction: the three to five most important steering metrics almost always carry several definitions at once. That is noticed only when a system outputs two of them side by side. That is the normal state, not an outlier, because in grown organisations multiple systems, multiple areas and decades of own language use meet.

The reflex most organisations react with is the same: collect more data, buy more tools, migrate the next platform. That treats a symptom. The context that makes a metric unambiguous does not arise by itself – it must be built explicitly before an AI system can use it.

The context layer: the four layers that make data truly AI-ready

Between the raw database and a trustworthy AI answer lie four layers. If one is missing, the system guesses at exactly that point, as quietly as with the Lift slide back then.

The first layer is the glossary: the binding definition of every important term, with formula, boundary and example, maintained in one place instead of in individual employees’ heads. The second layer is the semantic layer, the technical link between the glossary entry and the calculation in the database. It ensures that the same question in the data platform and in the AI system leads to the same number.

The third layer is the ontology or knowledge graph: the relationships between terms. It ensures a system knows that a customer margin belongs to a customer, a product and a period and is not an isolated number. The fourth layer is agent context: the rules that give an AI agent at query time which slice of the first three layers applies for this role, this user and this question.

Most projects that get stuck have built the fourth layer without the first three: a capable AI agent on a database that never had a glossary. The result is a system that answers quickly and cannot tell anyone why the answer is correct.

What it looks like when the layers are in place shows a prototype for management accounting we tested this summer. If someone asks for “contribution margin III”, the system delivers the binding definition with evidence from the metrics catalogue. If someone asks only for “contribution margin”, there comes no number but a follow-up question with the stored meanings to choose from. That looks unspectacular in a demo and is the whole difference in operations: the system does not guess, it asks. How the four layers interact in detail, from glossary care to agent context, is in the [Context Layer Whitepaper v3.0](https://beyond-chaotic-analytics.ch/docs/Whitepaper_Context Layer_v3.0.pdf).

Governance and roles: who decides the sequence

No organisation makes all data AI-ready at once; the question is always where to start. That is not a technical decision. It belongs at the level that owns the metric: the business-unit lead, the head of controlling, the executive team. Governance here means first prioritising by criteria, not setting up a new department.

Three questions carry this decision. First: Which metric will your AI system answer most often once it is productive? Second: Which metric already has several definitions in circulation today? Third: Which wrong answer would be most expensive, measured in bad decisions, rework or lost meeting time? The metric that ranks high on all three questions gets a glossary, a semantic layer and an agent context first. Everything else waits – with a cadence someone leads with date and outcome, not with a principle decision that disappears in the minutes.

The measure that remains

Das teuerste Problem in Unternehmen ist nicht Datenmangel. Es ist Entscheidungschaos.

Wir verwandeln Chaos in belastbare Entscheidungen, die Geld schützen, Tempo schaffen und umgesetzt werden.

Making data AI-ready does not mean cleaning every field before a system goes live. It means building four things for the terms with the greatest weight: definition, calculation, relationship and access rule – so that the same answer still holds in a year and someone in the house can explain why. The question that should stand before every signature under an AI project: Does the same question next week lead to the same number, and who in the room can explain the derivation?

If you want to check which of your metrics would fail this test today: write to me, or leave a comment with the metric you are least sure about. If you prefer a structured approach, the Clarity Audit offers a 90-minute inventory with four artefacts, including the first draft of a glossary entry for the metric that costs the most.

Teilen

Über den Autor

Adrian Bourcevet

Adrian Bourcevet

Experte für Analytics, Daten und KI. Unterstützt Unternehmen dabei, aus Daten wertvolle Erkenntnisse zu gewinnen.

Related articles

ask