A data product is not a dataset with an owner

Tables become data products, those responsible become product owners - often just new vocabulary for an old reality. What really distinguishes a data product from a dataset: purpose, consumer, boundary, contract and lifecycle.

← Back to overview
Roman Unterstöger

Roman Unterstöger

21 Min. Lesezeit

askbeyond chaotic analytics
·Teilen

Many companies have started to name their data differently.

Tables become data products. Those responsible become product owners. Central data platforms become data marketplaces.

That sounds like progress.

But sometimes it's just new vocabulary for an old reality.

A team continues to produce a data set. Another team should use him. Someone is entered as the owner. A description goes into the data catalog. And suddenly it says:

Data Product: Customer 360.

The only problem is:

Nobody knows exactly who this product was built for. Nobody can say what problem it solves. Changes are made without knowing the consumers. If the data is incorrect, the search for the responsible team begins. And after the go-live, hardly anyone is interested in whether the data set is actually used.

This is not a product.

This is a dataset with better marketing.

The difference is crucial. Because “Data as a Product” does not work by declaring data products. It only works when we actually start treating data like products - exactly the principle behind Data Mesh.

And to do that we first have to answer a surprisingly simple question:

A product does not start with the producer

Imagine a company is developing a new application.

The development team probably wouldn't start with the question:

Which database tables do we already have?

It would ask:

When it comes to data, we surprisingly often do the opposite.

We start with the sources.

There is a customer table in the ERP. In the CRM too. There are also orders, invoices, service cases and perhaps data from an online shop.

This data is extracted, transformed and made available in a platform.

Technically everything works.

Then the result is given a name:

Customer Data Product

But we haven't created a product yet.

We processed data.

A product only comes into being through the relationship between a problem, a consumer and a benefit.

This applies to software.

And it applies to data as well.

Who actually is your customer?

This question is self-evident for a classic product.

Surprisingly not with a data product.

Let's take a seemingly simple data product:

Customer 360.

Who is its consumer?

Marketing?

Then you might be interested in:

  • Which campaigns did a customer receive?
  • What products does he own?
  • When did he last react?
  • What is the expected customer lifetime value?
  • Which channel can he be contacted via?

The sales?

Then other information becomes important:

  • Who is the account owner?
  • Which opportunities are open?
  • Which contracts are expiring?
  • Which contacts decide?
  • What activities took place recently?

Finance?

Then “customer” could mean something different:

  • Who is the invoice recipient?
  • Which company owns the claim?
  • Which invoices are outstanding?
  • How high is the credit risk?
  • What payment terms apply?

Three consumers.

A seemingly identical object.

Three different meanings of “Customer 360.”

This is exactly where product thinking begins.

Not with the question:

What customer data can we provide?

Rather:

This fundamentally changes the architecture – and often also the domain boundaries. This is exactly what domain-oriented data model design is all about.

A data product needs a purpose

Geld weg. Zeit weg. Autorität weg. Er stoppt den Entscheidungs-Kollaps.

Wir bauen die Struktur, mit der Entscheidungen klarer, sicherer und umsetzbar werden.

A dataset can exist because data exists.

A data product should exist because someone wants to achieve something with it.

That sounds banal. In practice, it is precisely this distinction that separates useful data products from data graveyards.

A data product with the description

Contains harmonized customer information from CRM, ERP and web shop.

describes its technical development.

But it says almost nothing about its usefulness.

More interesting would be:

Provides marketing and customer service with a consistent view of active customers, their products, interactions and communication permissions.

Now more questions immediately arise:

These are not trivial, detailed questions.

These are the questions through which data becomes a product.

Ownership is necessary. But ownership is not yet product thinking.

The most common abbreviation is:

We need an owner for every data product.

Correct.

But incomplete.

Because writing a name next to a dataset does not solve the responsibility problem.

The crucial question is:

And perhaps the most unpleasant question:

Can this owner also reject requests?

Responsibility without the right to make decisions is not ownership.

It's escalation management.

If a data product owner is held responsible for quality but cannot influence the source systems or enforce quality rules, he or she does not have true ownership.

If he is responsible for availability but has no influence on the platform or operations, neither.

If five departments require different definitions and the owner is not allowed to make a decision, you will not have an owner.

You will have a moderator.

This can be a useful role.

But we should call them that too.

In the decision-making architecture, this is the same difference: ownership without the right to make decisions remains escalation. This is exactly what the Ownership Blueprint addresses - and, organizationally, the question of why approved decisions are not implemented.

Good data products have a limit

One of the most difficult problems is not the question of what goes into a data product.

But what doesn't belong there.

Because as soon as data is thought of as a product, the desire to make a product as universal as possible quickly arises.

Customer 360 is intended to support marketing.

And sales.

And finance.

And risk.

And customer service.

And AI.

And future use cases that we don't even know about today.

That sounds efficient.

In reality, a product is often created without a clear consumer.

The more requirements we pack in, the more difficult any change becomes.

A definition of “customer” suddenly has to work for six areas.

A change to a field requires coordination with twelve consumers.

Quality requirements contradict each other.

Currency requirements vary.

Access rights are becoming more complicated.

And at some point no one really owns the product anymore.

A good data product therefore needs more than just one purpose.

There also needs to be a limit.

The product must be able to say:

I am responsible for that.

And just like that:

Not for this.

This is not a defect.

This is product design.

Data does not become a product because it is in the marketplace

Data marketplaces make sense.

Data catalogs make sense.

A good search makes sense.

But a marketplace doesn't produce good data products.

It merely makes visible what already exists.

If you put bad, poorly documented, or meaningless data in a beautiful marketplace, you won't own any data products afterwards.

You have a more conveniently searchable data graveyard.

A [data catalog](/en/blog/data catalog/) remains indispensable for discoverability and governance - but it does not replace purpose, ownership or product promises.

The problem can be easily compared to an app store.

An app store makes applications discoverable.

But no one would claim that an application is good because it is listed there.

It has to work.

It has to be understandable.

She has to solve a problem.

She needs someone in charge.

She needs to be looked after.

Users need to know what to expect.

Updates must not arbitrarily destroy functions.

And when it is no longer needed, it should eventually disappear.

Why should Data Products have lower requirements?

A data product must be discoverable

At first this sounds like a pure catalog problem.

But it's not.

Findability starts with the naming.

Imagine an analyst is looking for sales data.

He finds:

F_SALES_V3

REV_AGG_MONTHLY

FIN_04

SALES_CURATED

BOOKINGS_FINAL_FINAL

Technically these names may be understandable.

From a consumer's perspective, they are almost worthless.

A product must be discoverable in the language of its users.

This applies to names, descriptions, terms, synonyms and technical contexts.

Anyone looking for “revenue” shouldn’t need to know that Finance internally speaks of net_revenue_recognized.

Anyone searching for “customer” should be able to understand why there may be multiple customer terms.

And anyone who finds a product should recognize:

Findability therefore doesn’t just mean:

I can find the dataset.

Rather:

I can see if this dataset solves my problem.

That's a significant difference.

A data product must be understandable

Imagine you get the following table:

customer_id revenue status date
84721 182,400 A 2026-07-31

The data is clean.

No NULL values.

Data types are correct.

Pipeline successful.

Is this a good product?

No idea.

Data can be technically perfect and technically almost useless.

A data product therefore needs more than just schema documentation.

It needs meaning.

And this is exactly where the connection to semantics, business glossaries, knowledge graphs and AI becomes interesting - and to the context layer: meaning must come from the company, not from the model. We therefore model relationships and terms explicitly, for example in Knowledge Graph in the company.

An LLM can do a great job of explaining to you what revenue likely means.

But “probably” is not an acceptable definition when it comes to business metrics.

The meaning must come from within the company.

Not from the model.

A data product must convey this meaning.

A data product must be trustworthy

“Data quality” is often treated like a binary property.

Data is either good or bad.

In reality, quality depends on the intended use.

For a monthly management report, data with a 24-hour lag may be completely sufficient.

They may be worthless for fraud detection.

A customer address can be good enough for a marketing analysis.

Not for the delivery of a physical product.

A data product should therefore not simply claim:

Quality: 98%

What does that mean?

More interesting are concrete expectations:

Trust does not come from data being error-free.

They never will be.

Trust comes from consumers knowing what to expect.

A data product needs a contract

As soon as other teams use a data product productively, a dependency arises.

And dependencies need stability.

Suppose a data product delivers:

customer_id

country

segment

annual_revenue

Three teams build reports on it.

A machine learning model uses segment.

An operational process uses country.

Then the producing team decides:

We'll clean up once.

country becomes country_code.

annual_revenue is removed.

segment gets new values.

Technically perhaps an improvement.

Possibly a production problem for four consumers.

This is exactly why data contracts are interesting.

Not as another governance document that no one reads.

But rather as an explicit agreement between producer and consumer.

This changes a pipeline of

Here is data.

to

You can rely on that.

This is a product promise - the same claim that Evidence-first AI places on the evidence base: first contract and evidence, then use and formulation.

Self-service does not start with access

Many companies want self-service analytics.

So they give more people access to more data.

The result is not automatically self-service.

Sometimes it's just self-service confusion.

True self-service means:

A user finds a suitable data product.

He understands what it means.

He knows how current it is.

He recognizes his quality.

He can understand where it comes from.

He knows how to gain access.

He can use it without calling three data engineers first.

And if something is wrong, he knows who is responsible.

That's a high bar.

But that should be exactly the aim.

Technically, zero-copy approaches like Delta Sharing and the logic behind sharing data instead of copying help – but only if the shared object is already a product with a promise, not just another share.

A data product is not self-service capable if its owner has to answer ten messages every week:

Then the owner is not a product owner.

He is the human documentation of the product.

Usage is not a side effect

This is where product thinking differs particularly clearly from classic data platform thinking.

When it comes to pipeline, project and migration, success is often:

running.

delivered.

completed.

That's not enough for one product.

A product that no one uses is not successful just because it technically works.

This is why data product teams should know:

This can be uncomfortable.

Maybe it turns out that hardly anyone uses the “strategic customer data product” in which six months were invested.

Perhaps analysts continue to export data from the old system.

Maybe there are five parallel revenue datasets.

Maybe the official product has the highest data quality, but no one understands its 180 columns.

These are not adoption problems that can be solved with training.

This is product information.

They tell you something about your product.

And sometimes a data product should die

This is also part of product thinking.

Data platforms have a remarkable ability to preserve things forever.

Tables.

Views.

Pipelines.

Dashboards.

APIs.

Datasets.

After all, someone could still need everything.

Many companies know the result:

Nobody knows what is official anymore.

Old and new versions coexist.

Consumers accidentally use outdated data.

Costs rise.

Dependencies become opaque.

And no one dares to delete anything.

Products have a life cycle.

Data Products should also have one:

Idea → Development → Publication → Usage → Further Development → Deprecation → Shutdown

If a product no longer has consumers, two serve the same purpose, or one is replaced, these questions should be allowed:

That too is ownership.

Not every dataset has to be a data product

Perhaps that is the most important consequence.

If we declare every dataset a data product, the term loses its meaning.

A temporary staging table is not a product.

An intermediate technical result of a pipeline is not a product.

A copy from a source system is not automatically a product.

A dataset that is only needed internally by a single process may also not be one.

And that's perfectly fine.

Not all data has to be products.

The product idea becomes valuable precisely because we use it where several consumers need a reliable, understandable and long-term maintained data interface.

Otherwise we will mainly create administrative costs.

Suddenly 4,000 tables need 4,000 owners.

4,000 descriptions.

4,000 quality definitions.

4,000 products in the marketplace.

This is not democratization of data.

This is bureaucracy with metadata.

Data products are changing the boundary between business and IT

Traditionally, the division of labor is often something like this:

The business defines requirements.

IT builds data solutions.

Business consumes results.

At Data Products, this separation works poorly.

Many properties of a data product are neither purely technical nor purely functional.

What does “customer” mean?

Technically.

How is this definition calculated from six systems?

Technically.

What quality is acceptable?

Technically.

How is it monitored automatically?

Technically.

Who can see the data?

Professional, legal and technical.

How quickly do they have to be available?

Business requirement with technical and financial consequences.

A real data product therefore forces both sides to share responsibility.

And perhaps that is exactly one of the biggest advantages of the concept.

Not the new architecture.

Not the Marketplace.

Not the new role model.

But the realization:

Data only has value when producers understand what consumers do with it.

That's exactly why the hybrid business analytics platform is not an end in itself: it connects sources in such a way that decisions become reliable - not in such a way that there are as many tables as possible in the marketplace.

Data Products are particularly relevant for AI

With generative AI, this discussion takes on a new dimension.

Previously, poor metadata and unclear semantics could be compensated for by having an experienced analyst know who to ask.

An AI agent does not automatically possess this informal organizational knowledge.

He sees data.

Descriptions.

Metadata.

relationships.

Policies.

If there are five different revenue datasets, he has to decide which one is relevant.

When two different definitions of “active customer” exist, it needs context.

If a key figure can only be used for certain countries, this rule must be explicit.

If data is out of date, timeliness must be visible.

And if a source is not authoritative, the LLM cannot simply choose from five plausible tables the one whose column names best fit the question.

This is exactly where data product thinking becomes the basis for enterprise AI.

Not because AI necessarily requires “data products” as an architectural pattern.

But because AI needs explicit answers to questions that people have often solved informally:

A clearly described data product answers exactly these questions - and thus provides the prerequisites for [company context not in the prompt](/en/blog/company context-belongs-not-in-the-prompt/) to end up in a testable model. For the specific point of use, the same applies as for AI Use Cases for Decision Acceleration: link to the decision type, not to the showcase.

A simple test for your data products

Open your data catalog or marketplace.

Choose one of your most important data products.

And without calling the owner, try to answer the following questions:

  1. Who consumes this product?

Not “the company”.

Specific roles, teams or applications.

  1. What decisions or processes does it support?

Not “analytics”.

Concrete benefit.

  1. What does the content mean technically?

Not just table and column descriptions.

Terms and rules.

  1. What quality can I expect?

Not “high”.

Measurable expectations and known boundaries.

  1. How current is the data?

And does this topicality correspond to the use case?

  1. What can change without consumers being informed?

Is there a contract or at least defined change rules?

  1. Who can make technical decisions about the product?

Not only: Who gets the ticket?

  1. Who actually uses it?

Not: Who could use it?

  1. What happens if it fails tomorrow?

Are there critical dependencies?

  1. When would you turn off this product?

If the answer is “never,” you probably don’t have lifecycle management.

If you can't answer most of these questions, it's no big deal.

But maybe you don't have a data product yet.

Maybe you have a dataset.

And that's exactly where you should start.

The product is not the table

This is ultimately the error in thinking behind many data product initiatives.

We look for the technical object that we can declare a product.

A table.

A view.

A scheme.

An API.

A dashboard.

But the product is not necessarily one of these things.

These things are interfaces to the product.

The actual product is a reliable promise to a consumer:

For this purpose, we make this data available to you with this meaning, this quality, this timeliness and these rules - and someone takes responsibility for ensuring that this promise is kept.

This is much more demanding than entering an owner in a data catalog.

But that’s exactly why the term “product” is useful in the first place.

If you just want to publish data, you don't need product thinking.

If you want other teams to be able to build decisions, applications, analytics and AI on top of this data, yes.

Because a dataset answers the question:

What data do we have?

A good data product answers a much more important question:

Teilen

Über den Autor

Roman Unterstöger

Roman Unterstöger

Enterprise AI & Decision Architecture. Verankert Entscheidungen operativ: Cadence, Governance und verbindliche Umsetzungsroutinen in SAP- und Analytics-Umgebungen.

Related articles

ask