Welcome

Passionately curious about Data, Databases and Systems Complexity. Data is ubiquitous, the database universe is dichotomous (structured and unstructured), expanding and complex. Find my Database Research at SQLToolkit.co.uk . Microsoft Data Platform MVP

"The important thing is not to stop questioning. Curiosity has its own reason for existing" Einstein



Tuesday, 18 August 2026

Why Trust Matters more than Discovery: Microsoft Purview Unified Catalog

In the first article in this series, I explored the challenge of visibility and how Microsoft Purview Unified Catalog helps organisations answer a simple but surprisingly difficult question: what data do we actually have?

For many organisations, solving this problem represents a significant milestone. Years of system growth, acquisitions, departmental solutions and technology change often create an environment where information exists but remains difficult to locate. Valuable datasets sit within platforms that few people know about, reports are recreated because earlier versions cannot be found, and knowledge about key information assets becomes concentrated within small groups of specialists.

Improving discovery removes many of those barriers, but it also introduces a new challenge. Once users can locate information more easily, their attention naturally shifts away from finding data and towards understanding it. Very rarely does somebody discover a dataset and immediately begin using it without asking further questions. Instead, they want to know whether the dataset is trusted, who owns it, how it is being maintained and whether it is suitable for the decision, analysis or report they are working on.

In practice, this is the point at which governance becomes far more interesting.

Most organisations do not struggle because they lack information. They struggle because they lack confidence in the information they have. Discovery helps people locate data, but trust determines whether that data is actually used.

Why Data Discovery is only the beginning

Many of the frustrations people experience with data are not caused by technology. They arise because the information lacks sufficient context.

Imagine an analyst searching a catalogue and finding three datasets that appear to contain customer information. All three are current. All three appear relevant. All three contain similar attributes and similar record counts. Discovery has succeeded because the analyst can see that the data exists. Unfortunately, discovery alone does not help determine which dataset should be used.

The questions that follow are usually business questions rather than technical questions.

Which dataset represents the approved source?

Which business area owns it?

What does "customer" actually mean within this context?

How frequently is it updated?

What transformations have been applied since the data was first collected?

Without those answers, users often fall back on familiar behaviour. They email colleagues, consult subject matter experts or continue using whichever data source they trusted previously. The catalogue exists, but confidence has not yet been established.

This is why governance programmes that focus exclusively on discovery often struggle to deliver their full value. Visibility is important, but visibility without context rarely creates trust.

Where Data Curation fits

One of the less discussed aspects of governance is curation. The term itself sounds administrative, which probably explains why it receives less attention than topics such as AI, analytics or compliance. In reality, curation sits at the heart of helping organisations bridge the gap between technical information and business understanding.

Most data assets are created within technology environments. Their names reflect system requirements, integration patterns or development conventions. To the engineers who build and maintain them, those names often make perfect sense. To everyone else, they can be cryptic, ambiguous or completely meaningless.

A curated asset looks different because it includes the information people actually need in order to understand it. Business descriptions explain what the asset represents. Ownership information identifies accountability. Classifications provide context about sensitivity and usage. Associated business terms explain how the asset fits within the language of the organisation.

This process transforms a technical asset into something that can be interpreted and trusted by a much wider audience.

The objective is not simply to document data. It is to create enough context that somebody encountering a dataset for the first time can understand its purpose and relevance without needing to find the person who created it.

The Business Glossary: Creating a Common Language

One of the most valuable governance capabilities within Microsoft Purview is the Business Glossary.

At first glance, a glossary sounds relatively straightforward. Many organisations assume it is simply a dictionary of approved business terms. In practice, its role is significantly more important than that.

Every organisation has terminology that appears obvious until people are asked to define it. Terms such as customer, employee, supplier, resident, contract or revenue are often assumed to have a consistent meaning. Governance workshops frequently reveal the opposite. Different teams use the same words while referring to slightly different concepts. Those differences may be perfectly reasonable within local contexts, but they become problematic when information is shared across departments, reports or analytical models.

A customer services team may define an active customer differently from the sales function. Finance may calculate revenue differently from operational reporting. Legal, risk and compliance teams may use terminology that reflects regulatory requirements rather than business reporting needs.

These are not necessarily disagreements. More often they are examples of organisational complexity becoming visible.

The Business Glossary provides a mechanism for governing this complexity. Within Microsoft Purview, glossary terms can be organised into domains, assigned owners and stewards, enriched with definitions and related terms, and connected directly to assets within the Unified Catalog. This relationship is particularly important because it links business language to the datasets, reports and information products that rely upon it.

When users search the catalogue, they are not simply looking at technical metadata. They can also see the business terminology associated with assets and understand how those assets relate to agreed organisational definitions. Rather than existing as a separate governance artefact that few people reference, the glossary becomes embedded within the discovery experience itself.

This is often where trust begins. People are far more likely to use information when they understand both what it contains and how the organisation expects it to be interpreted.

Why Lineage builds confidence

Even when terminology is clear and ownership is established, there is usually another question users want answered.

How did this data get here?

Most people consume information at the end of a process. They see a dashboard, a report, a model or, increasingly, an AI-generated response. What they do not see is the journey that information has taken through source systems, integrations, transformation processes and analytical platforms before reaching its final destination.

Understanding that journey is the role of data lineage.

Lineage provides visibility into how information moves through an organisation. Rather than viewing a dataset as an isolated asset, users can see its relationship to upstream systems, transformation processes and downstream consumers. This creates a much richer understanding of where information originated and what happened to it along the way.

The significance of lineage becomes particularly obvious when trust is challenged. If a figure changes unexpectedly, lineage helps explain why. If an upstream source system is modified, lineage can help identify which reports, dashboards and analytical processes may be affected. If two datasets appear similar, lineage may reveal that they originate from different systems and have undergone different transformations.

In other words, lineage provides evidence rather than assumption.

Within Microsoft Purview, lineage is captured automatically through integration with supported technologies and services. Data movement, transformation and processing activities within platforms such as Azure Data Factory, Microsoft Fabric, SQL environments and other supported services can be visualised as connected information flows. Instead of relying on manually maintained diagrams that quickly become outdated, organisations gain a dynamic view of how information actually moves through the estate.

For governance teams this increases visibility. For business users it often increases trust because they can see how a reported value is connected to its source.

Trust is what turns Data into value

Discovery remains a critical part of data governance. Organisations cannot govern information that they cannot find, which is why visibility, cataloguing and discovery capabilities provide such an important foundation.

However, discovery alone does not solve the larger challenge.

People create value from data when they are willing to use it. They use it when they understand it. They trust it when they have confidence in its meaning, ownership and provenance.

Business glossaries help establish shared language. Curation provides business context. Lineage explains how information was created and how it moves across the organisation. Together, these capabilities transform a catalogue from a searchable inventory into a trusted source of organisational knowledge.

Finding data is important. Being confident enough to use it is what ultimately matters.



No comments:

Post a Comment

Note: only a member of this blog may post a comment.