Welcome

Passionately curious about Data, Databases and Systems Complexity. Data is ubiquitous, the database universe is dichotomous (structured and unstructured), expanding and complex. Find my Database Research at SQLToolkit.co.uk . Microsoft Data Platform MVP

"The important thing is not to stop questioning. Curiosity has its own reason for existing" Einstein



Showing posts with label Data Governance. Show all posts
Showing posts with label Data Governance. Show all posts

Wednesday, 19 August 2026

Anthropic training programmes

Over the last week I have completed three Anthropic training programmes to expand my use of AI tools from Copilot and Gemini. While the courses focus on Claude, the value goes far beyond learning a particular AI tool. I completed:

  • Claude 101
  • AI Fluency Framework & Foundations
  • Claude Code in Action
Claude 101 provides a solid grounding in how to work effectively with large language models. It covers prompt design, structuring requests, and understanding where AI can genuinely add value versus where human judgement remains essential.

AI Fluency Framework & Foundations takes a broader view. Rather than concentrating on technology alone, it explores how individuals and organisations can develop the skills, mindset and practices needed to adopt AI successfully. It reinforces an important lesson about becoming AI-enabled is as much about people and ways of working as it is about the tools themselves.

Claude Code in Action was particularly interesting from a practical perspective. It demonstrates how AI can support software development workflows, automate repetitive tasks, assist with code generation and review, and help teams move faster while maintaining quality. Even for those of us who are not full-time developers, it offers valuable insight into how AI is changing the way technical teams work.

The skills I learned on a this tool cover the tool usage but are more about people, skills, governance, and trusted data.



Tuesday, 18 August 2026

Why Trust Matters more than Discovery: Microsoft Purview Unified Catalog

In the first article in this series, I explored the challenge of visibility and how Microsoft Purview Unified Catalog helps organisations answer a simple but surprisingly difficult question: what data do we actually have?

For many organisations, solving this problem represents a significant milestone. Years of system growth, acquisitions, departmental solutions and technology change often create an environment where information exists but remains difficult to locate. Valuable datasets sit within platforms that few people know about, reports are recreated because earlier versions cannot be found, and knowledge about key information assets becomes concentrated within small groups of specialists.

Improving discovery removes many of those barriers, but it also introduces a new challenge. Once users can locate information more easily, their attention naturally shifts away from finding data and towards understanding it. Very rarely does somebody discover a dataset and immediately begin using it without asking further questions. Instead, they want to know whether the dataset is trusted, who owns it, how it is being maintained and whether it is suitable for the decision, analysis or report they are working on.

In practice, this is the point at which governance becomes far more interesting.

Most organisations do not struggle because they lack information. They struggle because they lack confidence in the information they have. Discovery helps people locate data, but trust determines whether that data is actually used.

Why Data Discovery is only the beginning

Many of the frustrations people experience with data are not caused by technology. They arise because the information lacks sufficient context.

Imagine an analyst searching a catalogue and finding three datasets that appear to contain customer information. All three are current. All three appear relevant. All three contain similar attributes and similar record counts. Discovery has succeeded because the analyst can see that the data exists. Unfortunately, discovery alone does not help determine which dataset should be used.

The questions that follow are usually business questions rather than technical questions.

Which dataset represents the approved source?

Which business area owns it?

What does "customer" actually mean within this context?

How frequently is it updated?

What transformations have been applied since the data was first collected?

Without those answers, users often fall back on familiar behaviour. They email colleagues, consult subject matter experts or continue using whichever data source they trusted previously. The catalogue exists, but confidence has not yet been established.

This is why governance programmes that focus exclusively on discovery often struggle to deliver their full value. Visibility is important, but visibility without context rarely creates trust.

Where Data Curation fits

One of the less discussed aspects of governance is curation. The term itself sounds administrative, which probably explains why it receives less attention than topics such as AI, analytics or compliance. In reality, curation sits at the heart of helping organisations bridge the gap between technical information and business understanding.

Most data assets are created within technology environments. Their names reflect system requirements, integration patterns or development conventions. To the engineers who build and maintain them, those names often make perfect sense. To everyone else, they can be cryptic, ambiguous or completely meaningless.

A curated asset looks different because it includes the information people actually need in order to understand it. Business descriptions explain what the asset represents. Ownership information identifies accountability. Classifications provide context about sensitivity and usage. Associated business terms explain how the asset fits within the language of the organisation.

This process transforms a technical asset into something that can be interpreted and trusted by a much wider audience.

The objective is not simply to document data. It is to create enough context that somebody encountering a dataset for the first time can understand its purpose and relevance without needing to find the person who created it.

The Business Glossary: Creating a Common Language

One of the most valuable governance capabilities within Microsoft Purview is the Business Glossary.

At first glance, a glossary sounds relatively straightforward. Many organisations assume it is simply a dictionary of approved business terms. In practice, its role is significantly more important than that.

Every organisation has terminology that appears obvious until people are asked to define it. Terms such as customer, employee, supplier, resident, contract or revenue are often assumed to have a consistent meaning. Governance workshops frequently reveal the opposite. Different teams use the same words while referring to slightly different concepts. Those differences may be perfectly reasonable within local contexts, but they become problematic when information is shared across departments, reports or analytical models.

A customer services team may define an active customer differently from the sales function. Finance may calculate revenue differently from operational reporting. Legal, risk and compliance teams may use terminology that reflects regulatory requirements rather than business reporting needs.

These are not necessarily disagreements. More often they are examples of organisational complexity becoming visible.

The Business Glossary provides a mechanism for governing this complexity. Within Microsoft Purview, glossary terms can be organised into domains, assigned owners and stewards, enriched with definitions and related terms, and connected directly to assets within the Unified Catalog. This relationship is particularly important because it links business language to the datasets, reports and information products that rely upon it.

When users search the catalogue, they are not simply looking at technical metadata. They can also see the business terminology associated with assets and understand how those assets relate to agreed organisational definitions. Rather than existing as a separate governance artefact that few people reference, the glossary becomes embedded within the discovery experience itself.

This is often where trust begins. People are far more likely to use information when they understand both what it contains and how the organisation expects it to be interpreted.

Why Lineage builds confidence

Even when terminology is clear and ownership is established, there is usually another question users want answered.

How did this data get here?

Most people consume information at the end of a process. They see a dashboard, a report, a model or, increasingly, an AI-generated response. What they do not see is the journey that information has taken through source systems, integrations, transformation processes and analytical platforms before reaching its final destination.

Understanding that journey is the role of data lineage.

Lineage provides visibility into how information moves through an organisation. Rather than viewing a dataset as an isolated asset, users can see its relationship to upstream systems, transformation processes and downstream consumers. This creates a much richer understanding of where information originated and what happened to it along the way.

The significance of lineage becomes particularly obvious when trust is challenged. If a figure changes unexpectedly, lineage helps explain why. If an upstream source system is modified, lineage can help identify which reports, dashboards and analytical processes may be affected. If two datasets appear similar, lineage may reveal that they originate from different systems and have undergone different transformations.

In other words, lineage provides evidence rather than assumption.

Within Microsoft Purview, lineage is captured automatically through integration with supported technologies and services. Data movement, transformation and processing activities within platforms such as Azure Data Factory, Microsoft Fabric, SQL environments and other supported services can be visualised as connected information flows. Instead of relying on manually maintained diagrams that quickly become outdated, organisations gain a dynamic view of how information actually moves through the estate.

For governance teams this increases visibility. For business users it often increases trust because they can see how a reported value is connected to its source.

Trust is what turns Data into value

Discovery remains a critical part of data governance. Organisations cannot govern information that they cannot find, which is why visibility, cataloguing and discovery capabilities provide such an important foundation.

However, discovery alone does not solve the larger challenge.

People create value from data when they are willing to use it. They use it when they understand it. They trust it when they have confidence in its meaning, ownership and provenance.

Business glossaries help establish shared language. Curation provides business context. Lineage explains how information was created and how it moves across the organisation. Together, these capabilities transform a catalogue from a searchable inventory into a trusted source of organisational knowledge.

Finding data is important. Being confident enough to use it is what ultimately matters.



Friday, 14 August 2026

AI is Forcing Organisations to ask Data Governance questions they have avoided for years

When organisations begin exploring generative AI, the early conversations are usually focused on technology. Attention naturally turns towards copilots, agents, large language models, prompt engineering and how existing processes might be automated. The assumption is often that success will depend on choosing the right tools and identifying the right use cases.



Those discussions are important, but they rarely remain the centre of attention for long.

As AI initiatives move beyond experimentation and into real business scenarios, the conversation often shifts in an unexpected direction. Questions begin to emerge about ownership, trust, definitions and accountability. Teams discover that information which appeared well understood within individual departments becomes considerably more difficult to explain when it is surfaced across the organisation through a single AI-powered experience.

This is creating an interesting situation. Many organisations believe they are encountering AI challenges when, in reality, they are encountering long-standing governance challenges that have remained largely hidden until now.

For years it has been possible for businesses to operate successfully despite inconsistencies in the way data is managed. Different departments develop their own reporting processes, terminology and working practices. Over time these approaches become embedded into everyday operations. Finance may calculate a measure one way, while another business unit calculates it differently. Multiple systems may contain records relating to the same customer, product or asset. Ownership may be understood informally without being clearly defined.

None of these situations are unusual. In fact, they are common in organisations of every size and sector.

What has changed is that generative AI is exposing these inconsistencies in ways that traditional reporting platforms rarely did. Information that once remained within the boundaries of a specific application, report or team is increasingly being brought together and presented through a single interface. As soon as that happens, differences in meaning, ownership and interpretation become much more visible.

One of the more striking developments over the past year has been how quickly discussions about AI become discussions about data governance. An organisation may start by exploring how employees can use Copilot more effectively, only to find itself debating which definition of a business term should be treated as authoritative. A workshop intended to focus on automation can quickly become a conversation about data ownership. Questions about whether users can trust AI-generated responses often lead directly to questions about where underlying information originated and how it is managed.

These are not new concerns. Governance professionals have been dealing with them for decades. The difference is that they are no longer confined to governance programmes.

AI is bringing them into boardrooms, project teams and business conversations that might previously never have engaged with governance at all.

The issue is not that AI is creating poor governance. Rather, it is making gaps in governance more difficult to ignore.

A useful parallel can be found in the idea of technical debt. Most organisations understand that technology decisions made years ago can create future complexity. Shortcuts that seem reasonable at the time often require greater effort to address later. Data governance follows a similar pattern. Business definitions are left undocumented because everyone believes they share the same understanding. Ownership remains informal because responsibilities appear obvious. Metadata is treated as a technical concern rather than a business asset. Lineage documentation is postponed because delivery deadlines take priority.

Individually, these decisions rarely feel significant. Collectively, they create an environment where information becomes harder to understand, trust and govern over time.

Historically, organisations could continue operating with this ambiguity because people compensated for it. Experienced employees knew which reports to trust and who to contact when figures did not align. Unwritten knowledge often filled the gaps that formal governance processes had not addressed.

Generative AI changes that dynamic because it lacks this organisational context. It relies on information being discoverable, understandable and consistent. When definitions vary between teams, when ownership is unclear or when information carries little context, those weaknesses become more apparent. The technology is simply revealing what has always been there.

This is one reason metadata has suddenly become a much more strategic conversation. Business glossaries, catalogues, classifications, stewardship models and lineage are often viewed as traditional governance disciplines. Increasingly, they are becoming recognised as fundamental enablers for AI adoption. Organisations are realising that it is difficult to scale AI responsibly when basic questions about information cannot be answered consistently.

The organisations making the strongest progress with AI are not always the ones investing the most heavily in AI technology itself. More often, they are organisations that have a reasonable understanding of their information landscape. They know which data matters to the business, who is accountable for it, how it is defined and where it comes from. They have established enough structure and context to create confidence in the information being consumed.

That confidence matters because successful AI adoption is ultimately a trust exercise. Users need confidence that information is accurate, that responses can be explained and that decisions can be justified. Without trust, adoption slows regardless of how capable the underlying technology may be.

Perhaps the most interesting outcome of the current AI wave is that it is forcing organisations to revisit some of the fundamentals of information management. After years of being viewed as a compliance activity or a specialist discipline, data governance is finding itself at the centre of conversations about innovation, productivity and business transformation.

The irony is that many organisations began their AI journey expecting to focus primarily on technology. Instead, they are being asked to confront questions about data that have existed for years. They have questions about ownership, meaning, accountability,  and trust. Those are governance questions, and they are becoming increasingly difficult to avoid.

AI may not have been designed to improve data governance, but it is proving remarkably effective at showing organisations where governance needs attention. In many cases, the most valuable insight generated by AI is not contained within a response or recommendation. It is the realisation that understanding data remains one of the most important prerequisites for using it effectively.

Thursday, 6 August 2026

Governance should travel with the Data: Microsoft Purview Data Sharing

One of the more frustrating characteristics of traditional data governance is that it often becomes less effective at the exact moment information starts to create value. Organisations invest significant effort in cataloguing information, defining business terminology, assigning ownership, improving quality and establishing governance controls. Within a governed environment, confidence begins to grow because people understand what data exists, how it should be used and who is responsible for it. The challenge emerges when that information needs to be shared.

Historically, sharing data has often created a new copy of the problem alongside a new copy of the data. Information is extracted from a source system, copied into another platform and then distributed to a different department, partner or project team. While the immediate requirement has been satisfied, something important is frequently lost along the way. The governance context associated with that information does not always travel with it.

Definitions become disconnected from the business glossary. Ownership becomes less obvious. Security controls may differ between environments. Lineage becomes harder to follow. Before long, multiple versions of the same information exist across different locations, each carrying slightly different assumptions about how it should be managed.

Many organisations have spent years attempting to reduce the number of data silos within the business, only to discover that sharing data can be one of the fastest ways to create new ones.

This creates an important governance challenge.

Data only delivers value when people can access and use it. Restricting access to everything is rarely a practical solution. At the same time, uncontrolled sharing introduces risk, duplication and inconsistency. Effective governance requires a balance between enabling access and maintaining control.

This is where Microsoft Purview Data Sharing becomes particularly valuable.

Rethinking what it means to Share Data

When people hear the phrase "data sharing", they often assume it means moving data from one place to another.

Modern governance increasingly takes a different approach.

Rather than distributing copies of information throughout the organisation, the goal is to provide controlled access to trusted datasets while keeping the information where it already resides. Instead of sharing the asset itself, organisations increasingly focus on sharing access to the asset.

The distinction is more significant than it first appears.

A copied dataset immediately begins diverging from its source. Updates may occur in one location but not another. Governance controls may evolve independently. Users consuming the copied data may no longer have visibility of ownership, quality metrics or policy controls that existed within the original environment.

Providing access to the governed source avoids many of these problems because users remain connected to the same underlying asset.

The data and governance remains in one place. The organisation maintains a single version of the truth.

Why Data Sharing depends on Discovery and Trust

By the time an organisation begins sharing information more broadly, much of the work described in the previous articles has already become important.

Sharing trusted data assumes that trusted data has first been identified.

The Unified Catalog helps users discover available datasets. Business Glossary terms ensure there is a common understanding of what those datasets represent. Lineage helps explain where information originated and how it has been transformed. Data Quality provides confidence that the information meets expected standards.

Without those foundations, sharing simply distributes uncertainty more widely.

In many respects, governance capabilities become more valuable as data sharing increases because a larger audience needs confidence in the information being consumed. Users receiving access to data rarely have the benefit of local knowledge from the teams that originally created it. Governance provides the context that helps those users understand what they are accessing.

For this reason, sharing should not be viewed as a standalone capability. It is the outcome of many governance disciplines working together.

Governance should not end at the department boundary

One of the more common challenges within large organisations is that governance maturity often varies significantly between business areas.

Some domains have established ownership, clear definitions and strong stewardship practices. Others remain heavily dependent on local knowledge and informal processes. Sharing information across those boundaries can expose differences that were previously invisible.

A finance team may confidently share information with another department, only to discover that business definitions are interpreted differently elsewhere. Regulatory requirements understood within one team may not be obvious to another. Information classified as sensitive in one context may be treated differently in another.

These situations are rarely caused by poor intentions. More often they reflect the complexity of modern organisations.

A well-governed sharing model helps address this challenge because governance travels alongside the data. Users receiving access are not simply obtaining rows and columns. They are gaining access to ownership information, classification context, business definitions and other governance artefacts that help explain how the information should be interpreted and managed.

This creates a more consistent experience for both providers and consumers of data.

The Role of Data Policy

Sharing information safely ultimately depends on policy.

In the previous article, I explored how Data Policy moves governance from documentation into operational control. Data Sharing builds on the same principle.

Not everybody should have access to every dataset. Access decisions should reflect business need, sensitivity, ownership and organisational policy. As governance programmes mature, manually managing these decisions becomes increasingly difficult, particularly when information needs to be shared across departments, projects or external organisations.

Microsoft Purview helps organisations apply governance controls consistently through integrated policy management. Data can be shared according to rules that reflect the classifications, ownership structures and governance requirements already established elsewhere within the platform.

The important point is that governance is not being recreated at the point of sharing. It is being reused.

The policies that protect information within the organisation continue to provide value when that information is shared more broadly.

Supporting collaboration without creating new silos

The tension between collaboration and control has existed for as long as organisations have managed information.

Too much restriction reduces value because people struggle to access the information they need. Too little control increases risk and often leads to duplication, inconsistency and fragmented data landscapes.

Successful governance programmes recognise that the objective is not to prevent sharing. The objective is to make sharing safe, consistent and transparent.

This is particularly important as organisations invest in data products, cross-functional analytics and AI-driven initiatives. These capabilities depend on information flowing between teams, domains and systems. Value is increasingly created through reuse rather than isolation.

Data Sharing supports this by helping organisations move away from copying information and towards sharing governed access to trusted assets.

The Final Step in the Governance Journey

The earlier articles in this series focused on establishing visibility, building trust, operationalising governance and understanding how governance performance can be measured. Data Sharing brings those capabilities together by allowing governance to support one of the most important objectives of any data programme: enabling people to use information confidently across organisational boundaries.

This is an important distinction because governance is often perceived as something that restricts access to data. In reality, good governance should achieve the opposite. By creating trust, clarity and accountability, it becomes easier for organisations to share information safely rather than harder.

Ultimately, the success of a governance programme should not be measured by how effectively it controls data. It should be measured by how effectively it enables trusted use of data across the organisation.

Because the real value of governance is not found in catalogues, glossaries, dashboards or policies.

It is found in the confidence that allows people to use information, collaborate with others and make better decisions without creating new risks or new silos along the way.


Finding what your Organisation already knows: Microsoft Purview Unified Catalog

Most organisations have no shortage of data. What they lack is a clear view of which of it is relevant, current and trustworthy. Across almost every organisation there are databases, reports, data warehouses, data lakes, spreadsheets, business applications and operational systems containing information that somebody, somewhere, relies upon every day. New platforms arrive, legacy systems remain, departments develop their own solutions and the information landscape gradually expands year after year.

The challenge is rarely the absence of data. More often, the challenge is knowing what already exists. This becomes particularly visible whenever a new initiative begins. A project team starts looking for customer data. An analyst needs information to support a reporting requirement. An AI initiative requires access to trusted business information. The data almost certainly exists somewhere within the organisation, but locating it often becomes an exercise in networking rather than discovery. Emails are sent. Teams messages are exchanged. Conversations take place with individuals who have accumulated knowledge about particular systems over many years. Eventually the data is found, but the process raises an uncomfortable question. Why was finding it so difficult in the first place?

Many organisations have become accustomed to a culture of data by request. Access to information frequently depends on knowing who to ask rather than knowing where to look. Knowledge becomes concentrated within particular teams and individuals, creating operational dependencies that often remain invisible until those people move roles, leave the organisation or become unavailable. This is one of the problems Microsoft Purview Unified Catalog is designed to address.

The difference between knowing data exists and being able to find it

When people first hear the term data catalogue, they often imagine a searchable inventory of assets. That description is not wrong, but it is incomplete. A catalogue only has value if it remains current, accurate and connected to reality. Historically, many organisations attempted to maintain data inventories through spreadsheets, documents and manually curated repositories. These often delivered some value initially, but keeping them aligned with constantly changing technology estates proved difficult. Systems changed, databases evolved and new projects appeared long before documentation could be updated.

Microsoft approached the challenge differently. At the foundation of the Purview governance platform sits the Data Map, a service that scans connected data sources on a scheduled or on-demand basis and collects metadata from across the estate. Whether information resides within Azure, Fabric, SQL Server, Databricks, Power BI or a growing list of supported technologies, the Data Map provides the automated discovery capability that allows Purview to understand what exists within the environment. Importantly, what is collected is metadata rather than the data itself: Purview builds a picture of the estate without copying or exposing the underlying content.

This distinction is important because the Unified Catalog is not the scanning engine itself. The Data Map performs the discovery. The Unified Catalog turns that discovery into something users can explore, search and understand. Without the Data Map, the catalogue would quickly become another manually maintained inventory. Without the catalogue, the information collected by the Data Map would remain difficult for most users to consume. The value comes from the relationship between the two.

From technical metadata to business understanding

Discovering an asset is only the beginning of the journey. Knowing that a database table exists tells a technical user something useful, but it often tells a business user very little. A name, a schema and a collection of columns rarely explain whether a dataset is trusted, who owns it, how it is used or whether it should be used at all.

This is where the Unified Catalog begins to move beyond traditional metadata management. The Unified Catalog brings together technical information and business context within a single discovery experience. Datasets can be associated with business terms, classifications, ownership information, descriptions, lineage and governance information. Rather than presenting users with a list of technical assets, it starts to answer the questions people naturally ask when looking for data.

What does this dataset contain?

Who owns it?

Is it approved for reporting?

How does it relate to other assets?

Where did the information originate?

Can it be trusted?

These are fundamentally business questions rather than technical questions, which is why discoverability has become such an important governance capability. People rarely struggle to search for information. They struggle to determine whether the information they have found is the right information.

The Unified Catalog in Microsoft Purview

Within Microsoft Purview, the Unified Catalog serves as the central discovery experience for governed data assets. Users can search for datasets using business language rather than system names. They can explore information by domain, classification, glossary term or data product. Ownership information, lineage relationships and governance context are surfaced alongside technical metadata, helping users understand not only where data exists but also how it fits within the broader information landscape.

The catalogue is also more than a search box. Assets are organised into governance domains owned by the business, and packaged as data products that bundle related datasets with a described purpose, an accountable owner and terms of use. Alongside this sits data quality and health reporting, so stewards can see where definitions are missing, ownership is unclear or quality rules are failing, and consumers can request access to a product through a governed workflow rather than an email.

The introduction of the Unified Catalog is particularly significant because Microsoft is increasingly positioning it as the primary discovery and governance experience across the Microsoft data ecosystem. As organisations adopt Microsoft Fabric, OneLake, Purview and other platform services, the need for a common discovery layer becomes increasingly important. The catalogue provides a way of connecting data consumers with information assets without requiring detailed knowledge of the underlying technologies.

In many respects, the Unified Catalog represents a shift in governance thinking. Historically, governance initiatives often focused on controlling data. Increasingly, organisations are recognising that understanding and discoverability are equally important. Information that cannot be found, understood or trusted delivers little value regardless of how well it is protected.

Why this matters in the Age of AI

The renewed interest in data catalogues is not happening by accident. Generative AI is changing how people expect to interact with information. Employees increasingly assume that organisational knowledge should be discoverable, understandable and available at the point of need. They are less willing to navigate multiple systems, departments and processes simply to locate information that they believe already exists somewhere within the organisation.

At the same time, AI systems themselves depend heavily on context. Data without ownership, definitions or appropriate metadata becomes harder to interpret consistently. Many organisations are discovering that successful AI adoption is closely linked to their ability to organise and describe information in a way that makes sense beyond the boundaries of individual systems. An assistant grounded in an undocumented estate will answer confidently from whichever copy of the data it reaches first — and nobody will be able to say whether that copy was the right one. What appears to be an AI challenge often turns out to be a discoverability challenge.

More than a catalogue

The strongest data governance programmes are not built around catalogues. They are built around understanding. The value of Microsoft Purview Unified Catalog is not that it creates another inventory of information assets. Its value lies in helping organisations connect people with data more effectively, reducing reliance on undocumented individual knowledge and making information easier to discover, understand and trust. For many organisations, that represents a significant cultural shift. The goal is no longer to request information from the people who know where it lives. The goal is to create an environment where discovery becomes a normal part of working with data because in most organisations, the problem is not that valuable information is missing. The problem is that nobody realised it was already there.

 



Tuesday, 21 July 2026

Cabinet Level AI and How Britain’s Strategic Shift Changes the Data & AI Governance Landscape

The elevation of the Artificial Intelligence portfolio into the UK Cabinet marks a defining moment in British technology policy. With Kanishka Narayan promoted to attend Cabinet as Minister for AI, the message from Whitehall is unmistakable: artificial intelligence is no longer just a subset of digital policy or a niche driver of economic tech hubs. It is now a core pillar of national strategy, alongside economic growth, defense, and public infrastructure.

This structural shift signals that the UK intends to actively shape the global AI trajectory rather than merely adapt to it. However, accelerating AI innovation is only half the battle. Bringing dedicated ministerial oversight into the top room of government fundamentally alters how businesses, builders, and policymakers must approach data governance.

Opening the Floodgates for Innovation

For tech builders and investors, a dedicated Cabinet seat brings much needed political capital and decision-making speed. Historically, technology portfolios in government have wrestled with fragmented mandates across separate departments. Placing AI leadership directly within the Cabinet Office streamline policy across government bodies, offering clear advantages:
  •  Infrastructural Investment: Delivering state of the art AI requires significant physical infrastructure from data centre capacity and grid access to supercomputing networks. Centralized ministerial authority helps unblock planning hurdles and lower energy-access barriers for compute providers.
  • Public Sector Transformation: AI deployment is moving beyond private-sector start ups. Direct ministerial drive allows the government to integrate AI solutions across healthcare, transportation, and public administration, turning the state into an early anchor client for domestic innovation.
  •  Global Influence: As international debates rage over technological sovereignty, safety standards, and intellectual property, having a high level AI Minister ensures Britain has a direct, unified voice in shaping cross-border regulations.
Yet, innovation does not happen in a vacuum. The speed at which a nation can deploy advanced systems is directly bounded by the strength and reliability of its data foundations.

The Heightened Need for Agile Data Governance

It is a tech adage that holds truer than ever in the generative era. An AI model is only as safe, effective, and unbiased as the data used to train and run it.

As the UK ramps up its AI ambitions, the regulatory spotlight will inevitably shine brighter on data pipelines. Rather than viewing compliance as a friction point, modern organizations must recognize governance as an essential enabler of sustainable innovation.



 1. Moving Beyond Generic Privacy Compliance
Standard GDPR compliance is no longer enough when feeding complex foundational models or automated decision engines. Organizations now face intricate queries regarding copyright, consent for machine learning uses, synthetic data generation, and systemic bias. Cabinet level prioritization will drive clearer regulatory frameworks, forcing companies to prove where their data originated and how it was processed.
2. Trust as a Competitive Differentiator
Public trust remains fragile. High profile data leaks, hallucinated outputs, or opaque automated decisions can derail enterprise initiatives overnight. Clear, transparent data governance protocols, including rigorous lineage tracking and auditability, provide the legal certainty required to deploy AI models safely at scale.
3. Fostering Regulatory Sandboxes

A centralized AI strategy enables government regulators to expand "regulatory sandboxes" controlled environments where businesses can test frontier models against real-world datasets without triggering immediate penalty risks. This gives enterprises a safe arena to experiment while establishing clear benchmarks for safety, security, and privacy compliance.

Striking the Balance: What Businesses Should Do Next

The creation of a Cabinet-level AI minister reflects a broader truth: you cannot separate the thrill of innovation from the rigor of oversight. As the UK government aligns its resources to build, attract, and scale world-leading technology, the private sector must prepare its data architecture for stricter scrutiny and faster deployment cycles.

Organizations looking to capitalize on this shift should focus on three immediate priorities

 1. Audit Data Provenance: Ensure training data and operational pipelines have clear, documented chains of ownership and consent.
 2. Implement Human-in-the-Loop Governance: Establish cross-functional AI oversight teams combining legal, engineering, and product leaders.
 3. Design for Interoperability: Build data architectures flexible enough to adapt as national standards and international compliance rules evolve.

The British government has signaled its commitment to shaping the future of AI. Now, the responsibility falls on organizations to build the trustworthy, data-driven foundations required to lead in it.



References & Further Reading

  1. GOV.UK Official Announcement: Minister of State (Minister for Artificial Intelligence) Role & Profile — Official ministerial appointment details for Kanishka Narayan MP across the Cabinet Office and the Department for Business, Innovation, Science and Trade.

  2. Bloomberg / The Straits Times: Burnham Picks Narayan as First British AI Minister to Attend Cabinet (July 2026) — Coverage on the elevation of the AI portfolio to Cabinet level, the restructuring of UK tech departments, and national AI infrastructure strategy.

  3. ETIH EdTech Innovation Hub: Kanishka Narayan Named UK AI Minister Under Andy Burnham (July 2026) — Analysis of the UK government's strategic focus on AI innovation, industrial policy, and global competitiveness.

  4. Department for Science, Innovation and Technology (DSIT): AI Safety Institute & Sovereign AI Strategy Frameworks — Policy documentation outlining UK guidelines for AI safety standards, regulatory sandboxes, and enterprise data governance.

Monday, 20 July 2026

Why Fellowship Matters when Championing Data, AI, and Community Leadership

Reaching a milestone in one’s career is always an opportunity for reflection. Looking back on my journey as a Fellow of the British Computer Society (FBCS), I am reminded of why I joined this community in the first place and what driving tech leadership truly means.

Building professional communities since 2019, my focus has consistently been on the critical intersection where innovation meets responsibility. Over the years, championing robust Data and AI Governance has moved from a niche technical necessity to an urgent strategic priority. As models become more complex and integrated into everyday business and societal decisions, ensuring our data foundations are solid, ethical, and trustworthy is essential.

Sharing knowledge and mentoring others through these shifts isn't just a professional duty. It is at the core of real leadership.

To me, Fellowship is about using expertise to create impact that lasts. It’s about building resilient frameworks, empowering the next generation of technologists, and ensuring that as technology advances rapidly, it does so on a foundation of integrity and public trust.

Thank you to everyone who has been part of this community-building journey so far. Here’s to continuing the work, pushing boundaries in AI governance, and fostering spaces where impactful ideas can thrive.


Thursday, 16 July 2026

Governing the Governance programme: Microsoft Purview Data Estate Insights

One of the oddities of data governance is that organisations often spend considerable time measuring the things they are governing and very little time measuring governance itself.

Governance teams are regularly asked for evidence that data quality has improved, that ownership is becoming clearer, that sensitive information is being identified correctly or that users are finding the information they need more easily. These are entirely reasonable questions. The challenge is that most governance programmes are built around activities rather than outcomes. Assets are catalogued, business terms are defined, stewardship models are introduced and policies are agreed, but understanding whether those efforts are changing organisational behaviour can be much harder than expected.

At the outset of a governance programme this is rarely a significant concern. The early focus tends to be on establishing foundations. Organisations need visibility into their information landscape, which is why discovery and cataloguing become priorities. They need shared business language and greater confidence in the information they consume, making glossary management, curation and lineage important investments. Eventually attention turns towards quality, access management and policy enforcement as governance moves from documentation into operational practice.

As programmes mature, however, a different conversation starts to emerge. The question is no longer whether governance activities are taking place. The question becomes whether those activities are making a measurable difference.

This is where Microsoft Purview Data Estate Insights occupies a distinctive place within the wider governance platform.

Unlike Unified Catalog, Data Map, Business Glossary or Data Policy, Data Estate Insights is not primarily concerned with helping users discover or manage individual assets. Its purpose is to provide visibility into the governance capability itself. In many ways it serves as the executive view of governance, bringing together information that allows governance leaders, data owners and steering committees to understand how the programme is evolving over time.

One of the first areas this reveals is the health and coverage of the governance estate. Earlier articles in this series discussed how Purview uses Data Map to discover and scan information across connected systems. Once those capabilities are established, leadership teams inevitably become interested in the broader picture. As governance programmes mature, attention tends to move away from individual datasets and towards broader questions of coverage, adoption and visibility. Governance leaders want to understand whether the organisation has a comprehensive view of its information landscape, whether discovery processes are operating successfully and where gaps still exist. They are also interested in how the estate is changing over time, particularly as new technologies, business domains and information assets are brought into scope.

These questions are often more valuable than the underlying asset count because they provide insight into adoption. A catalogue containing thousands of assets may sound impressive, but its value is limited if significant parts of the organisation remain disconnected from governance processes. Visibility into coverage helps organisations understand not only the scale of their estate but also the extent to which governance has reached different parts of the business.

The same principle applies to business understanding.

Many governance programmes invest heavily in business glossaries and curation. Definitions are agreed, business terms are documented and stewardship responsibilities are established. The real challenge is understanding whether that effort is changing the way information is managed. Data Estate Insights provides visibility into how technical assets are being connected to business terminology, showing where business context is being applied and where gaps remain.

This is particularly useful because it highlights a common mistake within governance programmes. It is relatively easy to measure the number of glossary terms created. It is much harder to assess whether those terms are being adopted consistently across the organisation. A glossary only delivers value when it becomes connected to the information people actually use. Looking at curation coverage often reveals far more about governance maturity than simply counting definitions.

Another important perspective comes from classification and sensitivity analysis.

Governance teams rarely have the capacity to focus on every dataset equally. Some information carries greater operational value, greater regulatory significance or greater risk than others. Understanding where sensitive information exists across the estate helps governance leaders concentrate effort where it is most needed. The discussion stops being about individual assets and starts becoming a broader consideration of organisational risk, accountability and prioritisation.

For many leadership teams, this is where governance begins to intersect with wider business concerns. Discussions about metadata and stewardship gradually evolve into conversations about compliance exposure, information handling and operational resilience. Having visibility into classification trends provides context that is difficult to obtain through traditional governance reporting alone.

Data quality introduces a similar challenge.

Most organisations focus considerable effort on improving the quality of critical data. Rules are implemented, measurements are established and remediation activities are undertaken. While individual quality scores can be useful, executive audiences are rarely interested in isolated measurements. What matters is whether quality is improving across business-critical domains and whether governance interventions are producing sustained results.

Data Estate Insights helps provide that perspective by surfacing quality trends across the estate. Rather than treating quality as a series of isolated issues, governance leaders can observe patterns, identify areas where improvements are being sustained and recognise where challenges continue to emerge. This makes it easier to prioritise investment and demonstrate progress using evidence rather than anecdotal feedback.

Perhaps the most valuable aspect of Data Estate Insights is that it changes the nature of governance conversations. Instead of focusing exclusively on governance activities, organisations gain a clearer view of governance outcomes. Discovery coverage, glossary adoption, stewardship engagement, classification visibility and quality trends all contribute to a broader understanding of governance maturity.

This is important because mature governance programmes are rarely judged by the number of policies they create or the number of assets they catalogue. They are judged by whether trust in data is increasing, whether accountability is becoming clearer and whether the organisation is becoming more confident in the way it uses information.

Throughout this series, the focus has gradually moved from discovering data to understanding it, governing it and operationalising standards around it. Data Estate Insights represents the next logical step because it provides a way of understanding whether those efforts are producing tangible results. Rather than simply governing information assets, organisations gain the ability to measure, manage and continually improve the governance programme itself.

For many governance leaders, that is the point at which governance stops being a project and starts becoming a capability.






Wednesday, 15 July 2026

Microsoft MVP 2026 renewal 9th Year

Feeling humbled and honoured to be recognised as a Microsoft MVP. Grateful for the compassion, collaboration, and innovation that define this community and for the inspiring people who make Data Governance, AI Governance, and Responsible AI such meaningful fields to work in.

Award Category: Data Platform

Technology Areas: Microsoft Purview - Data Governance, Fabric Analytics

Credly badge

Thank you to everyone who shares knowledge, mentors others, and builds with purpose. Here’s to continuing the journey with curiosity, integrity, and a touch of creativity.



There is a great map showing all MVPs for Microsoft Purview.



Tuesday, 30 June 2026

AI Governance is a Hollow Framework Without Data Governance

The Hard Truth: We are trying to govern the outputs of frontier AI without establishing strict control over the inputs.

Imagine a near-future scenario: a frontier AI developer launches its next-generation model family. Within days, researchers uncover a zero-day jailbreak vulnerability that allows the model to map and exploit critical software vulnerabilities with unprecedented autonomy. In a scramble, the federal government issues an unprecedented emergency directive, forcing the developer to suspend global API access under the banner of national security.

While this sounds like a techno-thriller, the current geopolitical trajectory suggests this crisis is an inevitability. When governments eventually panic and react to high-risk algorithmic outputs, they will find that treating commercial AI models like sudden tactical threats is an unsustainable way to regulate technology.

AI models do not generate safety risks out of thin air; they learn them from data. Reactive government bans and real-time output filters are panic buttons. True thought leadership in this space requires looking upstream.

The Missing Link: Why Data Governance is AI Governance

Effective risk management for frontier models cannot rely on real-time safeguards alone. True resilience requires structural data governance built across three distinct operational pillars:

1. Data Provenance and Vulnerability Tracing

If a model can be steered into identifying critical software infrastructure vulnerabilities, we must ask: What specific datasets allowed it to map these exploits? Data governance mandates a transparent, verifiable ledger of training data. Regulators and developers must be able to audit what a model actually "knows" long before it is deployed to the public.

2. Dynamic Data Retention as a Defense Layer

When developers scramble to mitigate active exploits, they rely heavily on short-term telemetry retention policies to analyze user prompt interactions and track malicious behavior. Knowing exactly how user data is ingested, logged, and securely monitored is the only way to detect non-universal, highly sophisticated jailbreaks in real time.

3. Access Control and Data Sovereignty

Enforcing geographical or citizenship-based restrictions on a cloud-native, globally distributed API environment is a logistical nightmare. Without ironclad data access governance—restricting who can query the model and where that telemetry is stored—preventing unauthorized cross-border interaction with advanced reasoning systems is practically impossible.

Four Critical Questions for Tech Sovereignty

As the boundary between commercial technology and national security blurs, organizations and global regulators must confront the deeper systemic questions facing the ecosystem:

  • Who defines the threshold? Who determines when an advanced reasoning capability crosses the line from a massive commercial benefit to an existential national security threat?

  • What are the standards of validation? What transparent, independent, and technically grounded benchmarks must exist before a governing body can disrupt commercial ecosystems?

  • How do we prevent total fragmentation? If strict export controls dictate who can use the best models, how do we avoid a fractured digital world where access to advanced reasoning is determined entirely by geographical alignment?

  • What role does international cooperation play? When the regulatory actions of one nation can disable access for businesses worldwide, how do we build international institutions capable of managing global technological externalities?

Moving From Friction to Resilience

If we continue to treat AI safety as a series of sudden regulatory halts and reactive software patches, we will paralyze market innovation without actually making the digital estate any safer.

Responsible AI is the destination, but we cannot get there without two non-negotiable operational tracks:

  1. AI Governance: Providing the systemic oversight, legal compliance, and risk frameworks needed to manage model deployment.

  2. Data Governance: Securing the upstream integrity, tracing, and access controls of the information that shapes those models in the first place.

Reactive regulations are a sign of a system in deep friction. True leadership demands that we look upstream, securing the data infrastructure today so we can safely innovate the AI capabilities of tomorrow.



Sources & Further Reading (Alternative Options)

  • White House Policy: "Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence" — focusing on the mandates for safety testing and red-teaming for frontier models.

  • Geopolitical Precedents: Bureau of Industry and Security (BIS) guidelines on advanced computing and semiconductor export controls to showcase how the U.S. government actually restricts technology infrastructure.

  • Technical Frameworks: The NIST AI Risk Management Framework (AI RMF), which details the industry-standard pillars for measuring and governing AI risk, mapping beautifully to your data governance argument.

Saturday, 27 June 2026

The End of the Governance Silo: Building a Unified AI & Data Strategy

There’s a pattern emerging across organizations adopting AI. They stand up an “AI Governance” function. They build a new ethics board. They create new policies for models, prompts, and outputs. And yet, at the same time, they leave Data Governance exactly where it was separate, disconnected, and often treated as a legacy concern. It feels progressive. It looks sensible. But in reality, it creates something far more dangerous, The Governance Silo and with it comes a hidden cost the Silo Tax:

  • Slower deployment
  • Conflicting rules
  • And, most critically, gaps in accountability and control

In truth, AI governance is not a separate discipline. It never has been. AI is not a new domain to govern. It is an extension of the data ecosystem you already have and when those two worlds are separated, governance doesn’t just weaken it fractures.

The Dangerous Illusion of AI Governance as a Separate Discipline

The instinct to separate AI governance often comes from a good place. AI introduces new risks: bias, explainability, ethical use, automated decision-making. These feel different from traditional data concerns like quality, ownership, and classification. But this separation ignores a fundamental truth that AI is entirely dependent on data. Without strong data governance covering lineage, quality, ownership, and control AI governance simply cannot function effectively. You cannot explain an AI decision if you cannot explain the data that shaped it. You cannot ensure fairness in outputs if you cannot trust the inputs. You cannot manage AI risk if the data pipeline itself is opaque and yet, many organizations are trying to do exactly that.

The Transparency Gap: When AI Works… But No One Knows Why

Imagine an AI model making the “right” decision. It performs well. It delivers value. The business is happy. But then comes a challenge from a regulator, a customer, or an internal audit. Why did the model make that decision? This is where the governance silo breaks down. AI governance demands explainabilityBut explainability depends on data lineage knowing where data came from, how it was transformed, and how it was used. Without that lineage, the organization is left with a model that work but cannot be trusted and in an AI-driven world, that is not a technical issue. It’s a business risk. The real question is no longer Does the model perform? It is Can we prove why it behaves the way it does?



The Feedback Loop: When AI Starts Creating Its Own Data

AI doesn’t just consume data. It creates it. Predictions, classifications, synthetic datasets, generated content all of these become new data assets flowing back into the organization and this is where the second major risk emerges. If that AI-generated data is not governed, catalogued, classified, and controlled it begins to operate outside the governance perimeter.

Over time, this creates feedback loops:

  • Models trained on outputs from previous models
  • Synthetic data reinforcing hidden biases
  • Decisions based on increasingly distorted sources

Unchecked, these loops can degrade accuracy, amplify bias, and erode trust in AI systems. This is the point where governance stops being about compliance and becomes about control of reality itself. because if you lose control of your data, you lose control of your AI.

The Blueprint for a Unified Governance Model

So what does a better model look like? Not two parallel governance structures. Not another layer of oversight. But a single, joined-up governance system that treats data and AI as one continuous pipeline. In practice, that means three fundamental shifts.

1. A Shared Language Across Data and AI

The simplest problems are often the most damaging. If your Data team defines “sensitive data” differently to your AI team. If “accuracy” means something different in a model than it does in a dataset. You don’t have governance. You have misalignment. A unified governance model starts with a shared taxonomy, common definitions, classifications, and standards that flow consistently from data creation through to AI output. This is what eliminates conflicting rules and the friction they create.

2. A Single Source of Truth for Data and AI Assets

Most organizations already have a data catalog. Few have one that extends into AI. A unified model requires a single, integrated metadata layer where:

  • Data is tagged, classified, and owned
  • AI datasets are labelled as “AI-ready” or “restricted”
  • Lineage connects data sources directly to model outputs

This creates visibility across the entire pipeline from ingestion to decision and that visibility is what enables trust because governance is not about documentation. It is about knowing what is happening, in real time, across your data and AI ecosystem.

3. One Governance Body, Not Two

The final and often most overlooked shift is organizational. Many organizations create separate AI ethics boards alongside existing data governance councils. This is a mistake. Effective governance requires joined-up decision making, where:

  • Data sources are assessed alongside model outputs
  • Ethical considerations are evaluated across the full lifecycle
  • Accountability is defined end-to-end

A cross-functional governance council bringing together business, data, AI, risk, and compliance is already the established model for governing enterprise data.  The answer is not to create another council. It’s to evolve the one you already have.

From Silos to Systems: A Shift in Thinking

The organizations that struggle with AI governance are often those still thinking in layers:

  • Data layer
  • AI layer
  • Governance layer

But in reality, these are not separate stacks. They are one system.

Data flows into models.
Models generate outputs.
Outputs become new data.

And governance must sit across that entire loop. This is why leading organizations are moving toward a single governance umbrella one that integrates data and AI governance to create consistency, transparency, and enforceable controls because in a world of continuous data and continuous automation, governance can no longer be fragmented. It has to be continuous too.

Conclusion: The Road to Scalable AI

There’s a tendency in AI discussions to focus on the models, the algorithms, the tools and the capabilities. But that’s not where success will be determined. The organizations that win the AI race will not be those with the most advanced models. They will be the ones with the most trusted, controlled, and governed data pipelinesBecause ultimately AI is the car. Data Governance is the road. And no matter how powerful the car is you cannot win a race on a road full of potholes.