Welcome

Passionately curious about Data, Databases and Systems Complexity. Data is ubiquitous, the database universe is dichotomous (structured and unstructured), expanding and complex. Find my Database Research at SQLToolkit.co.uk . Microsoft Data Platform MVP

"The important thing is not to stop questioning. Curiosity has its own reason for existing" Einstein



Showing posts with label Data Quality. Show all posts
Showing posts with label Data Quality. Show all posts

Tuesday, 21 July 2026

Cabinet Level AI and How Britain’s Strategic Shift Changes the Data & AI Governance Landscape

The elevation of the Artificial Intelligence portfolio into the UK Cabinet marks a defining moment in British technology policy. With Kanishka Narayan promoted to attend Cabinet as Minister for AI, the message from Whitehall is unmistakable: artificial intelligence is no longer just a subset of digital policy or a niche driver of economic tech hubs. It is now a core pillar of national strategy, alongside economic growth, defense, and public infrastructure.

This structural shift signals that the UK intends to actively shape the global AI trajectory rather than merely adapt to it. However, accelerating AI innovation is only half the battle. Bringing dedicated ministerial oversight into the top room of government fundamentally alters how businesses, builders, and policymakers must approach data governance.

Opening the Floodgates for Innovation

For tech builders and investors, a dedicated Cabinet seat brings much needed political capital and decision-making speed. Historically, technology portfolios in government have wrestled with fragmented mandates across separate departments. Placing AI leadership directly within the Cabinet Office streamline policy across government bodies, offering clear advantages:
  •  Infrastructural Investment: Delivering state of the art AI requires significant physical infrastructure from data centre capacity and grid access to supercomputing networks. Centralized ministerial authority helps unblock planning hurdles and lower energy-access barriers for compute providers.
  • Public Sector Transformation: AI deployment is moving beyond private-sector start ups. Direct ministerial drive allows the government to integrate AI solutions across healthcare, transportation, and public administration, turning the state into an early anchor client for domestic innovation.
  •  Global Influence: As international debates rage over technological sovereignty, safety standards, and intellectual property, having a high level AI Minister ensures Britain has a direct, unified voice in shaping cross-border regulations.
Yet, innovation does not happen in a vacuum. The speed at which a nation can deploy advanced systems is directly bounded by the strength and reliability of its data foundations.

The Heightened Need for Agile Data Governance

It is a tech adage that holds truer than ever in the generative era. An AI model is only as safe, effective, and unbiased as the data used to train and run it.

As the UK ramps up its AI ambitions, the regulatory spotlight will inevitably shine brighter on data pipelines. Rather than viewing compliance as a friction point, modern organizations must recognize governance as an essential enabler of sustainable innovation.



 1. Moving Beyond Generic Privacy Compliance
Standard GDPR compliance is no longer enough when feeding complex foundational models or automated decision engines. Organizations now face intricate queries regarding copyright, consent for machine learning uses, synthetic data generation, and systemic bias. Cabinet level prioritization will drive clearer regulatory frameworks, forcing companies to prove where their data originated and how it was processed.
2. Trust as a Competitive Differentiator
Public trust remains fragile. High profile data leaks, hallucinated outputs, or opaque automated decisions can derail enterprise initiatives overnight. Clear, transparent data governance protocols, including rigorous lineage tracking and auditability, provide the legal certainty required to deploy AI models safely at scale.
3. Fostering Regulatory Sandboxes

A centralized AI strategy enables government regulators to expand "regulatory sandboxes" controlled environments where businesses can test frontier models against real-world datasets without triggering immediate penalty risks. This gives enterprises a safe arena to experiment while establishing clear benchmarks for safety, security, and privacy compliance.

Striking the Balance: What Businesses Should Do Next

The creation of a Cabinet-level AI minister reflects a broader truth: you cannot separate the thrill of innovation from the rigor of oversight. As the UK government aligns its resources to build, attract, and scale world-leading technology, the private sector must prepare its data architecture for stricter scrutiny and faster deployment cycles.

Organizations looking to capitalize on this shift should focus on three immediate priorities

 1. Audit Data Provenance: Ensure training data and operational pipelines have clear, documented chains of ownership and consent.
 2. Implement Human-in-the-Loop Governance: Establish cross-functional AI oversight teams combining legal, engineering, and product leaders.
 3. Design for Interoperability: Build data architectures flexible enough to adapt as national standards and international compliance rules evolve.

The British government has signaled its commitment to shaping the future of AI. Now, the responsibility falls on organizations to build the trustworthy, data-driven foundations required to lead in it.



References & Further Reading

  1. GOV.UK Official Announcement: Minister of State (Minister for Artificial Intelligence) Role & Profile — Official ministerial appointment details for Kanishka Narayan MP across the Cabinet Office and the Department for Business, Innovation, Science and Trade.

  2. Bloomberg / The Straits Times: Burnham Picks Narayan as First British AI Minister to Attend Cabinet (July 2026) — Coverage on the elevation of the AI portfolio to Cabinet level, the restructuring of UK tech departments, and national AI infrastructure strategy.

  3. ETIH EdTech Innovation Hub: Kanishka Narayan Named UK AI Minister Under Andy Burnham (July 2026) — Analysis of the UK government's strategic focus on AI innovation, industrial policy, and global competitiveness.

  4. Department for Science, Innovation and Technology (DSIT): AI Safety Institute & Sovereign AI Strategy Frameworks — Policy documentation outlining UK guidelines for AI safety standards, regulatory sandboxes, and enterprise data governance.

Saturday, 27 June 2026

The End of the Governance Silo: Building a Unified AI & Data Strategy

There’s a pattern emerging across organizations adopting AI. They stand up an “AI Governance” function. They build a new ethics board. They create new policies for models, prompts, and outputs. And yet, at the same time, they leave Data Governance exactly where it was separate, disconnected, and often treated as a legacy concern. It feels progressive. It looks sensible. But in reality, it creates something far more dangerous, The Governance Silo and with it comes a hidden cost the Silo Tax:

  • Slower deployment
  • Conflicting rules
  • And, most critically, gaps in accountability and control

In truth, AI governance is not a separate discipline. It never has been. AI is not a new domain to govern. It is an extension of the data ecosystem you already have and when those two worlds are separated, governance doesn’t just weaken it fractures.

The Dangerous Illusion of AI Governance as a Separate Discipline

The instinct to separate AI governance often comes from a good place. AI introduces new risks: bias, explainability, ethical use, automated decision-making. These feel different from traditional data concerns like quality, ownership, and classification. But this separation ignores a fundamental truth that AI is entirely dependent on data. Without strong data governance covering lineage, quality, ownership, and control AI governance simply cannot function effectively. You cannot explain an AI decision if you cannot explain the data that shaped it. You cannot ensure fairness in outputs if you cannot trust the inputs. You cannot manage AI risk if the data pipeline itself is opaque and yet, many organizations are trying to do exactly that.

The Transparency Gap: When AI Works… But No One Knows Why

Imagine an AI model making the “right” decision. It performs well. It delivers value. The business is happy. But then comes a challenge from a regulator, a customer, or an internal audit. Why did the model make that decision? This is where the governance silo breaks down. AI governance demands explainabilityBut explainability depends on data lineage knowing where data came from, how it was transformed, and how it was used. Without that lineage, the organization is left with a model that work but cannot be trusted and in an AI-driven world, that is not a technical issue. It’s a business risk. The real question is no longer Does the model perform? It is Can we prove why it behaves the way it does?



The Feedback Loop: When AI Starts Creating Its Own Data

AI doesn’t just consume data. It creates it. Predictions, classifications, synthetic datasets, generated content all of these become new data assets flowing back into the organization and this is where the second major risk emerges. If that AI-generated data is not governed, catalogued, classified, and controlled it begins to operate outside the governance perimeter.

Over time, this creates feedback loops:

  • Models trained on outputs from previous models
  • Synthetic data reinforcing hidden biases
  • Decisions based on increasingly distorted sources

Unchecked, these loops can degrade accuracy, amplify bias, and erode trust in AI systems. This is the point where governance stops being about compliance and becomes about control of reality itself. because if you lose control of your data, you lose control of your AI.

The Blueprint for a Unified Governance Model

So what does a better model look like? Not two parallel governance structures. Not another layer of oversight. But a single, joined-up governance system that treats data and AI as one continuous pipeline. In practice, that means three fundamental shifts.

1. A Shared Language Across Data and AI

The simplest problems are often the most damaging. If your Data team defines “sensitive data” differently to your AI team. If “accuracy” means something different in a model than it does in a dataset. You don’t have governance. You have misalignment. A unified governance model starts with a shared taxonomy, common definitions, classifications, and standards that flow consistently from data creation through to AI output. This is what eliminates conflicting rules and the friction they create.

2. A Single Source of Truth for Data and AI Assets

Most organizations already have a data catalog. Few have one that extends into AI. A unified model requires a single, integrated metadata layer where:

  • Data is tagged, classified, and owned
  • AI datasets are labelled as “AI-ready” or “restricted”
  • Lineage connects data sources directly to model outputs

This creates visibility across the entire pipeline from ingestion to decision and that visibility is what enables trust because governance is not about documentation. It is about knowing what is happening, in real time, across your data and AI ecosystem.

3. One Governance Body, Not Two

The final and often most overlooked shift is organizational. Many organizations create separate AI ethics boards alongside existing data governance councils. This is a mistake. Effective governance requires joined-up decision making, where:

  • Data sources are assessed alongside model outputs
  • Ethical considerations are evaluated across the full lifecycle
  • Accountability is defined end-to-end

A cross-functional governance council bringing together business, data, AI, risk, and compliance is already the established model for governing enterprise data.  The answer is not to create another council. It’s to evolve the one you already have.

From Silos to Systems: A Shift in Thinking

The organizations that struggle with AI governance are often those still thinking in layers:

  • Data layer
  • AI layer
  • Governance layer

But in reality, these are not separate stacks. They are one system.

Data flows into models.
Models generate outputs.
Outputs become new data.

And governance must sit across that entire loop. This is why leading organizations are moving toward a single governance umbrella one that integrates data and AI governance to create consistency, transparency, and enforceable controls because in a world of continuous data and continuous automation, governance can no longer be fragmented. It has to be continuous too.

Conclusion: The Road to Scalable AI

There’s a tendency in AI discussions to focus on the models, the algorithms, the tools and the capabilities. But that’s not where success will be determined. The organizations that win the AI race will not be those with the most advanced models. They will be the ones with the most trusted, controlled, and governed data pipelinesBecause ultimately AI is the car. Data Governance is the road. And no matter how powerful the car is you cannot win a race on a road full of potholes.


Friday, 12 June 2026

The Foundations of Intelligence: Why Your AI is Only as Good as Your DAMA Score

There is a quiet but critical misconception at the heart of today’s AI boom. Organizations believe they are investing in artificial intelligence. In reality, they are investing in data and often, that data isn’t ready. AI is a sophisticated engine. But it doesn’t run on innovation, hype, or vendor capability. It runs on data. And if that data is incomplete, inconsistent, poorly understood, or ethically questionable, the outcome isn’t just suboptimal it’s dangerous.

We are starting to see this play out at scale. AI projects stall, models produce biased outputs, and trust erodes. The narrative often focuses on the technology, but the root cause is rarely the model itself. It is almost always the data. Or more precisely: the absence of effective data governance. The uncomfortable truth is this for most AI failures are not AI failures at all. They are data governance failures in disguise. Frameworks like DAMA-DMBOK2 have spent years defining what good looks like in data management. What has changed is not the principles, but the stakes. In a reporting world, weak data might produce a misleading dashboard. In an AI-driven world, it can drive automated decisions at scale. This is why the conversation needs to shift from AI readiness to something far more grounded: data maturity.


The Four DAMA Pillars That Actually Matter for AI

DAMA-DMBOK outlines eleven knowledge areas, but when it comes to AI, four stand out as foundational. These are not optional capabilities. They are prerequisites.

1. Data Quality: Where AI Success Begins (and Ends)

For decades, organizations have lived with the idea of good enough data.

Reports can tolerate missing fields. Dashboards can work around anomalies. Humans are remarkably good at compensating for imperfect information. AI is not. An AI model does not “interpret” data in context—it learns patterns from it. If those patterns are flawed, biased, or inconsistent, the model will embed those flaws into its outputs. Worse, once learned, these patterns are incredibly difficult to remove. Dimensions like accuracy, completeness, and consistency are no longer operational concerns; they are existential ones.  The principle of garbage in, garbage out has never been more relevant. Even the most advanced models will produce unreliable results if the data they are trained on is flawed. This is not theoretical. Organizations are already seeing AI initiatives fail due to poor data quality, with research indicating that only a small fraction of companies believe their data is sufficiently ready for AI. Data Quality is not just a pillar. It is the foundation.

2. Metadata Management: The Missing Layer of Intelligence

If data quality determines whether AI works, metadata determines whether it makes sense. Metadata is often misunderstood as technical documentation, schemas, tables, field names. But for AI, it is far more than that. It is context. AI needs to understand:

  • What the data represents (business meaning)
  • Where it came from (lineage)
  • How it should be used (rules, classifications)
  • When it was last updated (timeliness)

Without this context, even the most advanced models become guesswork engines.

This is particularly critical for large language models interacting with enterprise data. These models are powerful, but they struggle with ambiguity and organizational nuance. Without metadata, they cannot distinguish between similar concepts, interpret domain-specific language, or validate the “truth” of a data point. Metadata effectively becomes the translation layer between human intent and machine interpretation. And yet, it is one of the most neglected areas in AI initiatives. Many organizations rush into model development while overlooking metadata strategy only to discover later that their AI cannot scale beyond experimentation. There is a growing recognition that metadata is not just supportive it is determinative. Without it, AI initiatives falter, regardless of model sophistication. 

3. Data Architecture: Designing for Machines, Not Just Reports

Traditional data architectures were designed for people.

Data warehouses centralised structured data for reporting and dashboards slow, stable, and human-interpreted. But AI does not consume data in the same way. It requires real-time access, integration across sources, and the ability to handle both structured and unstructured information. This is where modern architectural patterns come into play. Concepts like Data Fabric and Data Mesh, both explored within DAMA, represent a shift from centralisation to connectivity. Instead of moving data into a single repository, these approaches focus on making data accessible, governed, and usable wherever it resides. A data fabric, for example, creates a unified layer across distributed systems, enabling real-time integration and governance without physically moving data. This matters because AI thrives on:

  • Diverse data sources
  • Real-time signals
  • Context-rich environments

Traditional warehouses, designed for retrospective analysis, struggle to meet these demands. Modern architectures are not just technical upgrades, they are enablers of AI capability. If data cannot flow, AI cannot function.

4. Data Security and Ethics: The Line You Cannot Cross

The final pillar is where data governance transitions into AI governance. AI models do not inherently understand privacy, consent, or regulatory boundaries. They will learn from whatever data they are given. If that data includes sensitive, restricted, or biased information, the consequences can be severe. DAMA has long emphasised data security, privacy, and stewardship. In the AI era, these are no longer compliance exercises—they are ethical imperatives. Regulations like GDPR are not just legal constraints; they define the boundaries of what is acceptable in data usage. If an organization does not have clarity over data ownership, access rights, and usage permissions, it cannot claim to be operating ethical AI. More broadly, this is about trust. Without governance, organizations risk:

  • Embedding bias into automated decisions
  • Exposing sensitive data through AI outputs
  • Losing control over how data is used and reused

Strong governance ensures that AI is not only effective, but also accountable, transparent, and fair. 

The Real Question: How AI-Ready Are You?

For the C-suite, the implication is clear.

AI readiness is not about how many models you have deployed. It is not about how advanced your platform is. It is not even about how much data you hold.

It is about how well that data is governed.

Frameworks like DAMA-DMBOK provide a structured way to assess this. They define maturity across areas like quality, metadata, architecture, and security. And that maturity directly correlates to AI risk. If your organization is:

  • Immature in data quality → expect unreliable AI outcomes
  • Weak in metadata → expect confusion and inconsistency
  • Fragmented in architecture → expect scalability issues
  • Unclear on governance → expect ethical and regulatory risk

In other words, your DAMA maturity is your AI readiness. This is not theoretical. Research consistently shows that organizations struggle to make AI work not because of technology limitations, but because they lack the data foundations to support it. 

Final Thought: The Age of Data Governance Has Arrived

We are entering a phase where data governance is no longer a background function. It is becoming the defining capability of successful AI organizations. The companies that succeed with AI will not be those with the most advanced models. They will be those with the most disciplined data practices, those who understand that intelligence is not created by algorithms, but enabled by trust in data. AI is not a shortcut around governance. It is the ultimate test of it.

Wednesday, 10 June 2026

Microsoft Purview May 2026 Announcements Explained

May 2026 was one of the most important release moments for Microsoft Purview in recent years. It marked a clear shift from foundational governance tooling into operational, AI-era data governance at scaleHere is a quick summary of what tools became General Availability (GA).


AI governance and security
  • Data security and compliance protections for Microsoft Agent 365 (GA) 
  • Expanded Purview capabilities to govern AI activity, including agent-based workloads and AI interactions 

Data governance (data quality maturity)

  • Standalone data asset data quality scans (GA) 
  • Incremental data quality scans (GA)
  • Configurable data quality thresholds (GA) 

Data security posture management (DSPM)

  • New unified Data Security Posture Management experience (GA rollout in May 2026) 
This wasn’t just feature updates. Microsoft has effectively:
  • Turned Purview into the control plane for AI governance
  • Matured data quality into an operational, measurable discipline
  • Shifted data security from reactive controls to proactive posture management

The conversation as now switched from talking about implementing governance to talking about running governance continuously. This places governance in the age of AI. The most significant announcement in May wasn’t a single feature but was the integration of Purview with Microsoft Agent 365.

At GA, this introduces:

  • Centralised visibility of AI agents interacting with enterprise data
  • Data loss prevention and sensitivity enforcement applied to AI usage
  • Auditability and compliance over AI-driven actions 

This is a fundamental shift. Previously, governance focused on:

  • Data at rest
  • Data in motion
  • Human access patterns

Now, governance must deal with:

  • Autonomous agents accessing and acting on data
  • AI-generated outputs and derived data
  • Decisions made without direct human interaction

Purview is now positioned to govern these.

Data Quality

The data governance updates might look incremental, but they  are actually  significant. With May’s GA releases:

  • Data quality can be measured continuously (incremental scans)
  • Thresholds can be defined and enforced consistently
  • Data assets can be assessed independently at scale 

This moves data quality from periodic profiling exercises to always-on monitoring aligned to business expectations. For organizations, this means:

  • Data quality becomes a control, not an insight
  • Ownership becomes enforceable (through thresholds)
  • Governance shifts closer to operational accountability

This aligns strongly with what many frameworks (DAMA, DCAM) have always pushed. That Data Quality must be actively managed and not passively reported.

Data Security Posture Management (DSPM)

The new DSPM experience reaching GA is arguably the most strategic element of the May release. It introduced:

  • Unified visibility across traditional and AI-driven data environments
  • Risk-driven prioritisation of data security issues
  • Guided workflows to turn insights into action

It also extends beyond Microsoft-native data with integration with third-party data sources and tools and a single view of sensitive data across the estate. This matters because most organizations struggle with:

  • Fragmented visibility
  • Too many alerts, not enough prioritisation
  • Governance that stops at reporting

DSPM changes the conversation to what matters most, and what do we fix first? There was a subtle but important shift: governance of everything, not just Microsoft. 

Another key theme in May’s updates was expanding governance beyond Microsoft workloads. Examples include:

  • Visibility into third-party AI tools and environments 
  • Integration across broader ecosystems and data sources 

This is critical for real-world governance because the reality is:

  • Data does not live in one platform
  • AI is not limited to one vendor
  • Risk spans the entire digital estate

Purview is increasingly positioned as the normalising layer across that complexity. For organizations like those in housing, local government, or financial services (your typical audience), these updates directly address four growing risks:

1. AI adoption without governance

Agents and copilots are being deployed faster than policies can keep up.

→ Purview now provides policy enforcement and visibility at the AI layer.

2. Lack of data ownership and accountability

Data quality issues remain hidden until failure.

→ Thresholds and continuous scanning make ownership measurable.

3. Fragmented security controls

Tools exist, but there is no unified posture view.

→ DSPM provides a single, prioritised risk lens.

4. Increasing regulatory pressure

Frameworks are evolving faster than implementation capability. Purview now supports continuous compliance monitoring, not point-in-time audit.

The strategic takeaway shows a clear direction from Microsoft that Governance is no longer a framework or a project. It is an always-on operational capability. Purview is evolving into:
  • The execution layer for governance
  • The control point for AI and data risk
  • The bridge between business intent and technical enforcement

For organizations, the implication is equally clear:

  • Governance must move from design to operation
  • Ownership must move from assumed to measurable
  • Risk must move from identified to actively managed
The organizations that succeed with these updates won’t be the ones that deploy Purview fastest. They’ll be the ones that:
  • Define clear ownership and accountability first
  • Align governance to business outcomes, not tools
  • Use Purview to operationalise, not define their governance model

These announcements reinforce that Technology does not create governance. It makes it visible and enforces it.

Reference

What's new in Microsoft Purview | Microsoft Learn

Sunday, 28 December 2025

What Responsible AI Actually Means for Data Leaders in 2026

Responsible AI has become a buzzword, but for data leaders it’s a practical discipline. It is not just about lofty principles or glossy frameworks. It is about ensuring that models behave predictably, ethically, and transparently. That requires more than good intentions. It requires operational governance. Data quality, lineage, access control, and policy enforcement are not side notes; they are the mechanisms that make responsible AI real.

The challenge is that many organisations still treat responsible AI as a compliance checkbox. They focus on documentation rather than behaviour, and on principles rather than practice. But responsible AI is not something you declare—it’s something you operationalise. It lives in your data pipelines, your monitoring processes, your access controls, and your governance culture.

For 2026, the organisations that thrive will be those that embed responsible AI into their data strategy. This means aligning governance with the lifecycle of AI systems, from data sourcing to model deployment to ongoing monitoring. It means treating transparency as a design requirement, not an afterthought.

Responsible AI isn’t a brake on innovation, it’s the steering mechanism. Without it, organisations risk building systems they cannot explain, defend, or trust. With it, AI becomes a strategic advantage rather than a liability.



Monday, 7 April 2025

Why Fabric’s New Health & Quality Capabilities Matter More Than Ever

At FabCon this year, Microsoft doubled down on something many of us in data governance have been saying for a long time: trustworthy data doesn’t happen by accident. It is engineered, monitored, and continuously improved. The newly announced health, quality, and observability capabilities in Microsoft Fabric signal a decisive shift away from reactive firefighting and toward proactive, platform‑level assurance.

For organisations scaling AI, analytics, and operational data products, this matters. Data Quality and Observability are no longer “nice to have”; they are the minimum viable conditions for responsible, repeatable, and compliant data use.

Below is a concise, actionable breakdown of what these new capabilities mean—and how to turn them into immediate value across your estate.

1. Treat Data Health as a First‑Class Operational Signal

Fabric’s expanded health management capabilities give teams something they’ve historically lacked: a unified, platform‑native view of data system health. Instead of stitching together logs, alerts, and manual checks, you now get:

- Integrated telemetry across pipelines, workloads, and storage  

- Early‑warning indicators for degradation, drift, or failure  

- Operational insights that connect system behaviour to business impact  

This elevates data health from a technical afterthought to a governance‑aligned operational metric. For leaders, it means you can finally answer the question: “Is our data estate healthy enough to trust today’s decisions?”

Action: Establish a weekly “Data Health Review” ritual—short, structured, and tied to business outcomes. Treat it like you would a security posture review.

2. Use Data Quality as a Contract, Not a Cleanup Exercise

The new Fabric capabilities reinforce a principle I advocate in every governance programme: quality must be defined, measured, and enforced at the point of creation.

With Fabric’s enhanced quality tooling, teams can now:

- Define expectations (validity, completeness, timeliness) as part of the data product  

- Monitor quality continuously, not periodically  

- Surface issues directly to producers and consumers  

- Build trust signals into downstream AI and analytics workloads  

This shifts quality from reactive cleansing to proactive assurance , a contract between producers and consumers.

Action: Publish a lightweight “Quality Contract” template for all critical data products. Keep it simple: purpose, expectations, checks, and escalation paths.

3. Make Observability the Backbone of AI Governance

As AI workloads scale, observability becomes the difference between responsible innovation and uncontrolled risk. Fabric’s new observability features support:

- Traceability from source to model  

- Lineage‑aware debugging  

- Impact analysis when upstream changes occur  

- Evidence trails for audits, compliance, and Responsible AI reviews  

This is not just operational hygiene, it is AI governance in practice. You cannot assure fairness, accuracy, or safety in AI systems without deep visibility into the data that feeds them.

Action: Integrate Fabric observability outputs into your Responsible AI lifecycle checkpoints—especially model validation and change‑control reviews.

In summary the message from FabCon is clear: health, quality, and observability are now strategic capabilities, not technical chores. For organisations building modern data estates and especially for those embracing AI, where these features are the foundation of trust.









Saturday, 29 June 2024

The Growth of Microsoft Purview

What is Microsoft Purview is a question I often get asked. The answer is never what people expect. Over the last couple of years the Microsoft Purview solution has been growing in capability with many different applications being brought together under one umbrella term, Microsoft Purview. It is now described as having three areas.

  • Data security for information and cybersecurity teams 
  • Data governance for data consumers data engineers and data officers
  • Risk and compliance for risk compliance and legal teams

When we talk about Microsoft Purview it helps to know what the business problem is to identify which areas of the product are required.

I put together a diagram to help map out all the applications to date. In the diagram  it is clear to see the extent of the applications that Microsoft have brought together and added over the last year, to the suite of tools. The documentation refers to the 3 high level areas Risk and Compliance (shown in blue), Data Governance (shown in purple) and Security (shown in orange and green).  I have depicted the AI hub in green rather than orange, because it covers a different conceptual area of Responsible AI and that compliance protection for generative AI apps.

When I talk about data governance I am looking at apps such as data state health, roles and responsibilities for the data estate, the data catalogue, classification, lineage, all the new business domain areas such as a business glossary and the really new data quality set of tools to control and monitor quality and the health of the data. 

Thus, the new reimaged Microsoft Purview experience provides a holistic view of your data and enables better automated data management. 



Saturday, 24 June 2023

Data Quality inspiration, erudition and expedience

Excited to share my session on Data Quality inspiration, erudition and expedience at 17:45 at Data Toboggan Cool Runnings Saturday 24 June 2023.

This session covers

Great data quality is a key output of data systems. We often get wrapped up in the technological, like Microsoft Fabric, Synapse and Purview. This session covers the what why and how of data quality, to help provide a deeper understanding and roadmap for adoption.



Friday, 9 June 2023

Fabric an enabler for better data quality

 I was reading this article

How Microsoft Fabric aims to beat Amazon and Google in the cloud war and it sparked some thoughts on data quality.

Technical architectures are always changing but over the last few weeks we have witnessed a pivotal change with the introduction of Microsoft Fabric, creating a paradigm shift.  One lake has advantages of cost saving, transparency, flexibility, data governance and data quality.  Data governance and data quality are two of the areas that I feel need the most work in the data stack. Both are heavily reliant on people and how they perceive its importance. New style data governance is an enabler through distributed  teams in the business. Increased data quality is high on the list of areas that need improvement, but that has not seen any significant change with the move to cloud based services. A tool in Fabric called shortcuts helps enterprises with that single virtualized data lake across multi-clouds and is a stepping stone to better data quality.

Previously Synapse combined services into a single place for data lake and data warehouse with integrated Microsoft Purview to help with providing a greater holistic view of data. Fabric, as a SaaS service, goes one step further to provide a single place for data management and data governance on OneLake. It is enabling improved consistency and trustworthiness of data.  Bridging the gap between BI and AI  which brings with it a unique opportunity to improve data quality. Good quality data you can trust is the foundation stone of successful business growth.

There is a vast amount of documentation currently available to help us learn about Fabric. A few ket resources are below.

https://aka.ms/fabriccommunity

https://aka.ms/fabricblog

https://aka.ms/learn-fabric

https://aka.ms/fabricicons

Why Data Quality is important for your business

Original article I wrote was Published on LinkedIn 

Data Quality

#data volumes are increasing significantly. It is hard to keep track of the data used in a business and to know how accurate the data is, that is being used for business decisions. Accurate data is critical to the success of a business. Ensuring that data comes from the right source, is not duplicated, does not have syntax issues or is incorrectly updated requires good data management.

What is data quality?

DAMA UK defines data quality as ‘the planning , implementation , and control of activities that apply quality management techniques to data, in order to assure it is fit for consumption and meets the needs of data consumers’ So good data quality needs accurate data that meets the needs of a business.

There is also an international standard for data quality, ISO 8000. This states ‘Quality is actually the conformance of characteristics to requirements and, thus, any item of data can be of high quality for one use but not for another use that has differing requirements.’ The ISO8000 identifies that data quality has syntactic (format), semantic (meaning) and pragmatic (usefulness) characteristics.

Why is data quality important

Good data quality enables strong business decisions to be made leading to better business outcomes. These business decisions based on data can lead to greater profitability, helping improve situations and can be the starting place for predictive analytics using AI systems. Poor data quality often occurs due to human error and being able to put checks and balances in place to know the current state of the data is important.

Measuring data quality

It is important for every business to understand the level of their data quality maturity. This enables trustworthy decisions to be made and can even impact data ethical understanding. There are 6 main dimensions to consider.

  • Completeness – this metric addresses the requirements that all data sets and data items have all the relevant information. The measure is about that missing information, thinking about the proportion of data received, incomplete, missing, and data loss.
  • Accuracy – does the data reflect the data set real world, is it truthful containing the correct data entries.
  • Uniqueness – is a single view of the data. This metric looks at the extent of the duplication of data, with consideration of how data is controlled.
  • Validity – does the data match the rules such as syntax (format, type, range) of its definition.
  • Consistency – do the data sets match across data stored in two or more records. Does the pattern \ frequency of data match.
  • Timeliness – the degree the data represents reality at a point in time. Is the data available when required.

DAMA provide details on how to calculate these measure to provide those KPI’s to business. A useful technique to help provide a view on the current state of the data, is data profiling. This can look at things like counting nulls in data, the max/min value, max/min length, frequency distribution and data type and format.

Data enhancements

The main take away is to always consider the quality of the data you use and embed checks into the business processes. Always have a dashboard showing the current state of business data quality. If you are looking for a framework to help guide your thought processes the UK government has created a data management framework which is worth reviewing. In conclusion start with a review of the data quality of the core data sources you use in your data catalog and record the current state.

Key next steps to improve Data Quality

To start on the path of immediately improving the data quality, for the organisation and enable better analytics:

  • Understand what data exists
  • Assign a relevant data owner to the data sources
  • For each business use case understand the lineage of the data used for reports and dashboards.

Here are some helpful resources to assist with the above steps!

Creating a Data Catalog with Microsoft Purview - Cloud Adoption Framework

Improving Data Lineage with Microsoft Purview - Cloud Adoption Framework

Data quality considerations - Cloud Adoption Framework