Welcome

Passionately curious about Data, Databases and Systems Complexity. Data is ubiquitous, the database universe is dichotomous (structured and unstructured), expanding and complex. Find my Database Research at SQLToolkit.co.uk . Microsoft Data Platform MVP

"The important thing is not to stop questioning. Curiosity has its own reason for existing" Einstein



Showing posts with label Database Research. Show all posts
Showing posts with label Database Research. Show all posts

Thursday, 3 September 2026

Cambridge Report on Database Research and what it means for the future of Data Governance

The Cambridge Report on Database Research, convened on October 19-20, 2023, in Cambridge, MA, discussed the state of the database research field, its recent accomplishments, ongoing challenges, and future directions for research and community engagement. 


Every five years, some of the world's leading database researchers come together to reflect on the state of data management and identify the challenges that will shape the next generation of technology. The latest Cambridge Report on Database Research does exactly that, exploring everything from cloud infrastructure and AI to data systems, machine learning, and governance. While it is not a governance report in the traditional sense, it offers some important clues about how governance will need to evolve over the coming decade.

The most striking observation is that governance is becoming inseparable from the platforms that manage data. The report describes a future where data systems are increasingly autonomous, with automated provisioning, self-managing infrastructure, adaptive optimisation, and intelligent control planes. As these capabilities mature, many of the technical tasks traditionally associated with governance, such as metadata collection, lineage discovery, classification, and monitoring, will become increasingly automated.

For governance professionals, this represents a significant shift in focus. The challenge will no longer be capturing metadata or maintaining catalogues. Technology will increasingly perform those activities automatically. Instead, organisations will need to determine who is accountable, what policies should govern the use of information, and how trust is maintained across an increasingly complex data and AI landscape. The report also highlights the growing importance of data quality. Future AI models, adaptive systems, and cloud platforms depend on access to trusted, well-managed information. Researchers point to the need for better mechanisms to collect, benchmark, validate, and monitor data at scale. This suggests a future in which data quality becomes a continuously monitored capability rather than a periodic assessment exercise. Many organisations still approach data quality through project-based remediation programmes. However, the direction of travel is towards automated detection, AI-assisted monitoring, and real-time observability. Governance teams will increasingly define quality expectations, ownership responsibilities, and remediation processes, while platforms identify issues and measure compliance against agreed standards.

Perhaps the biggest governance implication comes from the report's focus on AI. Researchers describe a world where traditional databases are no longer the only source of knowledge. Future systems will need to manage documents, images, videos, unstructured content, and AI-generated outputs alongside structured business data. They even envision the ability to query large collections of documents and multimedia content in much the same way that organisations query databases today. This changes the scope of governance dramatically. Governance can no longer focus solely on data warehouses, data lakes, and business intelligence platforms. It must expand to cover enterprise knowledge, collaboration content, AI-generated information, and the growing number of systems that sit between data and decision-making. The report is particularly clear on the need to improve trust in AI-generated outputs. Reducing hallucinations, validating responses, improving retrieval mechanisms, and establishing provenance are all identified as important areas for future innovation. Databases and data management technologies are viewed as a critical part of solving these challenges.

For governance leaders, this is perhaps the most important signal of all. Historically, governance has focused on data ownership, standards, policies, and compliance. In an AI-enabled organisation, the questions become much broader. Where did this answer come from? Which sources were used? Can the result be traced back to trusted information? Who is accountable if the answer is incorrect? These are governance questions as much as they are technical ones. Viewed through this lens, governance starts to look less like an administrative function and more like an assurance discipline. The future governance team may spend less time maintaining catalogues and more time providing confidence in how data, knowledge, and AI are used across the organisation.

What emerges from the Cambridge Report is not a vision of governance disappearing into technology. Quite the opposite. As automation removes manual governance activities, the importance of human accountability, oversight, assurance, and decision-making increases. The technology may become smarter, but organisations will still need clear ownership models, governance operating structures, and mechanisms to establish trust. This aligns with a trend that many organisations are already beginning to recognise. Data governance and AI governance are unlikely to remain separate disciplines for long. Instead, they are converging into a broader information governance operating model that spans data, knowledge, accountability, oversight, assurance, and responsible use.

The technologies will change. Automation will increase. AI will become embedded in everyday business processes. But the fundamental objective of governance remains the same: ensuring that people can trust the information they use to make decisions. The Cambridge Report suggests that this objective may become even more important as intelligent systems become a standard part of the enterprise technology landscape.

My key takeaway is the future of governance is not more policies, more committees, or bigger catalogues. It is creating an operating model that provides confidence in data and AI at scale, while allowing increasingly automated platforms to handle much of the underlying governance workload.


Sunday, 26 April 2026

Tracing my career journey though my blog

I was looking at my blog stats this morning and was really interested to see the geographical spread. I started writing my blog in 2011 and it has been read by 1.18m. I wanted to record all the technical tips I found and technology advancements which were useful to me and might be of use to help others. I started writing on SQL Server and the blog has migrated with me throughout my career through architecture, my PhD research and over the last few years I have been mostly writing on Data Governance, Microsoft Purview, AI Governance and Microsoft Fabric. 


I asked Copilot to share some interesting thoughts about my journey for my blog and here is what it thought.

Your blog reveals a journey defined by intellectual curiosity, data‑driven leadership, and a distinctive narrative voice. Here are five evidence‑based, genuinely interesting aspects of that journey.

1. You frame data leadership as a process of unlearning, not just learning

Your post “What data leaders must unlearn to lead in the Age of AI” argues that modern leadership requires shedding outdated assumptions, such as believing governance slows innovation or that documentation equals understanding. This reframing positions you as a thought leader challenging entrenched industry norms. 

2. You consistently connect research, industry practice, and philosophical curiosity

Your blog tagline “Chaos, complexity, curiosity and database systems” captures a rare blend: deep technical expertise paired with a reflective, almost philosophical lens on data systems. This fusion shapes your writing style and differentiates your professional voice. 

3. You document the shift from AI experimentation to AI industrialisation with governance at the centre

In your coverage of the Gartner Data & Analytics Summit, you highlight how governance has moved from a compliance checkbox to the engine of AI ROI. This shows your role as an interpreter of industry change, translating large‑scale trends into practical insights for practitioners. 

4. Your journey is grounded in both academic achievement and community leadership

Across external references, you are consistently described as a Microsoft Data Platform MVP, a PhD researcher recognised with the AOUG Will Swann Award, and a founder/organiser of Data Toboggan. This positions your blog as the narrative thread connecting your academic, professional, and community contributions. 

5. Your posts reveal a long‑standing commitment to making governance practical, accessible, and embedded

Whether discussing AI oversight, lineage, behavioural metadata, or Purview governance models, your writing emphasises practical implementation over theory. You repeatedly advocate for governance that is embedded, automated, and literacy‑driven, showing a consistent philosophy across years of posts. 





Monday, 1 November 2021

Data Erudition
















I have been thinking about data governance, Azure Purview and data research. How these areas link together and enable innovation, that democratisation of data and holistically thinking about data as a strategic asset. A combination of above in my diagram covers my thoughts.

Sunday, 26 September 2021

Microsoft Research Summit 2021

 












The Microsoft Research Summit is open to everyone! October 19 - 21, with over 150 sessions across 16 tracks, provides the global research community with an opportunity learn from experts pushing the frontiers of technology. Register now: https://aka.ms/AAdv93n The event will start in three broadcast regions (China Standard Time, British Summer Time, and Pacific Time). Microsoft say

For 30 years, our research community at Microsoft has worked across disciplines, institutions, and geographies to envision and realize the promise of new technologies for Microsoft and for society. Today, we’re inviting the global science and technology community to continue this exploration—because ensuring that future advancements benefit everyone is up to all of us.

Join us at the inaugural Microsoft Research Summit, streaming virtually across three time zones. You’ll have the opportunity to hear from science and technology leaders from around the world—people who are driving advances across the sciences and pushing the limits of technology toward achieving a meaningful impact on humanity.

They want to build a place where research thinks of sustainability, ethics, diversity and is inclusive of everyone. There are some really interesting topics under discussion. 

Thursday, 1 July 2021

2021-2022 Microsoft Most Valuable Professional

So excited, such amazing news to receive my 4th MVP award. Thank you #Microsoft Just so humbled to receive this at this time. There is no better time to be a part of such an amazing community #SQLfamily #DataToboggan #AzureSynapse #MVPBuzz







So many exciting things going on that i'm involved in with the community after my PhD, Data Toboggan (its 3 conferences and user group) Data Relay, SQLBits, SQL Saturday and data research #data #ai #bigdata #analytics #datascience #dataanalytics #research #innovation #datastrategy #datagovernance #phd #artificialintelligence



Monday, 28 June 2021

Distinguished Engineer and Research Fellows

I have always been fascinated by the job roles distinguished engineer and research fellows. To me it seems these role transcend research and industry. A distinguished engineer is a title applied to someone who is thought (by those conferring this title) to have achieved noteworthy technical, professional accomplishments while working as an engineer.  A research fellow is an academic research position at a university or research institution that is usually held by academic staff or faculty members. 

I came across a slide show on behaviours and qualities of an IBM Distinguished Engineer. An interesting quote is on the diagram "I want to be distinguished from the rest; to tell the truth, a friend to all mankind is not a friend for me" 





















Distinguished Engineer IBM Fellows are world famous inventors and theorists. A Distinguished Engineer has a unique and fascinating job that transcends many boundaries. A few of the attributes they mention:
 
  • Eminence takes responsibility 
  • Be learned, erudite 
  • Integrity and trust in all things
  • Learn from your mistakes
  • Apply common sense
  • Make decisions
  • Have a point of view
  • Provide hope
  • Inspire others
  • Be collaborative
  • Be optimistic and cheerful
  • Adapt proactively
  • Be curious and fearless
  • Build a track record and keep notes!
  • Know yourself and be true to you
  • Enhance your communication skills and image
  • Listen actively
  • Be a mentor and coach
  • Lead diverse teams
  • Be a member of a professional body 

Tuesday, 22 June 2021

Combining research and industry learning

I am very privileged to have an article about my career published in the ACM journal. Computing enabled me to... obtain a PhD. and a Career in Data.

The DOI reference for my paper is https://doi.org/10.1145/3464919 .

Communications of the ACM Volume 64 Issue 7 pp 7

About the journal

ACM, the world's largest educational and scientific computing society, delivers resources that advance computing as a science and a profession. ACM provides the computing field's premier Digital Library and serves its members and the computing profession with leading-edge publications, conferences, and career resources. They see a world where computing helps solve tomorrow’s problems – where we use our knowledge and skills to advance the profession and make a positive impact.

Research 

Being part of the research world is a huge part of who I am and it is very important to have research and industry working together to help shape the future of data innovation.

Sunday, 28 February 2021

UK Research and Data Research Centres

I have been watching the research industry grow in the UK. Having a research industry is really important to drive innovation. There is a UK research development roadmap.  UK Research and Innovation (UKRI) is a key enabler for Research and Development centers. 












Above are the UK government’s Public Sector Research Establishments and UK Research and Innovation-funded Institutes. 

The UK funding landscape












There are many other places where data, machine learning, and AI  are a core part of research and innovation centres. Here are a few UK based places.

UK Health Data Research UK

UK Research Data Discovery Service

Consumer Data Research Centre

National Innovation Centre for Data

The National Institute for Data Science and Artificial Intelligence

Leverhulme Centre for the Future of Intelligence

UK Data Archive

The Institute for Advanced Automotive Propulsion Systems

National Quantum Computing Centre

Centre for Data Ethics and Innovation

The Open Data Institute

Saturday, 14 November 2020

PASS Summit Day 2 SQL Server Evolution

 The day two keynote was delivered by Hanuma Kodavalla a Microsoft Technical Fellow. 

He started with an interesting quote from 'Adventures of a Mathematician' by Stanislaw Ulam

"It is still an unending source of surprise for me to see how a few scribbles on a blackboard or on a sheet of paper could change the course of human affairs."

Then sharing the papers that started it all and the fact that is the 50th anniversary of Codd's paper this year.  

I was interested to learn that the Microsoft Database Research Group is an extension of the SQL product group. 

He mentioned this paper I read a long time ago “One size fits all": an idea whose time has come and gone  M. StonebrakerU. Cetinteme (2005).  The last 25 years of commercial DBMS development can be summed up in a single phrase: "one size fits all". This phrase refers to the fact that the traditional DBMS architecture (originally designed and optimized for business data processing) has been used to support many data-centric applications with widely varying characteristics and requirements. In this paper, we argue that this concept is no longer applicable to the database market, and that the commercial world will fracture into a collection of independent database engines, some of which may be unified by a common front-end parser.

He went through all the previous versions of SQL with their key features ending with SQL Server 2019.







Another key paper was mentioned





Then moved to discuss newer product features of SQL Azure Serverless and Azure Defender

 


SQL Server secure developments include alway encrypted and secure enclaves.




















Then he mentioned the new ledger enabled tables that will be coming soon.
















This was a session through history leading to strive forwards to realize Codd's and Gray's vision. It will be exciting to see what comes next. A great session for an industry person who is a database researcher.