Welcome

Passionately curious about Data, Databases and Systems Complexity. Data is ubiquitous, the database universe is dichotomous (structured and unstructured), expanding and complex. Find my Database Research at SQLToolkit.co.uk . Microsoft Data Platform MVP

"The important thing is not to stop questioning. Curiosity has its own reason for existing" Einstein



Showing posts with label Azure Cosmos DB. Show all posts
Showing posts with label Azure Cosmos DB. Show all posts

Thursday, 13 October 2022

Microsoft Ignite 2022


The Microsoft Ignite 2022 keynote on 12 October shared many innovations and there were lots of announcements. The Microsoft Book of News October 12 - 14, 2022  details these  https://news.microsoft.com/ignite-2022-book-of-news/

Satya talked about the digital imperative and the world's computer - Azure. The main 5 themes are to do more with less.









The technology world is changing rapidly and Satya mentioned that Gartner predicts that by 2025, 70 % of applications will be made by no code/low code tools, up from 25% in 2020. enabling hand drawn forms for example to be converted into apps with  AI

Microsoft and Databricks deepen partnership for modern, cloud-native analytics

https://techcommunity.microsoft.com/t5/azure-data-blog/microsoft-and-databricks-deepen-partnership-for-modern-cloud/ba-p/3640280  

Microsoft and Databricks have partnered to build a foundation in the Microsoft Intelligent Data Platform by integrating their hallmark capabilities to create an integrated solution for our customers.



Distributed PostgreSQL comes to Azure Cosmos DB

https://devblogs.microsoft.com/cosmosdb/distributed-postgresql-comes-to-azure-cosmos-db/

Azure Cosmos DB for PostgreSQL, a new Generally Available service to build cloud-native relational applications. Azure now offers its own single database service that supports both relational and NoSQL workloads. You can build cloud-native applications for relational and non-relational data using Cosmos DB.











Introducing the Microsoft Intelligent Data Platform Partner Ecosystem

https://techcommunity.microsoft.com/t5/azure-data-blog/introducing-the-microsoft-intelligent-data-platform-partner/ba-p/3640279

There was a launch of a powerful new Partner Ecosystem for the Microsoft Intelligent Data Platform deliver category-leading and cloud-native data and AI solutions integrated with the Microsoft Intelligent Data Platform to complement capabilities to address diverse scenarios.













Data Governance

Data Governance was mentioned at every single opportunity in many sessions. Data Governance looking at security and compliance, data management and responsible democratisation.  Announcements were Microsoft Purview Business workflows (GA), Business metamodel (Preview), Improved root cause analysis and traceability with SQL Dynamic lineage 













More details in the book of news shares:

  • Improved root cause analysis and traceability with SQL Dynamic lineage (now generally available) and fine-grained lineage (in preview) on Power BI datasets. You can do thorough root cause analysis from a single location in Microsoft Purview.
  • Metamodels that will enable customers to define organization, departments, data domains and business processes on their technical data. This feature is in preview.
  • then the Machine learning-based classifications will make detection of human names and addresses simple and scalable in user data. This feature is in preview.

These features with help with big data management, adding context on top of  data, improved classification with AI, manual process lineage and scorecard insights.  You can read about it below

Add business context to your hybrid data estate with Microsoft Purview

https://techcommunity.microsoft.com/t5/security-compliance-and-identity/add-business-context-to-your-hybrid-data-estate-with-microsoft/ba-p/3651989

Customize retention and deletion to help meet your specific business requirements

https://techcommunity.microsoft.com/t5/security-compliance-and-identity/customize-retention-and-deletion-to-help-meet-your-specific/ba-p/3613111

Read the article to learn more about the announcements in the Data Lifecycle and Records Management space that help organizations manage the lifecycle of data.

What's New in Microsoft Purview Compliance Manager

https://techcommunity.microsoft.com/t5/security-compliance-and-identity/what-s-new-in-microsoft-purview-compliance-manager/ba-p/3643375

Compliance manager helps

  • Eliminating blind spots with the right set of security, compliance, and privacy controls
  • Safeguarding critical data from external and internal threats
  • Identifying risks and addressing regulatory compliance requirements

There has been additional automated controls for Microsoft Priva and App Governance announced.

Power BI

Power BI updates Do more with enterprise self-service business intelligence  https://powerbi.microsoft.com/en-us/blog/microsoft-ignite-2022-do-more-with-enterprise-self-service-business-intelligence/ 

with video summary https://www.youtube.com/watch?v=PxGVcFb-zv0&t=28s









Tuesday, 25 May 2021

Microsoft Build May 2021

The Microsoft Build conference is running 25-27 May 2021. It is the digital event to expand your skills. The aim is to innovate for the challenges of tomorrow.

Satya Nadella talked again about how the world will be transformed through tech intensity and the importance of the environment. Microsoft aim to be the platform for platform creators and are releasing 100+ new updates during Build. The next generation of apps will be proactive rather than reactive.  We are at a pivotal time, multi cloud, multi edge, people centred which enables us to address opportunities to empower us and empower the world. It was another inspiring keynote.


There was a raft of announcements of new technology and enhancements to applications.


The general availability of Azure Cosmos DB Serverless was announced along with other Azure Cosmos DB enhancements.


 

Azure SQL ledger capability adds tamper-evident capabilities to Azure SQL Databases, available in Preview. Azure SQL Ledger is for sensitive systems enabling rich analytics and is enterprise ready.

Pytorch Enterprise is on Azure 

Azure Database for PostgreSQL has various features


Azure Purview now supports Azure Database for MySQL and Azure Database for PostgreSQL as a source for metadata, classification and lineage extraction






You can read about all the announcements in Harness the power of data and AI in your applications with Azure

Power BI has announced various features which can be read here Posts categorized: Announcements

My favourite announcement is Power BI in Jupyter notebooks . The new package lets you embed Power BI reports, dashboards, dashboard tiles, report visuals or Q&A in Jupyter notebooks easily.

The Microsoft Build 2021 Book of News covers all MS Build announcements. 




Thursday, 6 May 2021

Innovate Today with Azure SQL

Microsoft organised a digital event to explain how to build an effective cloud database management strategy that responds to today’s changing business requirements and tomorrow’s opportunities. The event built on the premise that the first adopted step has been a straightforward “lift and shift” to virtual machines (VMs). The Azure SQL digital event was on 4th May 2021: Innovate today with Azure SQL

Azure services power out-of-this-world solutions

A.I. Intelligent by default, AzureML and Cognitive Services

Hybrid Operational freedom, Azure Arc

Infrastructure Linux and Windows VMs

Data Choose the database that meets your workload’s needs

Apps App service, Azure Kubernetes services (AKS)

Tools Developer productivity, Azure DevOps

Fuelled by the best database for your workload

  • Azure PostgreSQL
  • Azure MySQL and MariaDB
  • Azure SQL Family
  • Azure Cosmos DB
  • Azure Cache for Redis



Wednesday, 7 November 2018

PASS Summit 2018 Keynote Day 1












The first keynote of PASS summit was delivered by Rohan Kumar entitled SQL Server and Azure Data Services: Harness the ultimate hybrid platform for data and AI





Customer priorities for a modernized data estate are: modernizing on-premises, modernizing to cloud, build cloud native apps and unlocking insights.






The announcements follow:

SQL Server 2019
SQL Server 2019 Public Preview  is a great way to celebrate the 25th anniversary of SQL Server

There is the introduction of big data clusters which combines Apache Spark and Hadoop into a single data platform called SQL Server. This combines the power of Spark with SQL Server over the relational and non-relation data sitting in SQL Server, HDFS and other systems like Oracle, Teradata, CosmosDB.

There are new capabilities around performance, availability and security for mission critical environments along with capability to leverage hardware innovations like persistent memory and enclaves.

Hadoop, ApacheSpark, Kubernetes and Java are native capabilities in the database engine.

Accelerated data recovery (ADR) was demonstrated and is incredible. It is at public preview.  The benefits of ADR are
  • Fast and consistent Database Recovery
  • Instantaneous Transaction rollback
  • Aggressive Log Truncation

Azure HDInsight 4.0

HDInsight 4.0 is now available in public preview.

There are several Apache Hadoop 3.0 innovations. Hive LLAP (Low Latency Analytical Processing known as Interactive Query in HDInsight) delivers ultra-fast SQL queries. The Performance metrics provide useful insight.

Integration with Power BI direct Query, Apache Zeppelin, and other tools. To learn more HDInsight Interactive Query with Power BI.

Data quality and GDPR compliance enabled by Apache Hive transactions
Improved ACID capabilities handle data quality (update/delete) issues at row level. This means that GDPR compliance requirements can now be meet with the ability to erase the data at row level. Spark can read and write to Hive ACID tables via Hive Warehouse Connector.

Apache Hive LLAP + Druid = single tool for multiple SQL use cases

Druid is a high-performance, column-oriented, distributed data store, which is well suited for user-facing analytic applications and real-time architectures. Druid is optimized for sub-second queries to slice-and-dice, drill down, search, filter, and aggregate event streams. Druid is commonly used to power interactive applications where sub-second performance with thousands of concurrent users are expected.

Hive Spark Integration
Apache Spark gets updatable tables and ACID transactions with Hive Warehouse Connector

There are several Apache Hadoop 3.0 innovations. Hive LLAP (Low Latency Analytical Processing called Interactive Query in HDInsight) for ultra-fast SQL queries. The Performance metrics provide useful insight.

Integration with Power BI Direct Query, Apache Zeppelin, and other tools. To learn more watch HDInsight Interactive Query with Power BI.

Better data quality and GDPR compliance enabled by Apache Hive transactions
Improved ACID capabilities handle data quality (update/delete) issues at row level. GDPR compliance requirements can now be meet with the ability to erase the data at row level. Spark can read and write to Hive ACID tables via Hive Warehouse Connector

Apache Hive LLAP + Druid = single tool for multiple SQL use cases

Druid is a high-performance, column-oriented, distributed data store, which is suited for user-facing analytic applications and real-time architectures. Druid is optimized for sub-second queries to slice-and-dice, drill down, search, filter, and aggregate event streams. Druid is commonly used to power interactive applications where sub-second performance with thousands of concurrent users are expected.

Hive Spark Integration
Apache Spark gets updatable tables and ACID transactions with Hive Warehouse Connector.



















Apache HBase and Apache Phoenix
Apache HBase 2.0 and Apache Phoenix 5.0 get new performance and stability features and all of the above have enterprise grade security.

Azure
Azure event hubs for Kafka is generally available
Azure Data Explorer is in public preview.

Azure Databricks Delta is in public preview
  • Connect data scientist and engineers
  • Prepare and clean data at massive scales
  • Build/train models with pre-configured ML

Azure Cosmos DB multi master replication was demoed with a drawing app, Azure Cosmos DB PxDraw
Azure SQL DB Managed Instances will be at General Availability (GA) on Dec 1st. This provides Availability Groups managed by Microsoft.

Power BI
















The new Dataflows is an enabler for self-service data prep in Power BI

Power BI Desktop November Update
  • Follow-up questions for Q&A explorerIt is possible to ask follow-up questions inside the Q&A explorer pop-up, which take into account the previous questions you asked.
  • Copy and paste between PBIX files
  • New modelling view makes it easier to work with large models.
  • Expand and collapse matrix row headers


Friday, 14 September 2018

Azure Cosmos DB multi-model database

Azure Cosmos DB has to be one of my favorite databases due to the breadth of available database types, its choice of consistency models and elastic scale out.

An introduction can be read here.

A definition for each of these types of databases is given.








Key-value
A key-value pair (KVP) is a set of two linked data items: a key, which is a unique identifier for some item of data, and the value, which is either the data that is identified or a pointer to the location of that data. Key-value pairs are frequently used in lookup tables, hash tables and configuration files.
https://searchenterprisedesktop.techtarget.com/definition/key-value-pair

Column
A column-oriented DBMS (or columnar database management system) is a database management system (DBMS) that stores data tables by column rather than by row.
https://en.wikipedia.org/wiki/Column-oriented_DBMS

Document
Document stores, also called document-oriented database systems, are characterized by their schema-free organization of data.That means records do not need to have a uniform structure, i.e. different records may have different columns. The types of the values ​​of individual columns can be different for each record. Columns can have more than one value (arrays). Records can have a nested structure. E.g. MongoDB
https://db-engines.com/en/article/Document+Stores

Graph
A graph database, also called a graph-oriented database, is a type of NoSQL database that uses graph theory to store, map and query relationships. Every node in a graph database is defined by a unique identifier, a set of outgoing edges and/or incoming edges and a set of properties expressed as key/value pairs.
https://whatis.techtarget.com/definition/graph-database


The five consistency levels offer predictable low latency guarantees and multiple well-defined relaxed consistency models.


Consistency Levels and guarantees

Consistency Level
Guarantees
Strong
Linearizability. Reads are guaranteed to return the most recent version of an item.
Bounded Staleness
Consistent Prefix. Reads lag behind writes by at most k prefixes or t interval
Session
Consistent Prefix. Monotonic reads, monotonic writes, read-your-writes, write-follows-reads
Consistent Prefix
Updates returned are some prefix of all the updates, with no gaps
Eventual
Out of order reads



There is a useful capacity planer that looks at request units throughput per second, request unit consumption and the amount of data storage needed by your application.


Monday, 7 May 2018

Microsoft Build Azure Cosmos DB




















Microsoft Build is underway sharing many useful features. The Azure Cosmos DB API is a versatile tool with a number of options. There are some quickstart tutorials and samples for these.

 Azure Cosmos DB now has multi-master write support. Multi-master in Azure Cosmos DB provides single-digit millisecond latency to write data and availability with built-in flexible conflict resolution support. There are some good examples in the article to help understand this functionality better.






















Azure Operational Data Services includes Azure SQL DB; PostgreSQL; MySQL; Redis Cache; and Cosmos DB.




Saturday, 5 May 2018

Azure CosmosDB Change Feed

The Azure CosmosDB change feed can provide a persistent log of records within an Azure CosmosDB container. You can learn about this from this concise presentation.







Tuesday, 10 April 2018

Leverage data for building























The leverage data to build intelligent apps presentation gives an insightful overview of the Microsoft Data Platform and how to innovate with analytics and AI. 

Tuesday, 3 April 2018

Cosmos DB SQL query cheat sheet

The new Azure Cosmos DB: SQL Query Cheat Sheet helps you write queries for SQL API data by displaying common database queries, keywords, built-in functions, and operators in an easy to print PDF reference sheet. Reference information for the MongoDB API, Table API, and Gremlin/Graph API are also included.





Tuesday, 13 March 2018

Azure Cosmos DB Data Explorer

A new tool is available to use. Data Explorer provides a rich and unified experience for inserting, querying, and managing Azure Cosmos DB data within the Azure portal. The data explorer brings together 3 tools, Document Explorer, Query Explorer, and Script Explorer.


Friday, 23 February 2018

Become an Azure Cosmos DB hero











There is an Azure Cosmos DB Technical Training Series available to sign up for. It has 7 parts covering a range of topics including a technical deep dive.

  • Technical overview of Azure Cosmos DB
  • Build real-time personalized experiences with AI and serverless technology
  • Using Graph API and Table API with Azure Cosmos DB 
  • Build or migrate your Mongo DB app to Azure Cosmos DB
  • Understanding Operations of Cosmos DB
  • Build Serverless Apps with Azure Cosmos DB and Azure Functions
  • Apply real-time analytics with Azure Cosmos DB and Spark

The training series has interactive Q&A throughout.

Monday, 22 January 2018

Migration from the Relational world to Graph
















I came across this useful blog SQL2Gremlin which translates the Northwind dataset. This was used as a sample database in older versions of SQL Server. The blog post explains the Apache TinkerPop's Gremlin graph traversal language using typical patterns found when querying data with SQL. The SQL examples make use of the T-SQL syntax.

This blog was helpful when looking at Azure Cosmos DB (Microsoft’s globally distributed multi-model database service). The  Gremlin console on the Azure portal is explained in the documentation, Azure Cosmos DB: create, query, and traverse a graph in the Gremlin console. The tutorial creates and queries vertices and edges, updates a vertex property, queries vertices, traverses the graph, and drops a vertex.
















The Gremlin console runs on Linux, Mac, and Windows. It can be downloaded from the Apache TinkerPop site.


Apache TinkerPop is a graph computing framework for both graph databases (OLTP) and graph analytic systems (OLAP).

Thursday, 2 November 2017

PASS Summit 2017 Day 2 Keynote






















I attended the Day 2 Keynote at PASS Summit presented by Rimma Nehme on Globally Distributed Databases Made Simple. This was an amazing presentation. It was presented seamlessly, explaining the technicalities of CosmosDB and how the globally distributed database works from the ground up.

Rimma raised the question, do we need another database? Databases need to meet the data needs for today and the future. Data is global, with large volumes of data being created every 60 seconds, which are continually growing and data is interconnected. The balance is shifting in the type of data and we need to have data globally next to users for processing, meaning the architecture needs to be different. 



CosmosDB was originally call Project Florence and was named as such because it is the place where the renaissance began. It was built in the cloud database for global distribution, with a fully resource governed stack and schema agnostic service. A single system image is used for all globally distributed resources. 

The resource model may have a database account / database that may span clusters and regions. The database is scaled out in terms of containers. It is designed to scale throughput and storage independently. There are two parts to the design. The physical system design is:






















The partitioning system design is:


The design is to enable elastically scalable storage, throughput, anywhere, anytime.

Resource governance cannot be an afterthought. The request unit/sec (RU) is the normalized currency.

There are 5 well-defined consistency models in Azure Cosmos DB  with clear trade offs: strong; bounded-stateless; sessions; consistent prefix and eventual.


There is native support for multiple data models with more coming in the future.







The talk continued to cover how indexing works in depth and the key points to remember about Cosmosdb are:


The talk concluded with a great quote “It is not the strongest of the species that survives, nor the most intelligent that survives. It is the one that is most adaptable to change.”

The slides can be downloaded.

Sunday, 1 October 2017

Microsoft for the Modern Data Estate

The Microsoft Ignite session on the modern data estate was full of announcements. All around us, data is driving digital transformation. Modernize with SQL Server 2017 on Linux and Windows and Azure Data Services; deliver modern intelligent applications using technologies like Azure Cosmos DB and Azure Database for PostgreSQL.


The world is changing.  We need to help invest in the future without being tied to the past. AI is a fundamental pillar to leverage that. If businesses invests in data they outperform other companies. 

Data doesn’t need to leave the database for data science to take place.
























The cloud first approach breeds faster innovation and SQL Server 2017 is proof of that.


Announcements
SQL Server 2017 on Linux, Docker , and Windows server
Supports for graph data and queries
Advances Machine Learning with R & Python
Native T-SQL scoring
Adaptive Query processing and Automatic Plan Correction
Vulnerability assessment for GDPR - preview
Intelligent insights into performance – preview
Support for Graph data and queries –GA
Adaptive query processing – GA
Native scoring and support for Azure Machine Learning - GA

SQL Database and Database Migration Service
Migration to the cloud is easy with this new service. 
Azure SQL Database is the intelligent cloud database for app developers. It learns and adapts, scales on the fly, enables multi-tenant SaaS apps, works in your environment, secures and protects. The systems of intelligence on SQL Database is shared.


Globally Distributed Applications

Announcing Azure functions for Azure Cosmos DB to build apps faster with a serverless infrastructure. 

Uncovering insights with big data and advanced analytics

The new Azure Data Factory allows easy modelling of diverse data integration scenarios. You can now with the preview service, easily move your SQL Server Integration Services (SSIS) workloads to cloud. There is also a data movements as a service with 30+ connectors. Azure SQL Data Warehouse has a compute- optimized tier and unlimited columnar storage in preview. The last announcement was the Power BI Report Server.



Saturday, 1 July 2017

Journey to the cloud

I have been reading a very interesting free ebook Enterprise Cloud Strategy by Barry Briggs and Eduardo Kassner. 

There are various types of modernisation discussed. 


It is based on Based on “Gartner Identifies Five Ways to Migrate Applications to the Cloud”, Gartner Inc., 2011. 

The overview of cloud migration principles discussed are summarised:























There is new a model that truly adopts cloud taking into account SaaS, PaaS and IaaS. This new architecture principle means considering the best placement of workloads, always looking at SaaS first.


This new way of thinking follows this placement decision tree which helps you make the correct decision on whether to use SaaS, PaaS, IaaS or private cloud.

On the data side, the paper talks about dividing your data into several categories related to risk. This is really about creating your own business data DNA. 'Data analytics' and 'BI and analytics'  cloud architectural blueprints are shared.

An interesting change for roles is the evolution of roles for the cloud. Microsoft suggest the data roles of Data Scientist and Information Architect emerge from the DBA and Statistician.