Welcome

Passionately curious about Data, Databases and Systems Complexity. Data is ubiquitous, the database universe is dichotomous (structured and unstructured), expanding and complex. Find my Database Research at SQLToolkit.co.uk . Microsoft Data Platform MVP

"The important thing is not to stop questioning. Curiosity has its own reason for existing" Einstein



Showing posts with label Deep Learning. Show all posts
Showing posts with label Deep Learning. Show all posts

Saturday, 2 March 2019

SQL Bits 2019 Keynote






















What an amazing SQLBits in Manchester. Four days packed full of leading edge data technology covering

  • SQL Server 2019 Big Data
  • Azure SQL Managed Database
  • Power BI
  • Kubernetes
  • Machine Learning
  • Python
  • Spark

This year SQLBits 2019 had a keynote.  It was nice for the event to have a keynote again. The theme Data Never Rests.  The Microsoft Data Platform Product group who spoke were Buck Woody, Bob Ward, Anna Thomas, Alain Dormehl, Adam Saxton and Patrick LeBlanc. An amazing set of speaks and fun keynote. They shared details of the evolution of the data platform to enable people to keep their skills up to date. The keynote is available to watch . There were several major announcements.


SQL Server 2019 will RTM in second half of the year. SQL Server 2019 CTP2.3 is available now with
  • Big data cluster enhancements
  • Accelerated database recovery
  • Performance enhancements
  • Graph data enhancements
  • SSAS enhancements

SQL Server 2019 is a modern innovation and there are various forms of the product.
  • On Premises
  • SQL Server Azure VM (IaaS)
  • Azure SQL DB Managed Instance (PaaS)
  • Azure SQL Data Warehouse


Azure SQL Database Hyperscale can autoscale up to 100TB and scale compute and storage independently.
During the keynote they showed Azure SQL Database Hyperscale where a 50TB database was restored in just under 8 minutes. That is nice accelerated database recovery.

Data virtualization and big data clusters is a game changing view with SQL Server 2019 big data clusters, data lake scale, machine learning and AI. Multiple data sources can be connected using external table, through the compute pool using Polybase connectors at the source.  Data persistence using multiple data sources is stored in shards of the data pool for SQL Server 2019 big data clusters data mart.



SQL 2019 will send push down predicated queries to other data platforms via Polybase to join SQL data with Oracle, Mongodb and CosmosDB data in one place efficiently.

SQL notebooks in Azure Data Studio is an awesome new feature. 

There is documentation to read and new courses for learning.

aka.ms/DataAccessGuide

and a Summary of All Exams and Certifications Launched in January, 2019!

aka.ms/DataEngCerts







Thursday, 13 September 2018

AI the art of the possible

During SQL Saturday Cambridge I attended a session by Terry McCann on using AI to write a session submission to SQL Saturday. It was a great session and I would recommend you attending it if you get a chance.

In the new data world it is important to understand the difference between AI, Machine Learning and Deep Learning.


























Then breaking this down further the differences between how machine learning works and deep learning is shown here. Deep learning is really a black box.




















Image : https://www.upwork.com/hiring/for-clients/log-analytics-deep-learning-machine-learning/

Terry mentioned a free book to read to learn more Neural Networks and Deep Learning. 

He mentioned also the book Harry Potter and the Portrait of what Looked Like a Large Pile of Ash which was written by an AI bot.  I hadn't come across this before but it lets you see the art of the possible in the future.

Image: https://imgur.com/gallery/gkLFz


Saturday, 7 July 2018

The Future State - Serendipitous Data Management

Gone are the days where companies can survive on existing products and services. The need to continually innovate to stay ahead in a fluid world, requires a change in direction. Many articles have been written, in both academic research and Industry, to try to predict what will be the future state of data technology and what will be this year's trends.

Currently research meets industry in a rebirth of industry-based research teams consisting of organisational only teams or industry collaborating with universities. Guzdial shares his thoughts in the Communicationsof ACM March 2018 journal that "for the majority of new computer science PhD's, the research environment in industry is currently more attractive". Particularly the need within industry to continually innovate cries out for more research divisions in industry. Part of this change is due to the rapid expansion of emerging technology but also the realization, of what data science and artificial intelligence (AI) can add to a business. Data science requires collaboration between people, teams and organisations as interdisciplinary skills are needed to solve today’s problems.

There is an emerging trend whereby more research institutes have been created or existing ones hiring more staff. Microsoft have created a new organization, Microsoft Research AI (MSR AI), to pursue game-changing advances in artificial intelligence. The research team combines advances in machine learning with innovations in language and dialog, human computer interaction, and computer vision to solve some of the toughest challenges in AI.

AI machine learning intelligence, based on big data, is a complex problem to solve, to empower people for the future. In the current world there is the need for collaboration. Greengard in the Communications of ACM March 2018 journal, raised a concern that "mountains of data produce incremental gains, and coordinating all the research groups and silos is a complex endeavour".  Managing data is complex and the key areas that I think will define the next revolution are in the graph.


Telling stories from the data is increasingly important in this ever-changing holistic environment. Skills need to be developed in this area as communicating the meaning of data is crucial. Aiming for improvement in business, science, robotics, space and health can initially appear through intelligent automation and can produce further actionable insights. 

Data visualization is a key component to telling the story and seeing anomalies. Parameswaran discussed at SIGMOD 2018, that it is the scale that brings databases and visualisation together. He highlighted two problem areas, too many tuples and too many visualisations. It is an interesting point to consider how to address the excessive data points and how to appropriately find the right visualization for the data, to gain insight at speed.  

Innovation is key to the next step. I believe that is by making beneficial discoveries by design through scientific experiments from quality data in a continuous and autonomous fashion. I call this Serendipitous Data Management. This improvement and innovation will come from having sound practices for big data management that enable actionable data insights at speed.

Another trend I am seeing in research and industry is looking at how data is processed in centralised data lakes and moving that processing to the edge, particularly for IOT at the moment. As well as this increasing security, if the data can remain at source, it also reduces the volume of data transit which is currently unsustainable. How to consolidate these distributed data sources and produce analysis across disparate systems is an interesting challenge to solve. In conclusion the system built on data creates a rapidly changing landscape of which I see as the key components in defining revolutionary changes to society and culture. 

Wednesday, 19 April 2017

Microsoft DataAmp – SQL Server 2017

The DataAmp webcast was packed full of announcements. The Webcast was delivered by Scott Guthrie and Joseph Sirosh. The SQL Server product delivering intelligence, trust and flexibility.

Microsoft confirmed that the next version of SQL Server is SQL Server 2017 and will be available simultaneously on Windows, Linux and Docker. Download the SQL Server 2017 datasheet.  It will be the first RDBMS to deliver AI with data. There is a convergence of cloud, data and intelligence. Delivering AI with data: the next generation of Microsoft’s data platform blog shares more information.  SQL Server 2017 Community Technology Preview 2.0 now available.

There were so many new features announced only a few are mentioned below.  There are adaptive query processing improvements which will enhance the performance of workloads. There is a You Tube video SQL Server 2017: Adaptive Query Processing discussing this. Threat detection is now in Azure SQL Database and is straight forward to configure. 



Hybrid Cloud just got easier to adopt with the new Azure migration resources and tools. To help with SQL Server migrations to the cloud features such as Service Broker, SQLAgent, Profiler etc. are now available in Azure. There is a new data migration service for automatic migration for SQL Server, Oracle and MySQL in Azure .

SQL Graph
Storing and analyzing graph data relationships. This includes full CRUD support to create nodes and edges and T-SQL query language extensions to provide multi-hop navigation using join-free pattern matching.  The SQL Server engine integration enables querying across SQL tables and graph data.

SQL Server on Linux
 

The official Microsoft repository for SQL Server in Docker containers is here.

Here are a few videos to help get you started with Linux:

Analytics 

SQL Server is the first commercial database to include Deep Learning algorithms, with the announcement of the Microsoft Cognitive Services general availability of the FACE API and Computer Vision API.
Azure Data Lake Services now have petabyte scale.
Azure Analysis Services became generally available. You Tube video: SQL Server 2017: BI enhancements.  
You can use Python for advanced analytics, You Tube video: SQL Server 2017: Advanced Analytics with Python


Azure DocumentDB

Azure DocumentDB is globally distrubuted and offer limitless scale of throughput and storage. It can be used for things such as IoT applications that need low response times and need to handle massive amounts of reads and writes.



Cortana Intelligence Solution Templates
You can now quickly build Cortana Intelligence Solutions from preconfigured solutions, reference architectures and design patterns. Some are released with more to follow.

The really important thing that I am excited about is the flexibility of choice within the SQL Server product.




Joseph Sirosh concluded comparing the industrial revolution with the intelligence revolution of today.