Welcome

Passionately curious about Data, Databases and Systems Complexity. Data is ubiquitous, the database universe is dichotomous (structured and unstructured), expanding and complex. Find my Database Research at SQLToolkit.co.uk . Microsoft Data Platform MVP

"The important thing is not to stop questioning. Curiosity has its own reason for existing" Einstein



Showing posts with label Predictive Analytics. Show all posts
Showing posts with label Predictive Analytics. Show all posts

Saturday, 7 July 2018

The Future State - Serendipitous Data Management

Gone are the days where companies can survive on existing products and services. The need to continually innovate to stay ahead in a fluid world, requires a change in direction. Many articles have been written, in both academic research and Industry, to try to predict what will be the future state of data technology and what will be this year's trends.

Currently research meets industry in a rebirth of industry-based research teams consisting of organisational only teams or industry collaborating with universities. Guzdial shares his thoughts in the Communicationsof ACM March 2018 journal that "for the majority of new computer science PhD's, the research environment in industry is currently more attractive". Particularly the need within industry to continually innovate cries out for more research divisions in industry. Part of this change is due to the rapid expansion of emerging technology but also the realization, of what data science and artificial intelligence (AI) can add to a business. Data science requires collaboration between people, teams and organisations as interdisciplinary skills are needed to solve today’s problems.

There is an emerging trend whereby more research institutes have been created or existing ones hiring more staff. Microsoft have created a new organization, Microsoft Research AI (MSR AI), to pursue game-changing advances in artificial intelligence. The research team combines advances in machine learning with innovations in language and dialog, human computer interaction, and computer vision to solve some of the toughest challenges in AI.

AI machine learning intelligence, based on big data, is a complex problem to solve, to empower people for the future. In the current world there is the need for collaboration. Greengard in the Communications of ACM March 2018 journal, raised a concern that "mountains of data produce incremental gains, and coordinating all the research groups and silos is a complex endeavour".  Managing data is complex and the key areas that I think will define the next revolution are in the graph.


Telling stories from the data is increasingly important in this ever-changing holistic environment. Skills need to be developed in this area as communicating the meaning of data is crucial. Aiming for improvement in business, science, robotics, space and health can initially appear through intelligent automation and can produce further actionable insights. 

Data visualization is a key component to telling the story and seeing anomalies. Parameswaran discussed at SIGMOD 2018, that it is the scale that brings databases and visualisation together. He highlighted two problem areas, too many tuples and too many visualisations. It is an interesting point to consider how to address the excessive data points and how to appropriately find the right visualization for the data, to gain insight at speed.  

Innovation is key to the next step. I believe that is by making beneficial discoveries by design through scientific experiments from quality data in a continuous and autonomous fashion. I call this Serendipitous Data Management. This improvement and innovation will come from having sound practices for big data management that enable actionable data insights at speed.

Another trend I am seeing in research and industry is looking at how data is processed in centralised data lakes and moving that processing to the edge, particularly for IOT at the moment. As well as this increasing security, if the data can remain at source, it also reduces the volume of data transit which is currently unsustainable. How to consolidate these distributed data sources and produce analysis across disparate systems is an interesting challenge to solve. In conclusion the system built on data creates a rapidly changing landscape of which I see as the key components in defining revolutionary changes to society and culture. 

Tuesday, 27 March 2018

Machine Learning


Predictive analytics uses various statistical techniques such as, machine learning to analyze collected data for patterns or trends to forecast future events. Machine learning uses predictive models that learn from existing data to forecast future behaviors, outcomes, and trends.

Machine Learning libraries enable data scientists to use dozens of algorithms, each with their strengths and weaknesses. Download the machine learning algorithm cheat sheet to help identify how to choose a machine learning algorithm.















Thursday, 1 March 2018

An Introduction to HDInsight

 
I attended a great session at SQLBits 2018 covering the basics of HDInsight by Edinson Medina. He introduced his talk explaining the term Big Data and that it is too complex for analysis in traditional databases. There are 2 types of processing batch processing, to shape the data for analysis and real time processing to capture streams of data for low latency querying.

Hadoop is described on the Hortonworks site as "Apache Hadoop is an open source software platform for distributed storage and distributed processing of very large data sets on computer clusters built from commodity hardware."

A Hadoop cluster looks like



















The underlying structure uses map reduce. The Tez engine is a newer faster engine for map reduce. The model is explained in the paper on "Analyzing performance of Apache Tez and MapReduce with hadoop multinode cluster on Amazon cloud"



HDInsight is 100% Apache Hadoop, but powered by the cloud.



There are many tools within the Hadoop ecosystem.

Hive
A meta data service that projects tabular schemas over folders and enables the folders to be queried as tables using a SQL like query.

Pig (an ETL Tool)
Performs a series of transformations to data relations based on Pig Latin statements.

OoZie
A workflow engine for actions in a Hadoop cluster supporting parallel work streams.

Scoop
A database integration service which enables bi-directional data transfer between an Hadoop cluster and databases via JDBC.

HBase
A low latency NoSQL database built on Hadoop modeled on Googles's BigTable. HBase stores files on HDFS.

Storm
An event processor for data streams such as real time monitoring and for event aggregation and logging. It defines a streaming topology that consists of spouts and bolts.

Spark
A fast general purpose computation engine that supports in memory operations. It is a unified stack for interactive, streaming and predictive analysis.

Ambari
A management platform for provisioning, managing, monitoring and securing Apache Hadoop clusters.

Zepplin notebooks 
A multi-purposed web-based notebook which brings data ingestion, data exploration, visualization, sharing and collaboration features to Hadoop and Spark.

Friday, 17 November 2017

Big Data LDN 2017


















I attended Big Data London 15-16 Nov 2017 with leading data & analytics experts showcasing their tools to help with delivering data-driven strategy. The conference showcased the fourth industrial revolution report which explains what the UK’s data leaders think about the state of the UK data economy.

A summary of things I found interesting during the two day event are summarized here.

Machine Learning is such a topical discussion point, but it is not that difficult to get started. An area to initially look at is co-occurrence and recommendation.  Co-occurrence helps you find behaviours and you can use that to find recommendations in areas such as textual analysis and intrusion detection.


Machine learning was described as the integration between analytics and operations. The three questions to ask were: what algorithm, what tools and what process. 90% of machine learning success is in data logistics (being able to handle lots of data types), not learning.

The CDO’s playbook was launched. The Chief Data Officer is a rapidly expanding role and this book offers practical advice on what this role is, how it fits into to other c-suite roles and provides actionable tips.

There are many challenges when dealing with citizen data. At the heart of audiences is

- single view of the customer
- deeper engagement
- supported intelligence
- relationship management

The main challenge is data quality and having a high enough quality of data to provide insight.

Citizens want to be data scientists and be able to dive into the data with ease. This self-service model can have challenges. Better governance, data management and operational efficiency are required together with the rise of managed service to remove the complexities of running these services.

The keynote on day 2, machine learning, AI and the future of big data analytics by Dr Amr 
Awadallah, Co-founder of Cloudera, talked about a history of waves.

- wave 1 automation of knowledge transfer
- wave 2 automation of food
- wave 3 automation of discovery
- wave 4 making and moving stuff (Industrial revolution)
- wave 5 automation of processes (IT revolution)
- wave 6 automation of decisions.

We are in wave 6 which is about collecting data and leveraging data to make decisions. It is different from the BI wave where humans made decisions. The new wave is learning how decisions are made and automating them. Things to consider for success are

- build a data driven culture
- develop the right team and skills
- be agile/lean in development
- leverage DevOps for production
- right size data governance 

There were discussions about data narrative and telling a story to the audience. The five steps learnt for better storytelling

- identify the right data
- choose the right visualizations
- calibrate visuals to your message
- remove unnecessary noise
- focus attention on what’s important

Matt Aslett talked on pervasive intelligence: the future of big data, machine learning and IoT, the details of which have been published in a report. He discussed trends and implications of the AI automation spectrum. It will bring about fundamental and wide ranging positive societal implication that will change the way we live, work, play, transact and travel. He mentioned a risk of having a small number of platform oriented companies that control the forces of production for generating value from data. The 4sight report on the future of IT is coming soon and sounds an interesting read.

Deep learning demystified explained why neural networks, that are not new, have only just come to the fore. It was because they were originally thought of as part of a failed experiment. In fact, it was that they did not use enough data. For supervised learning it works well with very large data sets. The key things to think of when considering deep learning are that it
- must have large data, a minimum of 10 million labels of data
- what level of accuracy do you need?
- can something simple work? – start with classical models such as linear models

There is a deep learning institute to learn more. 

The conference was useful and provided a wide range of discussions on high level data topics.

Thursday, 19 October 2017

Machines that learn to see and move: The future of artificial intelligence

















I attended the Institute for Mathematical Innovation (IMI) public lecture by Professor Andrew Blake, Research Director at The Alan Turing Institute on 18 October. Professor  Blake is a pioneer in the development of algorithms that make it possible for computers to behave as seeing machines. Before joining the Institute in 2015, Professor Blake held the position of Microsoft Distinguished Scientist and Laboratory Director at the Microsoft Research Lab in Cambridge, and he has been on the faculty at Oxford University. He is a part of a new startup FiveAI.

The session abstract:

Neural networks have taken the world of computing in general and artificial intelligence (AI) in particular by storm.

But in the future, AI will need to revisit these generative models which are used to make predictions. There are several reasons for this – system robustness, precision issues, transparency, and the high cost of labelling data.

This is particularly true for perceptual AI, needed for autonomous vehicles, where the need for simulators and the need to confront novel situations, will demand further development of generative, probabilistic models. 

He talked about the empirical detector and generative model. At the moment it is the era of deep learning and neural networks, that sit within the empirical detector area. A black box area of big data and optimal predictive power. The generative model is analysis by synthesis and comes with an ‘explanation’, like a model. It starts with a hypothesis, typically probabilistic. Professor Blake believes the generative model will come back as perceptual models need this. This is

  • to simulate labelled data
  • for data fusion - to increase reliability
  • to make detailed interpretations 
  • for online simulation - to explain hard to read situations
This was a very insightful lecture and very interesting to see the mention of analysis by synthesis. 

Friday, 6 October 2017

Machina Summit.AI













I attended IPExpo Europe 4-5 October in ExCel in London with the specific attendance at the Machina Summit.AI.

The opening keynote was by Professor Brian Cox OBE on ‘Where IT & Physics Collide’.  The talk interlinked big data, quantum mechanics and quantum computing. The whistle top tour mentioned the Sloan Digital Sky Survey, which are the most detailed three-dimensional maps of the universe; general relativity; history of space and time; the theory of cosmology; and quantum mechanics ending with quantum theory and predicting the distribution of galaxies. This was an amazing talk and gave a glimpse of the interconnected future.

This was followed by Brad Anderson, Corporate Vice President of Microsoft on ‘Business as usual in a digital war zone’. We live in turbulent times with a 300% increase in user account attacks this year, 96% of malware is automated polymorphic which costs business $15 million. Attacks happen in increasing waves and old defences never stand up against these attacks. In this intelligent war you need an intelligent graph. He introduced the Microsoft Azure Active Directory service as the new control plane. There is the need to eliminate false positives, classify email and guarantee data never leaves the browser and be able to use a real time evaluation engine.

A few other talks covered the practice of monitoring with machine data. There are 2 types of monitoring, transitional IT and the new data driven IT. For the latter there is the need to rethink and improve how IT operates using machine learning to be proactive. Organizational silos and increasing quality are things that need to be broken down to be able to address the velocity data in a more agile way to produce actionable insights.

Conrad Wolfram, Strategic Director, Wolfram Research talked about ‘Enterprise computation: the next frontier in AI and data science’ Todays data challenge is about accessibility of data, personalisation of data and providing insightful answers. Data Science is multi paradigm and machine learning does not have all the answers. Computation is required for everyone with smart automation and computational thinking is needed for everyone. Data science needs to be personalised, multifaceted but unified.

The day 2 keynote was given by Stuart Russell, Professor of Electrical Engineering and Computer Science, University California Berkeley on ‘Human-Compatible AI’. He discussed what is coming soon. Basic language understanding with web-scale question answering and intelligent assistants for health, education, finances and life (not chatbots!!). Robots for unstructured tasks (home, construction, agriculture) and new tools for economics, management and scientific research. He discussed the premise that eventually AI systems will make better* decisions than humans. Well *taking into account more information and looking further into the future. He argued that for the case of super intelligent AI, that you can’t switch off the machine and AI will never succeed.

Other sessions discussed the journey of chaos and how everything fails all the time. To address this there is the need to consider that every journey begins with a single step. There is the inevitable question to consider skills versus knowledge and that is practice.

Microsoft talked about their 'AI and Analytics in the Enterprise'. There is now a need to look at more than the rear view mirror, to see what happened. There is a convergence of cloud, data and AI. With that Microsoft have created an AI platform that is fast and agile, with AI built in and enterprise proven for on-premises to edge to create insights. The evolution of the data state takes into account increasing data volumes, new data sources and types and open source languages. There are 3 stages between the heterogeneous sources and providing apps and insights.
  • Ingest – data orchestration and monitoring
  • Store – Data Lake and storages
  • Machine learning – preparations and train ( Hadoop / spark / SQL and ML) then model and serve (on-prem, Cloud, IoT).

In summary the 2 day conference provided great insight into many new technical areas and raised thought provoking questions about the future of data and AI. 

Wednesday, 19 April 2017

Microsoft DataAmp – SQL Server 2017

The DataAmp webcast was packed full of announcements. The Webcast was delivered by Scott Guthrie and Joseph Sirosh. The SQL Server product delivering intelligence, trust and flexibility.

Microsoft confirmed that the next version of SQL Server is SQL Server 2017 and will be available simultaneously on Windows, Linux and Docker. Download the SQL Server 2017 datasheet.  It will be the first RDBMS to deliver AI with data. There is a convergence of cloud, data and intelligence. Delivering AI with data: the next generation of Microsoft’s data platform blog shares more information.  SQL Server 2017 Community Technology Preview 2.0 now available.

There were so many new features announced only a few are mentioned below.  There are adaptive query processing improvements which will enhance the performance of workloads. There is a You Tube video SQL Server 2017: Adaptive Query Processing discussing this. Threat detection is now in Azure SQL Database and is straight forward to configure. 



Hybrid Cloud just got easier to adopt with the new Azure migration resources and tools. To help with SQL Server migrations to the cloud features such as Service Broker, SQLAgent, Profiler etc. are now available in Azure. There is a new data migration service for automatic migration for SQL Server, Oracle and MySQL in Azure .

SQL Graph
Storing and analyzing graph data relationships. This includes full CRUD support to create nodes and edges and T-SQL query language extensions to provide multi-hop navigation using join-free pattern matching.  The SQL Server engine integration enables querying across SQL tables and graph data.

SQL Server on Linux
 

The official Microsoft repository for SQL Server in Docker containers is here.

Here are a few videos to help get you started with Linux:

Analytics 

SQL Server is the first commercial database to include Deep Learning algorithms, with the announcement of the Microsoft Cognitive Services general availability of the FACE API and Computer Vision API.
Azure Data Lake Services now have petabyte scale.
Azure Analysis Services became generally available. You Tube video: SQL Server 2017: BI enhancements.  
You can use Python for advanced analytics, You Tube video: SQL Server 2017: Advanced Analytics with Python


Azure DocumentDB

Azure DocumentDB is globally distrubuted and offer limitless scale of throughput and storage. It can be used for things such as IoT applications that need low response times and need to handle massive amounts of reads and writes.



Cortana Intelligence Solution Templates
You can now quickly build Cortana Intelligence Solutions from preconfigured solutions, reference architectures and design patterns. Some are released with more to follow.

The really important thing that I am excited about is the flexibility of choice within the SQL Server product.




Joseph Sirosh concluded comparing the industrial revolution with the intelligence revolution of today.