The hierarchy of needs for data science could help you be more effective with AI and machine learning.
Chaos, complexity, curiosity and database systems. A place where research meets industry
Welcome
Passionately curious about Data, Databases and Systems Complexity. Data is ubiquitous, the database universe is dichotomous (structured and unstructured), expanding and complex. Find my Database Research at SQLToolkit.co.uk . Microsoft Data Platform MVP
"The important thing is not to stop questioning. Curiosity has its own reason for existing" Einstein
"The important thing is not to stop questioning. Curiosity has its own reason for existing" Einstein
Saturday, 27 January 2018
Data Scientist Skills
There are an avalanche of skills required to become a data scientist. I came across this useful diagram.
The hierarchy of needs for data science could help you be more effective with AI and machine learning.
The hierarchy of needs for data science could help you be more effective with AI and machine learning.
Monday, 22 January 2018
Migration from the Relational world to Graph
I came across this useful blog SQL2Gremlin which translates the Northwind dataset. This was used as a sample database in older versions of SQL Server. The blog post explains the Apache TinkerPop's Gremlin graph traversal language using typical patterns found when querying data with SQL. The SQL examples make use of the T-SQL syntax.
This blog was helpful when looking at Azure Cosmos DB (Microsoft’s globally distributed multi-model database service). The Gremlin console on the Azure portal is explained in the documentation, Azure Cosmos DB: create, query, and traverse a graph in the Gremlin console. The tutorial creates and queries vertices and edges, updates a vertex property, queries vertices, traverses the graph, and drops a vertex.

The Gremlin console runs on Linux, Mac, and Windows. It can be downloaded from the Apache TinkerPop site.
Apache TinkerPop is a graph computing framework for both graph databases (OLTP) and graph analytic systems (OLAP).
Wednesday, 17 January 2018
Field Guide to Data Science
I read this really useful guide about data science. The Field Guide to Data Science was created to help organizations of all types and missions understand how to make use of data as a resource
More details about understanding the DNA of data can be found here.
More details about understanding the DNA of data can be found here.
Saturday, 30 December 2017
Diagrams help explain complex data
Diagrams are useful to help explain qualitative data. There are various diagramming tools that can help with understanding complex systems. These are the main tools that I use
The diagrams above were provided by the Open University and guidelines for constructing such diagrams are explained.
The Open University guide to using diagrams can be seen in this video.
- Rich Pictures
- Spray Diagrams
- Systems Map
- Influence Diagram
- Multiple Cause Diagram
The diagrams above were provided by the Open University and guidelines for constructing such diagrams are explained.
The Open University guide to using diagrams can be seen in this video.
Wednesday, 20 December 2017
Continual Change and Complexity
This year has been an entire year of change for me, that will continue into the new year. Continual change is the way of the new world. With data and AI being embedded into every realm of technology, we can expect more frequent and smaller changes on a day to day basis. I have enjoyed researching immensely and being able to apply that research to understanding the complexity of real world database problems.
As the holidays approach I wish you all a very Merry Christmas and Happy a New Year.
As the holidays approach I wish you all a very Merry Christmas and Happy a New Year.
Thursday, 14 December 2017
The DevOps Model
TechNet UK had a live stream of talks back in September and this was a diagram that they shared. I think it is a helpful picture describing the process.
Friday, 17 November 2017
Big Data LDN 2017
I
attended Big Data London 15-16 Nov 2017 with leading data &
analytics experts showcasing their tools to help with delivering data-driven
strategy. The conference showcased the fourth industrial revolution report which
explains what the UK’s data leaders think about the state of the UK data
economy.
A
summary of things I found interesting during the two day event are summarized
here.
Machine
Learning is such a topical discussion point, but it is not that difficult to get
started. An area to initially look at is co-occurrence and recommendation. Co-occurrence helps you find behaviours and you
can use that to find recommendations in areas such as textual analysis and
intrusion detection.
Further reading: Chapter 4. Co-occurrence and Recommendation
Machine learning
was described as the integration between analytics and operations. The three
questions to ask were: what algorithm, what tools and what process. 90% of machine
learning success is in data logistics (being able to handle lots of data types),
not learning.
The CDO’s playbook was launched. The Chief Data
Officer is a rapidly expanding role and this book offers practical advice on
what this role is, how it fits into to other c-suite roles and provides
actionable tips.
There are many challenges
when dealing with citizen data. At the heart of audiences is
- single view
of the customer
- deeper engagement
- supported intelligence
- relationship
management
The main
challenge is data quality and having a high enough quality of data to provide
insight.
Citizens want
to be data scientists and be able to dive into the data with ease. This self-service
model can have challenges. Better governance, data management and operational efficiency
are required together with the rise of managed service to remove the complexities
of running these services.
The keynote on
day 2, machine learning, AI and the future of big data analytics by Dr Amr
Awadallah, Co-founder of Cloudera, talked about a history of waves.
- wave 1 automation of
knowledge transfer
- wave 2 automation of
food
- wave 3 automation of
discovery
- wave 4 making and
moving stuff (Industrial revolution)
- wave 5 automation of
processes (IT revolution)
- wave 6 automation of
decisions.
We are in wave
6 which is about collecting data and leveraging data to make decisions. It is
different from the BI wave where humans made decisions. The new wave is learning
how decisions are made and automating them. Things to
consider for success are
- build a
data driven culture
- develop the
right team and skills
- be agile/lean
in development
- leverage DevOps
for production
- right size
data governance
There were discussions about data narrative and telling a story to the audience. The five steps
learnt for better storytelling
- identify the
right data
- choose the
right visualizations
- calibrate
visuals to your message
- remove unnecessary
noise
- focus attention
on what’s important
Matt Aslett
talked on pervasive intelligence: the future of big data, machine learning and
IoT, the details of which have been published in a report. He discussed trends and implications of the AI automation spectrum. It will bring about fundamental and wide ranging
positive societal implication that will change the way we live, work, play,
transact and travel. He mentioned a risk of having a small number of platform
oriented companies that control the forces of production for generating value
from data. The 4sight report on the future of IT is coming soon and sounds an interesting
read.
Deep learning
demystified explained why neural networks, that are not new, have only just
come to the fore. It was because they were originally thought of as part of a
failed experiment. In fact, it was that they did not use enough data. For
supervised learning it works well with very large data sets. The key things to
think of when considering deep learning are that it
- must have
large data, a minimum of 10 million labels of data
- what level
of accuracy do you need?
- can something
simple work? – start with classical models such as linear models
There is a deep learning
institute to learn more.
The conference was useful and provided a wide range of discussions on high level data topics.
Subscribe to:
Posts (Atom)






