Welcome

Passionately curious about Data, Databases and Systems Complexity. Data is ubiquitous, the database universe is dichotomous (structured and unstructured), expanding and complex. Find my Database Research at SQLToolkit.co.uk . Microsoft Data Platform MVP

"The important thing is not to stop questioning. Curiosity has its own reason for existing" Einstein



Showing posts with label Graph Databases. Show all posts
Showing posts with label Graph Databases. Show all posts

Friday, 16 November 2018

Big Data LDN Day 2





















The Fourth Industrial Revolution Report – Download for Free

Keynote

The day 2 keynote was given by Michael Stonebraker, Turing Prize winner, IEEE John von Neumann Medal Holder, Professor at MIT, Co-founder of Tamr entitled Big Data, Disruption and the 800 pound gorilla in the corner.

A few vignettes were mentioned Hamiltons, Dewitts and Amadeus.  Hadoop (meaning map reduce) was started to be used in 2010 and Google stopped using it in 2011. Hadoop now means a HDFS file system.  Cloudera's big problem is that no one wants map reduce. Map reduce is not used for anything.

The data warehouse is yesterdays problem. BI is simple SQL. Data Science has complex problems and it is a different skill set. It is based on deep learning, machine learning and linear algebra, nothing to do with SQL. Deep learning is all the rage but you need vast amounts of training data.  It is not possible to explain why the black box gives certain recommendations so it is not good when data providence is required. 

Big velocity is a big problem over time. Pattern matching and CEP (Complex Event Processing)
like Storm is not competitive. Don’t run Oracle but instead run Mongodb, Cassandra or Redis. NoSQL means no standards and no ACID. ACID is a good idea. NoSQL means you always give up something as per CAP theorem. Declarative languages are a great idea.

Data discovery is a big problem. You spend 90% of your time finding and cleaning data. Then 10% finding and cleaning the errors. Very little time is spent doing data integration. It is a data integration challenge.

Graphs

Jim Webber from Neo4J gave an insightful talk about how useful graphs are to solve problems and predict outcomes. There was some great examples of how to use graphs. He talked about triad closure and strong and week ties. Also mentioning a couple of papers to read

Effects of organizational support on organizationalcommitment  Fakhraei M, Imami R, Manuchehri S (2015)
Semi-Supervised Classification with Graph ConvolutionalNetworks Thomas N. Kipf , Max Welling  (2017) 

and a free ebook

Free Book: Graph Databases By Ian Robinson, Jim Webber, and Emil Eifrém

It is important to have semantic domain knowledge for inference and understanding in graphs as graphs depend on the context. Graphml convolutional network graph will be the data structure for AI.

The Joy of Data
The closing session of the event was delivered by Dr Hannah Fry, Associate Professor in the mathematics of Cities – UCL. This was an amazing session exploring what visualization and insights can be achieved from understanding the data.

She started the talk with the strange Wikipedia phenomenon that all routes lead to philosophy. So by clicking the first proper link on a page you will eventually end up on the philosophy page. 

There are 2 parallel universes where people click the link and the mathematical universe.  Data is the bridge.

She showed how data could be used to investigate why the bicycle transport scheme in London was seeing all the bikes ending up in the wrong place. Vans have to go round moving bikes into the right places during the day. This was the result of people liking to cycle down the hills but not up.

Another example showed that Islington station was a bottleneck which caused a cascading problem because it has a lack of transport routes from there. There were many other interesting examples and how gossip can pay by using network science to track the problem down.

Big Data LDN had some amazing sessions and insightful content. Big Data LDN will be back next year 13-14 Nov 2019.

Saturday, 7 July 2018

The Future State - Serendipitous Data Management

Gone are the days where companies can survive on existing products and services. The need to continually innovate to stay ahead in a fluid world, requires a change in direction. Many articles have been written, in both academic research and Industry, to try to predict what will be the future state of data technology and what will be this year's trends.

Currently research meets industry in a rebirth of industry-based research teams consisting of organisational only teams or industry collaborating with universities. Guzdial shares his thoughts in the Communicationsof ACM March 2018 journal that "for the majority of new computer science PhD's, the research environment in industry is currently more attractive". Particularly the need within industry to continually innovate cries out for more research divisions in industry. Part of this change is due to the rapid expansion of emerging technology but also the realization, of what data science and artificial intelligence (AI) can add to a business. Data science requires collaboration between people, teams and organisations as interdisciplinary skills are needed to solve today’s problems.

There is an emerging trend whereby more research institutes have been created or existing ones hiring more staff. Microsoft have created a new organization, Microsoft Research AI (MSR AI), to pursue game-changing advances in artificial intelligence. The research team combines advances in machine learning with innovations in language and dialog, human computer interaction, and computer vision to solve some of the toughest challenges in AI.

AI machine learning intelligence, based on big data, is a complex problem to solve, to empower people for the future. In the current world there is the need for collaboration. Greengard in the Communications of ACM March 2018 journal, raised a concern that "mountains of data produce incremental gains, and coordinating all the research groups and silos is a complex endeavour".  Managing data is complex and the key areas that I think will define the next revolution are in the graph.


Telling stories from the data is increasingly important in this ever-changing holistic environment. Skills need to be developed in this area as communicating the meaning of data is crucial. Aiming for improvement in business, science, robotics, space and health can initially appear through intelligent automation and can produce further actionable insights. 

Data visualization is a key component to telling the story and seeing anomalies. Parameswaran discussed at SIGMOD 2018, that it is the scale that brings databases and visualisation together. He highlighted two problem areas, too many tuples and too many visualisations. It is an interesting point to consider how to address the excessive data points and how to appropriately find the right visualization for the data, to gain insight at speed.  

Innovation is key to the next step. I believe that is by making beneficial discoveries by design through scientific experiments from quality data in a continuous and autonomous fashion. I call this Serendipitous Data Management. This improvement and innovation will come from having sound practices for big data management that enable actionable data insights at speed.

Another trend I am seeing in research and industry is looking at how data is processed in centralised data lakes and moving that processing to the edge, particularly for IOT at the moment. As well as this increasing security, if the data can remain at source, it also reduces the volume of data transit which is currently unsustainable. How to consolidate these distributed data sources and produce analysis across disparate systems is an interesting challenge to solve. In conclusion the system built on data creates a rapidly changing landscape of which I see as the key components in defining revolutionary changes to society and culture. 

Monday, 7 May 2018

Microsoft Build Azure Cosmos DB




















Microsoft Build is underway sharing many useful features. The Azure Cosmos DB API is a versatile tool with a number of options. There are some quickstart tutorials and samples for these.

 Azure Cosmos DB now has multi-master write support. Multi-master in Azure Cosmos DB provides single-digit millisecond latency to write data and availability with built-in flexible conflict resolution support. There are some good examples in the article to help understand this functionality better.






















Azure Operational Data Services includes Azure SQL DB; PostgreSQL; MySQL; Redis Cache; and Cosmos DB.




Friday, 23 February 2018

Become an Azure Cosmos DB hero











There is an Azure Cosmos DB Technical Training Series available to sign up for. It has 7 parts covering a range of topics including a technical deep dive.

  • Technical overview of Azure Cosmos DB
  • Build real-time personalized experiences with AI and serverless technology
  • Using Graph API and Table API with Azure Cosmos DB 
  • Build or migrate your Mongo DB app to Azure Cosmos DB
  • Understanding Operations of Cosmos DB
  • Build Serverless Apps with Azure Cosmos DB and Azure Functions
  • Apply real-time analytics with Azure Cosmos DB and Spark

The training series has interactive Q&A throughout.

Monday, 22 January 2018

Migration from the Relational world to Graph
















I came across this useful blog SQL2Gremlin which translates the Northwind dataset. This was used as a sample database in older versions of SQL Server. The blog post explains the Apache TinkerPop's Gremlin graph traversal language using typical patterns found when querying data with SQL. The SQL examples make use of the T-SQL syntax.

This blog was helpful when looking at Azure Cosmos DB (Microsoft’s globally distributed multi-model database service). The  Gremlin console on the Azure portal is explained in the documentation, Azure Cosmos DB: create, query, and traverse a graph in the Gremlin console. The tutorial creates and queries vertices and edges, updates a vertex property, queries vertices, traverses the graph, and drops a vertex.
















The Gremlin console runs on Linux, Mac, and Windows. It can be downloaded from the Apache TinkerPop site.


Apache TinkerPop is a graph computing framework for both graph databases (OLTP) and graph analytic systems (OLAP).

Friday, 6 October 2017

Machina Summit.AI













I attended IPExpo Europe 4-5 October in ExCel in London with the specific attendance at the Machina Summit.AI.

The opening keynote was by Professor Brian Cox OBE on ‘Where IT & Physics Collide’.  The talk interlinked big data, quantum mechanics and quantum computing. The whistle top tour mentioned the Sloan Digital Sky Survey, which are the most detailed three-dimensional maps of the universe; general relativity; history of space and time; the theory of cosmology; and quantum mechanics ending with quantum theory and predicting the distribution of galaxies. This was an amazing talk and gave a glimpse of the interconnected future.

This was followed by Brad Anderson, Corporate Vice President of Microsoft on ‘Business as usual in a digital war zone’. We live in turbulent times with a 300% increase in user account attacks this year, 96% of malware is automated polymorphic which costs business $15 million. Attacks happen in increasing waves and old defences never stand up against these attacks. In this intelligent war you need an intelligent graph. He introduced the Microsoft Azure Active Directory service as the new control plane. There is the need to eliminate false positives, classify email and guarantee data never leaves the browser and be able to use a real time evaluation engine.

A few other talks covered the practice of monitoring with machine data. There are 2 types of monitoring, transitional IT and the new data driven IT. For the latter there is the need to rethink and improve how IT operates using machine learning to be proactive. Organizational silos and increasing quality are things that need to be broken down to be able to address the velocity data in a more agile way to produce actionable insights.

Conrad Wolfram, Strategic Director, Wolfram Research talked about ‘Enterprise computation: the next frontier in AI and data science’ Todays data challenge is about accessibility of data, personalisation of data and providing insightful answers. Data Science is multi paradigm and machine learning does not have all the answers. Computation is required for everyone with smart automation and computational thinking is needed for everyone. Data science needs to be personalised, multifaceted but unified.

The day 2 keynote was given by Stuart Russell, Professor of Electrical Engineering and Computer Science, University California Berkeley on ‘Human-Compatible AI’. He discussed what is coming soon. Basic language understanding with web-scale question answering and intelligent assistants for health, education, finances and life (not chatbots!!). Robots for unstructured tasks (home, construction, agriculture) and new tools for economics, management and scientific research. He discussed the premise that eventually AI systems will make better* decisions than humans. Well *taking into account more information and looking further into the future. He argued that for the case of super intelligent AI, that you can’t switch off the machine and AI will never succeed.

Other sessions discussed the journey of chaos and how everything fails all the time. To address this there is the need to consider that every journey begins with a single step. There is the inevitable question to consider skills versus knowledge and that is practice.

Microsoft talked about their 'AI and Analytics in the Enterprise'. There is now a need to look at more than the rear view mirror, to see what happened. There is a convergence of cloud, data and AI. With that Microsoft have created an AI platform that is fast and agile, with AI built in and enterprise proven for on-premises to edge to create insights. The evolution of the data state takes into account increasing data volumes, new data sources and types and open source languages. There are 3 stages between the heterogeneous sources and providing apps and insights.
  • Ingest – data orchestration and monitoring
  • Store – Data Lake and storages
  • Machine learning – preparations and train ( Hadoop / spark / SQL and ML) then model and serve (on-prem, Cloud, IoT).

In summary the 2 day conference provided great insight into many new technical areas and raised thought provoking questions about the future of data and AI. 

Wednesday, 10 May 2017

Azure Cosmos DB

Another database annoucment today from Microsoft. The first globally distributed, multi-model database service. Microsoft describe Azure Cosmos DB as containing a write optimized, resource governed, schema-agnostic database engine that natively supports multiple data models: key-value, documents, graphs, and columnar.





















Azure Cosmos DB is an evolutionary leap for DocumentDB which states it contains APIs for accessing data including MongoDB, DocumentDB SQL, Gremlin (preview), and Azure Tables (preview).



Azure Cosmos DB contains a write optimized, resource governed, schema-agnostic database engine that natively supports multiple data models: key-value, documents, graphs, and columnar. 

The Key Capabilities
  • Turnkey global distribution
  • Multiple data models and popular APIs for accessing and querying data
  • Elastically scale throughput and storage on demand, worldwide
  • Build highly responsive and mission-critical applications
  • Ensure "always on" availability
  • Write globally distributed applications, the right way
  • Money back guarantees
  • No database schema/index management
  • Low cost of ownership

Wednesday, 19 April 2017

Microsoft DataAmp – SQL Server 2017

The DataAmp webcast was packed full of announcements. The Webcast was delivered by Scott Guthrie and Joseph Sirosh. The SQL Server product delivering intelligence, trust and flexibility.

Microsoft confirmed that the next version of SQL Server is SQL Server 2017 and will be available simultaneously on Windows, Linux and Docker. Download the SQL Server 2017 datasheet.  It will be the first RDBMS to deliver AI with data. There is a convergence of cloud, data and intelligence. Delivering AI with data: the next generation of Microsoft’s data platform blog shares more information.  SQL Server 2017 Community Technology Preview 2.0 now available.

There were so many new features announced only a few are mentioned below.  There are adaptive query processing improvements which will enhance the performance of workloads. There is a You Tube video SQL Server 2017: Adaptive Query Processing discussing this. Threat detection is now in Azure SQL Database and is straight forward to configure. 



Hybrid Cloud just got easier to adopt with the new Azure migration resources and tools. To help with SQL Server migrations to the cloud features such as Service Broker, SQLAgent, Profiler etc. are now available in Azure. There is a new data migration service for automatic migration for SQL Server, Oracle and MySQL in Azure .

SQL Graph
Storing and analyzing graph data relationships. This includes full CRUD support to create nodes and edges and T-SQL query language extensions to provide multi-hop navigation using join-free pattern matching.  The SQL Server engine integration enables querying across SQL tables and graph data.

SQL Server on Linux
 

The official Microsoft repository for SQL Server in Docker containers is here.

Here are a few videos to help get you started with Linux:

Analytics 

SQL Server is the first commercial database to include Deep Learning algorithms, with the announcement of the Microsoft Cognitive Services general availability of the FACE API and Computer Vision API.
Azure Data Lake Services now have petabyte scale.
Azure Analysis Services became generally available. You Tube video: SQL Server 2017: BI enhancements.  
You can use Python for advanced analytics, You Tube video: SQL Server 2017: Advanced Analytics with Python


Azure DocumentDB

Azure DocumentDB is globally distrubuted and offer limitless scale of throughput and storage. It can be used for things such as IoT applications that need low response times and need to handle massive amounts of reads and writes.



Cortana Intelligence Solution Templates
You can now quickly build Cortana Intelligence Solutions from preconfigured solutions, reference architectures and design patterns. Some are released with more to follow.

The really important thing that I am excited about is the flexibility of choice within the SQL Server product.




Joseph Sirosh concluded comparing the industrial revolution with the intelligence revolution of today.