Welcome

Passionately curious about Data, Databases and Systems Complexity. Data is ubiquitous, the database universe is dichotomous (structured and unstructured), expanding and complex. Find my Database Research at SQLToolkit.co.uk . Microsoft Data Platform MVP

"The important thing is not to stop questioning. Curiosity has its own reason for existing" Einstein



Showing posts with label Big Data. Show all posts
Showing posts with label Big Data. Show all posts

Monday, 9 August 2021

A Summer Retrospective: a bygone era


A few weeks in rural France is just the place to contemplate life and take a step back into a bygone era. An era where there are no phones, no internet and no television. Life can quite easily pass you by and you could go for weeks not speaking to a sole. The fruit on the trees ripen in the orchard, the birds waiting for the perfect moment to swoop and eat the fruit. The roads are mostly empty with the occasional car or logging lorry passing by. Cycling is heaven with the roads to yourself.

This rural area, 214 million years ago, had all life within 300 miles of Rochechouart wiped out when a meteorite, around one of the 15 largest ever to come crashing down on earth. The geological signs of he creator are still present today. This bygone era was also rife with conflict. From the last battle of Richard the 1st - the Lionheart, who laid siege to the Chateau of Chalus-Chabrol, located at the border between Aquitaine and the French kingdom, to the hideouts in the forest of the Maquis du Limousin, who were one of the largest groups of French resistance fighters in the Second World War. The village of Oradour-sur-Glane remains an empty ruin as a memorial for the massacre of its inhabitants.

In this backdrop, technology seems a lifetime away. I can't stress enough the tremendous benefits of taking a technological break for your mental health. You can dream and innovative without the interruption of everyday life.

The age of cloud computing, big data and the algorithm requires a 360-degree perspective. A socio-technical perspective is critical. Reflecting on the changes to the earth, made to this unique landscape from space, you realize that data is in the environment. It is not possible to be an expert in all areas as data is the environment. Data is history, is in the maps, is used in conflict resolution and is used for impact analysis. Data is completely inseparable from life and it drives life, not only business. The  choice of tools available help you navigate through data are vast.

The question is can one truly ever master the entirety of life. Data is life, the past, the present and the future. To truly be a master it requires collaboration, communication and control, as data weaves its interconnected complexity throughout life. A holistic view of this diverse scientific area is required to provide a sustainable future. There is no one best practice that can help navigate this web of graph vertices and edges.

To that end I summise that taking a technological break enables the mind to contemplate and blue sky thinking roam free. Happy Summer break. 

Wednesday, 16 October 2019

After Data Relay


Data Relay has finished for the 2019 season. This year I was involved as Head of Marketing and was the Bristol event owner again. I enjoyed being a part of the Data Relay team. It is great to be able to put on an event with free training across the breadth of the Microsoft Data Platform.  

I also presented on “The big data and database management paradigm shift”  The abstract

With the emergence of big data this has created a new paradigm for data and database management.  Structured and unstructured data must now work together to produce actionable insights. This session will share details on how to embrace this new era of database management, to prepare you for SQL Server 2019.

The slides summary 




















References



Thursday, 5 September 2019

The Mirage and Metamorphosis of Data and AI



I have just taken a break to rejuvenate my creative juices. It was a time to reflect and innovate. We are often so busy in our day to day lives we don't stop and reflect. I spent my time reading and catching up on bleeding edge technology. I am always fascinated to see what is coming next, what problems researchers are trying to address and how Data and AI could be utilised to benefit industry and the world around us.

The role I enjoy the most is as a Data and AI philosopher providing thought leadership. We are at an exciting time in history to witness and contribute to the mirage and metamorphosis of Data and AI. My explorations find exciting challenges in diversity and Data and AI at the centre of most things we want to achieve. Research is increasingly needed in industry to achieve business success due to the increasing complexity within industry and the world around us. We need to move away from agile for certain tasks to enable complexity to be understood and use systems thinking techniques.  My findings on the future mirage and metamorphosis of Data and AI are a complex interconnected world around Data and AI and mastering that complexity is the key to success.



  











References
The data and AI market landscape 2019: The next wave of hybrid emerges
https://www.zdnet.com/article/the-data-and-ai-market-landscape-2019-the-next-wave-of-hybrid-emerges/
Part I: A Turbulent Year: The 2019 Data & AI Landscape
https://mattturck.com/data2019/
Part II: Major Trends in the 2019 Data & AI Landscape
https://mattturck.com/2019trends/
Navigating AI hype in search of success, Oliver Pickup (Sunday Times 12 May 2019)
The real big-data problem and why only machine learning can fix it
https://siliconangle.com/2019/08/09/real-big-data-problem-machine-learning-can-fix-mitcdoiq-startupoftheweek/
Big Data is just Data
https://buckwoody.wordpress.com/2019/08/26/big-data-is-just-data/
Maximising the AI opportunity
https://info.microsoft.com/rs/157-GQE-382/images/UK-DIGTRNS-CNTNT-content-MGC0003240.pdf
The Data Ethics Framework principles
https://www.gov.uk/government/publications/data-ethics-framework/data-ethics-framework

Friday, 2 August 2019

The Big Data Problem

The article The real big-data problem and why only machine learning can fix it and video from the MIT CDO conference, Cambridge, MA contains an interesting discussion on why ETL and MDM don't scale and why placing a schema later doesn't deliver usable data. The key is using machine learning to classify and prep data.



Thursday, 20 June 2019

The Common Data Model

The Common Data Model (CDM) is a standard and extensible collection of schemas (entities, attributes, relationships) that represents business concepts and activities with well-defined semantics, to facilitate data interoperability. 

The Common Data Model, that was announced at Ignite, as part of the Open Data Initiative, a jointly-developed vision by Microsoft, Adobe and SAP. CDM is already supported in the Common Data Service, Dynamics 365, PowerApps, Power BI, and upcoming Azure data services. This data is continually developing at Microsoft. I do wonder what consistency there will be between the Microsoft Common Data Model and the Splunk Common Information Model






























Saturday, 2 March 2019

SQL Bits 2019 Keynote






















What an amazing SQLBits in Manchester. Four days packed full of leading edge data technology covering

  • SQL Server 2019 Big Data
  • Azure SQL Managed Database
  • Power BI
  • Kubernetes
  • Machine Learning
  • Python
  • Spark

This year SQLBits 2019 had a keynote.  It was nice for the event to have a keynote again. The theme Data Never Rests.  The Microsoft Data Platform Product group who spoke were Buck Woody, Bob Ward, Anna Thomas, Alain Dormehl, Adam Saxton and Patrick LeBlanc. An amazing set of speaks and fun keynote. They shared details of the evolution of the data platform to enable people to keep their skills up to date. The keynote is available to watch . There were several major announcements.


SQL Server 2019 will RTM in second half of the year. SQL Server 2019 CTP2.3 is available now with
  • Big data cluster enhancements
  • Accelerated database recovery
  • Performance enhancements
  • Graph data enhancements
  • SSAS enhancements

SQL Server 2019 is a modern innovation and there are various forms of the product.
  • On Premises
  • SQL Server Azure VM (IaaS)
  • Azure SQL DB Managed Instance (PaaS)
  • Azure SQL Data Warehouse


Azure SQL Database Hyperscale can autoscale up to 100TB and scale compute and storage independently.
During the keynote they showed Azure SQL Database Hyperscale where a 50TB database was restored in just under 8 minutes. That is nice accelerated database recovery.

Data virtualization and big data clusters is a game changing view with SQL Server 2019 big data clusters, data lake scale, machine learning and AI. Multiple data sources can be connected using external table, through the compute pool using Polybase connectors at the source.  Data persistence using multiple data sources is stored in shards of the data pool for SQL Server 2019 big data clusters data mart.



SQL 2019 will send push down predicated queries to other data platforms via Polybase to join SQL data with Oracle, Mongodb and CosmosDB data in one place efficiently.

SQL notebooks in Azure Data Studio is an awesome new feature. 

There is documentation to read and new courses for learning.

aka.ms/DataAccessGuide

and a Summary of All Exams and Certifications Launched in January, 2019!

aka.ms/DataEngCerts







Thursday, 10 January 2019

Cloudera vision and strategy

Today the joint vision for the new Cloudera was shared. It was interesting to hear their strategy going forward. I was expecting to hear something revolutionary and new but seems very much the same as other companies at the moment.

Here is a summary of the points.

They will be the only provider to run across all cloud providers Azure, AWS, Google Cloud, IBM and Oracle. Both companies had the same vision to make the impossible possible, to transform data into clear and actionable insights and be committed to open source to give flexibility to its customers.
Cloudera want to

  • Invest in real time streaming at the edge
  • Be enterprise grade
  • Cloud native
  • A data warehouse
  • Provide AI industrialization
  • To deliver the industries first enterprise data cloud

They are developing the next generation platform called the Cloudera Data Platform. It will consist of


100% open source
The best of HDP3 + CDH 6
Hybrid and multi-cloud
Unified, from the edge to AI
Supported through till at least January 2022
Provide predictable and flexible migration paths
To separate compute and storage using technologies like Kubernetes
Have a consistent security ecosystem




There are two application changes:

The Cloudera Data Science workbench will now work with HDP.












HDF to work with CDH

Cloudera talked about the industrialization of AI which requires strategy, people and organization, security,governance and compliance and technology for an enterprise grade AI operation.

Cloudera have launched a new machine learning powered platform by Kubernetes. It is in preview.


Friday, 4 January 2019

From the Edge to AI

Hortonworks completed their merger with Cloudera to make them the second largest open source software company in the world. The company is now called Cloudera. The combined platform will enable enterprises to create greater value from data with:

  • The right data analytics, running on data anywhere
  • Strong enterprise-grade and enterprise-wide data security, governance and management
  • Flexibility to choose among multi and hybrid clouds

There is a virtual event on 10 January from the edge to AI to hear about their vision and direction.

Friday, 16 November 2018

Big Data LDN Day 2





















The Fourth Industrial Revolution Report – Download for Free

Keynote

The day 2 keynote was given by Michael Stonebraker, Turing Prize winner, IEEE John von Neumann Medal Holder, Professor at MIT, Co-founder of Tamr entitled Big Data, Disruption and the 800 pound gorilla in the corner.

A few vignettes were mentioned Hamiltons, Dewitts and Amadeus.  Hadoop (meaning map reduce) was started to be used in 2010 and Google stopped using it in 2011. Hadoop now means a HDFS file system.  Cloudera's big problem is that no one wants map reduce. Map reduce is not used for anything.

The data warehouse is yesterdays problem. BI is simple SQL. Data Science has complex problems and it is a different skill set. It is based on deep learning, machine learning and linear algebra, nothing to do with SQL. Deep learning is all the rage but you need vast amounts of training data.  It is not possible to explain why the black box gives certain recommendations so it is not good when data providence is required. 

Big velocity is a big problem over time. Pattern matching and CEP (Complex Event Processing)
like Storm is not competitive. Don’t run Oracle but instead run Mongodb, Cassandra or Redis. NoSQL means no standards and no ACID. ACID is a good idea. NoSQL means you always give up something as per CAP theorem. Declarative languages are a great idea.

Data discovery is a big problem. You spend 90% of your time finding and cleaning data. Then 10% finding and cleaning the errors. Very little time is spent doing data integration. It is a data integration challenge.

Graphs

Jim Webber from Neo4J gave an insightful talk about how useful graphs are to solve problems and predict outcomes. There was some great examples of how to use graphs. He talked about triad closure and strong and week ties. Also mentioning a couple of papers to read

Effects of organizational support on organizationalcommitment  Fakhraei M, Imami R, Manuchehri S (2015)
Semi-Supervised Classification with Graph ConvolutionalNetworks Thomas N. Kipf , Max Welling  (2017) 

and a free ebook

Free Book: Graph Databases By Ian Robinson, Jim Webber, and Emil Eifrém

It is important to have semantic domain knowledge for inference and understanding in graphs as graphs depend on the context. Graphml convolutional network graph will be the data structure for AI.

The Joy of Data
The closing session of the event was delivered by Dr Hannah Fry, Associate Professor in the mathematics of Cities – UCL. This was an amazing session exploring what visualization and insights can be achieved from understanding the data.

She started the talk with the strange Wikipedia phenomenon that all routes lead to philosophy. So by clicking the first proper link on a page you will eventually end up on the philosophy page. 

There are 2 parallel universes where people click the link and the mathematical universe.  Data is the bridge.

She showed how data could be used to investigate why the bicycle transport scheme in London was seeing all the bikes ending up in the wrong place. Vans have to go round moving bikes into the right places during the day. This was the result of people liking to cycle down the hills but not up.

Another example showed that Islington station was a bottleneck which caused a cascading problem because it has a lack of transport routes from there. There were many other interesting examples and how gossip can pay by using network science to track the problem down.

Big Data LDN had some amazing sessions and insightful content. Big Data LDN will be back next year 13-14 Nov 2019.

Wednesday, 14 November 2018

Big Data LDN day 1




















I attended Big Data LDN 13-14 November 2018.

The event was busy with vendor product session and technical sessions.  All the sessions were 30 mins so there was a quick turn around following each session. The sessions ran throughout the day with no break for lunch. 

One of the sessions discussed the fourth industrial revolution and the fact that it is causing a cultural shift. The areas of importance that were mentioned were
  • Skills
  • Digital Infrastructure
  • Search and resilience
  • Ethics and digital regulation
Two institutions were mentioned as leading the way. The Alan Turing Institute as the national institute for data science and artificial intelligence and the Ada Lovelace Institute, an independent research and deliberative body with a mission to ensure data and AI work for people and society.

Text Analytics
I attended an interesting session on text analysis. Text analytics process unstructured text to find patterns and relevant information to transform business.  It is far harder than image analysis due to

  • Obstalele - the quantity of data
  • Polymorphy of language
  • Polysemy of language – where words have many forms and meaning
  • Misspellings
Accuracy of sentiment analysis is hard. Sentiment analysis determines the degree of positive, negative or neutral expression. Some tools are bias. Topic modelling was a method discussed for latent dirichlet allocation (LDA). Topic modeling is a form of unsupervised learning that seeks to categorize documents by topic.

Governance

The changing face of governance has created a resurgence and rebirth of data governance. Data is important to classify, reuse and be trustworthy. A McKinsey survey about data integrity and trust of data was mentioned that the talked about defensive (single source of trust) and offensive (multi versions of the truth).

The Great Data Debate

The end of the first day Big Data LDN assembled a unique panel of some of the world’s leaders in data management for The Great Data Debate.

The panelists included
  • Dr Michael Stonebraker Turing Award winner the inventor of Ingres, Postgres, Illustra, Vertica, Streambase and now CTO of Tamr.
  • Dan Wolfson, Distinguished Engineer and Director of Data & Analytics, IBM Watson Media & Weather,
  • Raghu Ramakrishnan, Global CTO for Data at Microsoft
  • Doug Cutting co-creator of Hadoop
  • Chief Architect of Cloudera
  • Phillip Radley Chief Data Architect at BT
There is a growing challenge of complexity and agility in architecture. When data scientists start looking at the data, 80% of time is spent data cleaning and then a further 10% of the time cleaning errors from the data integration. Data scientists are data unifiers not data scientists.  There are two things to consider

  • How to do data unification with lots of tools
  • Everyone will move to the cloud at some point due to economic pressures.

Data lineage is important and privacy needs to be by design. It is possible to have self service for easy analytics but not for more complicated things. A question to also consider is why not clean data at source before migrating it. Democratizing data will require data that is always on and always clean.

There will be no one size fits all. Instead packages will come, such as SQL Server 2019 bundling tools outside such as Spark and HDFS. Going forward there is likely to be 

  • A database management regime in a large database management ecosystem. 
  • A need a best of breed of tools and a uniform lens to view all lineage, all data and all tasks.

The definition of what is a database is, has evolved over time.   There are a few things to consider going forward

  • Diversity of engines for storage and processing.  
  • Keep track of data meta systems after cleaning, data enrichment and provenance is important. 
  • Keep training data attached to the machine learning (ML)  model. 
  • Need enterprise catalog management. 
  • ML brings competitive advantage
  • Separate data from compute

It is a data unification problem in a data catalog era.

  

Wednesday, 7 November 2018

PASS Summit 2018 Keynote Day 1












The first keynote of PASS summit was delivered by Rohan Kumar entitled SQL Server and Azure Data Services: Harness the ultimate hybrid platform for data and AI





Customer priorities for a modernized data estate are: modernizing on-premises, modernizing to cloud, build cloud native apps and unlocking insights.






The announcements follow:

SQL Server 2019
SQL Server 2019 Public Preview  is a great way to celebrate the 25th anniversary of SQL Server

There is the introduction of big data clusters which combines Apache Spark and Hadoop into a single data platform called SQL Server. This combines the power of Spark with SQL Server over the relational and non-relation data sitting in SQL Server, HDFS and other systems like Oracle, Teradata, CosmosDB.

There are new capabilities around performance, availability and security for mission critical environments along with capability to leverage hardware innovations like persistent memory and enclaves.

Hadoop, ApacheSpark, Kubernetes and Java are native capabilities in the database engine.

Accelerated data recovery (ADR) was demonstrated and is incredible. It is at public preview.  The benefits of ADR are
  • Fast and consistent Database Recovery
  • Instantaneous Transaction rollback
  • Aggressive Log Truncation

Azure HDInsight 4.0

HDInsight 4.0 is now available in public preview.

There are several Apache Hadoop 3.0 innovations. Hive LLAP (Low Latency Analytical Processing known as Interactive Query in HDInsight) delivers ultra-fast SQL queries. The Performance metrics provide useful insight.

Integration with Power BI direct Query, Apache Zeppelin, and other tools. To learn more HDInsight Interactive Query with Power BI.

Data quality and GDPR compliance enabled by Apache Hive transactions
Improved ACID capabilities handle data quality (update/delete) issues at row level. This means that GDPR compliance requirements can now be meet with the ability to erase the data at row level. Spark can read and write to Hive ACID tables via Hive Warehouse Connector.

Apache Hive LLAP + Druid = single tool for multiple SQL use cases

Druid is a high-performance, column-oriented, distributed data store, which is well suited for user-facing analytic applications and real-time architectures. Druid is optimized for sub-second queries to slice-and-dice, drill down, search, filter, and aggregate event streams. Druid is commonly used to power interactive applications where sub-second performance with thousands of concurrent users are expected.

Hive Spark Integration
Apache Spark gets updatable tables and ACID transactions with Hive Warehouse Connector

There are several Apache Hadoop 3.0 innovations. Hive LLAP (Low Latency Analytical Processing called Interactive Query in HDInsight) for ultra-fast SQL queries. The Performance metrics provide useful insight.

Integration with Power BI Direct Query, Apache Zeppelin, and other tools. To learn more watch HDInsight Interactive Query with Power BI.

Better data quality and GDPR compliance enabled by Apache Hive transactions
Improved ACID capabilities handle data quality (update/delete) issues at row level. GDPR compliance requirements can now be meet with the ability to erase the data at row level. Spark can read and write to Hive ACID tables via Hive Warehouse Connector

Apache Hive LLAP + Druid = single tool for multiple SQL use cases

Druid is a high-performance, column-oriented, distributed data store, which is suited for user-facing analytic applications and real-time architectures. Druid is optimized for sub-second queries to slice-and-dice, drill down, search, filter, and aggregate event streams. Druid is commonly used to power interactive applications where sub-second performance with thousands of concurrent users are expected.

Hive Spark Integration
Apache Spark gets updatable tables and ACID transactions with Hive Warehouse Connector.



















Apache HBase and Apache Phoenix
Apache HBase 2.0 and Apache Phoenix 5.0 get new performance and stability features and all of the above have enterprise grade security.

Azure
Azure event hubs for Kafka is generally available
Azure Data Explorer is in public preview.

Azure Databricks Delta is in public preview
  • Connect data scientist and engineers
  • Prepare and clean data at massive scales
  • Build/train models with pre-configured ML

Azure Cosmos DB multi master replication was demoed with a drawing app, Azure Cosmos DB PxDraw
Azure SQL DB Managed Instances will be at General Availability (GA) on Dec 1st. This provides Availability Groups managed by Microsoft.

Power BI
















The new Dataflows is an enabler for self-service data prep in Power BI

Power BI Desktop November Update
  • Follow-up questions for Q&A explorerIt is possible to ask follow-up questions inside the Q&A explorer pop-up, which take into account the previous questions you asked.
  • Copy and paste between PBIX files
  • New modelling view makes it easier to work with large models.
  • Expand and collapse matrix row headers


Monday, 24 September 2018

SQL Server 2019, Big Data and AI


At Microsoft Ignite SQL Server 2019 was launched. An amazing product for the future combining SQL Server 2019 with big data and analytics. It is great to see the combining of multiple tools in once place, a one stop shop for large and small data, structured and unstructured and from multiple sources.

There are 3 major components to SQL Server 2019.




















The creation of a data virtualization layer that handles complexity of all data sources and format.  Enabling the integration of structured and unstructured data without moving the data.

The streamlining of data management with SQL Server 2019 big data clusters deployed in Kubernetes integrating HDFS and Spark. The architecture is explained in more depth here and looks like




The creation of a complete AI platform that can use Spark to analyse both structured and unstructured data anywhere, use SQL Server machine learning services and SparkML.




In summary SQL Server big data clusters allow you to deploy scalable clusters of SQL Server, Spark, and HDFS Docker containers running on Kubernetes.

Read More




Tuesday, 1 May 2018

Big Data Exploration

The big data landscape is growing and exploration of the data can help make better decisions. I came across this great infographic from IBM.


Friday, 20 April 2018

DataWorks Summit 2018




This was the first time I had attended the DataWorks summitIdeas. Insights. Innovation. for big data. I had the privilege to attend the Luminaries dinner on arrival at the conference. The dinner was held for the European data heroes award. The Hortonworks data heroes initiative recognizes the data visionaries, data scientists, and data architects transforming their businesses and organizations through Big Data.

Each day started with a set of keynotes.

Day 1 Opening Keynotes
The Single Most Important Formula for Business Success Scott Gnau - Hortonworks
Changing the Data Game with Open Metadata and Governance Mandy Chessell - IBM
Big Data Success In Practice: The Biggest Mistakes To Avoid Across The Top 5 Business Use Cases Bernard Marr - Bernard Marr & Co.
Munich Re: Driving a Big Data Transformation Andreas Kohlmaier - Munich Re

Scott Gnau opened his talk with an hypothesis “Data is your cloud is your business” Connecting disparate data to provide for real time information enables us to innovated fast. A data strategy is imperative, it needs to include governance, security and adopt rapid change. Data drives our lives everyday from smart edge devices to all businesses.





















He concluded with your data strategy is your cloud strategy is your business strategy if (A) =(B) and (B) = (C) then (A) =(C).

Bernard Marr then shared his insights about AI automating more things faster and the fourth industrial revolution.  He mentioned the top 5 business use cases as

  • Informing: to make better decisions
  • Understand: know you customers better
  • Improvement: customer value proposition
  • Automation: key business processes
  • Monetization: data as an asset
A couple of interesting points raised were about specialist data hunting units to find new data sources and automation requirements to improve operations.  Data diversity is key to improve analytics along with data governance.

Day 2 Keynotes
Renault: A Data Lake Journey Kamelia Benchekroun - Renault Group
Are You Ready For GDPR? Jamie Engesser - Hortonworks, Srikanth Venkat - Hortonworks Inc
Embracing GDPR to Improve Your Business Practices in the Digital Age Enza Iannopollo - Forrester Research
Driving High Impact Business Outcomes from Artificial Intelligence Frank Saeuberlich – Teradata

Day 2 Forester Enza Iannopollo discussed embracing GDPR to improve your business practices in the digital age. Privacy by design and by default requires new business processes to be established and cultural change to happen. GDPR requires compliance across the organization and with external partners. The compliance strategies are only as good as your risk assessment and mitigation. The classification of data is a key place to start. Concluding the sessions with a quote
“Good Data protection normally enables you to do more things with data, not less” Tim Gough Head of Data Protection Guardian News and Media

Data Steward Studio (DSS) was launched at the conference. It is one of several services available for Hortonworks DataPlane Service; it provides a suite of capabilities that allows users to understand and govern data across enterprise data lakes.





Sunday, 1 April 2018

Literature Map


When you start any research project, you need to set the research in the context of the current literature. This will establish a framework for the importance of the study. This document was the starting place for organizing the literature of interest in my research.

Thesis Title: A Study in Best Practices and Procedures for the Management of Database Systems



Friday, 16 March 2018

Big Data LDN Keynotes

The 2018 opening  keynotes of Big Data LDN have been announced.

Jay Kreps and Michael Stonebraker will be delivering the two opening keynotes.

Jay Kreps, opens the event on day 1, Tuesday 13th November. The Ex-Lead Architect for Data Infrastructure at LinkedIn, Co-creator of Apache Kafka and Co-founder & CEO of Confluent will take to the stage in the keynote theatre at 09:30.

Michael Stonebraker, the Turing Prize winner, IEEE John von Neumann Medal Holder, Co-founder of Tamr and Professor at MIT will address the keynote theatre at 09:30 on day 2, Wednesday 14th November.

Friday, 2 March 2018

There are revised patterns available for big data advanced analytical capabilities using the Azure Databricks platforms with Azure Machine Learning.




The new capabilities will enable advance analytics to be carried out using Azure Machine Learning. The different types of data requirements and consumption are integrated using CosmosDB.

Databricks in Azure

Databricks is a big data unified analytics platform that harness the power of AI. It is built on top of Spark, serverless and is highly elastic cloud based. Azure Databricks is in preview currently. This new Azure service aims to accelerate innovation by enabling data science with a high-performance analytics platform that’s optimized for Azure. It has native integration with other Azure services such as Power BI, SQL Data Warehouse, Cosmos DB as well as from enterprise-grade Azure security, including Active Directory integration, compliance, and enterprise-grade SLAs. More information can be found in these two links

A technical overview of Azure Databricks
https://azure.microsoft.com/en-gb/blog/a-technical-overview-of-azure-databricks/

Introduction to Azure Databricks
https://channel9.msdn.com/Events/Connect/2017/T257

Databricks is a collaborative workspace.

























Databricks have an ebook Simplifying Data Engineering to Accelerate Innovation which covers

  • The three primary keys to better data engineering
  • How to build and run faster and more reliable data pipelines
  • How to reduce operational complexity and total cost of infrastructure ownership
  • 5 examples of enterprises building reliable and highly performant data pipelines