Welcome

Passionately curious about Data, Databases and Systems Complexity. Data is ubiquitous, the database universe is dichotomous (structured and unstructured), expanding and complex. Find my Database Research at SQLToolkit.co.uk . Microsoft Data Platform MVP

"The important thing is not to stop questioning. Curiosity has its own reason for existing" Einstein



Showing posts with label Machine Learning. Show all posts
Showing posts with label Machine Learning. Show all posts

Tuesday, 1 March 2022

Start your learning journey in Azure AI

To start your learning journey into Azure AI with a helping hand – Powered by Women in AI, the first episode of this three part series can now be viewed on YouTube following the live broadcast Monday, February 28, 2022 4:30 PM to 5:30 PM GMT.

Develop proficiency with the most in-demand skills and advance your career in 30 Days. There is 10 hours of learning paths to complete that set you up for the certification. Follow the link below to take part in the 30 days learn challenge.

https://aka.ms/Azure-AI30DTL

Part one zero to hero






Monday, 28 February 2022

Azure AI Fundamentals: Get Started with AI on Azure


We have an exciting session coming up on Monday 28 February at 4.30PM (GMT) to enable you to get started on your learning journey. Register for the first in a series of three sessions at meetup.

Series Description:

In honour of IWD which takes place in March, this Azure AI series will be delivered by some of our female tech experts. This series is designed to help you get started on a Microsoft Azure AI learning journey. Not matter your previous knowledge of AI, this actionable set of sessions will take you through what is available in Azure AI, Computer Vision and Conversational AI and give you space to ask a set of expert’s questions about each topic.

We will only scratch the surface of what is available on Microsoft learn and will encourage you to continue your learning between weekly sessions in the series. This series partners with the Microsoft Certified: Azure AI Fundamentals certification – if you follow this learning path you will be well on your way to taking this exam.

You can be eligible for 50 percent off the cost of a Microsoft Certification exam by completing your challenge within 30 days. See Terms and Conditions for eligibility details. 10 hours of learning paths to complete that set you up for the certification. Follow the link below to take part in the 30 days learn

https://aka.ms/Azure-AI30DTL

Learn Module: https://aka.ms/learnlive-20220228A

Thursday, 9 April 2020

Spark AI Summit 2020

The largest data and machine learning conference, Spark and AI Summit brings together over 7,500 engineers, scientists, developers, analysts and leaders from around the world to San Francisco every year. Over four days, we shape the future of big data, analytics and AI as we share knowledge, hear from thought leaders and train on open-source technologies like Apache Spark, Delta Lake, MLflow, Koalas, TensorFlow and PyTorch. This year the virtual event for data teams has free general admission, 22-26 June.


Friday, 4 October 2019

Two New Datasets to Improve Natural Language Understanding Models

Google’s PAWS data set helps AI models capture word order and structure. Yuan Zhang, Research Scientist and Yinfei Yang, Software Engineer, Google Research posted: Read more

Word order and syntactic structure have a large impact on sentence meaning — even small perturbations in word order can completely change interpretation. For example, consider the following related sentences:

Flights from New York to Florida.
Flights to Florida from New York.
Flights from Florida to New York.

All three have the same set of words. However, 1 and 2 have the same meaning — known as paraphrase pairs — while 1 and 3 have very different meanings — known as non-paraphrase pairs. The task of identifying whether pairs are paraphrase or not is called paraphrase identification, and this task is important to many real-world natural language understanding (NLU) applications such as question answering Read on

Thursday, 22 August 2019

Microsoft ML for Apache Spark

Microsoft Research announce a new version Microsoft ML for Apache Spark, an open-source and distributed ML and microservice library. v0.18 brings Vowpal Wabbit on Spark, Speech to Text & more!

Microsoft Machine Learning for Apache Spark (MMLSpark) is an ecosystem of enhancements that expand the Apache Spark distributed computing library to tackle problems in Deep Learning. It enables sending streaming data to Power BI.
Website: http://aka.ms/spark Paper: http://aka.ms/spark-paper



Global AI Nights 2019


The Global AI Night is a free evening event organized in London by community people, who are passionate about Artificial Intelligence on the Microsoft Azure. It is at the The Microsoft Reactor in London on Thursday September 5, 2019 5:45 PM – 10:00 PM Register here.

Friday, 2 August 2019

The Big Data Problem

The article The real big-data problem and why only machine learning can fix it and video from the MIT CDO conference, Cambridge, MA contains an interesting discussion on why ETL and MDM don't scale and why placing a schema later doesn't deliver usable data. The key is using machine learning to classify and prep data.



Thursday, 25 July 2019

SQL Server 2019 Workshop Lab

SQL Server 2019 is a modern data platform designed to tackle the challenges of today's data professional.






















There is a new self-paced free lab is available to learn some of the concepts and how to solve modern data challenges using a hands-on lab approach.

SQL Server 2019 provides many new capabilities including:

  • Data Virtualization with Polybase and Big Data Clusters to reduce the need for data movement
  • Intelligent Performance to boost query performance with no application changes
  • Security enhancements such as Always Encrypted and Data Classification
  • Mission Critical Availability including Availability Groups on Kubernetes and Accelerated Database Recovery
  • Modern Development capabilities including Machine Learning Services and Extensibility with Java and the language of your choice
  • SQL Server on the platform of your choice with compatibility including Windows, Linux, Docker, Kubernetes, and Arm64 (Azure SQL Database Edge)

Monday, 6 May 2019

Google AI training data set

Google has released an AI training data set with 5 million images and 200,000 landmarks. The open-sourced Google-Landmarks-v2 contains a larger landmark recognition corpus. Google has also launched two new challenges Landmark Recognition 2019 and Landmark Retrieval 2019 on Kaggle.


Tuesday, 30 April 2019

Azure Open Datasets

Azure Open Datasets are curated public datasets that can be used to add scenario-specific features to machine learning solutions for more accurate models. Open Datasets are on Microsoft Azure and are available to Azure Databricks, Machine Learning service, and Machine Learning Studio. Access to the datasets is through the APIs and other products, such as Power BI and Azure Data Factory.



Saturday, 2 March 2019

SQL Bits 2019 Keynote






















What an amazing SQLBits in Manchester. Four days packed full of leading edge data technology covering

  • SQL Server 2019 Big Data
  • Azure SQL Managed Database
  • Power BI
  • Kubernetes
  • Machine Learning
  • Python
  • Spark

This year SQLBits 2019 had a keynote.  It was nice for the event to have a keynote again. The theme Data Never Rests.  The Microsoft Data Platform Product group who spoke were Buck Woody, Bob Ward, Anna Thomas, Alain Dormehl, Adam Saxton and Patrick LeBlanc. An amazing set of speaks and fun keynote. They shared details of the evolution of the data platform to enable people to keep their skills up to date. The keynote is available to watch . There were several major announcements.


SQL Server 2019 will RTM in second half of the year. SQL Server 2019 CTP2.3 is available now with
  • Big data cluster enhancements
  • Accelerated database recovery
  • Performance enhancements
  • Graph data enhancements
  • SSAS enhancements

SQL Server 2019 is a modern innovation and there are various forms of the product.
  • On Premises
  • SQL Server Azure VM (IaaS)
  • Azure SQL DB Managed Instance (PaaS)
  • Azure SQL Data Warehouse


Azure SQL Database Hyperscale can autoscale up to 100TB and scale compute and storage independently.
During the keynote they showed Azure SQL Database Hyperscale where a 50TB database was restored in just under 8 minutes. That is nice accelerated database recovery.

Data virtualization and big data clusters is a game changing view with SQL Server 2019 big data clusters, data lake scale, machine learning and AI. Multiple data sources can be connected using external table, through the compute pool using Polybase connectors at the source.  Data persistence using multiple data sources is stored in shards of the data pool for SQL Server 2019 big data clusters data mart.



SQL 2019 will send push down predicated queries to other data platforms via Polybase to join SQL data with Oracle, Mongodb and CosmosDB data in one place efficiently.

SQL notebooks in Azure Data Studio is an awesome new feature. 

There is documentation to read and new courses for learning.

aka.ms/DataAccessGuide

and a Summary of All Exams and Certifications Launched in January, 2019!

aka.ms/DataEngCerts







Wednesday, 16 January 2019

Monday, 14 January 2019

The AI Journey

The AI Journey is a interesting blog post that discusses the pragmatic approach to AI and use, the pattern for AI and the journey. 

The patterns seen are for virtual agents, ambient intelligence, AI assisted professionals, knowledge mining and autonomous systems More details are discussed here.

The question of where to start is being asked in many circles and BI is still the foundation. Without good quality data there is no AI. The largest hurdle I think that needs to be overcome is data ingest quality.


Thursday, 10 January 2019

Cloudera vision and strategy

Today the joint vision for the new Cloudera was shared. It was interesting to hear their strategy going forward. I was expecting to hear something revolutionary and new but seems very much the same as other companies at the moment.

Here is a summary of the points.

They will be the only provider to run across all cloud providers Azure, AWS, Google Cloud, IBM and Oracle. Both companies had the same vision to make the impossible possible, to transform data into clear and actionable insights and be committed to open source to give flexibility to its customers.
Cloudera want to

  • Invest in real time streaming at the edge
  • Be enterprise grade
  • Cloud native
  • A data warehouse
  • Provide AI industrialization
  • To deliver the industries first enterprise data cloud

They are developing the next generation platform called the Cloudera Data Platform. It will consist of


100% open source
The best of HDP3 + CDH 6
Hybrid and multi-cloud
Unified, from the edge to AI
Supported through till at least January 2022
Provide predictable and flexible migration paths
To separate compute and storage using technologies like Kubernetes
Have a consistent security ecosystem




There are two application changes:

The Cloudera Data Science workbench will now work with HDP.












HDF to work with CDH

Cloudera talked about the industrialization of AI which requires strategy, people and organization, security,governance and compliance and technology for an enterprise grade AI operation.

Cloudera have launched a new machine learning powered platform by Kubernetes. It is in preview.


Thursday, 6 December 2018

Microsoft Connect() 2018

Microsoft shared yet more innovations at Connect()2018 which considered ubiquitous computing and the fact that technology is transforming business.



Watch the keynote and many other sessions.

They announced the general availability of Azure Machine Learning service, which enables developers and data scientists to efficiently build, train and deploy machine learning models.

Azure Kubernetes Service (AKS) virtual node public preview was announced for serverless Kubernetes. This new feature enabled you to elastically provision additional compute capacity in seconds.

Friday, 26 October 2018

Machine Learning on Azure

At Microsoft Ignite there were many data announcements. Azure AI is another such area that covers the next wave of innovation aimed at transforming business. There are 3 solution areas. 

Predictive models to optimise business process

These are a set of pretrained models for Azure Cognitive Services and ONNX (Open Neural Network Exchange) that enables model interoperability across frameworks. Machine Learning is available with Azure Databricks, Azure Machine Learning and Machine Learning VMs

















AI powered apps to integrate vision, speech and language

There are now services specifically designed to help build AI powered apps & agents.

Knowledge mining to uncover insight from documents
There is valuable information hidden in documents, forms, pdfs and images. Azure Cognitive Search (in preview) adds Cognitive Services on top of Azure Search. 

  



Reading
Azure AI – Making AI real for business

Thursday, 13 September 2018

AI the art of the possible

During SQL Saturday Cambridge I attended a session by Terry McCann on using AI to write a session submission to SQL Saturday. It was a great session and I would recommend you attending it if you get a chance.

In the new data world it is important to understand the difference between AI, Machine Learning and Deep Learning.


























Then breaking this down further the differences between how machine learning works and deep learning is shown here. Deep learning is really a black box.




















Image : https://www.upwork.com/hiring/for-clients/log-analytics-deep-learning-machine-learning/

Terry mentioned a free book to read to learn more Neural Networks and Deep Learning. 

He mentioned also the book Harry Potter and the Portrait of what Looked Like a Large Pile of Ash which was written by an AI bot.  I hadn't come across this before but it lets you see the art of the possible in the future.

Image: https://imgur.com/gallery/gkLFz


Saturday, 7 July 2018

The Future State - Serendipitous Data Management

Gone are the days where companies can survive on existing products and services. The need to continually innovate to stay ahead in a fluid world, requires a change in direction. Many articles have been written, in both academic research and Industry, to try to predict what will be the future state of data technology and what will be this year's trends.

Currently research meets industry in a rebirth of industry-based research teams consisting of organisational only teams or industry collaborating with universities. Guzdial shares his thoughts in the Communicationsof ACM March 2018 journal that "for the majority of new computer science PhD's, the research environment in industry is currently more attractive". Particularly the need within industry to continually innovate cries out for more research divisions in industry. Part of this change is due to the rapid expansion of emerging technology but also the realization, of what data science and artificial intelligence (AI) can add to a business. Data science requires collaboration between people, teams and organisations as interdisciplinary skills are needed to solve today’s problems.

There is an emerging trend whereby more research institutes have been created or existing ones hiring more staff. Microsoft have created a new organization, Microsoft Research AI (MSR AI), to pursue game-changing advances in artificial intelligence. The research team combines advances in machine learning with innovations in language and dialog, human computer interaction, and computer vision to solve some of the toughest challenges in AI.

AI machine learning intelligence, based on big data, is a complex problem to solve, to empower people for the future. In the current world there is the need for collaboration. Greengard in the Communications of ACM March 2018 journal, raised a concern that "mountains of data produce incremental gains, and coordinating all the research groups and silos is a complex endeavour".  Managing data is complex and the key areas that I think will define the next revolution are in the graph.


Telling stories from the data is increasingly important in this ever-changing holistic environment. Skills need to be developed in this area as communicating the meaning of data is crucial. Aiming for improvement in business, science, robotics, space and health can initially appear through intelligent automation and can produce further actionable insights. 

Data visualization is a key component to telling the story and seeing anomalies. Parameswaran discussed at SIGMOD 2018, that it is the scale that brings databases and visualisation together. He highlighted two problem areas, too many tuples and too many visualisations. It is an interesting point to consider how to address the excessive data points and how to appropriately find the right visualization for the data, to gain insight at speed.  

Innovation is key to the next step. I believe that is by making beneficial discoveries by design through scientific experiments from quality data in a continuous and autonomous fashion. I call this Serendipitous Data Management. This improvement and innovation will come from having sound practices for big data management that enable actionable data insights at speed.

Another trend I am seeing in research and industry is looking at how data is processed in centralised data lakes and moving that processing to the edge, particularly for IOT at the moment. As well as this increasing security, if the data can remain at source, it also reduces the volume of data transit which is currently unsustainable. How to consolidate these distributed data sources and produce analysis across disparate systems is an interesting challenge to solve. In conclusion the system built on data creates a rapidly changing landscape of which I see as the key components in defining revolutionary changes to society and culture. 

Sunday, 1 July 2018

Tutorial on Tree Based Modeling

I found a useful tutorial on tree based learning





















The tutorial includes

  • What is a Decision Tree? How does it work?
  • Regression Trees vs Classification Trees
  • How does a tree decide where to split?
  • What are the key parameters of model building and how can we avoid over-fitting in decision trees?
  • Are tree based models better than linear models?
  • Working with Decision Trees in R and Python
  • What are the ensemble methods of trees based model?
  • What is Bagging? How does it work?
  • What is Random Forest ? How does it work?
  • What is Boosting ? How does it work?
  • Which is more powerful: GBM or Xgboost?
  • Working with GBM in R and Python
  • Working with Xgboost in R and Python
  • Where to Practice ?

Tuesday, 27 March 2018

Machine Learning


Predictive analytics uses various statistical techniques such as, machine learning to analyze collected data for patterns or trends to forecast future events. Machine learning uses predictive models that learn from existing data to forecast future behaviors, outcomes, and trends.

Machine Learning libraries enable data scientists to use dozens of algorithms, each with their strengths and weaknesses. Download the machine learning algorithm cheat sheet to help identify how to choose a machine learning algorithm.