Welcome

Passionately curious about Data, Databases and Systems Complexity. Data is ubiquitous, the database universe is dichotomous (structured and unstructured), expanding and complex. Find my Database Research at SQLToolkit.co.uk . Microsoft Data Platform MVP

"The important thing is not to stop questioning. Curiosity has its own reason for existing" Einstein



Showing posts with label SQL Conferences. Show all posts
Showing posts with label SQL Conferences. Show all posts

Saturday, 2 March 2019

SQL Bits 2019 Keynote






















What an amazing SQLBits in Manchester. Four days packed full of leading edge data technology covering

  • SQL Server 2019 Big Data
  • Azure SQL Managed Database
  • Power BI
  • Kubernetes
  • Machine Learning
  • Python
  • Spark

This year SQLBits 2019 had a keynote.  It was nice for the event to have a keynote again. The theme Data Never Rests.  The Microsoft Data Platform Product group who spoke were Buck Woody, Bob Ward, Anna Thomas, Alain Dormehl, Adam Saxton and Patrick LeBlanc. An amazing set of speaks and fun keynote. They shared details of the evolution of the data platform to enable people to keep their skills up to date. The keynote is available to watch . There were several major announcements.


SQL Server 2019 will RTM in second half of the year. SQL Server 2019 CTP2.3 is available now with
  • Big data cluster enhancements
  • Accelerated database recovery
  • Performance enhancements
  • Graph data enhancements
  • SSAS enhancements

SQL Server 2019 is a modern innovation and there are various forms of the product.
  • On Premises
  • SQL Server Azure VM (IaaS)
  • Azure SQL DB Managed Instance (PaaS)
  • Azure SQL Data Warehouse


Azure SQL Database Hyperscale can autoscale up to 100TB and scale compute and storage independently.
During the keynote they showed Azure SQL Database Hyperscale where a 50TB database was restored in just under 8 minutes. That is nice accelerated database recovery.

Data virtualization and big data clusters is a game changing view with SQL Server 2019 big data clusters, data lake scale, machine learning and AI. Multiple data sources can be connected using external table, through the compute pool using Polybase connectors at the source.  Data persistence using multiple data sources is stored in shards of the data pool for SQL Server 2019 big data clusters data mart.



SQL 2019 will send push down predicated queries to other data platforms via Polybase to join SQL data with Oracle, Mongodb and CosmosDB data in one place efficiently.

SQL notebooks in Azure Data Studio is an awesome new feature. 

There is documentation to read and new courses for learning.

aka.ms/DataAccessGuide

and a Summary of All Exams and Certifications Launched in January, 2019!

aka.ms/DataEngCerts







Wednesday, 14 November 2018

Big Data LDN day 1




















I attended Big Data LDN 13-14 November 2018.

The event was busy with vendor product session and technical sessions.  All the sessions were 30 mins so there was a quick turn around following each session. The sessions ran throughout the day with no break for lunch. 

One of the sessions discussed the fourth industrial revolution and the fact that it is causing a cultural shift. The areas of importance that were mentioned were
  • Skills
  • Digital Infrastructure
  • Search and resilience
  • Ethics and digital regulation
Two institutions were mentioned as leading the way. The Alan Turing Institute as the national institute for data science and artificial intelligence and the Ada Lovelace Institute, an independent research and deliberative body with a mission to ensure data and AI work for people and society.

Text Analytics
I attended an interesting session on text analysis. Text analytics process unstructured text to find patterns and relevant information to transform business.  It is far harder than image analysis due to

  • Obstalele - the quantity of data
  • Polymorphy of language
  • Polysemy of language – where words have many forms and meaning
  • Misspellings
Accuracy of sentiment analysis is hard. Sentiment analysis determines the degree of positive, negative or neutral expression. Some tools are bias. Topic modelling was a method discussed for latent dirichlet allocation (LDA). Topic modeling is a form of unsupervised learning that seeks to categorize documents by topic.

Governance

The changing face of governance has created a resurgence and rebirth of data governance. Data is important to classify, reuse and be trustworthy. A McKinsey survey about data integrity and trust of data was mentioned that the talked about defensive (single source of trust) and offensive (multi versions of the truth).

The Great Data Debate

The end of the first day Big Data LDN assembled a unique panel of some of the world’s leaders in data management for The Great Data Debate.

The panelists included
  • Dr Michael Stonebraker Turing Award winner the inventor of Ingres, Postgres, Illustra, Vertica, Streambase and now CTO of Tamr.
  • Dan Wolfson, Distinguished Engineer and Director of Data & Analytics, IBM Watson Media & Weather,
  • Raghu Ramakrishnan, Global CTO for Data at Microsoft
  • Doug Cutting co-creator of Hadoop
  • Chief Architect of Cloudera
  • Phillip Radley Chief Data Architect at BT
There is a growing challenge of complexity and agility in architecture. When data scientists start looking at the data, 80% of time is spent data cleaning and then a further 10% of the time cleaning errors from the data integration. Data scientists are data unifiers not data scientists.  There are two things to consider

  • How to do data unification with lots of tools
  • Everyone will move to the cloud at some point due to economic pressures.

Data lineage is important and privacy needs to be by design. It is possible to have self service for easy analytics but not for more complicated things. A question to also consider is why not clean data at source before migrating it. Democratizing data will require data that is always on and always clean.

There will be no one size fits all. Instead packages will come, such as SQL Server 2019 bundling tools outside such as Spark and HDFS. Going forward there is likely to be 

  • A database management regime in a large database management ecosystem. 
  • A need a best of breed of tools and a uniform lens to view all lineage, all data and all tasks.

The definition of what is a database is, has evolved over time.   There are a few things to consider going forward

  • Diversity of engines for storage and processing.  
  • Keep track of data meta systems after cleaning, data enrichment and provenance is important. 
  • Keep training data attached to the machine learning (ML)  model. 
  • Need enterprise catalog management. 
  • ML brings competitive advantage
  • Separate data from compute

It is a data unification problem in a data catalog era.

  

Monday, 1 October 2018

SQL Relay 2018

I am speaking at SQL Relay 2018 in Bristol, Friday 12 October. My session is: Research skills for industry experts.

Abstract
It is becoming ever increasing the need to present analytic outcomes. Analytics are only ever as good, as the robustness of the data collection and analysis. This session will cover the raft of research skills that can be applied in industry to improve the quality of your investigative work.


Tuesday, 25 September 2018

SQLBits 2019

SQLBits 2019 has been announced. It is in the heart of Manchester. The last time it was in Manchester was in 2009. I am already excited about this next event. 




Saturday, 8 September 2018

SQL Relay: the travelling conference



Registration is open for SQL Relay. This is a conference with a difference. It tours round the UK for 5 consecutive days bringing international and MVP speakers to your local area. It is an amazing conference to attend with lots of sessions to choose from. This years events are at

Newcastle - Monday 8 October
https://sqlrelay2018-newcastle.eventbrite.co.uk/

Leeds - Tuesday 9 October
https://sqlrelay2018-leeds.eventbrite.co.uk/

Birmingham - Wednesday 10 October
https://sqlrelay2018-birmingham.eventbrite.co.uk/

Reading - Thursday 11 October
https://sqlrelay2018-reading.eventbrite.co.uk/

Bristol - Friday 12 October
https://sqlrelay2018-bristol.eventbrite.co.uk/

Sunday, 10 June 2018

24 Hours of PASS Summit Preview 2018

Free Microsoft Data Platform training is to be held 12 June 2018 starting at 12:00 UTC. It is a great opportunity to hear some amazing sessions in 24 hours, of 1 hour back to back content.

Topics covered in this edition include Performance Tuning, Azure Data Lake, Digital Storytelling, Advanced R, Power BI and more! Join in on sessions from Kendra Little, Melissa Coates, Rob Sewell, Brent Ozar, Dejan Sarka, Devin Knight, Mico Yuk and many more experts in their fields.

Saturday, 9 June 2018

ACM SIGMOD/PODS Conference 2018

It is that time of year again when the annual ACM SIGMOD/PODS Conference. To be held 10 -15 June 2018 in Houston. The ACM SIGMOD/PODS Conference is a leading international forum for database researchers, practitioners, developers, and users to explore cutting-edge ideas and results, and to exchange techniques, tools, and experiences.

Microsoft has contributed to the programming of SIGMOD with several researchers serving on committees and inclusion in workshops, research sessions, industry sessions, demo sessions, and poster sessions.


Monday, 9 April 2018

Advice and guidance on becoming a speaker or volunteer

I watched this great session giving 'advice and guidance on becoming a speaker or volunteer' from SQLBits this year. 


I felt humbled when I listened to the SQLBits session recording as I am named as an absolute legend for attending all 16 SQLBits and helping for over 8 years. I had never spoken, never presented or been involved in the public facing side of the conference. It is such a great feeling helping the conference be successful, helping others enjoy what working with data brings and being a part of the sqlfamily. Thanks to SQLBits for enabling me to be a part of such an amazing event for all of these years.


Saturday, 24 March 2018

SQL Relay Session Submission Open



SQLRelay session submission is open https://sessionize.com/sqlrelay2018/. Please submit a session to this travelling conference. Speakers can present at a single or multiple events, it's up to you.

We are running 5 events in the same week in 2018, Monday to Friday, covering 5 different cities within the UK.

    Mon 8 Oct - Newcastle
    Tue 9 Oct - Leeds
    Wed 10 Oct - Birmingham
    Thu 11 Oct - Reading
    Fri 12 Oct - Bristol

We cover a broad range of topics at different levels, from SQL Server DBA to Advanced Analytics in Azure, taking in all aspects of the Microsoft Data Platform. 


Friday, 16 March 2018

Big Data LDN Keynotes

The 2018 opening  keynotes of Big Data LDN have been announced.

Jay Kreps and Michael Stonebraker will be delivering the two opening keynotes.

Jay Kreps, opens the event on day 1, Tuesday 13th November. The Ex-Lead Architect for Data Infrastructure at LinkedIn, Co-creator of Apache Kafka and Co-founder & CEO of Confluent will take to the stage in the keynote theatre at 09:30.

Michael Stonebraker, the Turing Prize winner, IEEE John von Neumann Medal Holder, Co-founder of Tamr and Professor at MIT will address the keynote theatre at 09:30 on day 2, Wednesday 14th November.

Friday, 2 March 2018

The Magic of Data


It was the 10th anniversary of SQLBits this year, with the conference tag line being, the magic of data. The conference was held at Olympia, London between 21-24 February. I am proud to have attended every conference since its inception and that I have been a helper for the last 8 years. At the start of each conference it is always an interesting challenge to understand the new venue layout and what things we can do to make this the best SQLBits conference ever for the attendees and speakers.

There were the usual two days of expert instructor led training days. I looked after a PowerBI day with Adam Saxton and Python day with Dejan Sarka. The Python training included Machine Learning for in-database SQL Server. It was a very helpful to have an overview of data mining, machine learning and statistics for data scientists. Having an appreciation of the maths and algorithms is important in this new diverse data world.

Friday arrived with a mix of general sessions running in multiple tracks. I initially attended a session by Mark Wilcock on Text Analytics Of A Bank's Public Reports. The recording can be seen here. The text analysis in R that was demonstrated was very similar to the type of qualitative analysis I undertook in my PhD. The session an introduction to HDInsight was a great starting point for managing big data.

The date that every data person has in their head this year, is 25 May 2018. That is the GDPR deadline. The big question for Microsoft was understanding the telemetry data collected and its pipeline to ensure that they comply to GDPR. It was great to hear about all the work they have done to address GDPR for the data platform.

I attended more sessions in data science and SQL Graph. Graph databases are very useful in certain scenarios. The on-premises SQL Server 2017 graph engine is different to that of the graph API in Cosmos DB and has different syntax. There are many new features still to come for SQL Graph.

Other very interesting sessions were on performance tuning with the tiger tool box, R in PowerBI, the flexibility of SQL Server 2017, inside the classic machine learning algorithms with Professor Mark Whitehorn, and a session on don't cross the streams, a closer look at Stream Analytics by Johan Ludvig BrattÃ¥s. That concluded the breath of topics I  covered in this years conference. The conference covers an amazing breadth and depth of topics from database management, development, BI, data management and data science. My lightning talk experience from this year is shared here.

The rest of my time was spent mingling and sharing data experiences. It was an honor to have been able to be a part of the conference again.

Wednesday, 28 February 2018

SQLBits Lightning Talks















This year at SQLBits I was asked early Friday afternoon if I would put my name forward for the lightning talks. As a helper for 8 years I wanted to be helpful and agreed stating I had no laptop, access to material or pre-prepared slides. I quickly thought of the only topic I could reasonably cover with very little preparation, that being ‘how I became a doctor’.  As a first-time speaker it was quite daunting to stand up in front of lots of people. Without slides you have nothing to distract the attendees, at the session, from watching you.

I received a mail late afternoon confirming I had been chosen to present. I was originally down to room monitor another session at the same time. This resulted in my frantic search to find another helper who would swap room monitoring sessions. Luckily another helper kindly agreed.
Friday evening, knowing I had to get up early Saturday morning to help and monitor a few sessions, it left no time to prepare. I left the party early Friday evening to pack and to write some semblance of order for a 5 minute presentations.

I felt quite embarrassed to speak alongside others who had prepared slides, those who had presented before or those who had given lightning talks before. I was surprised with the first run though live I finished 10 seconds early, thus keeping to the strict 5 minute time limit.


After the end of the presentation the judges gave their constructive feedback to the room. I knew before they provided feedback that the lack of slides and laptop was against me obtaining great feedback. As another lightning speaker said, this was all a very daunting experience for a first-time speaker and wouldn’t encourage new speakers.





I would say standing on the stage takes courage and all the lightning talk presenters should be proud they did it. Now after returning home after a very exhilarating and exhausting week of helping, learning and networking, I have created the set of slides  for the presentation in case anyone wanted more information. 




Saturday, 24 February 2018

SQL Relay 2018

If you enjoyed SQLBits come to SQLRelay in October 8-12. We are visiting 5 venues around the UK. More details will follow in due course. We look forward to seeing you there,



Saturday, 17 February 2018

SQLBits 2018 The Magic of Data



It is nearly time for the 2018 edition of SQLBits. The four day conference is at Olympia, London. It is the leading data professional conference in Europe. It offers the op3portunity to learn, network, develop and share your data knowledge. There is a great shift in the provision of data management allowing operational and predictive insight and embedded AI on any platform. This conference offers attendees the chance to experience new paradigms and to access an array data related events.  
  •          4 days of world class training
  •         Over 170 specialist sessions
  •          More than 10 hours of best practice sharing with your peers
  •          Networking events with the Microsoft Product Group
  •          Introductions to software vendors and consultants
  •          The infamous SQLBits Party
The conference has grown from strength to strength over the years and it is privilege to be a conference helper for an 8th year.

Wednesday, 1 November 2017

PASS Summit 2017 Day 1 Keynote

I attended PASS Summit 2017 which was my second year of attendance. I enjoyed the conference enormously. It is enjoyable being immersed in data and being with people who are enthusiastic in the field.

The Day 1 Keynote "Microsoft for the Modern Data Estate" was presented by Rohan Kumar. Data is driving transformation. Data, Cloud and AI are the three most disruptive trends of our time. The modern data estate, enables simplicity and common sense. It takes any data from any source, structured or unstructured data and large or small data. The modern data estate provides a seamless infrastructure between on premises, private and public cloud, enabling a hybrid set up that hides the dichotomy of these disparate systems. Seamless flexibly and a choice of engines.

New features in SQL Server 2017

There are many changes to SQL Server 2017. SQL Server 2017 has industry leading performance and security now on Linux and Docker. The key engine changes 
  • Support for graph data and queries
  • Advanced Machine Learning with R and Python
  • Native T-SQL scoring
  • Adaptive Query Processing and Automation Plan Correction

SQL Server 2017 will enable deployment in seconds on Linux and Windows containers and has special pricing for SQL Server on Linux and Red Hat Enterprise Linux.


















New Features Azure SQL Database

Azure SQL Database offers intelligent DBaaS, privacy and trust, seamless and compatibility and competitive TCO. There is seamless migration to the cloud with the cloud first approach breading faster innovations. The list of changes presented


Azure Data Factory now provides a managed environment for SQL Server Integration Services (SSIS) packages and easily move your SSIS workloads to cloud.

There was the announcement made for a new tool called, Microsoft SQL Operations Studio, a free lightweight modern data operations tools for SQL everywhere.
















These are but some of the changes coming to the products.


Saturday, 14 October 2017

SQL Relay 2017














SQL Relay took place between 9 – 13 October 2017. It was the end of my first year of being on the organising committee which has been great fun. 

The relay begin in Reading, moving to Nottingham, Leeds, Birmingham and ending Bristol. This year I helped out on site at 2 events Reading and Bristol. The event brought 4 tracks, 3 general tracks and a workshop track, to each venue. With only 1 hour in the morning before the event starts to set up, it is all hands on deck to get all the attendees registered and the event starting on time. We were lucky to have so many amazing volunteers who helped during the days and without sponsors and speakers we wouldn’t have been able to run the events. I became the event lead for Bristol and it  was nice to be able to bring the event back to the city this year. We enabled around 1000 people to be trained, enabled the SQL community to grow and for people to learn something new. It is a privilege to be a part of this unique event.  

Friday, 6 October 2017

Machina Summit.AI













I attended IPExpo Europe 4-5 October in ExCel in London with the specific attendance at the Machina Summit.AI.

The opening keynote was by Professor Brian Cox OBE on ‘Where IT & Physics Collide’.  The talk interlinked big data, quantum mechanics and quantum computing. The whistle top tour mentioned the Sloan Digital Sky Survey, which are the most detailed three-dimensional maps of the universe; general relativity; history of space and time; the theory of cosmology; and quantum mechanics ending with quantum theory and predicting the distribution of galaxies. This was an amazing talk and gave a glimpse of the interconnected future.

This was followed by Brad Anderson, Corporate Vice President of Microsoft on ‘Business as usual in a digital war zone’. We live in turbulent times with a 300% increase in user account attacks this year, 96% of malware is automated polymorphic which costs business $15 million. Attacks happen in increasing waves and old defences never stand up against these attacks. In this intelligent war you need an intelligent graph. He introduced the Microsoft Azure Active Directory service as the new control plane. There is the need to eliminate false positives, classify email and guarantee data never leaves the browser and be able to use a real time evaluation engine.

A few other talks covered the practice of monitoring with machine data. There are 2 types of monitoring, transitional IT and the new data driven IT. For the latter there is the need to rethink and improve how IT operates using machine learning to be proactive. Organizational silos and increasing quality are things that need to be broken down to be able to address the velocity data in a more agile way to produce actionable insights.

Conrad Wolfram, Strategic Director, Wolfram Research talked about ‘Enterprise computation: the next frontier in AI and data science’ Todays data challenge is about accessibility of data, personalisation of data and providing insightful answers. Data Science is multi paradigm and machine learning does not have all the answers. Computation is required for everyone with smart automation and computational thinking is needed for everyone. Data science needs to be personalised, multifaceted but unified.

The day 2 keynote was given by Stuart Russell, Professor of Electrical Engineering and Computer Science, University California Berkeley on ‘Human-Compatible AI’. He discussed what is coming soon. Basic language understanding with web-scale question answering and intelligent assistants for health, education, finances and life (not chatbots!!). Robots for unstructured tasks (home, construction, agriculture) and new tools for economics, management and scientific research. He discussed the premise that eventually AI systems will make better* decisions than humans. Well *taking into account more information and looking further into the future. He argued that for the case of super intelligent AI, that you can’t switch off the machine and AI will never succeed.

Other sessions discussed the journey of chaos and how everything fails all the time. To address this there is the need to consider that every journey begins with a single step. There is the inevitable question to consider skills versus knowledge and that is practice.

Microsoft talked about their 'AI and Analytics in the Enterprise'. There is now a need to look at more than the rear view mirror, to see what happened. There is a convergence of cloud, data and AI. With that Microsoft have created an AI platform that is fast and agile, with AI built in and enterprise proven for on-premises to edge to create insights. The evolution of the data state takes into account increasing data volumes, new data sources and types and open source languages. There are 3 stages between the heterogeneous sources and providing apps and insights.
  • Ingest – data orchestration and monitoring
  • Store – Data Lake and storages
  • Machine learning – preparations and train ( Hadoop / spark / SQL and ML) then model and serve (on-prem, Cloud, IoT).

In summary the 2 day conference provided great insight into many new technical areas and raised thought provoking questions about the future of data and AI. 

Wednesday, 23 August 2017

SQL Relay 2017 Registration is Open

SQL Relay is returning for its 8th year. The UK roadshow sees the relay carry the Data Platform learning baton between 5 cities.



SQL Relay features top quality Microsoft Data Platform content from Microsoft, and nationally and internationally renowned speakers. There are four tracks; three primary tracks providing in depth tips and tricks for your SQL Server and Azure environment, there is also a morning and afternoon in-depth workshop, spaces are limited on these so make sure you register your place below. With over 1000 registrations on the last SQL Relay, reserve your place quickly.

Come and join the awesome UK roadshow SQL training events in October to learn, grow and adapt. Registration is open and free.

Reading             
Nottingham      
Leeds                   
Bristol        

Tuesday, 1 November 2016

Data Intelligence

I had the amazing opportunity to attend PASS Summit 2016, the largest Microsoft SQL Server event I the world. The event provided the opportunity to meet many international experts and engage with Microsoft engineers in every field.

As a first time attendee there was a lot of logistics to understand to get the most from the event.  I was amazed by the number of Europeans who attended the conference, many of whom I know as a helper for many years at SQLBits. PASS Summit is the pinnacle of the year and I can say I gained much from this event which otherwise would not have been possible.

The first summit keynote delivered by Joseph Sirosh who presented types of A.C.I.D. intelligence with various patterns, intelligent DB, intelligent lake and deep intelligence. A.C.I.D. intelligence being Algorithms, Cloud, IoT and Data. Intelligence is now in every piece of software with applications that continually learn from the data and subsequent information.  This pushes intelligence to where the data lives.

The intelligent database incorporates the new functionality of R Services, provides an operating system of choice (Windows or Linux) for any data deployed anywhere.  The SQL Server 2016 functionality is extended with the hybrid transaction and analytical processing (HTAP) solution which the In-Memory OLTP, In-Memory Analytics, In-Memory Azure SQL Database (launched 15 November) combined with Polybase enable fast querying of structured and unstructured data. Polybase can connect to all data sources such as MongoDB, Hadoop, Teradata, Oracle.  Adding machine learning to the suite of tools add benefits such as real time fraud detection.  DocumentDB properties were also discussed highlighting the blazing fast performance and global replication.

The intelligence lake enables the handling of petabytes of data through algorithms and the extensible data lake. Azure analysis services is available at public preview and Azure SQL Data Warehouse with its parallel processing and scale out was offered as an exclusive one month free trial. There was a great demo by Julie Koesmarno on Azure cognitive services with U-SQL which provided sentiment analysis of War and Peace.

The final part of the key note presented deep learning which looked at many real life examples of learning everywhere from collecting data reviewing whether power lines looked in a good state of repair to face detection to medical research detecting cancer cells.


The keynote was truly inspirational. There were many other amazing sessions with a vast amount of information on diverse topics which I will share in separate posts.

Sunday, 8 May 2016

SQLBits in Space










SQLBits XV was held between 4 -6 May 2016 at the Exhibition Centre in Liverpool. It was the official UK launch event of SQL Server 2016 which will RTM 1st June. There were lots of amazing sessions held for the first time in domes.

The keynote was delivered by Joseph Sirosh, the corporate vice president of the Microsoft Data Group. His keynote entitled the unreasonable effectiveness of data. A paper was written by Alon Halevy, Peter Norvig, and Fernando Pereira of the same title.  Joseph Sirosh mentioned the future effectiveness of data and the Sloan Digital Sky Survey, the 1st astronomy datascope. A take away was that there are many new data services and R should be the language to learn. 

There were many sessions covering the new features of SQL Server 2016 on both the BI and administration side. A highlight for me was the training day on data science by Buck Woody and using the Cortana Intelligence suite. 

The Cortana Analytics Suite big data and advances analytics process



The Azure IaaS and PaaS Services are embedded into these services.

  

U-SQL is another new language that allows you to query unstructured data. Michale Rys delivered a very informative session on Azure Data Lake and U-SQL. The traditional data warehousing approach is

  The new Data Lake approach





 The slides are http://www.slideshare.net/MichaelRys

 There are many new features in SQL Server 2016 and many new data features in Azure.