Welcome

Passionately curious about Data, Databases and Systems Complexity. Data is ubiquitous, the database universe is dichotomous (structured and unstructured), expanding and complex. Find my Database Research at SQLToolkit.co.uk . Microsoft Data Platform MVP

"The important thing is not to stop questioning. Curiosity has its own reason for existing" Einstein



Friday, 9 November 2012

Pass Summit 2012: Polybase What, Why, How.

[The original material links in this article are no longer available on the web and have been removed.] This was a talk in 2012 that I attended.]

This session was very interesting and shared the technical workings of Polybase. It was delivered by David DeWitt, Director at Microsoft from work in the Gray Systems Lab.  

He set the scene explaining the Hadoop ecosystem.













Then defined the main components.











This is a two universe world of both structured and unstructured data. The 2 alternative solutions are Sqoop which has limitations and Polybase which is a superior alternative.


 As stated on the Gray Systems Lab site "the goal of the Polybase project is to allow SQL Server PDW users to execute queries against data stored in Hadoop, specifically the Hadoop distributed file system (HDFS). Polybase is agnostic on both the type of the Hadoop cluster (Linux or Windows)"




The sessions discussed the approach and drawbacks of each method. There are several phases of delivery planned.
  • Phase 1: in PDW next year
  • Phase 2: working on
  • Phase 3: thinking about
 David DeWitt ended the presentation with
  • The world has changed
  • Map reduce really is not the right tool
  • Polybase for PDW is a first step.

Thursday, 8 November 2012

The Data Lifecycle: Turning Data Into Business Value





The second PASS summit keynote was delivered by Quentin Clark, Corporate Vice President, Data Platform Group. The keynote took a journey through the data lifecycle. The data lifecycle being broken down into 5 areas.
  • Collaborate
  • Operationalize
  • Manage
  • Discover and refine
  • Visualize
The lifecycle starts with managing all the data. Combining any data wherever it lives, be it relational and non-relational, cloud & on-premises and model architecture. Then from that data finding all the relevant information and ways to unlock the value of the data with models and analysis. After which creating visualizations of the data and allowing collaboration on the data. The data requires governance and control,  transforming insights into repeatable business process. Self service is the key element being delivered by the new tools.  The demos showed many of the available tools including Data Explorer http://www.microsoft.com/en-us/sqlazurelabs/labs/dataexplorer.aspx and how AlwaysOn allows you to add Azure as an availability Replica in SSMS which adds scalability.

The session ended with the message that Microsoft are providing a complete platform solution to let you find what the data is trying to say and embracing the new value of data.

Wednesday, 7 November 2012

Pass Summit 2012 Day One Keynote: Accelerate Insight on any Data

Today is the official start to the 2012 PASS summit. The PASS Summit 2012 Day One Keynote session was delivered by Microsoft corporate vice president of the Data Platform Group (DPG) Ted Kummert. The keynote was packed with announcements on changes in the database engines.


Service Pack 1

The first announcement made at PASS is the release of SQL Server 2012 SP1. Some of the new features include
  • Cross-Cluster Migration of AlwaysOn Availability Groups for OS Upgrade
  • Selective XML Index to improve performance
  • DBCC SHOW_STATISTICS works with SELECT permission
  • New dynamic management function which returns statistics properties (sys.dm_db_stats_properties)
  • Express now comes with the complete SQL Server Management Studio (SSMS)
  • SlipStream Full Installation
  • Management Object Support Added for Resource Governor DDL
To read more read what's new in SQL Server 2012 SP1 http://bit.ly/YKmft4

Ted Kummert continued saying the world of data is changing with superabundant supplies. It is approaching a tipping point with volume, variety, velocity, hardware innovation, software innovation and cloud.  There is also a change in architectural assumptions with big data providung new insights with new sources of data.  There is the need to accelerate business process and insight, manage any data, any size, anywhere and enable pervasive insight.








Hekaton

The second announcement is project code name ‘Hekaton’ to accelerate in-memory for business process and insight. Project Hekaton brings an in-memory engine process to the transactional world. After converting tables to then in-memory engine there is almost a 10x performance increase. It allows recompiles of stored procedures so they will run in memory,  which could also offer a 30 times performance increase.
This is in the next release of sql.
The Hekaton AMR tool is designed to help identify hot spots in the database application and provide assistance to migrate things such as tables and stored procedures.

HDinsight

For Non-relational data Microsoft has released to CTP HDinsight Server. Microsoft’s Apache Hadoop based solutions for Windows Server and Windows Azure. http://www.microsoft.com/en-us/download/details.aspx?id=35397

SQL Server 2012 Parallel Data Warehouse
Microsoft announced the new SQL Server 2012 Parallel Data Warehouse. Queries that previously took 20 minutes to run only take 20 seconds. It offers an up to 50x performance gain with an optimized architecture. Due out in H1 2013.

Polybase

Polybase is a breakthrough in data processing. It integrates Hadoop data and relational data to allow combined TSQL queries with an optimized architecture. It will allow future expansion to other data sets. Polybase will unify the relational and non-relational world. Built for big data and coming in H1 2013.

Microsoft’s strategy is to design an Enterprise platform for  integrating data. We have to learn one new concept and that is about joining up data coming from multiple sources.

Other
Additional annoucements mentioned
  • Clustered column store index with updatable tables
  • Power View and PowerPivot fully integrated in Excel, now a complete BI tool
  • Fully interactive maps inside Excel
  • DAX queries on top of molap cubes

In conclusion data is bringing about lots of changes which adds richness  to the solutions.

PASS is running the first ever Business Intelligence conference in spring 2013 in Chicago.



Wednesday, 3 October 2012

Day 3 PASS SQL Rally Nordic

Day 3 launched straight into sessions. The session designing Hybrid Systems for the Enterprise by Buck Woody was based on the ebook  Building Hybrid Applications in the Cloud on Windows Azure http://www.tinyurl.com/9ffms3t . The session provided an excellent guide for creating a cloud architecture. The session finished quoting Plato “The beginning is the most important part of the work.”

The next couple of sessions were very useful, exploring under the hood with the sessions using statistics and relational algebra. The sessions were Optimizing column stores with statistical analysis by Thomas Kejser and the SQL Server 2012 Query Optimiser by Conor Cunningham.

Column store was originally written about in papers in the 1970’s and has only now been incorporated into SQL Server. Mathematical techniques were discussed, low cardinality, correlation of columns, entropy, mutual information leading to the big unsolved question of computer science  P=NP. More information is here http://blog.kejser.org/2012/07/27/what-is-the-best-sort-order-for-a-column-store/

The Query Optimiser session gave a flying visit to how the optimiser works. The optimizer finds a good plan rather than optimal plan as it may take to days to find the absolute optimal plan. SQL Server does a good job at selecting a good plan. The relational algebra trees were explained from the basics to logical , physical  tree concepts answering the question what is a query?  More details on previous talks on this subject matter http://sqlbits.com/Speakers/Conor_Cunningham

The last 2 sessions I attended covered Windows 2012 Infrastructure for SQL Server and Cloud Ready Data Services.

All in all an excellent conference with great sessions and lots of exchanging SQL views with like-minded people.

Tuesday, 2 October 2012

Day 2 PASS SQL Rally Nordic

The main conference started with a keynote delivered by Kamal Hathi BI on Big Data. This started with a historical look at database\ data progression. With social and web analytics and live data fields, how do we predict future outcomes? Utilising advanced analytics on big data this can be visually displayed with the creation of a storyboard connecting the many data sets such as SQL Server, Hadoop, Excel etc. with the Microsoft toolset. There is no one single tool but a collection of tools for the job.

Other sessions I attended were from the DBA track which covered extended events, clustering and memory management for SQL Server 2012.

An interesting session shared technical information on building highly scalable and available cloud applications. Key consideration for the architectural design
  • Data Warehousing is not a good fit for the cloud
  • Scale out not Scale up
  • Everything has a limit
  • Design for failure
  • Design for continuity
  • Optimise for density
The key to a successful cloud, distributed computing system, is all in the architectural design. A phase from the session “Telemetry is Life” left me thinking about patterns and practices.  The end of another  great day.

Monday, 1 October 2012

PASS SQL Rally Nordic 2012 Day 1


The PASS SQL Rally Nordic from October 1-3 2012 in Copenhagen started with 3 pre-conference seminars. I attended a session covering the Microsoft Business Intelligence (BI) Presentation layer taken by Stacia Misner. This excellent session covered the theory of the BI Maturity Model from TDWI by Wayne Eckerson, a roadmap to analytical completion and collaborative decision making. Then this was mapped to the BI stack of SharePoint, SQL Server Reporting Services, PowerView, PowerPivot, Performance Point and Excel.

Reading Material from the session
Book on Competing on Analytics: The New Science of Winning  Thomas H. Davenport and Jeanne G. Harris

Friday, 21 September 2012

24 hours of PASS

24 hours of PASS provided 24 back to back hours of free SQL Server Training. The event ran from Thursday 20 September 2012 until Friday 21 September 2012. There were some amazing sessions again http://www.sqlpass.org/24hours/fall2012/ . This event gave you a glimpse into the level of forthcoming technical content to be covered at the PASS Summit 2012.

I was very interested in the session delivered by Denny Lee on the Introduction to Microsoft’s Big Data Platform and Hadoop Primer. He defined Big Data as 4V's volume, velocity, variability and variety utilising techniques and technologies that make handling big data at extreme scale economical. Big Data is allowing us to ask new sets of questions. He discussed Scale up and Scale out commoditized distribution.

Hadoop
 
He moved onto discussing what is Hadoop explaining that a lot of data is machine generated these days and the data is loaded first then modelled after.  The infrastructure allows automatic distribution and replication across nodes. Hadoop was based on the Yahoo Nutch project in 2003 and renamed to Hadoop in 2005. A reference book to look at is Tom White's Hadoop: The Definitive Guide.

To find out more about Microsoft's product  https://www.hadooponazure.com/ .