{"id":2464,"date":"2016-09-30T00:00:00","date_gmt":"2016-09-30T00:00:00","guid":{"rendered":""},"modified":"2019-01-23T16:49:48","modified_gmt":"2019-01-23T16:49:48","slug":"strata-word-cloud-2012-vs-2016-data-lakes-spark-real-time-and-other-trends","status":"publish","type":"post","link":"https:\/\/cazenasite.com\/?p=2464","title":{"rendered":"The Strata Word Cloud 2012 vs 2016: Data Lakes, Spark, Real-Time and other Trends"},"content":{"rendered":"<p><img decoding=\"async\" src=\"\/wp-content\/uploads\/2019\/01\/strata_comparison_sm1.png\" style=\"float: right; max-width: 320px; margin-left: 10px;\"\/><\/p>\n<p>This week Cazena made a major announcement, un-coincidentally timed with Strata + Hadoop World. We seriously enhanced our Data Lake as a Service, which is based on Cloudera Enterprise, runs on Microsoft Azure or AWS, and includes many new features for data science. &nbsp;Read more <a href=\"http:\/\/www.cazena.com\/news\/dlaas-on-azure\" target=\"_blank\">here<\/a>. It\u2019s been exciting to see the momentum in the Big Data as a Service category and I loved sharing the news at Strata. Walking through buzzy hum of expo floor conversations, I overheard the same terms over and over. It began to feel like floating through a word cloud. That gave me an idea \u2013 and with it, I uncovered some surprising trends.&nbsp;<\/p>\n<p>I scraped the session descriptions from the 2012 schedule in Santa Clara, the first year Strata merged with Hadoop World. Then I did the same for session descriptions in 2016 in NYC. (I know, analysts, I probably already have sample problems. But this is just for fun.) Then I put the text into a word cloud generator (Tagul is cool!), did some lightweight data massaging and output the graphics. It\u2019s not super scientific, but it\u2019s an interesting back-of-the-napkin look at trends and passes the gut test for industry regulars.<\/p>\n<p><strong>What\u2019s New? Data Lakes and Spark. <\/strong>Big data technology innovation moves quickly. Neither of these terms appeared in any session descriptions in 2012. The new in-memory engine Spark made a huge impact this year, having only debuted in 2014. Now it\u2019s often considered a must-have (and Cazena has it \u201cas a Service,\u201d BTW.) But perhaps the biggest splash was made by the \u201cdata lake,\u201d a Hadoop-based repository for storing raw data. The term was around in 2012, but didn\u2019t make any Strata session descriptions. Now data lakes are all over the place, and at the peak of the hype cycle this year. Cazena also sees lots of interest, inspiring the recent updates to our cloud-based Data Lake as a Service.&nbsp;<\/p>\n<p><strong>What\u2019s Hot? Cloud, Real-time, Stream, IOT, AI, Machine Learning and\u2026Data Warehousing?!<\/strong>&nbsp;These terms ranked in 2012, but were used much more frequently in 2016. Clearly, cloud has come a long way in four years, and everything has gotten bigger and faster. The desire for \u201creal-time,\u201d a term that barely ranked in 2012, has driven technology innovation in streaming, especially in Kafka, a stream processing platform. Analytics have gotten much more sophisticated, with more talk about machine learning and artificial intelligence. Unsurprisingly, Internet of Things (IOT) get a lot more lip service. While IOT only had a couple of dedicated sessions this year, many referenced the potential and challenge of IOT sensor data.<\/p>\n<p>The biggest surprise was more sessions mentioning data warehousing at Strata this year, way more than four years ago. That\u2019s likely due to the conference expanding topically, but also the desire to apply new technologies, like cloud and Hadoop, to aging data warehouse architectures. Cazena knows this challenge well. Many clients are augmenting aging appliances with Big Data as a Service, by migrating data science or other workloads to the cloud. That\u2019s why Cazena\u2019s Data Mart\/Warehouse as a Service leverages technologies like Pivotal Greenplum and Amazon Redshift, runs on Microsoft Azure and AWS, supports standard BI\/analytics tools and includes the Cazena Gateway for secure integration with on-premises systems. Just sayin\u2019.<\/p>\n<p><strong>What\u2019s Still Cool? Big Data, Hadoop, Data Science, Pipelines.&nbsp;<\/strong>The conference\u2019s focus on Big Data and Hadoop remains clear \u2013 the relative use of terms was consistent year over year. I was a bit surprised to find the same pattern with the terms data science and data scientist. It feels like much more emphasis on this role recently. That may be because there are more robust tools emerging to help (wink, wink) \u2013 and because this conference has focused on the data science since it\u2019s early days. On a smaller scale, the concept of data \u201cpipeline\u201d was also used about equally this year and four years ago.<\/p>\n<p><strong>What Terms Are Fading? Social, MapReduce<\/strong> (?!) While IOT is the big data example du jour now, social data was the big thing four years ago. There are definitely social analytics success stories, often for consumer brands, but attention seems to have shifted to other, shinier objects for now. Interestingly, \u201cMapReduce\u201d appeared often in 2012 Strata session descriptions and was notably absent this year. It\u2019s clearly still very much in use, but perhaps it\u2019s that there more tools, services and platforms abstracting that complex technology. Like\u2026you guessed it, Cazena\u2019s Data Lake as a Service.<\/p>\n<p>The meta message? The tech world changes quickly. That\u2019s why Big Data as a Service is transformational. Instead of tracking these rapid-fire changes, figuring out what\u2019s significant, testing, training and deploying, now you can simply plug your favorite tools and datasets into a cloud service like Cazena. We bring the benefits of new technologies to you, without all the complexity. It may be early days for the Big Data as a Service term\u2026but as you can see, things change fast around here.&nbsp;<\/p>\n","protected":false},"excerpt":{"rendered":"<p>This week Cazena made a major announcement, un-coincidentally timed with Strata + Hadoop World. We seriously enhanced our Data Lake as a Service, which is based on Cloudera Enterprise, runs on Microsoft Azure or AWS, and includes many new features for data science. &nbsp;Read more here. It\u2019s been exciting to see the momentum in the [&hellip;]<\/p>\n","protected":false},"author":9,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[5],"tags":[45,46],"class_list":["post-2464","post","type-post","status-publish","format-standard","hentry","category-blog","tag-bi-analytics","tag-technical"],"_links":{"self":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/posts\/2464","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/users\/9"}],"replies":[{"embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=2464"}],"version-history":[{"count":0,"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/posts\/2464\/revisions"}],"wp:attachment":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=2464"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=2464"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=2464"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}