{"id":2497,"date":"2020-05-01T00:00:00","date_gmt":"2020-05-01T00:00:00","guid":{"rendered":"https:\/\/live-cazena.pantheonsite.io\/451-research-matt-aslett-cloud-data-lakes\/"},"modified":"2020-06-04T01:55:01","modified_gmt":"2020-06-04T01:55:01","slug":"451-research-guest-blog-cloud-data-lakes","status":"publish","type":"post","link":"https:\/\/cazenasite.com\/451-research-matt-aslett-cloud-data-lakes\/","title":{"rendered":"451 Research&#8217;s Matt Aslett on Simplifying Cloud Data Lakes"},"content":{"rendered":"<p><em>Cazena is proud to publish\u00a0this exclusive insight\u00a0from <a href=\"https:\/\/451research.com\/analyst-team\/analyst\/Matt+Aslett\" target=\"_blank\" rel=\"noopener noreferrer\">Matt Aslett, 451 Research Vice President<\/a>\u00a0for Data, AI and Analytics.\u00a0<\/em><\/p>\n<p>Despite criticism, the data lake has emerged in the last decade as a solution to managing large volumes of data, and is relatively widely adopted. While simple on paper, the data lake is a lot more difficult to deliver in practice, with multiple moving parts that can fragment without significant manual intervention.<\/p>\n<p>Data processing is shifting to the cloud, but IaaS and PaaS data lake environments still require a lot of manual intervention on the part of the user. True SaaS data lake environments mask the complexity, enabling users to concentrate on data processing and analysis rather than infrastructure management and software orchestration.<\/p>\n<h3>Why data lakes?<\/h3>\n<p>As companies of all shapes and sizes recognize the increasing importance of data and take steps to become more data-driven, one of the key challenges has been how best to rapidly convert large volumes of data into business insight.<\/p>\n<p>It isn\u2019t enough just to store more data from a greater variety of sources \u2013 not just from enterprise applications but also increasingly sensors and the Internet of Things, as well as AI and machine learning. The important thing is how you add knowledge and insight to turn that data into wisdom and unlock its value.<\/p>\n<p>In the last decade, the concept of the data lake emerged as a new approach to storing, processing and analyzing large volumes of data \u2013 particularly the combination of structured and unstructured raw data in its native format \u2013 that can be accessed by multiple users for multiple purposes.<\/p>\n<p>Data from 451 Research\u2019s Voice of the Enterprise illustrates the growth in the adoption of data lake environments. More than 58% of respondents are using data lakes, to some extent, with a further 15% considering adoption.<\/p>\n<h3><em>451 Research:\u00a0Adoption of a Data Lake Environment<\/em><\/h3>\n<p><img decoding=\"async\" style=\"width: 700px; height: 330px;\" src=\"\/wp-content\/uploads\/2019\/11\/cazena_blog_1.png\" alt=\"451 Research - Data Lake Adoption\" \/><\/p>\n<p><em>Source: 451 Research<\/em><\/p>\n<p>The term \u2018data lake\u2019 is credited to Pentaho&#8217;s founder, James Dixon, who first used it in a <a href=\"https:\/\/jamesdixon.wordpress.com\/2010\/10\/14\/pentaho-hadoop-and-data-lakes\/\" target=\"_blank\" rel=\"noopener noreferrer\">2010 blog post<\/a> in which he compared it to a large body of water that is fed from various source streams and that can be accessed by multiple users for multiple purposes. Dixon contrasted the data lake with a data mart, which could be considered the equivalent of bottled water: cleansed and packaged for easy consumption.<\/p>\n<p>The idea was attractive but didn\u2019t explicitly explain what a data lake was technically, or how to build one. Many early adopters set off ingesting data from multiple sources into Hadoop in the hope of proving value further down the line.<\/p>\n<p>Without a clear idea of the requirements or use cases, however, instead of a data lake, several initial projects resulted in the creation of a data swamp \u2013 a single environment housing large volumes of raw data that couldn\u2019t be easily accessed for any purpose, let alone multiple purposes.<\/p>\n<p>A couple of years ago, 451 Research pointed out that what a \u2018data lake\u2019 failed to address, as an analogy, was how multiple users would access that data for multiple purposes, and offered an alternative analogy \u2013 the \u2018<a href=\"https:\/\/clients.451research.com\/reportaction\/79978\/Toc\" target=\"_blank\" rel=\"noopener noreferrer\">data treatment plant<\/a>\u2019 \u2013 arguing that industrial-scale processes were required to make data acceptable for multiple desired end uses and the multiple methods for accessing and processing data.<\/p>\n<p>That wasn\u2019t a battle we were ever likely to win, but while the term \u2018data lake\u2019 has taken off, greater emphasis is now being placed on those industrial-scale processes that are required to turn the data lake concept from theory into reality. This includes data governance functionality, including a data catalog, to create an inventory of what data is in the environment, and also self-service data preparation functionality via which users can discover, integrate, cleanse and enrich data to make it suitable for analysis.<\/p>\n<p>The definition of a data lake has expanded over time, beyond Hadoop, to address Hadoop-compatible cloud storage, and also potentially relational databases as part of what might be considered a logical data lake, as well as optional capabilities such as data virtualization, analytics acceleration and stream-based continuous data integration.<\/p>\n<h3><em>451 Research:\u00a0The Components of a Functioning Data Lake<\/em><\/h3>\n<p><img decoding=\"async\" style=\"width: 700px; height: 300px;\" src=\"\/wp-content\/uploads\/2019\/11\/451Research_datalakes_cazena_blog_2.png\" alt=\"451Research_datalakes_cazena_blog_2.png\" \/><\/p>\n<p><em>Source: 451 Research<\/em><\/p>\n<p>This is all well and good in terms of a theoretical architecture, but it&#8217;s a lot more difficult to deliver in practice. In particular, it is very brittle and has the potential to fragment if not held together with a lot of duct tape and manual intervention.<\/p>\n<p>One potential solution is the cloud, and we have seen an increasing shift toward data workloads being deployed into cloud environments, in part because they offer the potential to lower the configuration and management complexity.<\/p>\n<p>It is important to note, however, that all cloud-based data processing is not created equal. There are significant differences \u2013 in terms of the reduction in complexity \u2013 between IaaS, PaaS and SaaS environments.<\/p>\n<p>IaaS certainly reduces the requirement on the user to configure and manage the underlying infrastructure, but the user still adopts the financial risk of selecting the most appropriate infrastructure resources, and the onus remains on the user to configure and orchestrate the software environment, as well as the data processing jobs.<\/p>\n<p>PaaS reduces this configuration complexity, as well as the financial risk, but there is still a significant requirement on the user to assemble, configure and manage the relevant software services: many PaaS data lake offerings can best be thought of as \u2018blueprints\u2019 for assembling a data lake, rather than the finished article.<\/p>\n<p>For enterprises without the in-house skills to manage and assemble a data lake themselves, SaaS provides a more appropriate option. The service provider takes on the responsibility (both technical and financial) for configuring, managing and orchestrating not just the underlying infrastructure, but also the software environment, enabling users to focus their attention on performing data processing and analytics: converting large volumes of data into business insight.<\/p>\n<h3><em>451 Research: Cloud data processing options.<\/em><\/h3>\n<p><img decoding=\"async\" style=\"width: 700px; height: 359px;\" src=\"\/wp-content\/uploads\/2019\/11\/451Research_datalakes_cazena_blog_3.png\" alt=\"451Research_datalakes_cazena_blog_3.png\" \/><\/p>\n<p><em>Source: 451 Research<\/em><\/p>\n<h3>Conclusion &amp; Resources<\/h3>\n<p>There are multiple technology choices for assembling and running data lakes environments. To learn more about navigating the options and recent advances in cloud technologies, explore the webinar and resources to learn more.<\/p>\n","protected":false},"excerpt":{"rendered":"<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">451 Research Vice President Matt Aslett offers a timely update on cloud data lakes in this guest blog.&nbsp;Despite criticism, the data lake has emerged in the last decade as a solution to managing large volumes of data, and is relatively widely adopted. While simple on paper, the data lake is a lot more difficult to deliver in practice, with multiple moving parts that can fragment without significant manual intervention.&nbsp;Data processing is shifting to the cloud, but&nbsp;IaaS&nbsp;and&nbsp;PaaS&nbsp;data lake environments still require a lot of manual intervention on the part of the user. True SaaS data lake environments mask the complexity, enabling users to concentrate on data processing and analysis rather than infrastructure management and software orchestration.<\/p>\n","protected":false},"author":13,"featured_media":2992,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[5],"tags":[195,153],"class_list":["post-2497","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-blog","tag-analyst-insight","tag-cloud-data-lake"],"_links":{"self":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/posts\/2497","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/users\/13"}],"replies":[{"embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=2497"}],"version-history":[{"count":5,"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/posts\/2497\/revisions"}],"predecessor-version":[{"id":3188,"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/posts\/2497\/revisions\/3188"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/media\/2992"}],"wp:attachment":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=2497"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=2497"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=2497"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}