{"id":3559,"date":"2020-08-25T19:20:27","date_gmt":"2020-08-25T19:20:27","guid":{"rendered":"https:\/\/www.cazena.com\/?p=3559"},"modified":"2021-01-08T20:38:24","modified_gmt":"2021-01-08T20:38:24","slug":"what-is-a-cloud-data-lake","status":"publish","type":"post","link":"https:\/\/cazenasite.com\/?p=3559","title":{"rendered":"What is a Cloud Data Lake?"},"content":{"rendered":"<p>A <a href=\"https:\/\/jamesdixon.wordpress.com\/2010\/10\/14\/pentaho-hadoop-and-data-lakes\/\" target=\"_blank\" rel=\"noopener noreferrer\">data lake<\/a> is a generalized data processing platform that supports a wider variety of data and analytical processing <a href=\"https:\/\/cazenasite.com\/data-warehouses-vs-data-lakes-differences-blog\/\" target=\"_blank\" rel=\"noopener noreferrer\">above and beyond standard SQL data warehouses<\/a><em>. \u00a0<\/em>For over a decade, enterprises have invested heavily to build on-premises data lakes.\u00a0 However, over the past few years a new trend is emerging, the Cloud Data Lake.<\/p>\n<p>The Cloud Data Lake is a next-generation Data Lake hosted in the cloud that delivers more attractive price\/performance, a variety of analytical engines, best-of-breed tooling, all on virtually unlimited Cloud storage.<\/p>\n<p>Residing in public cloud environments such as <a href=\"https:\/\/aws.amazon.com\/big-data\/datalakes-and-analytics\/what-is-a-data-lake\/\" target=\"_blank\" rel=\"noopener noreferrer\">AWS<\/a> and <a href=\"https:\/\/azure.microsoft.com\/en-us\/solutions\/data-lake\/\" target=\"_blank\" rel=\"noopener noreferrer\">Microsoft Azure<\/a>, the Cloud Data Lake is more than just storage. The Cloud Data Lake is a complete analytical environment that supports a variety of analytical tools and languages (SQL, R, Python, Java, Scala, etc.), supporting a variety of workloads, from traditional analytics, BI, streaming event\/IoT processing, to advanced Machine Learning and AI processing.<\/p>\n<p>Compared to their on-premises counterparts, the Cloud Data Lake brings a set of distinctly different advantages across Storage, Compute, and Cost. With these advantages, however, come new challenges around skills required and the operational complexity of Cloud Data Lakes. This post will explore the advantages and challenges of this new analytical platform.<\/p>\n<h3>Storage<\/h3>\n<p>One of the big challenges with data lake deployments is the growth of data. Data is being created at astonishing rates and piles up quickly.<\/p>\n<p>With on-premises data lakes, you must regularly monitor the data growth within your data lake. As your data grows and approaches your capacity, you must add additional disk drives to existing hardware, or purchase additional compute and storage to expand your cluster even though you might not require the additional compute power.<\/p>\n<p><a href=\"https:\/\/www.cazena.com\/wp-content\/uploads\/2020\/08\/Onpremises-Data-Lake-Image-1.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-3572\" src=\"https:\/\/www.cazena.com\/wp-content\/uploads\/2020\/08\/Onpremises-Data-Lake-Image-1-1024x492.png\" alt=\"\" width=\"800\" height=\"384\" srcset=\"https:\/\/cazenasite.com\/wp-content\/uploads\/2020\/08\/Onpremises-Data-Lake-Image-1-1024x492.png 1024w, https:\/\/cazenasite.com\/wp-content\/uploads\/2020\/08\/Onpremises-Data-Lake-Image-1-300x144.png 300w, https:\/\/cazenasite.com\/wp-content\/uploads\/2020\/08\/Onpremises-Data-Lake-Image-1-768x369.png 768w, https:\/\/cazenasite.com\/wp-content\/uploads\/2020\/08\/Onpremises-Data-Lake-Image-1-1536x738.png 1536w, https:\/\/cazenasite.com\/wp-content\/uploads\/2020\/08\/Onpremises-Data-Lake-Image-1-2048x984.png 2048w\" sizes=\"auto, (max-width: 800px) 100vw, 800px\" \/><\/a><\/p>\n<p>With Cloud Data Lakes, storage is essentially infinite (it is serverless) when using a cloud vendor\u2019s low-cost object storage, such as S3 on AWS and Azure Data Lake Storage (ADLS) on Azure. These storage layers offer many 9s of durability as well as availability and have automatic geo-replication.\u00a0 With limitless capacity, there is no need to capacity plan for data growth.<\/p>\n<h3>Compute<\/h3>\n<p>The separation of storage from compute is significant as it increases the flexibility and capacity of your Cloud Data Lake.<\/p>\n<p><a href=\"https:\/\/www.cazena.com\/wp-content\/uploads\/2020\/08\/Cloud-Data-Lake-Image-1.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-3573\" src=\"https:\/\/www.cazena.com\/wp-content\/uploads\/2020\/08\/Cloud-Data-Lake-Image-1-1024x409.png\" alt=\"\" width=\"800\" height=\"319\" srcset=\"https:\/\/cazenasite.com\/wp-content\/uploads\/2020\/08\/Cloud-Data-Lake-Image-1-1024x409.png 1024w, https:\/\/cazenasite.com\/wp-content\/uploads\/2020\/08\/Cloud-Data-Lake-Image-1-300x120.png 300w, https:\/\/cazenasite.com\/wp-content\/uploads\/2020\/08\/Cloud-Data-Lake-Image-1-768x307.png 768w, https:\/\/cazenasite.com\/wp-content\/uploads\/2020\/08\/Cloud-Data-Lake-Image-1-1536x613.png 1536w, https:\/\/cazenasite.com\/wp-content\/uploads\/2020\/08\/Cloud-Data-Lake-Image-1-2048x818.png 2048w\" sizes=\"auto, (max-width: 800px) 100vw, 800px\" \/><\/a><\/p>\n<p>In a Cloud Data Lake, compute, the analytic engines such as Spark, Hive, Presto or Impala, can be on-demand and compute can be elastic.<\/p>\n<p>Analytics engines can be created on-demand for specific purposes. Different teams can spin up different compute engines for different workloads, such as ETL, ML, ad hoc analytics, etc., on the same shared data. It is not uncommon to spin up a Spark cluster for a few hours to process a pipeline of data then shutdown the compute resources when the pipeline has completed.<\/p>\n<p>The result is that you can provision infrastructure that is optimized for individual workloads, and since the infrastructure can be transient, the result is often is significantly reduced infrastructure cost.<\/p>\n<p>Analytics engines can be configured to increase compute on-demand, adding the power of compute elasticity to your data lake.\u00a0 Often this results in performance SLAs with reduced infrastructure cost over the long run.<\/p>\n<p>These scenarios are practically impossible with on-premises clusters.\u00a0 To increase compute in your data lake, you must add additional hardware.\u00a0 Unless you have it lying around, this means procuring the hardware and installing it; often, this process takes weeks to months. To avoid or at least reduce this delay you must perform regular compute capacity analysis and projections to stay ahead of the game.<\/p>\n<h3>Cost<\/h3>\n<p>With a Cloud Data Lake, you only pay for the compute that you use, if you are not using it, you can easily shut it down, and avoid the wasted expense.<\/p>\n<p>For on-premises data lakes, once you spend the money on new hardware, you own it. If it goes unused it is still a capitalized expense &#8212; one that you may be stuck with 3-5 years. This means that even if there are new options that better fit for your workloads, you can only adopt them during a hardware refresh.<\/p>\n<p>Software licensing costs are similar.\u00a0 With on-premises data lakes, you must buy software licenses and software support contracts, and if you find you are not using the software, that doesn\u2019t matter, you usually can\u2019t get your money back.\u00a0 With Cloud Data Lakes, software and services are usually billed hourly \u2013 if you are not using the service you don\u2019t have to pay for it.<\/p>\n<h3>Cloud Data Lake Challenges<\/h3>\n<p>Enterprises should evaluate the use of Cloud Data Lakes based on their architectural advantages described above. However, Cloud Data Lakes also present new challenges around their complexity and operational skills requirements.<\/p>\n<p>Common challenges for Cloud Data Lakes include long deployment cycles for production, integration issues with on-premises applications and users, ensuring security and compliance, data governance, and managing on-going costs.<\/p>\n<p>Just as an example, around security and compliance, you must put serious thought into securing your data, especially if you are planning to store sensitive data in the cloud data lake.\u00a0 Much like an on-premises solution, you should define encryption of data in motion as well as at rest.\u00a0 Also be sure to not expose your data and services to the internet.\u00a0 This means no public IP addresses and fully auditing access to data and services.<\/p>\n<p>A number of cloud-native or third-party security services and controls must be deployed, integrated and managed for your specific cloud data lake. And once the Cloud Data Lake is provisioned, the work doesn\u2019t stop there.\u00a0 You still need ongoing SecOps to fully secure your data, detect and protect from the ever-present threats to the system.<\/p>\n<p>Enterprises must be aware of these issues and work to address them. Many enterprises will need to augment their teams with new skills and expertise around the new Cloud Data Lake stack. Partners with expertise can help with skills. Also, new <a href=\"https:\/\/cazenasite.com\/cloud-data-lake-solution\">SaaS Cloud Data Lake offerings<\/a> are emerging that accelerate deployments with minimal operational complexity.<\/p>\n<h3>Conclusion<\/h3>\n<p>Cloud Data Lakes deliver significant flexibility at potentially significant cost savings versus traditional on-premises data lakes. Low-cost limitless storage capacity and on-demand flexible compute, where you pay for only the compute you use.\u00a0 Cloud Data Lakes are the next-generation enterprise data platform for all analytical workloads, including BI, ML, and data engineering. As enterprises evaluate Cloud Data Lakes, they must work to address their deployment and operational complexity.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>A data lake is a generalized data processing platform that supports a wider variety of data and analytical processing above and beyond standard SQL data warehouses. \u00a0For over a decade, enterprises have invested heavily to build on-premises data lakes.\u00a0 However, over the past few years a new trend is emerging, the Cloud Data Lake. The [&hellip;]<\/p>\n","protected":false},"author":9,"featured_media":3573,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[5],"tags":[153],"class_list":["post-3559","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-blog","tag-cloud-data-lake"],"_links":{"self":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/posts\/3559","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/users\/9"}],"replies":[{"embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=3559"}],"version-history":[{"count":15,"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/posts\/3559\/revisions"}],"predecessor-version":[{"id":3749,"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/posts\/3559\/revisions\/3749"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/media\/3573"}],"wp:attachment":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=3559"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=3559"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=3559"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}