{"id":2474,"date":"2017-06-19T00:00:00","date_gmt":"2017-06-19T00:00:00","guid":{"rendered":""},"modified":"2018-05-09T12:43:33","modified_gmt":"2018-05-09T12:43:33","slug":"hadoop-cloud-flexible-challenging","status":"publish","type":"post","link":"https:\/\/cazenasite.com\/?p=2474","title":{"rendered":"Hadoop in the Cloud: Flexible, but Challenging"},"content":{"rendered":"<p><img decoding=\"async\" src=\"\/wp-content\/uploads\/2018\/05\/Yellow-question.png\" style=\"float: right; max-width: 320px; margin-left: 10px;\"\/><\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\"><em><strong>&nbsp;Series:&nbsp;<\/strong><\/em><em style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-variant-caps: normal;\"><strong>The<\/strong><\/em><em style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-variant-caps: normal;\"><strong>&nbsp;Hidden Challenges of Putting Hadoop and Spark in Production<\/strong><\/em><\/p>\n<p><em>A recent Gartner survey estimates that only 14% of Hadoop deployments are in production. We\u2019re not surprised. We\u2019ve been in many conversations with companies that have been piloting Hadoop to bolster their analytic capabilities beyond relational databases. Common challenges fall into a few important categories, which we explore in this blog series:<\/em><\/p>\n<ul style=\"font-size: 16px; font-family: proxima-nova, Arial, sans-serif; font-style: normal; font-variant-caps: normal;\">\n<li><a href=\"http:\/\/www.cazena.com\/blog\/choosing-right-infrastructure-hadoop\" target=\"_blank\"><em>Infrastructure: Choosing and configuring servers for Hadoop<\/em><\/a><\/li>\n<li><a href=\"http:\/\/www.cazena.com\/blog\/hadoop-performance-sizing-and-scaling\" target=\"_blank\"><em>Performance Optimization: Scaling and tuning Hadoop for price-performance<\/em><\/a><\/li>\n<li><strong><em>To Cloud or Not: Selecting, configuring and new challenges<\/em><\/strong><\/li>\n<li><a href=\"http:\/\/www.cazena.com\/blog\/cloud-security-big-data-service\" target=\"_blank\"><em>Security: What to consider<\/em><\/a><\/li>\n<\/ul>\n<hr style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal; width: 1055px;\" \/>\n<p class=\"rtecenter\" style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\"><span style=\"font-size: 20px;\"><strong>Hadoop in the Cloud: Flexible, but Challenging<\/strong><\/span><\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">Many enterprises are asking tough \u2013 but reasonable \u2013 questions about Hadoop, Spark and many other new data technologies:<\/p>\n<p class=\"rteindent1\" style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">How do I get a functional production environment more quickly? What do I buy vs. build? How can services help?<\/p>\n<p class=\"rteindent1\" style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">How do I deal with the shifting open-source software landscape in a sustainable manner? Do I bet on the current promising processing engine \u2013 how do I hedge my bets?<\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">As our series introduction stated, one&nbsp;survey estimated&nbsp;that only 14% of Hadoop deployments are in production.That\u2019s because&nbsp;in most corporate environments, production-ready means that the cluster is secure and has the operational processes established for access, governance, compliance and service-level agreements (SLAs.) At the most basic level, companies must have a plan for how they will monitor infrastructure, software components&nbsp;and overall system health.<\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">But they also need to plan for processes like patching and upgrading, which is complex with an ecosystem like Hadoop, Spark and related technologies, evolving at a rapid pace. Ensuring that you keep a record of&nbsp;the hundreds&nbsp;of configurations that have been applied to each layer of the stack is mandatory if you want to recover from a disaster&#8230;<\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">This is especially challenging in enterprises, which is why we often hear of initial production Hadoop environments taking nine to 12&nbsp;months&nbsp;(at best!). It\u2019s the operational processes and maintenance that can turn into a much bigger headache than the software deployment.<\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\"><strong>Leveraging the cloud&nbsp;for Hadoop: Helpful, not a panacea<\/strong><\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">Deploying Hadoop and Spark in the cloud means that there is no long-term commitment to either the server type that you choose for your data nodes nor the number of nodes that you start with. There is little need to have a configuration that is configured beyond the project(s) that are currently funded and are in motion. This implies that sizing your cluster can be based on known variables for existing projects rather than trying to look out into the future and extrapolate usage. That\u2019s helpful.<\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">A second advantage is that you have flexibility if you get the initially sizing incorrect. Taking away a node as you have over provisioned is a viable option. Try doing that on premises! The cloud gives you a lot of flexibility when it comes to the infrastructure layer,&nbsp;<strong>but<\/strong>&nbsp;you still need to lay down the software in the right manner and keep that overall configuration operational.<\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\"><strong>Configuring and Managing Hadoop in the Cloud<\/strong><\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">Configuring Hadoop and Spark in the cloud, and keeping it operational, is time-consuming and where the \u201cas a service\u201d model starts to really show its&nbsp;value. If you\u2019ve not yet run your own production Hadoop environment, ask a friend. There\u2019s a growing understanding of the care and feeding of a cluster, cloud or not.&nbsp;<\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">That\u2019s why having <strong>automated processes<\/strong>,&nbsp;which&nbsp;can ensure that configurations can be repeated and verified, are a critical step for a&nbsp;production environment. You must have processes&nbsp;to conform to the various compliance attestations that production environments need to have\u2013and it\u2019s not easy&nbsp;to DIY:<\/p>\n<ul>\n<li style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">Can you guarantee that your cluster is Kerberos enabled and configured in the same way each time?<\/li>\n<li style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">Can you guarantee the cluster is configured appropriately with gateway nodes to ensure that you have control on all service access?<\/li>\n<li style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">What happens when one of the cloud VM\u2019s that supports a data node gets sick or dies all together?<\/li>\n<\/ul>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">And if all of that sounds scary or techie to you, are you really ready to DIY in the cloud?<\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">All of these operational responsibilities are necessary for Hadoop in the cloud. They are time&nbsp;consuming, and can take away from valuable staff time for analytics. All of this does not even touch on security which is a topic all by itself.<\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">But the good news is, there are services like Cazena\u2019s that can get you to production with <a href=\"\/what-is-cazena\">Hadoop in the cloud<\/a> quickly, without worrying about any of the things above. That\u2019s the trend for many of the data science, advanced analytics and business units we\u2019ve been interacting with. It didn&#8217;t&nbsp;take long for them to realize the value&nbsp;of a fully-managed platform. You get to the analytic tasks faster and dedicate more of your team to&nbsp;analytics, rather than dev-ops and keeping the lights on.<\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">&nbsp;<\/p>\n","protected":false},"excerpt":{"rendered":"<p><span style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">As part of our ongoing series on&nbsp;<\/span>Productionizing<span style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">&nbsp;Hadoop and Spark in the cloud, we explore performance optimization, and how companies scale and tune for the best performance. We also discuss what\u2019s required for production-grade deployments, often an underestimated part of the process.<br \/>\n&nbsp;<\/span><\/p>\n","protected":false},"author":12,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[5],"tags":[22,46],"class_list":["post-2474","post","type-post","status-publish","format-standard","hentry","category-blog","tag-data-science","tag-technical"],"_links":{"self":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/posts\/2474","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/users\/12"}],"replies":[{"embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=2474"}],"version-history":[{"count":0,"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/posts\/2474\/revisions"}],"wp:attachment":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=2474"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=2474"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=2474"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}