{"id":2473,"date":"2017-05-25T00:00:00","date_gmt":"2017-05-25T00:00:00","guid":{"rendered":""},"modified":"2018-05-09T13:36:03","modified_gmt":"2018-05-09T13:36:03","slug":"hadoop-performance-sizing-and-scaling","status":"publish","type":"post","link":"https:\/\/cazenasite.com\/?p=2473","title":{"rendered":"Hadoop Performance, Sizing and Scaling"},"content":{"rendered":"<p><img decoding=\"async\" src=\"\/wp-content\/uploads\/2018\/05\/horton-run.png\" style=\"float: right; max-width: 320px; margin-left: 10px;\"\/><\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\"><em><strong>Series<\/strong><strong>:&nbsp;<\/strong><\/em><em style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-variant-caps: normal;\"><strong>The<\/strong><\/em><em style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-variant-caps: normal;\"><strong>&nbsp;Hidden Challenges of Putting Hadoop and Spark in Production<\/strong><\/em><\/p>\n<p><em>A recent Gartner survey estimates that only 14% of Hadoop deployments are in production. We\u2019re not surprised. We\u2019ve been in many conversations with companies that have been piloting Hadoop to bolster their analytic capabilities beyond relational databases. Common challenges fall into a few important categories, which we explore in this blog series:<\/em><\/p>\n<ul style=\"font-size: 16px; font-family: proxima-nova, Arial, sans-serif; font-style: normal; font-variant-caps: normal;\">\n<li><a href=\"http:\/\/www.cazena.com\/blog\/choosing-right-infrastructure-hadoop\" target=\"_blank\"><em>Infrastructure: Choosing and configuring servers for Hadoop<\/em><\/a><\/li>\n<li><strong><em>Performance Optimization: Scaling and tuning Hadoop for price-performance<\/em><\/strong><\/li>\n<li><a href=\"http:\/\/www.cazena.com\/blog\/hadoop-cloud-flexible-challenging\" target=\"_blank\"><em>To Cloud or Not: Selecting, configuring and new challenges<\/em><\/a><\/li>\n<li><a href=\"http:\/\/www.cazena.com\/blog\/cloud-security-big-data-service\" target=\"_blank\"><em>Security: What to consider<\/em><\/a><\/li>\n<\/ul>\n<hr style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal; width: 1055px;\" \/>\n<p class=\"rtecenter\" style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\"><span style=\"font-size: 20px;\"><strong>Hadoop&nbsp;Performance, Sizing and Scaling<\/strong><\/span><\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">Closely related to infrastructure decisions for Hadoop and Spark are the questions about how to size your cluster. Since Hadoop and Spark are newer technologies, often with new workloads, it\u2019s frequently a big unknown. Most organizations don\u2019t have a lot of history (if any) to estimate cluster characteristics \u2013 and it\u2019s not just a matter of data volume estimates. We have seen cases where the initial cluster is grossly undersized due to rapid adoption (arguably a good thing) and other deployments were the cluster is drastically underutilized, often due to operational or business issues. It\u2019s a tough equation.<\/p>\n<p>Contrary to popular belief, managing Hadoop performance doesn&#8217;t&nbsp;automatically get easier if you deploy in the cloud vs. on-premises. In fact, the stakes can be even higher with metered services, where a poor performance configuration can translate to a big bill. There\u2019s no easy advice here. It\u2019s critical to do your research, consider many factors (which components of the Hadoop ecosystem, storage, compute, growth, adoption, etc.) and be ready to experiment.<\/p>\n<p>With the cluster in place, the next step is configuring Hadoop and optimizing it for production. Configuring the cluster has everything to do with the workload, which drives how you configure&nbsp;HDFS, or if you use&nbsp;HDFS&nbsp;at all. After storage, you then keep stepping up the stack and configure Yarn, Spark, Impala and other components. However this is not a one time process. The Hadoop ecosystem is rapidly expanding and evolving, so configuration optimization is an ongoing process.<\/p>\n<p>Let\u2019s use the&nbsp;Cloudera&nbsp;Impala example that I used in the&nbsp;<a href=\"http:\/\/www.cazena.com\/blog\/choosing-right-infrastructure-hadoop\" target=\"_blank\">last blog<\/a>, which described how Impala rapidly evolved to exploit more and more cores on a single machine. It would have been totally feasible to have configured a cluster 24 months ago that had lots of nodes with very few cores. However today, you can get better price-performance using fewer nodes that have more cores.<\/p>\n<p>Another example is the transition from&nbsp;MapReduce&nbsp;(M\/R) to Spark, which usually entails having nodes that have a lot more memory than the optimal node for a M\/R workload. Trying to do this on premise is close to impossible, unless you have unlimited budget and can depreciate hardware in a year or less.<\/p>\n<p><strong>The Cloud Helps (But Is Not a Panacea)<\/strong><\/p>\n<p>Leveraging the public cloud (AWS, Azure, etc) can mitigate these challenges. There is no long-term commitment to either the server type that you choose for your data nodes or the number of nodes that you start with. The cloud helps you truly leverage advances by the chip manufacturers, as well as advances by the cloud platforms, with minimal lag time. And with the cloud, there are no sunk costs due to hardware not being fully depreciated.<\/p>\n<p>Sizing of Hadoop clusters in the cloud can take a totally different approach to what would be traditionally done on-premises. There will obviously need to be some requirements gathering and planning to initially bring up a Hadoop platform that will meet current analytic needs. But there\u2019s room for flexibility, so it\u2019s almost more important to have processes that can monitor and constantly evolve the platform to provide the best price-performance.<\/p>\n<p>The cloud provides us many great tools to dynamically size and configure platforms, however it also adds in some questions. How do I secure this environment that is no longer behind my firewall? How do I monitor and manage this environment to the same level as my existing on-premises systems?<\/p>\n<p>We will explore the process of putting these distributed technologies into production in the next blog.<\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 12px; font-style: normal; font-variant-caps: normal;\"><em>Original artwork by Carlos Joaquin in collaboration with Cazena.<br \/>\nApache\u00ae,&nbsp;<a href=\"http:\/\/hadoop.apache.org\/\" target=\"_blank\">Apache Hadoop, Hadoop\u00ae<\/a>, and the yellow elephant logo are either registered trademarks or trademarks of the&nbsp;<a href=\"http:\/\/www.apache.org\/\" target=\"_blank\">Apache Software&nbsp;<\/a><a href=\"http:\/\/www.apache.org\/\" target=\"_blank\">Foundation<\/a>in&nbsp;the United States and\/or other countries.<\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p><span style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">As part of our ongoing series on Productionizing Hadoop and Spark in the cloud, we explore performance optimization, and how companies scale and tune for the best performance. We also discuss what\u2019s required for production-grade deployments, often an underestimated part of the process.<\/span><\/p>\n","protected":false},"author":12,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[5],"tags":[20,22,46],"class_list":["post-2473","post","type-post","status-publish","format-standard","hentry","category-blog","tag-cloudera","tag-data-science","tag-technical"],"_links":{"self":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/posts\/2473","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/users\/12"}],"replies":[{"embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=2473"}],"version-history":[{"count":0,"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/posts\/2473\/revisions"}],"wp:attachment":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=2473"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=2473"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=2473"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}