{"id":2472,"date":"2017-04-18T13:40:07","date_gmt":"2017-04-18T13:40:07","guid":{"rendered":""},"modified":"2018-05-09T13:38:17","modified_gmt":"2018-05-09T13:38:17","slug":"choosing-right-infrastructure-hadoop","status":"publish","type":"post","link":"https:\/\/cazenasite.com\/?p=2472","title":{"rendered":"Choosing the Right Infrastructure for Hadoop"},"content":{"rendered":"<p><img decoding=\"async\" src=\"\/wp-content\/uploads\/2018\/05\/horton-overload.png\" style=\"float: right; max-width: 320px; margin-left: 10px;\"\/><\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\"><em><strong>Series: The Hidden Challenges of Putting Hadoop and Spark in Production<\/strong><\/em><\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\"><em>A recent Gartner survey estimates that only 14% of Hadoop deployments are in production. We\u2019re not surprised. We\u2019ve been in many conversations with companies that have been piloting Hadoop to bolster their analytic capabilities beyond relational databases. These companies want to use new distributed processing frameworks, like Apache Spark, which allow them to process data more efficiently, especially data that is inconsistently or variably structured.<\/em><\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\"><em>The pilots typically involve finding a few existing servers within the datacenter and&nbsp;repurposing&nbsp;them to support Hadoop. There is often minimal infrastructure investment for trying out these new technologies, and sometimes, it\u2019s not even an official project.<\/em><\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\"><em>But the stories often continue similarly: The pilot shows promising results, and a series of meetings leads to a set of potential use cases. A project is established, the clock starts ticking, and the realization sets in: &nbsp;<\/em><\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\"><em>Building a production-grade configuration for Hadoop is a non-trivial exercise. Whether on-premises in a local data center, or in the public cloud, putting Hadoop in production requires expertise and experience to get right.<\/em><\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\"><em>Common challenges fall into a few important categories, which we\u2019ll explore in this blog series:<\/em><\/p>\n<ul style=\"font-size: 16px; font-family: proxima-nova, Arial, sans-serif; font-style: normal; font-variant-caps: normal;\">\n<li><strong><em>Infrastructure: Choosing and configuring servers for Hadoop<\/em><\/strong><\/li>\n<li><a href=\"http:\/\/www.cazena.com\/blog\/hadoop-performance-sizing-and-scaling\" target=\"_blank\"><em>Performance Optimization: Scaling and tuning Hadoop for price-performance<\/em><\/a><\/li>\n<li><a href=\"\/blog\/hadoop-cloud-flexible-challenging\" target=\"_blank\"><em>To Cloud or Not: Selecting, configuring and new challenges<\/em><\/a><\/li>\n<li><a href=\"http:\/\/www.cazena.com\/blog\/cloud-security-big-data-service\" target=\"_blank\"><em>Security: What to consider<\/em><\/a><\/li>\n<\/ul>\n<hr style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal; width: 1019px;\" \/>\n<p class=\"rtecenter\" style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\"><span style=\"font-size: 20px;\"><strong>Choosing the Right Infrastructure for Hadoop<\/strong><\/span><\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">First, you need to know what type of machines will support the Hadoop cluster. In the single-engine database world (Oracle,&nbsp;GPDB,&nbsp;Vertica, etc.), this is a constrained problem. Yet in the Hadoop world, it is not as simple. That\u2019s because the Hadoop ecosystem has multiple processing engines and components from multiple vendors and open source projects. The list is long: Apache,&nbsp;Hortonworks,&nbsp;Cloudera, Spark, Impala,&nbsp;MapReduce, etc. The different engines are optimized for different types of workloads.<\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">One company we spoke with started their pilot using&nbsp;MapReduce&nbsp;(M\/R) as the primary engine. However, they decided that&nbsp;Cloudera&nbsp;Spark\u2019s in-memory capabilities made it the faster, more appropriate processing engine for the problem they needed to solve. That\u2019s when they learned the hard way that the optimal infrastructure for a M\/R dominant&nbsp;Cloudera&nbsp;Hadoop configuration is&nbsp;<em>different<\/em>&nbsp;from a Spark dominant one due to the additional memory requirements. That meant they needed to reconfigure the existing hardware by buying new memory modules, complicating and slowing down the project.<\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\"><strong>Keeping up with a (Very) Rapidly Evolving Ecosystem<\/strong><\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">The underlying technologies in the Hadoop ecosystem are also rapidly evolving and forking-off as new projects start and vendors add in their own components. This makes things even harder to manage and plan for. Consider not long ago, in 2015, Impala (Cloudera\u2019s SQL engine for Hadoop), did not exploit multiple cores. That\u2019s a big deal when you\u2019re trying to optimize performance. Building a configuration with the best price for performance for that Impala release meant lots of machines with very few cores. But today&nbsp;Cloudera&nbsp;Impala exploits multiple cores \u2013 so the new configuration for maximum performance at the lowest cost would be fewer machines with more cores. Totally different. That\u2019s just one example of how rapidly technology is evolving.<\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">If you\u2019re trying to keep up with that in your datacenter, you\u2019re constantly optimizing. It becomes incredibly difficult to choose the right type of server for on-premises deployments and manage capacity planning. In enterprise environments, servers are typically expected to last a minimum of 3 years, with the norm being closer to 5 years. Without a crystal ball, many companies overbuy hardware. They get really powerful machines \u201cjust to be safe,\u201d even though that power and capacity might sit idle for years.<\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">In our <a href=\"http:\/\/www.cazena.com\/blog\/hadoop-performance-sizing-and-scaling\" target=\"_blank\">next blog<\/a>, we\u2019ll talk more about a closely related subject: Hadoop performance, sizing and scaling.<\/p>\n<p style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 12px; font-style: normal; font-variant-caps: normal;\"><em>Original artwork by Carlos Joaquin&nbsp;in collaboration with Cazena.<br \/>\nApache\u00ae, <a href=\"http:\/\/hadoop.apache.org\/\" target=\"_blank\">Apache Hadoop, Hadoop\u00ae<\/a>, and the yellow elephant logo are either registered trademarks or trademarks of the <a href=\"http:\/\/www.apache.org\/\" target=\"_blank\">Apache Software <\/a><a href=\"http:\/\/www.apache.org\/\" target=\"_blank\">Foundation<\/a>in the United States and\/or other countries.<\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p><span style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">In the first installment of our series on The Hidden Challenges of Putting Hadoop and Spark in Production, we explore the infrastructure selection process. This series was inspired after we read a recent Gartner survey, which estimates that only 14% of Hadoop deployments are in production. We\u2019re not surprised&#8230;.<\/span><br \/>\n&nbsp;<\/p>\n","protected":false},"author":12,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[5],"tags":[20,22,46],"class_list":["post-2472","post","type-post","status-publish","format-standard","hentry","category-blog","tag-cloudera","tag-data-science","tag-technical"],"_links":{"self":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/posts\/2472","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/users\/12"}],"replies":[{"embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=2472"}],"version-history":[{"count":0,"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/posts\/2472\/revisions"}],"wp:attachment":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=2472"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=2472"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=2472"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}