{"id":2471,"date":"2017-03-23T13:26:37","date_gmt":"2017-03-23T13:26:37","guid":{"rendered":""},"modified":"2019-01-29T11:49:32","modified_gmt":"2019-01-29T11:49:32","slug":"five-hurdles-impacting-data-science-teams-collaboration-and-efficiency","status":"publish","type":"post","link":"https:\/\/cazenasite.com\/?p=2471","title":{"rendered":"Five Hurdles Impacting Data Science Teams\u2019 Collaboration and Efficiency"},"content":{"rendered":"<p><img decoding=\"async\" src=\"\/wp-content\/uploads\/2019\/01\/pitfall_wall_sm.jpg\" style=\"float: right; max-width: 320px; margin-left: 10px;\"\/><\/p>\n<p><a href=\"http:\/\/www.cazena.com\/sandbox\" target=\"_blank\">Cazena\u2019s Data Science Sandbox Test Drive<\/a> has spurred many conversations in the last few weeks, at Strata + Hadoop event in San Jose, meetups and other events. There is an interesting common thread mentioned by the leaders of data science and advanced analytics groups: All are focused on how to make their team as productive as possible. The resources for these teams are notoriously hard to find. So, naturally, team leaders want to ensure that these scarce, highly-skilled workers have everything they need.<\/p>\n<p>Here are the common challenges I\u2019ve heard about recently that are impacting data science team\u2019s ability to collaborate efficiently and productively:<\/p>\n<p><strong>1) Working (too) locally<\/strong><\/p>\n<p>One of the things that these leaders point to as a major efficiency impact is members of their teams working locally on their laptops or desktops. Many data scientists are quite adept at quickly creating complex models by extracting data from various enterprise and non-enterprise systems. My initial instinct was that the \u201cdesktop constraint\u201d discussion was going to lead into a conversation about the challenges of volume and scale with limited local compute and storage resources \u2013 which it did \u2013 but there\u2019s another big issue that\u2019s second on their list of challenges.<\/p>\n<p><strong>2) No Common Environment for Sharing Data, Models and Code Snippets<\/strong><\/p>\n<p>The leaders want their teams to be able to iterate quickly and fail fast. To do this, data science teams need to be able to easily share data, models and code snippets. However, this is where the inefficiencies often start. Since most of these analysts are working locally, on their own desktops, sharing data involves copying or emailing files between team members.<\/p>\n<p>In one extreme case, we spoke to an organization, which had its analytics team scattered across three countries and two continents. The average size of the files they were passing back and forth was only 10GB, but when you were doing this multiple times a day &#8212; the team was typically spending at least an hour a day waiting for files to arrive. One analyst tried to make his day more efficient by scheduling his snack breaks during the times when he was waiting for a file from someone else on his team.<\/p>\n<p>Another common scenario that leaders face is multiple members of their team extracting the same data. This is inefficient at the team level as well as putting unneeded strain on the systems that they extract data from.<\/p>\n<p><strong>3) Serious Library and Version Incompatibilities<\/strong><\/p>\n<p>Working locally also results in library and version incompatibilities. In the company example above, each individual analyst had their own versions of local libraries, which meant that time was required to refactor code snippets they got from other team members, who all had their own, different local libraries.<\/p>\n<p>A vexing example of this was when an analyst sent a code snippet to another team member that referenced a library that the recipient did not have. The recipient then had to download the library which was a newer version to the one originally used. Then the analyst had to refactor the code to work with the newer version. This cycle only continued when that analyst made changes, then sent an updated code snippet back to the original team member, who then had to refactor it again to work with the older version. Even if you don\u2019t know anything about refactoring, that was clearly not an efficient process.<\/p>\n<p><strong>4) Slow and Inefficient processing (especially for high-volume data)<\/strong><\/p>\n<p>While sample sizes can always be adjusted, single-threaded processing is increasingly an issue for data science productivity. Many teams want to use distributed computing, and engines like Apache Spark \u2013 but the time and skills required for implementation are too daunting. Some are allowed \u201ctimeshare\u201d situations on larger clusters, but that can also lead to many manual processes and sometimes inefficient processing if the data can\u2019t be in close proximity to the processing engines.<\/p>\n<p><strong>5) Can\u2019t Leverage Cloud due to Security and Compliance<\/strong><\/p>\n<p>There\u2019s an obvious well-known solution to compute and storage constraints \u2013 aka the public cloud \u2013 but that\u2019s been off limits for many due to the security and compliance policies at their companies. Teams and their leaders haven\u2019t been able to take advantage of the cloud, because it\u2019s time-consuming to address the many requirements of enterprise security. That\u2019s precisely why Cazena\u2019s services are single-tenant and private for each enterprise team.<\/p>\n<p><b>How Cazena helps<\/b><\/p>\n<p>Cazena\u2019s Data Science Sandbox as a Service addresses all of these pitfalls \u2013 and a few more. The immediate value that many leaders see in the Cazena Service is a secure, centralized cloud platform, which allows datasets to be stored in a single place. It\u2019s much more than a secure fileshare though. Using a Cazena platform also gives teams a simple way to use a variety of distributed processing engines (e.g. Spark, MPP) \u2013 all in close proximity to their data for maximum efficiency.<\/p>\n<p>Having centralized tools means that the entire team uses a common version of R and Python. Libraries that the members need are added to a single place so that all analysts have access to the same set of libraries and more importantly, the version of the libraries is consistent. When a library is upgraded to a newer version, <b>all<\/b> members of the team automatically now need to use this new version.<\/p>\n<p>These features of the Cazena platform form the critical foundation that enable data science teams to be productive. Analytic leaders get most excited about those efficiency-enhancing functions. The fact that Cazena also gives teams the flexibility to use any analytic language <span style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">\u2013<\/span> and makes it easy to run analytics across full datasets \u2013 is just the cherry on the top.<\/p>\n<p>Do you have additional examples of simple things that make the daily life of a data scientist more <span style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">\u2013<\/span> or less <span style=\"font-family: proxima-nova, Arial, sans-serif; font-size: 16px; font-style: normal; font-variant-caps: normal;\">\u2013<\/span> productive? I am really interested in learning about those to see if there are additional things we can do to add value to the data science process.<\/p>\n<p>&nbsp;<\/p>\n<p><em>Lovan Chetty is Director of Product Management with Cazena, and&nbsp;can be reached at <a href=\"mailto:lovan@cazena.com\">lovan@cazena.com<\/a>.<\/em><\/p>\n<p>&nbsp;<\/p>\n","protected":false},"excerpt":{"rendered":"<p>There is an interesting theme mentioned by the leaders of data science and advanced analytics groups: All are focused on how to make their team as productive as possible. The resources for these teams are notoriously hard to find. &nbsp;So, naturally, team leaders want to ensure that these scarce, highly-skilled workers have everything they need to be efficient. Here are the most common pitfalls we hear about. Do you agree?<\/p>\n","protected":false},"author":12,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[5],"tags":[22],"class_list":["post-2471","post","type-post","status-publish","format-standard","hentry","category-blog","tag-data-science"],"_links":{"self":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/posts\/2471","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/users\/12"}],"replies":[{"embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=2471"}],"version-history":[{"count":0,"href":"https:\/\/cazenasite.com\/index.php?rest_route=\/wp\/v2\/posts\/2471\/revisions"}],"wp:attachment":[{"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=2471"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=2471"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/cazenasite.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=2471"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}