MapR Adds Apache Spark Stack to Hadoop Distribution
And partners with Databricks to provide support and work on Spark and Hadoop
This is a Press Release edited by StorageNewsletter.com on April 18, 2014 at 2:50 pmMapR Technologies, Inc. announced a partnership with Databricks and the addition of the Apache Spark technology stack to the MapR Distribution.
The Spark in-memory processing framework provides programming ease, and real-time processing advantages.
“It has become clear that Apache Spark offers a combination of high-performance in-memory data processing and multiple computation models that is well suited to serving as the basis of next-generation data processing platforms,” commented Matt Aslett, research director, data platforms and analytics, 451 Research. “MapR’s support for the complete Spark stack, combined with its partnership with Databricks, should give Hadoop users the confidence to start developing applications to take advantage of Spark’s performance and flexibility.“
Organizations are looking for ways to derive value from big data. Spark improves both performance and developer productivity:
-
Performance. Spark provides a general-purpose execution framework with in-memory pipelining to speed up end-to-end application performance. For many applications, this results in a 5-100x performance improvement.
-
Developer productivity. Spark jobs can require as little as 1/5th the number of lines of code. It provides a programming abstraction allowing developers to design applications as operations on data collections (known as RDDs, or Resilient Distributed Dataset). Developers can build these applications in multiple programming languages, including Java, Scala and Python, and the same code can be reused across batch, interactive and streaming applications.
Many organizations are running Spark in production in their MapR environments. Their Spark-based applications benefit from the dependability and performance of the MapR Distribution, and from the ability to process real-time operational data thanks to the direct access NFS interface.
With the inclusion of the Spark stack in the distribution, MapR customers can obtain 24×7 support for all projects in the Spark stack. In addition, both companies are joining forces to drive the roadmap and accelerate innovation on these projects. This will benefit customers and the broader Hadoop community over the coming years, starting with the upcoming release of Apache Spark 1.0.
“As the driving force behind Spark, Databricks is thrilled to enter into such a strategic partnership with MapR,” said Ion Stoica, CEO, Databricks. “We are looking forward to combine MapR’s enterprise-grade dependability and performance with Spark, the next generation big data engine, to enable enterprises to unlock deeper insights from their data, faster.“
“The open source community is developing tremendous technology innovations at a rapid pace,” said John Schroeder, CEO and cofounder, MapR. “MapR provides a future-proof investment for our customers with the most open distribution to give them flexibility to pick the right solution with the widest range of compute frameworks and libraries.“
The Spark stack includes five Apache open source projects. With the introduction of the Spark stack, the distribution includes more than 20 Apache open source projects. These projects support a range of use cases, including batch, interactive, streaming, graph and machine leaning. It is the only distribution with a monthly release cadence for Apache open source projects. This enables customers to upgrade components of the distribution without upgrading the entire cluster. It also enables faster adoption of open source innovation while reducing the risk of disruption to the production cluster services.