The Fifth Elephant 2015

A conference on data, machine learning, and distributed and parallel computing

Anatomy of RDD : A Deep dive into Spark RDD Data structure.

Submitted by Madhukara Phatak (@phatak-dev) on Wednesday, 6 May 2015

videocam_off

Technical level

Advanced

Section

Full Talk

Status

Submitted

Vote on this proposal

Login to vote

Total votes:  +14

Objective

RDD is the core abstraction of Apache Spark. So understanding RDD in depth is very
crucial to use spark very effectively. This talks aims to take audience a deep
dive into RDD to make them understand why it’s so powerful.

Description

This is an Advanced talks aimed toward people who already know Spark. This talk
tries to deconstruct RDD abstraction to peek inside. We will be discussing about

  • Immutability and Distribution
  • Partitions
  • Partition API’s like mapParittions, lookUp etc
  • Implementation of Laziness
  • RDD dependency hierarchy
  • Transformation and Action implementation
  • Caching implementation

All the above topics are discussed with real code.

Requirements

  • Prior experience of Working with Spark

Speaker bio

Madhukara phatatak is a Bigdata consultant @ Datamantra. He has been actively working in Hadoop,Spark and its ecosystem projects from last 5 years.

He was lead developer of Nectar, a ML library for hadoop.He also contributed to hadoop source code to improve cyclic checks in Jobcontrol api.With raise of Apache Spark, he with his team has open sourced courseera machine learning course examples on spark here. He blogs on spark here. Also he runs a Spark meetup group in Bangalore.

Links

Slides

http://www.slideshare.net/datamantra/anatomy-of-rdd

Comments

  • 0
    sharmila m 2 months ago

    56y54

Login with Twitter or Google to leave a comment