Big data analysis with Apache Spark

Jul 2015

13 Mon

14 Tue

15 Wed

16 Thu 08:30 AM – 06:35 PM IST

17 Fri 08:30 AM – 06:30 PM IST

18 Sat 09:00 AM – 06:30 PM IST

19 Sun

Make a submission

NIMHANS Convention center

Machine Learning, Distributed and Parallel Computing, and High-performance Computing are the themes for this year’s edition of Fifth Elephant.

The deadline for submitting a proposal is 15th June 2015

We are looking for talks and workshops from academics and practitioners who are in the business of making sense of data, big and small.

Track 1: Discovering Insights and Driving Decisions

This track is about general, novel, fundamental, and advanced techniques for making sense of data and driving decisions from data. This could encompass applications of the following ML paradigms:

Statistical Visualizations
Unsupervised Learning
Supervised Learning
Semi-Supervised Learning
Active Learning
Reinforcement Learning
Monte-carlo techniques and probabilistic programming
Deep Learning

Across various data modalities including multi-variate, text, speech, time series, images, video, transactions, etc.

Track 2: Speed at Scale

This track is about tools and processes for collecting, indexing, and processing vast amounts of data. The theme includes:

Distributed and Parallel Computing
Real Time Analytics and Stream Processing
MapReduce and Graph Computing frameworks
Kafka, Spark, Hadoop, MPI
Stories of parallelizing sequential programs
Cost/Security/Disaster Management of Data

HasGeek believes in open source as the binding force of our community. If you are describing a codebase for developers to work with, we’d like it to be available under a permissive open source license. If your software is commercially licensed or available under a combination of commercial and restrictive open source licenses (such as the various forms of the GPL), please consider picking up a sponsorship. We recognize that there are valid reasons for commercial licensing, but ask that you support us in return for giving you an audience. Your session will be marked on the schedule as a sponsored session.

Workshops

If you are interested in conducting a hands-on session on any of the topics falling under the themes of the two tracks described above, please submit a proposal under the workshops section. We also need you to tell us about your past experience in teaching and/or conducting workshops.

Hosted by

The Fifth Elephant

The Fifth Elephant - known as one of the best data science and Machine Learning conference in Asia - has transitioned into a year-round forum for conversations about data and ML engineering; data science in production; data security and privacy practices. more

All submissions

Previous Next

Big data analysis with Apache Spark

Submitted May 6, 2015

Section: Workshop Technical level: Beginner

Apache Spark is a new upcoming big data processing engine. It’s getting popular for it’s of ease of use and it’s unification of different big data work load. The objective this workshop is to get your hands dirty with it.

Outline

We will go over the following in the workshop

Evolution of Big data systems
Why Apache Spark?
Apache Spark architecture
Installing Spark
Working with Spark REPL
Working with RDD’s
RDD Map/Reduce
RDD examples in Scala
Using Spark for streaming data
Spark SQL

Requirements

General Programming Knowledge
General knowledge on big data systems like Hadoop

Need a laptop with Spark installed. I will share specific steps for installation near to the workshop.

Speaker bio

Madhukara phatatak is a Bigdata consultant @ Datamantra. He has been actively working in Hadoop,Spark and its ecosystem projects from last 5 years.

He was lead developer of Nectar, a ML library for hadoop.He also contributed to hadoop source code to improve cyclic checks in Jobcontrol api.With raise of Apache Spark, he with his team has open sourced courseera machine learning course examples on spark here. He blogs on spark here. Also he runs a Spark meetup group in Bangalore.

Comments

Jul 2015

13 Mon

14 Tue

15 Wed

16 Thu 08:30 AM – 06:35 PM IST

17 Fri 08:30 AM – 06:30 PM IST

18 Sat 09:00 AM – 06:30 PM IST

19 Sun

Make a submission

NIMHANS Convention center

Hosted by

The Fifth Elephant

The Fifth Elephant 2015

Track 1: Discovering Insights and Driving Decisions

Track 2: Speed at Scale

Commitment to Open Source

Workshops

Big data analysis with Apache Spark

Outline

Requirements

Speaker bio

Links

Comments