Explore Products

Best Sellers

New Releases

Books

Videos

Audiobooks

Learning Hub

Free Learning

Frank Kane's Taming Big Data with Apache Spark and Python: Real-world examples to help you analyze large datasets with Apache Spark

Frank Kane

$19.99 per month

3.8 (11 Ratings)

Paperback Jun 2017 296 pages 1st Edition

Frank Kane

$19.99 per month

3.8 (11 Ratings)

Paperback Jun 2017 296 pages 1st Edition

What do you get with a Packt Subscription?

Free for first 7 days. $19.99 p/m after that. Cancel any time!

Unlimited ad-free access to the largest independent learning library in tech. Access this title and thousands more!

50+ new titles added per month, including many first-to-market concepts and exclusive early access to books as they are being written.

Innovative learning tools, including AI book assistants, code context explainers, and text-to-speech.

Thousands of reference materials covering every tech concept you need to stay up to date.

Subscribe now

View plans & pricing

View table of contents

Preview Book

Download Code

Key benefits

Understand how Spark can be distributed across computing clusters
Develop and run Spark jobs efficiently using Python
A hands-on tutorial by Frank Kane with over 15 real-world examples teaching you Big Data processing with Spark

Description

Frank Kane’s Taming Big Data with Apache Spark and Python is your companion to learning Apache Spark in a hands-on manner. Frank will start you off by teaching you how to set up Spark on a single system or on a cluster, and you’ll soon move on to analyzing large data sets using Spark RDD, and developing and running effective Spark jobs quickly using Python. Apache Spark has emerged as the next big thing in the Big Data domain – quickly rising from an ascending technology to an established superstar in just a matter of years. Spark allows you to quickly extract actionable insights from large amounts of data, on a real-time basis, making it an essential tool in many modern businesses. Frank has packed this book with over 15 interactive, fun-filled examples relevant to the real world, and he will empower you to understand the Spark ecosystem and implement production-grade real-time Spark projects with ease.

Who is this book for?

If you are a data scientist or data analyst who wants to learn Big Data processing using Apache Spark and Python, this book is for you. If you have some programming experience in Python, and want to learn how to process large amounts of data using Apache Spark, Frank Kane’s Taming Big Data with Apache Spark and Python will also help you.

What you will learn

Find out how you can identify Big Data problems as Spark problems
Install and run Apache Spark on your computer or on a cluster
Analyze large data sets across many CPUs using Spark's Resilient Distributed Datasets
Implement machine learning on Spark using the MLlib library
Process continuous streams of data in real time using the Spark streaming module
Perform complex network analysis using Spark's GraphX library
Use Amazon s Elastic MapReduce service to run your Spark jobs on a cluster

What do you get with a Packt Subscription?

Free for first 7 days. $19.99 p/m after that. Cancel any time!

Unlimited ad-free access to the largest independent learning library in tech. Access this title and thousands more!

50+ new titles added per month, including many first-to-market concepts and exclusive early access to books as they are being written.

Innovative learning tools, including AI book assistants, code context explainers, and text-to-speech.

Thousands of reference materials covering every tech concept you need to stay up to date.

Subscribe now

View plans & pricing

Product Details

Publication date : Jun 30, 2017

Length: 296 pages

Edition : 1st

Language : English

ISBN-13 : 9781787287945

Category :

Data

Languages :

Python

Concepts :

Big Data

Tools :

Apache Spark

Frequently bought together

Python: End-to-end Data Analysis

May 2017 931 pages

Course

$89.99

Hands-On Data Science and Python Machine Learning

Jul 2017 420 pages

3.8 (4)

eBook

$24.99 ~~$35.99~~

Frank Kane's Taming Big Data with Apache Spark and Python

Jun 2017 296 pages

3.8 (11)

eBook

$24.99 ~~$35.99~~

Total $ 177.97

$89.99

$43.99

Total $ 177.97

7 Chapters

Getting Started with Spark

Getting set up - installing Python, a JDK, and Spark and its dependencies

Installing the MovieLens movie rating dataset

Run your first Spark program - the ratings histogram example

Summary

Spark Basics and Spark Examples

What is Spark?

The Resilient Distributed Dataset (RDD)

Ratings histogram walk-through

Key/value RDDs and the average friends by age example

Running the average friends by age example

Filtering RDDs and the minimum temperature by location example

Running the minimum temperature example and modifying it for maximums

Running the maximum temperature by location example

Counting word occurrences using flatmap()

Improving the word-count script with regular expressions

Sorting the word count results

Find the total amount spent by customer

Check your results and sort them by the total amount spent

Check your sorted implementation and results against mine

Summary

Advanced Examples of Spark Programs

Finding the most popular movie

Using broadcast variables to display movie names instead of ID numbers

Finding the most popular superhero in a social graph

Running the script - discover who the most popular superhero is

Superhero degrees of separation - introducing the breadth-first search algorithm

Accumulators and implementing BFS in Spark

Superhero degrees of separation - review the code and run it

Item-based collaborative filtering in Spark, cache(), and persist()

Improving the quality of the similar movies example

Summary

Running Spark on a Cluster

Introducing Elastic MapReduce

Setting up our Amazon Web Services / Elastic MapReduce account and PuTTY

Partitioning

Troubleshooting Spark on a cluster

More troubleshooting and managing dependencies

Summary

SparkSQL, DataFrames, and DataSets

Introducing SparkSQL

Executing SQL commands and SQL-style functions on a DataFrame

Using DataFrames instead of RDDs

Summary

Other Spark Technologies and Libraries

Introducing MLlib

Using MLlib to produce movie recommendations

Analyzing the ALS recommendations results

Using DataFrames with MLlib

Spark Streaming and GraphX

Summary

Where to Go From Here? – Learning More About Spark and Data Science

Recommendations for you

Python Machine Learning By Example

Jul 2024 518 pages

4.9 (8)

eBook

$24.99 ~~$36.99~~

Unlocking Data with Generative AI and RAG

Sep 2024 346 pages

eBook

$21.99 ~~$31.99~~

Building LLM Powered Applications

May 2024 342 pages

4.2 (22)

eBook

$27.98 ~~$39.99~~

Python for Algorithmic Trading Cookbook

Aug 2024 404 pages

4.6 (15)

eBook

$32.99 ~~$47.99~~

Machine Learning with PyTorch and Scikit-Learn

Feb 2022 774 pages

4.4 (95)

eBook

$29.99 ~~$43.99~~

RAG-Driven Generative AI

Sep 2024 334 pages

4.5 (13)

eBook

$24.99 ~~$35.99~~

Artificial Intelligence Engines

Nov 2024 217 pages

eBook

$6.98 ~~$9.99~~

AI Product Manager's Handbook

Nov 2024 484 pages

eBook

Hands-On Reinforcement Learning with Python

Jun 2018 318 pages

2.6 (18)

eBook

$20.98 ~~$29.99~~

LLM Engineer's Handbook

Oct 2024 522 pages

4.8 (13)

eBook

$47.99

People who bought this also bought

Machine Learning Engineering with Python

Aug 2023 462 pages

4.6 (36)

eBook

$27.98 ~~$39.99~~

Deep Learning with TensorFlow and Keras – 3rd edition

Oct 2022 698 pages

4.6 (45)

eBook

$27.98 ~~$39.99~~

Modern Generative AI with ChatGPT and OpenAI Models

May 2023 286 pages

4.2 (34)

eBook

$27.98 ~~$39.99~~

Generative AI with LangChain

Dec 2023 368 pages

4 (34)

eBook

$27.98 ~~$39.99~~

Causal Inference and Discovery in Python

May 2023 456 pages

4.5 (49)

eBook

$21.99 ~~$31.99~~

About the author

Frank Kane

Frank Kane has spent nine years at Amazon and IMDb, developing and managing the technology that automatically delivers product and movie recommendations to hundreds of millions of customers all the time. He holds 17 issued patents in the fields of distributed computing, data mining, and machine learning. In 2012, Frank left to start his own successful company, Sundog Software, which focuses on virtual reality environment technology and teaches others about big data analysis.

See other products by Frank Kane

FAQs

What is included in a Packt subscription?

A subscription provides you with full access to view all Packt and licnesed content online, this includes exclusive access to Early Access titles. Depending on the tier chosen you can also earn credits and discounts to use for owning content

How can I cancel my subscription?

To cancel your subscription with us simply go to the account page - found in the top right of the page or at https://subscription.packtpub.com/my-account/subscription - From here you will see the ‘cancel subscription’ button in the grey box with your subscription information in.

What are credits?

Credits can be earned from reading 40 section of any title within the payment cycle - a month starting from the day of subscription payment. You also earn a Credit every month if you subscribe to our annual or 18 month plans. Credits can be used to buy books DRM free, the same way that you would pay for a book. Your credits can be found in the subscription homepage - subscription.packtpub.com - clicking on ‘the my’ library dropdown and selecting ‘credits’.

What happens if an Early Access Course is cancelled?

Projects are rarely cancelled, but sometimes it's unavoidable. If an Early Access course is cancelled or excessively delayed, you can exchange your purchase for another course. For further details, please contact us here.

Where can I send feedback about an Early Access title?

If you have any feedback about the product you're reading, or Early Access in general, then please fill out a contact form here and we'll make sure the feedback gets to the right team.

Can I download the code files for Early Access titles?

We try to ensure that all books in Early Access have code available to use, download, and fork on GitHub. This helps us be more agile in the development of the book, and helps keep the often changing code base of new versions and new technologies as up to date as possible. Unfortunately, however, there will be rare cases when it is not possible for us to have downloadable code samples available until publication.

When we publish the book, the code files will also be available to download from the Packt website.

How accurate is the publication date?

The publication date is as accurate as we can be at any point in the project. Unfortunately, delays can happen. Often those delays are out of our control, such as changes to the technology code base or delays in the tech release. We do our best to give you an accurate estimate of the publication date at any given time, and as more chapters are delivered, the more accurate the delivery date will become.

How will I know when new chapters are ready?

We'll let you know every time there has been an update to a course that you've bought in Early Access. You'll get an email to let you know there has been a new chapter, or a change to a previous chapter. The new chapters are automatically added to your account, so you can also check back there any time you're ready and download or read them online.

I am a Packt subscriber, do I get Early Access?

Yes, all Early Access content is fully available through your subscription. You will need to have a paid for or active trial subscription in order to access all titles.

How is Early Access delivered?

Early Access is currently only available as a PDF or through our online reader. As we make changes or add new chapters, the files in your Packt account will be updated so you can download them again or view them online immediately.

How do I buy Early Access content?

Early Access is a way of us getting our content to you quicker, but the method of buying the Early Access course is still the same. Just find the course you want to buy, go through the check-out steps, and you’ll get a confirmation email from us with information and a link to the relevant Early Access courses.

What is Early Access?

Keeping up to date with the latest technology is difficult; new versions, new frameworks, new techniques. This feature gives you a head-start to our content, as it's being created. With Early Access you'll receive each chapter as it's written, and get regular updates throughout the product's development, as well as the final course as soon as it's ready.We created Early Access as a means of giving you the information you need, as soon as it's available. As we go through the process of developing a course, 99% of it can be ready but we can't publish until that last 1% falls in to place. Early Access helps to unlock the potential of our content early, to help you start your learning when you need it most. You not only get access to every chapter as it's delivered, edited, and updated, but you'll also get the finalized, DRM-free product to download in any format you want when it's published. As a member of Packt, you'll also be eligible for our exclusive offers, including a free course every day, and discounts on new and popular titles.

Frank Kane's Taming Big Data with Apache Spark and Python: Real-world examples to help you analyze large datasets with Apache Spark

What do you get with a Packt Subscription?

Key benefits

Description

Who is this book for?

What you will learn

Product Details

What do you get with a Packt Subscription?

Product Details

Packt Subscriptions

Frequently bought together

Table of Contents

Recommendations for you

Customer reviews

Filter reviews by

People who bought this also bought

About the author

FAQs