Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Tomasz Drabas, Denny Lee

Rating No ratings yet

Combine the power of Apache Spark and Python to build effective big data applications Key Features Perform effective data processing, machine learning, and analytics using PySpark Overcome challenges in developing and deploying Spark solutions using Python Explore recipes for efficiently combining Python and Apache Spark to process data Book Description Apache Spark is an open source framework for efficient cluster computing with a strong interface for data parallelism and fault tolerance. The PySpark Cookbook presents effective and time-saving recipes for leveraging the power of Python and putting it to use in the Spark ecosystem. You'll start by learning the Apache Spark architecture and how to set up a Python environment for Spark. You'll then get familiar with the modules available in PySpark and start using them effortlessly. In addition to this, you'll discover how to abstract data with RDDs and DataFrames, and understand the streaming capabilities of PySpark. You'll then move on to using ML and MLlib in order to solve any problems related to the machine learning capabilities of PySpark and use GraphFrames to solve graph-processing problems. Finally, you will explore how to deploy your applications to the cloud using the spark-submit command. By the end of this book, you will be able to use the Python API for Apache Spark to solve any problems associated with building data-intensive applications. What you will learn Configure a local instance of PySpark in a virtual environment Install and configure Jupyter in local and multi-node environments Create DataFrames from JSON and a dictionary using pyspark.sql Explore regression and clustering models available in the ML module Use DataFrames to transform data used for modeling Connect to PubNub and perform aggregations on streams Who this book is for The PySpark Cookbook is for you if you are a Python developer looking for hands-on recipes for using the Apache Spark 2.x ecosystem in the best possible way. A thoroug

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# PySpark Cookbook: Over 60 Recipes for Implementing Big Data Processing and Analytics Using Apache Spark and Python ## 【One-Line Pitch】 A practical recipe collection for Python developers who want to harness Apache Spark 2.x for big data processing, machine learning, and streaming analytics—covering everything from environment setup to cloud deployment. ## 【Book Arc】 - **Opening (~0%–9%)**: Introduces the Apache Spark architecture and walks through installing Spark with Python, including prerequisites like Java and Python, and setting up the GitHub repository with sample code and data. - **Early (~9%–28%)**: Covers advanced installation scenarios—building Spark from source, installing from binaries, setting up multi-node clusters, configuring Jupyter notebooks with Spark kernels, and using virtual machines like Cloudera for learning. - **Early–Middle (~28%–44%)**: Dives into RDDs (Resilient Distributed Datasets)—creating them from collections or files, applying transformations like map(), filter(), and union(), and understanding partitions and performance considerations. - **Middle (~44%–53%)**: Shifts to DataFrames as the primary abstraction—reading data from JSON and CSV, inferring or programmatically specifying schemas, and comparing RDD vs. DataFrame performance using Spark SQL's Catalyst optimizer. - **Middle–Late (~53%–100%)**: Covers data preparation for modeling, machine learning with MLlib and the ML module, structured streaming, graph processing with GraphFrames, and deploying applications to the cloud using spark-submit. ## 【Key Takeaways】 - **Environment setup is the first real challenge** (Early): The book provides shell scripts to check Java, Python, R, Scala, and Maven versions, and to automate Spark installation—saving hours of manual configuration. - **Building Spark from source gives you control** (Early): Using make-distribution.sh with flags like -Phadoop-2.7 and -Phive lets you create a custom Spark distribution tailored to your Hadoop and Hive needs. - **Multi-node clusters require careful planning** (Early): The book shows how to create hosts.txt files, set up SSH access, and automate deployment across servers—essential for moving beyond single-machine learning. - **RDDs are the foundation of Spark's data abstraction** (Middle): Creating RDDs with sc.parallelize() or sc.textFile(), then transforming them with map(), filter(), and union(), teaches the core functional programming model. - **DataFrames offer significant performance gains** (Middle): Running the same group-by query via RDD reduceByKey() vs. DataFrame GROUP BY shows how Spark SQL's WholeStageCodegen and Catalyst optimizer dramatically speed up execution. - **Schema management is critical for data quality** (Middle): Whether inferring schemas from JSON/CSV or specifying them programmatically with StructType and StructField, controlling data types prevents silent errors downstream. - **Pandas UDFs bridge Python and Spark** (Middle): Using @f.pandas_udf with scipy functions shows how to apply complex Python computations to Spark DataFrames at scale. ## 【Reading Tips】 - **Skim Chapter 1's installation scripts** if you already have Spark running—the shell functions are thorough but verbose; focus on the multi-node and Jupyter sections if you're setting up a cluster. - **Deep-read Chapter 2 on RDDs** even if you plan to use DataFrames—understanding transformations like map() and filter() builds intuition for how Spark processes data. - **Pay attention to the RDD vs. DataFrame performance comparison** in Chapter 2/3—it explains why DataFrames are now the recommended API and how Catalyst optimization works. - **The pandas_udf example is a must-try**—it demonstrates a pattern you'll reuse constantly for integrating Python libraries (like scipy) with Spark. - **The book's later chapters (ML, streaming, GraphFrames) are recipe-based**—skim the "Getting ready" sections and jump to the specific recipe you need rather than reading linearly. ## 【Coverage Limits】 This guide covers the book's opening through the DataFrame chapters (~53% of the book). The later sections on MLlib, the ML module, structured streaming, GraphFrames, and cloud deployment are summarized from the table of contents but not detailed from excerpts. ##
Excerpt 1
Create DataFrames from JSON and a dictionary using pyspark.sql Explore regression and clustering models available in the ML module Use DataFrames to transfor...
View in text
Excerpt 2
ill be moving the binaries to: it will either be /opt/spark (default) or your home directory if you use the -ns (or --nosudo) switch when calling the ./insta...
View in text
Excerpt 3
virtualbox- ​on- windows/ ​ On Linux: https://www.packtpub.com/books/content/installing-virtualbox-lin ux On Mac: https:/ ​/​www. ​youtube. ​com/ ​watch? ​v=...
View in text
Excerpt 4
eter) and that it will attempt to assign the right datatype to each column based on the content (the inferSchema parameter assigns strings by default). In co...
View in text
Excerpt 5
pDuplicates() [ 138 ] Preparing Data for Modeling Chapter 4 Well, it looks like we have two records with 'Id == 3'. Let's check whether they're the same: The...
View in text
Excerpt 6
ef labelEncode(label): return [int(label[0] == '>50K')] final_data = ( final_data .map(lambda row: labelEncode(row[0]) + [item for sublist in row[1:] for ite...
View in text
Excerpt 7
expectancy and assumes that a marginal effect of one of the features accelerates or decelerates a process failure. DecisionTreeRegressor, a counterpart of De...
View in text
Excerpt 8
oach: [ 241 ] Machine Learning with the ML Module Chapter 6 Discretizing continuous variables Sometimes, it is actually useful to have a discrete representat...
View in text
Tags
AI categories
Big DataProgrammingData
ISBN: 1788835360
Publisher: Packt Publishing
Publish Year: 2018
Language: English
Pages: 330
File Format: PDF
File Size: 11.1 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…