Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Jean-Georges Perrin

Rating No ratings yet

The Spark distributed data processing platform provides an easy-to-implement tool for ingesting, streaming, and processing data from any source. In Spark in Action, Second Edition, you’ll learn to take advantage of Spark’s core features and incredible processing speed, with applications including real-time computation, delayed evaluation, and machine learning. Unlike many Spark books written for data scientists, Spark in Action, Second Edition is designed for data engineers and software engineers who want to master data processing using Spark without having to learn a complex new ecosystem of languages and tools. You’ll instead learn to apply your existing Java and SQL skills to take on practical, real-world challenges. Key Features · Lots of examples based in the Spark Java APIs using real-life dataset and scenarios · Examples based on Spark v2.3 Ingestion through files, databases, and streaming · Building custom ingestion process · Querying distributed datasets with Spark SQL For beginning to intermediate developers and data engineers comfortable programming in Java. No experience with functional programming, Scala, Spark, Hadoop, or big data is required. About the technology Spark is a powerful general-purpose analytics engine that can handle massive amounts of data distributed across clusters with thousands of servers. Optimized to run in memory, this impressive framework can process data up to 100x faster than most Hadoop-based systems. Author Bio An experienced consultant and entrepreneur passionate about all things data, Jean-Georges Perrin was the first IBM Champion in France, an honor he’s now held for ten consecutive years. Jean-Georges has managed many teams of software and data engineers.

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A hands-on guide to Apache Spark for Java and SQL developers who want to process large-scale data without first learning Scala or a new big-data ecosystem. Best for beginning-to-intermediate engineers who prefer working code over theory. 【Book Arc】 - **Opening (~0%–20%)**: Establishes Spark as a general-purpose, in-memory analytics engine and frames the book's Java-first, SQL-friendly approach for developers coming from traditional software engineering. - **Early (~20%–40%)**: Introduces core Spark concepts and the Java APIs, building toward practical data processing on real datasets rather than abstract theory. - **Middle (~40%–60%)**: Covers ingestion from files, databases, and streaming sources, showing how to bring varied real-world data into Spark pipelines. - **Late (~60%–80%)**: Moves into querying distributed datasets with Spark SQL and building custom ingestion processes tailored to specific needs. - **Ending (~80%–100%)**: Extends toward advanced applications such as real-time computation, delayed evaluation, and machine learning, rounding out the processing toolkit. 【Key Takeaways】 - **Spark is a general-purpose, in-memory analytics engine** (Opening): it distributes work across clusters of thousands of servers and can run up to 100x faster than most Hadoop-based systems, which is why it matters for large-scale processing. - **This book targets engineers, not data scientists** (Opening): it deliberately avoids forcing readers into Scala or functional programming, letting Java and SQL skills carry the load. - **Java APIs are the primary teaching vehicle** (Early): examples are grounded in the Spark Java APIs using real-life datasets and scenarios, so the learning transfers directly to production code. - **Ingestion is a first-class concern** (Middle): the book covers pulling data from files, databases, and streaming sources, plus building custom ingestion processes when off-the-shelf connectors fall short. - **Spark SQL lets you query distributed datasets declaratively** (Late): applying existing SQL knowledge to cluster-scale data lowers the barrier for engineers without big-data backgrounds. - **Streaming and real-time computation extend batch processing** (Late): the platform handles continuous data flows, not just static datasets. - **Delayed evaluation and machine learning round out the feature set** (Ending): these advanced topics show Spark as more than a query engine. - **No prior Spark, Hadoop, Scala, or big-data experience is assumed** (Opening): the on-ramp is designed for developers comfortable in Java who are new to distributed systems. 【Reading Tips】 - **Deep-read the Java API examples**: they are the book's core value; typing them out against the real datasets will teach more than skimming. - **Skim conceptual framing if you already know Spark basics**: the opening material is orientation for newcomers, so experienced readers can move quickly to ingestion and SQL chapters. - **Treat ingestion and custom ingestion as the practical heart**: this is where most real-world friction lives, so slow down there. - **Use Spark SQL chapters to bridge existing skills**: if you know SQL, this is your fastest path into distributed querying. - **Save streaming and ML for last**: they build on the batch and SQL foundations, so sequence matters. 【Coverage Limits】 The excerpts describe the book's scope, audience, and feature list but do not include chapter-level detail, code specifics, or named examples; this guide therefore maps themes rather than concrete walkthroughs.
Excerpt 1
书名: Spark in Action, Second Edition (Jean-Georges Perrin) (Z-Library) (1) 作者: Jean-Georges Perrin The Spark distributed data processing platform provides an...
View in text
Tags
AI categories
Big DataDataJava
ISBN: 1617295523
Publish Year: 2020
Language: English
Pages: 600
File Format: PDF
File Size: 36.0 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…