Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Gerard Maas, Francois Garillot

Rating No ratings yet

Before you can build analytics tools to gain quick insights, you first need to know how to process data in real time. With this practical guide, developers familiar with Apache Spark will learn how to put this in-memory framework to use for streaming data. You’ll discover how Spark enables you to write streaming jobs in almost the same way you write batch jobs. Authors Gerard Maas and François Garillot help you explore the theoretical underpinnings of Apache Spark. This comprehensive guide features two sections that compare and contrast the streaming APIs Spark now supports: the original Spark Streaming library and the newer Structured Streaming API. • Learn fundamental stream processing concepts and examine different streaming architectures • Explore Structured Streaming through practical examples; learn different aspects of stream processing in detail • Create and operate streaming jobs and applications with Spark Streaming; integrate Spark Streaming with other Spark APIs • Learn advanced Spark Streaming techniques, including approximation algorithms and machine learning algorithms • Compare Apache Spark to other stream processing projects, including Apache Storm, Apache Flink, and Apache Kafka Streams

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A practical guide for Spark developers who want to process data in real time, comparing the original Spark Streaming library with the newer Structured Streaming API so you can pick the right tool and build reliable streaming jobs. Best for engineers already comfortable with Spark batch programming who now need to handle unbounded, continuously arriving data. 【Book Arc】 - **Opening (~0%–10%)**: Frames the core problem—extracting insight from infinite event streams—and introduces the two Spark streaming APIs the book contrasts. Establishes foundational concepts like event time vs. processing time and the uncertainty of unbounded input. - **Early (~10%–30%)**: Covers stream-processing fundamentals: timelines, windows (tumbling/sliding), the effect of time on aggregation, and architectural models (Lambda vs. Kappa). Positions Spark's abstraction layers (Core, SQL, libraries) as the platform for streaming. - **Early–Middle (~30%–45%)**: Explains Spark's fault tolerance and resilience—RDDs, driver/executor architecture, shuffle service, and high-availability modes—then moves into the Structured Streaming programming model: sources, transformations, sinks, output modes, and the lazy start() lifecycle. - **Middle (~45%–65%)**: Deepens Structured Streaming with practical detail: source and sink configuration (Kafka, socket, file, memory, console, foreach), event-time handling, windows, watermarks, and stateful operations. - **Late (~65%–90%)**: Shifts to Spark Streaming (the original DStream API), its integration with other Spark APIs, and advanced techniques including approximation algorithms and machine learning on streams. - **Ending (~90%–100%)**: Compares Spark to other stream processors (Storm, Flink, Kafka Streams) and points to community resources for staying current. 【Key Takeaways】 - **Event time vs. processing time is the central distinction** (Early): the book stresses that correlating, ordering, and aggregating events depends on which timeline you use, and that late/out-of-order data is a first-class concern. - **Unbounded input introduces uncertainty** (Early): streaming systems can't assume regular arrival, so peak-load planning and resource matching become core design challenges. - **Lambda vs. Kappa architectures are a real trade-off** (Early): Kappa is lighter and simpler, but the book argues Lambda may still be needed in some cases; batch datasets also serve as benchmarks for streaming iterations. - **Spark's fault tolerance is foundational, not streaming-specific** (Early–Middle): RDDs, driver/executor design, shuffle service, and HA modes underpin long-running streaming jobs. - **Structured Streaming reuses batch APIs but with key differences** (Middle): most Dataset/DataFrame transformations carry over, yet operations like a full count have no defined meaning on a stream and require writeStream.start(). - **Sources, sinks, and output modes form the streaming contract** (Middle): the book details Kafka, socket, file, memory, console, and foreach sinks, plus how output modes relate to aggregations. - **Typed Datasets can be slower than DataFrames** (Middle): closures in Dataset operations are opaque to the query planner, so the DataFrame API may optimize better. - **Spark Streaming (DStreams) remains relevant** (Late): the book covers its operation, integration with other Spark APIs, and advanced approximation/ML techniques alongside Structured Streaming. 【Reading Tips】 - Deep-read the early chapters on event time, windows, and architectures—these concepts recur throughout and are the hardest to retrofit later. - Skim the source/sink configuration catalogs (Kafka, file, foreach) on first pass; return to them as reference when implementing. - Pay close attention to the Structured Streaming programming model chapter—the lazy evaluation and start() lifecycle trip up newcomers. - If you only need Structured Streaming, treat the Spark Streaming (DStream) chapters as comparative context rather than required reading. - Use the final comparison chapter to decide whether Spark fits your use case versus Flink or Kafka Streams. 【Coverage Limits】 The excerpts are heavily weighted toward the table of contents and early conceptual chapters; the later Spark Streaming, advanced techniques, and comparison chapters are only partially represented, so this guide's late-stage coverage is inferred from chapter titles and brief fragments rather than full detail.
Excerpt 1
130 Configuration 131 Operations 132 The Rate Source 132 Options 133 11. Structured Streaming Sinks. . . . . . . . . . . . . . . . . . . . . . . . . . . . ....
View in text
Excerpt 2
ut. 6 | Chapter 1: Introducing Stream Processing Figure 1-1. Abstraction layers (horizontal) and libraries (vertical) offered by Spark As abstraction layers...
View in text
Excerpt 3
omputation of that program occurs as the data arrives. How‐ ever, if the network partitions separate some part of the cluster, the driver might be able to ke...
View in text
Excerpt 4
set and DataFrame APIs for transforming the streaming data. • Some common operations from the batch API do not make sense in streaming mode. • Sinks are the...
View in text
Excerpt 5
also offers a programmable interface that allows us to work with arbitrary external systems. Reliable Sinks The sinks considered reliable or production ready...
View in text
Excerpt 6
mapGroupsWithState requires the state handling function to produce a single record for each group processed at every trigger interval. This is fine when the...
View in text
Excerpt 7
|1.0 | |216|2018-08-06 00:13:20.687|1 |0.0 | |217|2018-08-06 00:13:21.687|1 |0.0 | |218|2018-08-06 00:13:22.687|0 |0.0 | |219|2018-08-06 00:13:23.687|0 |0.0...
View in text
Excerpt 8
r spark.streaming.minRememberDuration is provided as a time duration, the actual window will be computed as the ceiling(remember_duration/ batch_interval). F...
View in text
Tags
AI categories
Big DataDataProgramming
ISBN: 1491944242
Publisher: O'Reilly Media
Publish Year: 2019
Language: English
Pages: 452
File Format: PDF
File Size: 8.3 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…