Before you can build analytics tools to gain quick insights, you first need to know how to process data in real time. With this practical guide, developers familiar with Apache Spark will learn how to put this in-memory framework to use for streaming data. You’ll discover how Spark enables you to write streaming jobs in almost the same way you write batch jobs.
Authors Gerard Maas and François Garillot help you explore the theoretical underpinnings of Apache Spark. This comprehensive guide features two sections that compare and contrast the streaming APIs Spark now supports: the original Spark Streaming library and the newer Structured Streaming API.
• Learn fundamental stream processing concepts and examine different streaming architectures
• Explore Structured Streaming through practical examples; learn different aspects of stream processing in detail
• Create and operate streaming jobs and applications with Spark Streaming; integrate Spark Streaming with other Spark APIs
• Learn advanced Spark Streaming techniques, including approximation algorithms and machine learning algorithms
• Compare Apache Spark to other stream processing projects, including Apache Storm, Apache Flink, and Apache Kafka Streams
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
【One-Line Pitch】
A practical guide for Spark developers who want to process data in real time, comparing the original Spark Streaming library with the newer Structured Streaming API so you can pick the right tool and build reliable streaming jobs. Best for engineers already comfortable with Spark batch programming who now need to handle unbounded, continuously arriving data.
【Book Arc】
- **Opening (~0%–10%)**: Frames the core problem—extracting insight from infinite event streams—and introduces the two Spark streaming APIs the book contrasts. Establishes foundational concepts like event time vs. processing time and the uncertainty of unbounded input.
- **Early (~10%–30%)**: Covers stream-processing fundamentals: timelines, windows (tumbling/sliding), the effect of time on aggregation, and architectural models (Lambda vs. Kappa). Positions Spark's abstraction layers (Core, SQL, libraries) as the platform for streaming.
- **Early–Middle (~30%–45%)**: Explains Spark's fault tolerance and resilience—RDDs, driver/executor architecture, shuffle service, and high-availability modes—then moves into the Structured Streaming programming model: sources, transformations, sinks, output modes, and the lazy start() lifecycle.
- **Middle (~45%–65%)**: Deepens Structured Streaming with practical detail: source and sink configuration (Kafka, socket, file, memory, console, foreach), event-time handling, windows, watermarks, and stateful operations.
- **Late (~65%–90%)**: Shifts to Spark Streaming (the original DStream API), its integration with other Spark APIs, and advanced techniques including approximation algorithms and machine learning on streams.
- **Ending (~90%–100%)**: Compares Spark to other stream processors (Storm, Flink, Kafka Streams) and points to community resources for staying current.
【Key Takeaways】
- **Event time vs. processing time is the central distinction** (Early): the book stresses that correlating, ordering, and aggregating events depends on which timeline you use, and that late/out-of-order data is a first-class concern.
- **Unbounded input introduces uncertainty** (Early): streaming systems can't assume regular arrival, so peak-load planning and resource matching become core design challenges.
- **Lambda vs. Kappa architectures are a real trade-off** (Early): Kappa is lighter and simpler, but the book argues Lambda may still be needed in some cases; batch datasets also serve as benchmarks for streaming iterations.
- **Spark's fault tolerance is foundational, not streaming-specific** (Early–Middle): RDDs, driver/executor design, shuffle service, and HA modes underpin long-running streaming jobs.
- **Structured Streaming reuses batch APIs but with key differences** (Middle): most Dataset/DataFrame transformations carry over, yet operations like a full count have no defined meaning on a stream and require writeStream.start().
- **Sources, sinks, and output modes form the streaming contract** (Middle): the book details Kafka, socket, file, memory, console, and foreach sinks, plus how output modes relate to aggregations.
- **Typed Datasets can be slower than DataFrames** (Middle): closures in Dataset operations are opaque to the query planner, so the DataFrame API may optimize better.
- **Spark Streaming (DStreams) remains relevant** (Late): the book covers its operation, integration with other Spark APIs, and advanced approximation/ML techniques alongside Structured Streaming.
【Reading Tips】
- Deep-read the early chapters on event time, windows, and architectures—these concepts recur throughout and are the hardest to retrofit later.
- Skim the source/sink configuration catalogs (Kafka, file, foreach) on first pass; return to them as reference when implementing.
- Pay close attention to the Structured Streaming programming model chapter—the lazy evaluation and start() lifecycle trip up newcomers.
- If you only need Structured Streaming, treat the Spark Streaming (DStream) chapters as comparative context rather than required reading.
- Use the final comparison chapter to decide whether Spark fits your use case versus Flink or Kafka Streams.
【Coverage Limits】
The excerpts are heavily weighted toward the table of contents and early conceptual chapters; the later Spark Streaming, advanced techniques, and comparison chapters are only partially represented, so this guide's late-stage coverage is inferred from chapter titles and brief fragments rather than full detail.
omputation of that program occurs as the data arrives. How‐ ever, if the network partitions separate some part of the cluster, the driver might be able to ke...
set and DataFrame APIs for transforming the streaming data. • Some common operations from the batch API do not make sense in streaming mode. • Sinks are the...
also offers a programmable interface that allows us to work with arbitrary external systems. Reliable Sinks The sinks considered reliable or production ready...
mapGroupsWithState requires the state handling function to produce a single record for each group processed at every trigger interval. This is fine when the...
r spark.streaming.minRememberDuration is provided as a time duration, the actual window will be computed as the ceiling(remember_duration/ batch_interval). F...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
Stream Processing with Apache Spark Mastering Structured Streaming and Spark Streaming (Gerard Maas, Francois Garillot)(Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
Stream Processing with Apache Spark Mastering Structured Streaming and Spark Streaming (Gerard Maas, Francois Garillot)(Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment