Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Seymour, Mitch

Rating No ratings yet

Working with unbounded and fast-moving data streams has historically been difficult. But with Kafka Streams and ksqlDB, building stream processing applications is easy and fun. This practical guide explores the world of real-time data systems through the lens of these popular technologies and explains important stream processing concepts against a backdrop of interesting business problems. Mitch Seymour, senior data systems engineer at Mailchimp, introduces you to both Kafka Streams and ksqlDB so that you can choose the best tool for each unique stream processing project. Non-Java developers will find the ksqlDB path to be an especially gentle introduction to stream processing. In this book, you'll learn: Basic and advanced uses of Kafka Streams and ksqlDB How to transform, enrich, and process event streams How to build both stateless and stateful stream processing applications The different notions of time and the role it plays in stream processing How to to build event-driven microservices on top of continuous event streams Features, operational characteristics, deployment patterns, and configuration tips for both technologies

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A hands-on tour of real-time stream processing that teaches you to build, operate, and choose between Kafka Streams and ksqlDB using concrete business scenarios. Best for data engineers, analysts, and data scientists who want to move from batch thinking to continuous event processing. 【Book Arc】 - **Opening (~0%–15%)**: Frames the problem of unbounded, fast-moving data and lays the Kafka foundation — topics, partitions, consumer groups, and the commit log as the storage abstraction behind streams. - **Early (~15%–35%)**: Introduces Kafka Streams itself: its operational goals (scalability, reliability, maintainability), the processor topology/DAG model, tasks and threads, and the high-level DSL versus the lower-level Processor API. - **Middle (~35%–55%)**: Moves from stateless to stateful processing — stateless operators like map/mapValues/flatMapValues, custom serialization (including Avro and Schema Registry trade-offs), then stateful operations such as joins and aggregations. - **Late (~55%–80%)**: Deepens state management — persistent store layout, fault tolerance via changelog topics and standby replicas, rebalancing pitfalls, and techniques to control state size and deduplicate writes. - **Ending (~80%–100%)**: Shifts to ksqlDB as the SQL-based alternative, covering querying key-value stores (local vs. remote), windows and time semantics, and deployment/operational patterns for both technologies. 【Key Takeaways】 - **Streams are stored as append-only commit logs** (Early): Kafka models continuous data as an ordered, append-only log — the same abstraction found in databases and version control — which underpins how topics persist and replay events. - **Kafka Streams favors a simple deployment model** (Early): Unlike cluster-based engines such as Flink or Spark Streaming, it runs as a library with a friendlier learning curve and supports true event-at-a-time processing. - **Topologies are DAGs of source, stream, and sink processors** (Early): Designing the topology first, then implementing it in Java, is the recommended mental model for building any Kafka Streams app. - **Tasks and threads are the unit of parallelism** (Early): A task maps to a topic-partition; adding threads redistributes tasks across CPUs without changing task count, which matters for scaling. - **Prefer mapValues/flatMapValues over map/flatMap** (Middle): Value-only operators let Kafka Streams execute more efficiently and avoid unnecessary rekeying unless stateful operations downstream require it. - **Stateful processing trades simplicity for power** (Middle): Capturing state enables joins and aggregations but introduces concerns around fault tolerance, rebalancing, and state size that stateless pipelines avoid. - **Schema handling is a real design decision** (Middle): Embedding Avro schemas in each record avoids extra infrastructure, while Confluent Schema Registry yields smaller payloads and safer schema evolution at the cost of another service. - **ksqlDB offers a gentler on-ramp** (Late): Non-Java developers can express stream processing in SQL, making it a practical entry point before committing to the Kafka Streams API. 【Reading Tips】 - Deep-read the Kafka fundamentals chapter if you're new to topics, partitions, and consumer groups — later chapters assume this fluency. - Skim the code-heavy setup and serialization sections on a first pass; return to them when you actually implement a pipeline. - Treat the state management and rebalancing material as the hardest part — slow down there, since operational mistakes here cause real production pain. - If you're not a Java developer, jump to the ksqlDB chapters first to build intuition, then circle back to Kafka Streams. - Keep the GitHub repo open alongside the book; the tutorials are meant to be run, not just read. 【Coverage Limits】 This guide is synthesized from stratified excerpts covering roughly the first half of the book plus table-of-contents fragments; later ksqlDB chapters and detailed operational configurations are only partially represented, so specifics there are inferred from headings rather than full text.
Page 8
. . . . . . . . . . . . . . . . . . . . . . . . . . . 143 Introducing Our Tutorial: Patient Monitoring Application 144 Project Setup 146 Data Models 147 Time...
View in text
Excerpt 2
e to unexpected failure. Therefore, Kafka needs some way of maintaining the membership of each group, and redistributing work when necessary. To facilitate t...
View in text
Excerpt 3
ource method is an arbitrary name for this stream processor. In this case, we simply call this processor UserSource. We will refer to this name in the next l...
View in text
Excerpt 4
ller schema ID in each record instead of the entire schema. As shown in Figure 3-8, the benefit of the first approach is you don’t have to set up and run a s...
View in text
Excerpt 5
e joiner could be expressed using the following pseudocode: (scoreEvent, player) -> combine(scoreEvent, player); But we can do much better than that. It’s mo...
View in text
Excerpt 6
the purpose of aggregating and joining. Kafka Streams sup‐ ports a few different types of windows, so let’s take a look at each type to determine which imple...
View in text
Excerpt 7
nt has not checked out, then perform the aggregation logic. While tombstones are useful for keeping key-value stores small, there is another method we can us...
View in text
Excerpt 8
arameter again since the parameter types match the previous source processor. One thing to note about the preceding example is that you will see no mention o...
View in text
Tags
AI categories
DataBig DataBackend
ISBN: 1492062499
Publish Year: 2021
Language: English
Pages: 400
File Format: PDF
File Size: 9.1 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…