Big Data teaches you to build big data systems using an architecture that takes advantage of clustered hardware along with new tools designed specifically to capture and analyze web-scale data. It describes a scalable, easy-to-understand approach to big data systems that can be built and run by a small team. Following a realistic example, this book guides readers through the theory of big data systems, how to implement them in practice, and how to deploy and operate them once they're built.
Purchase of the print book includes a free eBook in PDF, Kindle, and ePub formats from Manning Publications.
About the Book
Web-scale applications like social networks, real-time analytics, or e-commerce sites deal with a lot of data, whose volume and velocity exceed the limits of traditional database systems. These applications require architectures built around clusters of machines to store and process data of any size, or speed. Fortunately, scale and simplicity are not mutually exclusive.
Big Data teaches you to build big data systems using an architecture designed specifically to capture and analyze web-scale data. This book presents the Lambda Architecture, a scalable, easy-to-understand approach that can be built and run by a small team. You'll explore the theory of big data systems and how to implement them in practice. In addition to discovering a general framework for processing big data, you'll learn specific technologies like Hadoop, Storm, and NoSQL databases.
This book requires no previous exposure to large-scale data analysis or NoSQL tools. Familiarity with traditional databases is helpful.
What's Inside
Introduction to big data systems
Real-time processing of web-scale data
Tools like Hadoop, Cassandra, and Storm
Extensions to traditional database skills
About the Authors
Nathan Marz is the creator of Apache Storm and the originator of the Lambda Architecture for big data systems. James Warren is an analytics architect with a background in machine learnin
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
【One-Line Pitch】
A practical guide to building web-scale data systems with the Lambda Architecture—an approach simple enough for a small team to run but scalable to any data size. Best for engineers and architects with traditional database experience who need to move beyond single-machine systems.
【Book Arc】
- **Opening (~0%–15%)**: Frames why traditional databases fail at web scale and introduces the Lambda Architecture as a response—three layers (batch, serving, speed) built on immutable, append-only master data. Solves the "how do I even think about this problem" stage.
- **Early (~15%–30%)**: Establishes the core properties a big data system must have—human-fault tolerance, low-latency reads and updates, horizontal scalability, generalization, extensibility, and ad hoc query support. Then defines the data model: rawness, immutability, perpetuity, and the fact-based model.
- **Middle (~30%–55%)**: Dives into the batch layer. Covers schema design with Apache Thrift, the limitations of serialization frameworks, why the master dataset needs simpler storage than a database, and how HDFS works (namenode/datanode, block replication). Also introduces MapReduce and higher-level abstractions like JCascalog.
- **Late (~55%–80%)**: Moves through the serving layer and into the speed layer—real-time processing, stream processing concepts, and micro-batch approaches including Trident. The running example SuperWebAnalytics.com ties batch and speed layer implementations together.
- **Ending (~80%–100%)**: Revisits the Lambda Architecture in depth, covering incremental batch processing, measuring and optimizing batch layer resource usage, and how the query layer merges batch and realtime views. Closes with a consolidated view of the full architecture.
【Key Takeaways】
- **The Lambda Architecture separates batch and speed concerns** (Early): batch layer recomputes views from all data; speed layer handles recent data with low latency; queries merge both. This lets you get correctness and freshness without compromising either.
- **Immutability and recomputation are the core robustness mechanism** (Early): because the master dataset is append-only and eternally true, human errors (bad deploys, corrupted values) can be recovered by recomputing views—no complex rollback needed.
- **Data is raw, immutable, and perpetual; views are derived** (Early): the fact-based model stores only what cannot be derived from anything else. One person's view can be another's data—the distinction depends on what you can compute.
- **Serialization frameworks enforce structure but not semantic validity** (Middle): Thrift checks types and required fields, but cannot enforce "ages must be non-negative." You need additional validation before writing to the master dataset.
- **The master dataset needs simpler storage than a traditional database** (Middle): many database features actively get in the way. HDFS-style distributed filesystems—with block chunking, replication, and a namenode lookup—fit the batch layer better.
- **MapReduce is general but low-level; higher abstractions matter** (Middle): MapReduce can compute any scalable function, but tools like JCascalog make it far easier to express those computations.
- **Micro-batch stream processing bridges batch and realtime** (Late): Trident and micro-batch topologies provide fault-tolerant, in-memory processing for the speed layer, finishing the SuperWebAnalytics.com example.
- **Scalability is horizontal across all layers** (Early): scaling means adding machines, not upgrading them. The architecture generalizes to financial systems, social analytics, scientific applications, and more.
【Reading Tips】
- **Deep-read the opening chapters on data properties and the Lambda Architecture equations**—they are the conceptual foundation everything else builds on. Skim the tool-specific chapters (Thrift, HDFS, JCascalog) if you already know those tools.
- **Follow the SuperWebAnalytics.com example** across batch and speed layer implementations. It is the thread that turns abstract architecture into working code.
- **Pay attention to the "why not a database" arguments** in the storage chapters. They explain design trade-offs that are easy to miss if you jump straight to implementation.
- **Treat the batch/speed/serving layer chapters as a progression**: read batch layer first, then serving, then speed. The query layer chapter at the end makes more sense after you understand what it is merging.
- **Use the architecture-in-depth chapter as a review** after finishing the implementation chapters—it consolidates incremental batch processing and resource optimization.
【Coverage Limits】
The excerpts cover the book's structure, core concepts, and several implementation chapters, but do not include detailed code listings or the full SuperWebAnalytics.com implementation. Specific deployment and operations guidance mentioned in the blurb is not covered in the available excerpts.
Excerpt 1
xtensions to traditional database skills About the Authors Nathan Marz is the creator of Apache Storm and the originator of the Lambda Architecture for big d...
f increasing data or load by adding resources to the system. The Lambda Architecture is horizontally scalable across all layers of the system stack: scaling...
(birthdate, (Web scraping) Potential number friends, …) ads Number of of friends) friends b Tom provides detailed profile information c Only the public infor...
monstrate how such a tool can be used for the batch layer. HDFS and Hadoop MapReduce are the two prongs of the Hadoop project: a Java framework for distribut...
umber of HDFS blocks and eases the demand on the namenode. Read Support for paral- The number of tasks in a MapReduce job is determined by the num- lel proce...
fly to the moon fly fly to the moon the fly to the moon to fly to the moon moon fly to the moon moon dog dog dog dog Figure 6.16 Illustration of a pipe dia-...
utput of two different generator predicates, AGE and GENDER. Because each instance of the variable must have the same value for any resulting tuples, JCascal...
nction(), "?c").out("?z"); where each field is incremented Although it’s a simple query, there’s considerable repetition because it must explicitly apply Inc...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
Big Data Principles and best practices of scalable realtime data systems (Nathan Marz, James Warren)(Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
Big Data Principles and best practices of scalable realtime data systems (Nathan Marz, James Warren)(Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment