Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Jay Kreps

Rating No ratings yet

Why a book about logs? That’s easy: the humble log is an abstraction that lies at the heart of many systems, from NoSQL databases to cryptocurrencies. Even though most engineers don’t think much about them, this short book shows you why logs are worthy of your attention. Based on his popular blog posts, LinkedIn principal engineer Jay Kreps shows you how logs work in distributed systems, and then delivers practical applications of these concepts in a variety of common uses—data integration, enterprise architecture, real-time stream processing, data system design, and abstract computing models. Go ahead and take the plunge with logs; you’re going love them. Learn how logs are used for programmatic access in databases and distributed systems Discover solutions to the huge data integration problem when more data of more varieties meet more systems Understand why logs are at the heart of real-time stream processing Learn the role of a log in the internals of online data systems Explore how Jay Kreps applies these ideas to his own work on data infrastructure systems at LinkedIn

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A short, opinionated field guide to the append-only log as the unifying abstraction behind databases, distributed consensus, and modern data pipelines. Read it if you build or operate data systems and want a mental model that connects replication, ETL, and stream processing into one story. 【Book Arc】 - **Opening (~0%–10%)**: Defines the log as an abstract, append-only, time-ordered sequence of records — not a text file — and argues it is the fundamental structure for recording what happened and when. Solves the "why should I care?" problem. - **Early (~10%–30%)**: Shows the log's roots in databases: write-ahead logging for crash recovery, log shipping for replication, and the duality between a changelog and a table. Also introduces consensus (Paxos, ZAB, RAFT) as log-building. - **Early–Middle (~25%–40%)**: Reframes data integration as the real bottleneck, introduces the "Maslow-like hierarchy" of data needs, and explains why event data breaks traditional ETL. - **Middle (~40%–55%)**: The LinkedIn case study: the N² pipeline problem, the move to a central log hub, and the birth of Kafka as a general-purpose pipeline. - **Late (~55%–80%)**: Practical architecture guidance — where transformations belong (producer, real-time post-processing, or load-time), and how a central log democratizes data access across an organization. - **Ending (~80%–100%)**: Extends the log concept to stream processing and abstract computing models; excerpts do not cover the final chapters in detail. 【Key Takeaways】 - **The log is an abstract data structure, not a log file** (Opening): an append-only sequence where entry numbers act as logical timestamps, decoupled from physical clocks — essential for distributed systems. - **Databases already run on logs** (Early): write-ahead logs guarantee atomicity and durability; each table or index is a projection of the log's history. Log shipping keeps replicas in sync. - **Tables and changelogs are dual** (Early): a log of changes can materialize a table, and a table's updates can be published as a changelog. This duality underpins stream processing. - **Consensus is really log-building** (Early): Paxos, ZAB, RAFT, and Viewstamped Replication all solve the same practical problem — maintaining a distributed, consistent log. The log is the more natural abstraction than a single-value register. - **Data integration is the unglamorous bottleneck** (Early–Middle): most organizations lack reliable, complete data flow but jump to advanced analytics. Without a solid base, a Hadoop cluster is "an expensive space heater." - **Point-to-point pipelines scale as O(N²)** (Middle): the LinkedIn experience shows that connecting every system to every other system is unbuildable. A central log hub reduces integration work to one connection per system. - **Transformations belong at the producer** (Middle): cleanup should happen before publishing to the log, be lossless and reversible; value-added processing happens as real-time post-processing on the raw log; destination-specific aggregation happens at load time. - **A central log democratizes data access** (Middle–Late): new systems need only one integration point, making search, real-time monitoring, and other capabilities feasible without rebuilding the entire pipeline. 【Reading Tips】 - **Deep-read the first two chapters**: the log definition and the database/consensus material are the conceptual foundation for everything else. - **Skim the LinkedIn history if you know Kafka**: the case study is illustrative but the architectural lessons (N² problem, central hub) are the real payload. - **Focus on the transformation-placement section**: it is the most actionable part for designing real pipelines. - **Treat it as a blog-post collection, not a textbook**: chapters are short and somewhat independent; read in order but don't expect deep proofs. - **Take away the mental model, not the implementation details**: the book is about why logs matter, not how to configure Kafka. 【Coverage Limits】 This guide is based on stratified excerpts covering roughly the first half of the book; the later chapters on stream processing and abstract computing models are only lightly represented. Specific implementation details of Kafka and the final chapters' arguments are not covered here.
Page 9
r 2014: First Edition Preface Conventions Used in This Book The following typographical conventions are used in this book: Italic Indicates new terms, URLs,...
View in text
Excerpt 2
l the problem of maintaining a distributed, consistent log. My suspicion is that our view of this is a little biased by the path of history, perhaps due to t...
View in text
Excerpt 3
up traditional data integration approaches because it tends to be several orders of magnitude larger than transactional data. The Explosion of Specialized Da...
View in text
Excerpt 4
or where a particular cleanup or transformation can reside: It can be done by the data producer prior to adding the data to the company-wide log. It can be d...
View in text
Excerpt 5
. As these processes are replaced with continuous feeds, we naturally start to move towards continuous processing to smooth out the processing resources need...
View in text
Excerpt 6
ystem that will let you retain the full log of the data you want to be able to reprocess and that allows for multiple subscribers. For example, if you will w...
View in text
Excerpt 7
ou can replay it to recreate the state of the source system. That is, if I have the log of changes, I can replay that log into a table in another database an...
View in text
Excerpt 8
e log to retain a complete copy of data and compact the log itself. This moves a significant amount of complexity out of the serving layer, which is system-s...
View in text
Tags
AI categories
DataBackendProgramming Language
ISBN: 1491909382
Publisher: O'Reilly Media
Publish Year: 2014
Language: English
Pages: 60
File Format: PDF
File Size: 3.1 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…