Why a book about logs? That’s easy: the humble log is an abstraction that lies at the heart of many systems, from NoSQL databases to cryptocurrencies. Even though most engineers don’t think much about them, this short book shows you why logs are worthy of your attention.
Based on his popular blog posts, LinkedIn principal engineer Jay Kreps shows you how logs work in distributed systems, and then delivers practical applications of these concepts in a variety of common uses—data integration, enterprise architecture, real-time stream processing, data system design, and abstract computing models.
Go ahead and take the plunge with logs; you’re going love them.
Learn how logs are used for programmatic access in databases and distributed systems
Discover solutions to the huge data integration problem when more data of more varieties meet more systems
Understand why logs are at the heart of real-time stream processing
Learn the role of a log in the internals of online data systems
Explore how Jay Kreps applies these ideas to his own work on data infrastructure systems at LinkedIn
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
【One-Line Pitch】
A short, opinionated field guide to the append-only log as the unifying abstraction behind databases, distributed consensus, and modern data pipelines. Read it if you build or operate data systems and want a mental model that connects replication, ETL, and stream processing into one story.
【Book Arc】
- **Opening (~0%–10%)**: Defines the log as an abstract, append-only, time-ordered sequence of records — not a text file — and argues it is the fundamental structure for recording what happened and when. Solves the "why should I care?" problem.
- **Early (~10%–30%)**: Shows the log's roots in databases: write-ahead logging for crash recovery, log shipping for replication, and the duality between a changelog and a table. Also introduces consensus (Paxos, ZAB, RAFT) as log-building.
- **Early–Middle (~25%–40%)**: Reframes data integration as the real bottleneck, introduces the "Maslow-like hierarchy" of data needs, and explains why event data breaks traditional ETL.
- **Middle (~40%–55%)**: The LinkedIn case study: the N² pipeline problem, the move to a central log hub, and the birth of Kafka as a general-purpose pipeline.
- **Late (~55%–80%)**: Practical architecture guidance — where transformations belong (producer, real-time post-processing, or load-time), and how a central log democratizes data access across an organization.
- **Ending (~80%–100%)**: Extends the log concept to stream processing and abstract computing models; excerpts do not cover the final chapters in detail.
【Key Takeaways】
- **The log is an abstract data structure, not a log file** (Opening): an append-only sequence where entry numbers act as logical timestamps, decoupled from physical clocks — essential for distributed systems.
- **Databases already run on logs** (Early): write-ahead logs guarantee atomicity and durability; each table or index is a projection of the log's history. Log shipping keeps replicas in sync.
- **Tables and changelogs are dual** (Early): a log of changes can materialize a table, and a table's updates can be published as a changelog. This duality underpins stream processing.
- **Consensus is really log-building** (Early): Paxos, ZAB, RAFT, and Viewstamped Replication all solve the same practical problem — maintaining a distributed, consistent log. The log is the more natural abstraction than a single-value register.
- **Data integration is the unglamorous bottleneck** (Early–Middle): most organizations lack reliable, complete data flow but jump to advanced analytics. Without a solid base, a Hadoop cluster is "an expensive space heater."
- **Point-to-point pipelines scale as O(N²)** (Middle): the LinkedIn experience shows that connecting every system to every other system is unbuildable. A central log hub reduces integration work to one connection per system.
- **Transformations belong at the producer** (Middle): cleanup should happen before publishing to the log, be lossless and reversible; value-added processing happens as real-time post-processing on the raw log; destination-specific aggregation happens at load time.
- **A central log democratizes data access** (Middle–Late): new systems need only one integration point, making search, real-time monitoring, and other capabilities feasible without rebuilding the entire pipeline.
【Reading Tips】
- **Deep-read the first two chapters**: the log definition and the database/consensus material are the conceptual foundation for everything else.
- **Skim the LinkedIn history if you know Kafka**: the case study is illustrative but the architectural lessons (N² problem, central hub) are the real payload.
- **Focus on the transformation-placement section**: it is the most actionable part for designing real pipelines.
- **Treat it as a blog-post collection, not a textbook**: chapters are short and somewhat independent; read in order but don't expect deep proofs.
- **Take away the mental model, not the implementation details**: the book is about why logs matter, not how to configure Kafka.
【Coverage Limits】
This guide is based on stratified excerpts covering roughly the first half of the book; the later chapters on stream processing and abstract computing models are only lightly represented. Specific implementation details of Kafka and the final chapters' arguments are not covered here.
Page 9
r 2014: First Edition Preface Conventions Used in This Book The following typographical conventions are used in this book: Italic Indicates new terms, URLs,...
l the problem of maintaining a distributed, consistent log. My suspicion is that our view of this is a little biased by the path of history, perhaps due to t...
up traditional data integration approaches because it tends to be several orders of magnitude larger than transactional data. The Explosion of Specialized Da...
or where a particular cleanup or transformation can reside: It can be done by the data producer prior to adding the data to the company-wide log. It can be d...
. As these processes are replaced with continuous feeds, we naturally start to move towards continuous processing to smooth out the processing resources need...
ystem that will let you retain the full log of the data you want to be able to reprocess and that allows for multiple subscribers. For example, if you will w...
ou can replay it to recreate the state of the source system. That is, if I have the log of changes, I can replay that log into a table in another database an...
e log to retain a complete copy of data and compact the log itself. This moves a significant amount of complexity out of the serving layer, which is system-s...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
I Heart Logs Event Data, Stream Processing, and Data Integration (Jay Kreps)(Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
I Heart Logs Event Data, Stream Processing, and Data Integration (Jay Kreps)(Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment