Austin Parker, Daniel Spoonhower, Jonathan Mace & Rebecca Isaacs Foreword by Ben Sigelman Distributed Tracing in Practice Instrumenting, Analyzing, and Debugging Microservices
(This page has no text content)
Austin Parker, Daniel Spoonhower, Jonathan Mace, and Rebecca Isaacs with Ben Sigelman Distributed Tracing in Practice Instrumenting, Analyzing, and Debugging Microservices Boston Farnham Sebastopol TokyoBeijing
978-1-492-05663-8 [LSI] Distributed Tracing in Practice by Austin Parker, Daniel Spoonhower, Jonathan Mace, and Rebecca Isaacs, with Ben Sigelman Copyright © 2020 Ben Sigelman, Austin Parker, Daniel Spoonhower, Jonathan Mace, and Rebecca Isaacs. All rights reserved. Printed in the United States of America. Published by O’Reilly Media, Inc., 1005 Gravenstein Highway North, Sebastopol, CA 95472. O’Reilly books may be purchased for educational, business, or sales promotional use. Online editions are also available for most titles (http://oreilly.com). For more information, contact our corporate/institutional sales department: 800-998-9938 or corporate@oreilly.com. Acquisitions Editor: John Devins Development Editor: Sarah Grey Production Editor: Katherine Tozer Copyeditor: Chris Morris Proofreader: JM Olejarz Indexer: Sue Klefstad Interior Designer: David Futato Cover Designer: Karen Montgomery Illustrator: Rebecca Demarest April 2020: First Edition Revision History for the First Edition 2020-04-13 First Release See http://oreilly.com/catalog/errata.csp?isbn=9781492056638 for release details. The O’Reilly logo is a registered trademark of O’Reilly Media, Inc. Distributed Tracing in Practice, the cover image, and related trade dress are trademarks of O’Reilly Media, Inc. The views expressed in this work are those of the authors, and do not represent the publisher’s views. While the publisher and the authors have used good faith efforts to ensure that the information and instructions contained in this work are accurate, the publisher and the authors disclaim all responsibility for errors or omissions, including without limitation responsibility for damages resulting from the use of or reliance on this work. Use of the information and instructions contained in this work is at your own risk. If any code samples or other technology this work contains or describes is subject to open source licenses or the intellectual property rights of others, it is your responsibility to ensure that your use thereof complies with such licenses and/or rights.
Table of Contents Foreword. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ix Introduction: What Is Distributed Tracing?. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . xiii 1. The Problem with Distributed Tracing. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1 The Pieces of a Distributed Tracing Deployment 3 Distributed Tracing, Microservices, Serverless, Oh My! 4 The Benefits of Tracing 6 Setting the Table 7 2. An Ontology of Instrumentation. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9 White Box Versus Black Box 10 Application Versus System 13 Agents Versus Libraries 15 Propagating Context 16 Interprocess Propagation 18 Intraprocess Propagation 20 The Shape of Distributed Tracing 23 Tracing-Friendly Microservices and Serverless 23 Tracing in a Monolith 25 Tracing in Web and Mobile Clients 27 3. Open Source Instrumentation: Interfaces, Libraries, and Frameworks. . . . . . . . . . . . . . 31 The Importance of Abstract Instrumentation 32 OpenTelemetry 34 OpenTracing and OpenCensus 43 OpenTracing 43 OpenCensus 48 iii
Other Notable Formats and Projects 53 X-Ray 53 Zipkin 54 Interoperability and Migration Strategies 54 Why Use Open Source Instrumentation? 57 Interoperability 58 Portability 58 Ecosystem and Implicit Visibility 59 4. Best Practices for Instrumentation. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61 Tracing by Example 61 Installing the Sample Application 62 Adding Basic Distributed Tracing 62 Custom Instrumentation 70 Where to Start—Nodes and Edges 71 Framework Instrumentation 72 Service Mesh Instrumentation 75 Creating Your Service Graph 76 What’s in a Span? 79 Effective Naming 79 Effective Tagging 80 Effective Logging 81 Understanding Performance Considerations 82 Trace-Driven Development 85 Developing with Traces 86 Testing with Traces 89 Creating an Instrumentation Plan 91 Making the Case for Instrumentation 91 Instrumentation Quality Checklist 93 Knowing When to Stop Instrumenting 95 Smart and Sustainable Instrumentation Growth 97 5. Deploying Tracing. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 99 Organizational Adoption 100 Start Close to Your Users 100 Start Centrally: Load Balancers and Gateways 101 Leverage Infrastructure: RPC Frameworks and Service Meshes 102 Make Adoption Repeatable 103 Tracer Architecture 104 In-Process Libraries 105 Sidecars and Agents 106 Collectors 107 iv | Table of Contents
Centralized Storage and Analysis 108 Incremental Deployment 109 Data Provenance, Security, and Federation 110 Frontend Service Telemetry 110 Server-Side Telemetry for Managed Services 114 6. Overhead, Costs, and Sampling. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 117 Application Overhead 118 Latency 118 Throughput 120 Infrastructure Costs 122 Network 122 Storage 123 Sampling 124 Minimum Requirements 124 Strategies 126 Selecting Traces 130 Off-the-Shelf ETL Solutions 131 7. A New Observability Scorecard. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 135 The Three Pillars Defined 136 Metrics 136 Logging 138 Distributed Tracing 139 Fatal Flaws of the Three Pillars 140 Design Goals 141 Assessing the Three Pillars 142 Three Pipes (Not Pillars) 144 Observability Goals and Activities 145 Two Goals in Observability 145 Two Fundamental Activities in Observability 146 A New Scorecard 148 The Path Ahead 152 8. Improving Baseline Performance. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 153 Measuring Performance 154 Percentiles 156 Histograms 158 Defining the Critical Path 160 Approaches to Improving Performance 163 Individual Traces 163 Biased Sampling and Trace Comparison 165 Table of Contents | v
Trace Search 167 Multimodal Analysis 169 Aggregate Analysis 171 Correlation Analysis 173 9. Restoring Baseline Performance. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 179 Defining the Problem 180 Human Factors 182 (Avoiding) Finger-Pointing 182 “Suppressing” the Messenger 183 Incident Hand-off 184 Good Postmortems 184 Approaches to Restoring Performance 185 Integration with Alerting Workflows 185 Individual Traces 186 Biased Sampling 187 Real-Time Response 189 Knowing What’s Normal 191 Aggregate and Correlation Root Cause Analysis 195 10. Are We There Yet? The Past and Present. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 201 Distributed Tracing: A History of Pragmatism 202 Request-Based Systems 202 Response Time Matters 202 Request-Oriented Information 203 Notable Work 203 Pinpoint 204 Magpie 204 X-Trace 206 Dapper 207 Where to Next? 208 11. Beyond Individual Requests. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 209 The Value of Traces in Aggregate 211 Example 1: Is Network Congestion Affecting My Application? 211 Example 2: What Services Are Required to Serve an API Endpoint? 211 Organizing the Data 212 A Strawperson Solution 212 What About the Trade-offs? 214 Sampling for Aggregate Analysis 214 The Processing Pipeline 215 Incorporating Heterogeneous Data 217 vi | Table of Contents
Custom Functions 217 Joining with Other Data Sources 218 Recap and Case Study 219 The Value of Traces in Aggregate 219 Organizing the Data 220 Sampling for Aggregate Analysis 220 The Processing Pipeline 221 Incorporating Heterogeneous Data 221 12. Beyond Spans. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 223 Why Spans Have Prevailed 223 Visibility 223 Pragmatism 224 Portability 224 Compatibility 225 Flexibility 225 Why Spans Aren’t Enough 225 Graphs, Not Trees 226 Inter-Request Dependencies 227 Decoupled Dependencies 228 Distributed Dataflow 229 Machine Learning 230 Low-Level Performance Metrics 231 New Abstractions 232 Seeing Causality 234 13. Beyond Distributed Tracing. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 237 Limitations of Distributed Tracing 238 Challenge 1: Anticipating Problems 239 Challenge 2: Completeness Versus Costs 240 Challenge 3: Open-Ended Use Cases 240 Other Tools Like Distributed Tracing 241 Census 241 A Motivating Example 242 A Distributed Tracing Solution? 243 Tag Propagation and Local Metric Aggregation 244 Comparison to Distributed Tracing 245 Pivot Tracing 246 Dynamic Instrumentation 246 Recurring Problems 247 How Does It Work? 247 Dynamic Context 248 Table of Contents | vii
Comparison to Distributed Tracing 248 Pythia 249 Performance Regressions 249 Design 251 Overheads 251 Comparison to Distributed Tracing 251 14. The Future of Context Propagation. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 253 Cross-Cutting Tools 253 Use Cases 254 Distributed Tracing 254 Cross-Component Metrics 255 Cross-Component Resource Management 255 Managing Data Quality Trade-offs 256 Failure Testing of Microservices 257 Enforcing Cross-System Consistency 258 Request Duplication 258 Record Lineage in Stream Processing Systems 259 Auditing Security Policies 259 Testing in Production 259 Common Themes 260 Should You Care? 260 The Tracing Plane 261 Is Baggage Enough? 262 Beyond Key-Value Pairs 264 Compiling BDL 265 BaggageContext 266 Merging 266 Overheads 266 A. The State of Distributed Tracing Circa 2020. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 269 B. Context Propagation in OpenTelemetry. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 275 Bibliography. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 281 Index. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 285 viii | Table of Contents
Foreword Human beings have struggled to understand production software for exactly as long as human beings have had production software. We have these marvelously fast machines, but they don’t speak our language and—despite their speed and all of the hype about artificial intelligence—they are still entirely unreflective and opaque. For many (many) decades, our efforts to understand production software ultimately boiled down to two types of telemetry data: log data and time series statistics. The time series data—also known as metrics—helped us understand that “something ter‐ rible” was happening inside of our computers. If we were lucky, the logging data would help us understand specifically what that terrible thing was. But then everything changed: our software needed more than just one computer. In fact, it needed thousands of them. We broke the software into tiny, independently operated services and distributed those fragmented services across the planet, atomized among the millions of comput‐ ers housed in massive datacenters. And with so many processes involved in every end-user request, the logs and statistics from individual machines told only a sliver of the story. It felt like we were flying blind. I started working on distributed tracing in early 2005. At the time, I was a 25-year-old software engineer working—somewhat grudgingly, if I’m being candid—on a far- flung service within the Google AdWords backend infrastructure. Like the rest of the company, I was trying to write software capable of withstanding a punishing load from the outside world (by this point, Google was already a verb and we had scaled well into uncharted territory for commodity hardware). We were running microservi‐ ces before that term had been invented, and when we needed some new abstraction layer or infrastructure, we were almost always forced to write it in-house (for one thing, GitHub hadn’t even been incorporated yet). ix
To make a long story short, we kept the ship afloat…but it was a mess. And nobody except for the old-timer super-geniuses (read: not me) had any clue where the bodies were buried or how it all actually fit together. That’s when I met Sharon Perl, truly by accident. She had been a research scientist at DEC’s Systems Research Center in the 1990s (i.e., when it was cool!) and came to Google in the very early days: 2001, if I remember correctly. In that short impromptu conversation with Sharon, I asked her what she was working on, and she rattled off a list of interesting systems software projects: a distributed blob store, a Google-scale identity service, a distributed lockservice…and then this thing called Dapper. Dapper was “a distributed tracing system,” whatever that was. Needless to say, I had never heard of a distributed tracing system—in 2005, hardly any nonacademic had—but it sounded fascinating. At the time, Dapper was just a prototype that Sharon codeveloped with Mike Burrows and Luiz Barroso. They had patched Google’s internal RPC subsystem and control-flow packages in order to propagate a few GUIDs alongside each request as the request bounced from service to service. It wasn’t fully operational, but an early proof-of-concept showed that the fundamentals were sound. For the first time, an ordinary Google engineer actually had some hope of understanding what happened to an individual web request in the 150 milliseconds it took to touch hundreds or thousands of distinct microservices. I was hooked. Here was something truly novel, powerful, and—from a personal standpoint—wildly understaffed! So I started to dig into the Dapper codebase, clean things up, round edges, and deal with more than my share of internal bureaucracy (among other things, Dapper had a daemon running with root privileges on every piece of production hardware at Google and, wisely, they put some process around that sort of thing). Suffice it to say, after a year or two of wall time and some phenom‐ enal work from a team of engineers, we were able to deploy Dapper across all of Goo‐ gle’s backend software. To the best of my knowledge, this was the first time any organization had run distributed tracing continuously for a production system at scale. …and so we deployed Dapper across all of Google, and we solved observability. If only! The truth is that Dapper was a point solution to some painful yet isolated problems. In the early days, it was hard to get people to even use it, much less benefit from it. My team’s KPI was the number of weekly logins, and I remember when the number hovered in the low double digits, month after month. We would dream up clever analytical features, deploy them, wait for the thundering herds of enthusiastic users, and then feel disappointed. Eventually we did find a way to increase usage and thus organizational value to Goo‐ gle, but it wasn’t with new analytical features or insightful visualizations. In fact, it was something really basic that only required a few hundred lines of code: one of my x | Foreword
colleagues integrated links to relevant Dapper traces into a tool that Google engineers already used many times a day. It turns out that some small fraction of people would click on those links, and sometimes they found something really valuable on the other side. That was it. Just a simple integration into an existing workflow. Beyond the data engi‐ neering and instrumentation challenges, distributed tracing is hard because it’s often thought of not just as a new set of telemetry but as a distinct, segregated product experience. No matter how compelling that product experience is, developers (like all people) are creatures of habit who do not want to learn a new tool to check proac‐ tively. Tracing data and insights must fit into the context of preexisting workflows and tasks to be done. This is the best way to give tracing-oriented insights the expo‐ sure needed to justify the investment in a fundamentally new data source. Distributed tracing is still in its infancy. Thinking back to the early days of the Dap‐ per project, when I was just ramping up on the codebase, I asked Luiz Barroso if he could spare 30 minutes to help me understand a few things. Luiz was already quite distinguished, but was (and remains) humble, friendly, and generous with his time, so he agreed. When I met with him, I must have sounded a bit naive, but I was also unfathomably excited about what I wanted to do to Dapper. I wanted to build in a just-in-time sampling mechanism, create a declarative programming language for user-defined queries that execute across application services, integrate kernel traces, and more. I asked him what he thought. Ever the voice of wisdom, he let me down easy and explained that simply getting Dapper into production would be a major accomplishment and would take years. “Start there,” he said. Luiz was right about that. Fifteen years later, much of our industry hasn’t gotten a whole lot further than that, at least in production. Distributed tracing is worth it, but it’s hard! Still, it’s a very young discipline, and the last section of this book provides a window on what’s still to come. In another 15 years, we will look back on distributed tracing circa 2020 as both critical and primitive. By understanding where the technol‐ ogy is going, we’ll be better able to position ourselves to adapt to the dynamic land‐ scape surrounding tracing and observability in general. Stepping back, it’s important to remember that nobody works with “just one microservice.” Our industry moved to microservices so that our dev teams could operate with inde‐ pendence, and to a certain extent, we got our wish—at least where continuous inte‐ gration/continuous deployment is concerned. But this “independence” was an illusion; in production, these microservices are in fact highly interdependent, and a failure or slowdown in one service propagates across the stack of microservices, leav‐ ing chaos and confusion (and many frantic Slack messages) in its wake. Foreword | xi
Distributed traces must be part of the solution to this problem. They are the only window we have into how the hundreds of services in deep, multilayered microser‐ vice architectures actually interact as they fulfill end-to-end user requests. They may be a relative newcomer to the telemetry world compared to time series stats and vanilla logs, but they are also the most vital when it comes to understanding the larger system. Without tracing data, we are reduced to guess-and-check across seas of disorganized logging data and metrics dashboards. Yet it’s not nearly as simple as adding distributed tracing. While healthy observability in distributed systems must involve distributed traces, we still need to figure out how. How do we make distributed tracing useful? How do we adopt it? How do we inte‐ grate it into our existing workflows and processes? And how do we future-proof these efforts? These are fascinating and challenging questions, and they are the subject of this book. We hope you enjoy it. — Ben Sigelman Cofounder and CEO of Lightstep and cocreator of Dapper xii | Foreword
Introduction: What Is Distributed Tracing? If you’re reading this book, you may already have some idea what the words dis‐ tributed tracing mean. You may also have no idea what they mean—for all we know, you’re simply a fan of bandicoots (the animal on the cover). We won’t judge, promise. Either way, you’re reading this to gain some insight into what distributed tracing is, and how you can use it to understand the performance and operation of your micro‐ services and other software. With that in mind, let’s start out with a simple definition. Distributed tracing (also called distributed request tracing) is a type of correlated log‐ ging that helps you gain visibility into the operation of a distributed software system for use cases such as performance profiling, debugging in production, and root cause analysis of failures or other incidents. It gives you the ability to understand exactly what a particular individual service is doing as part of the whole, enabling you to ask and answer questions about the performance of your services and your distributed system as a whole. That was easy—see you next book! What’s that? Why’s everyone asking for a refund? Oh… We’re being told that you need a little more than that. Well, let’s take a step back and talk about software, specifically distributed software, so that we can better understand the problems that distributed tracing solves. Distributed Architectures and You The art and science of developing, deploying, and operating software is constantly in flux. New advances in computing hardware and software have dramatically pushed the boundaries of what an application looks like over the years. While there’s an inter‐ esting digression here about how “everything old is new again,” we’ll focus on changes over the past two decades or so, for the sake of brevity. xiii
Prior to advances in virtualization and containerization, if you needed to deploy some sort of web-based application, you would need a physical server, possibly one dedicated to your application itself. As traffic increased to your application, you would either need to increase the physical resources of that server (adding RAM, for example) or you would need multiple servers that each ran their own copy of your application. With a monolithic server process, this horizontal scaling often led to unfavorable trade-offs in cost, performance, and organizational overhead. Running multiple instances of your server meant you were duplicating all functionality of the server, rather than scaling individual subcomponents independently. With traditional infra‐ structure, you were often forced to make a decision about how many minutes (or hours!) of degraded performance was acceptable while you brought additional capacity online—servers aren’t cheap to run, so why would you run at peak capacity if you didn’t need to? Finally, as the size and complexity of your application increased, along with the amount of developers who were working on it, testing and validating new changes became more difficult. As your organization grew, it became unreasona‐ ble for developers to understand a single codebase, not to mention the shape of the entire system. Increasingly smaller changes increased the odds of a ripple effect that led to total application failure as their impact radiated out from one component to another. Time marched on, however, and solutions to these problems were built. Software was created that abstracted away the details of physical hardware such as virtualization, allowing for a single physical server to be split into multiple logical servers. Docker and other containerization technologies extended this concept, providing a light‐ weight and user-friendly abstraction over heavier-weight virtual machines, moving the question of “who deploys this software” from operators to developers. The popu‐ larization of cloud computing and its notion of on-demand computing resources solved the problem of resource scaling, as it became possible to increase the amount of RAM or CPU cores for a given server at the click of a button. Finally, the idea of microservice architectures came about to address the complexity imposed by ever- larger and more complicated software-oriented businesses by structuring large appli‐ cations around loosely coupled independent services. Today, it’s arguable that most applications are distributed in some fashion, even if they don’t use microservices. Simple client-server applications themselves are distributed —consider the classic question of “A call to my server has timed out; was the response lost, or was the work not done at all?” Additionally, they may have a variety of dis‐ tributed dependencies, such as datastores that are consumed as a service offered by a cloud provider, or a whole host of third-party APIs that provide everything from ana‐ lytics to push notifications and more. xiv | Introduction: What Is Distributed Tracing?
1 [Sig19] Why is distributed software so popular? The arguments for distributed software are pretty clear: Scalability A distributed application can more easily respond to demand, and its scaling can be more efficient. If a lot of people are trying to log in to your application, you could scale out only the login services, for example. Reliability Failures in one component shouldn’t bring down the entire application. Dis‐ tributed applications are more resilient because they split up functions through a variety of service processes and hosts, ensuring that even if a dependent service goes offline, it shouldn’t impact the rest of the application. Maintainability Distributed software is more easily maintainable for a couple of reasons. Divid‐ ing services from each other can increase how maintainable each component is by allowing it to focus on a smaller set of responsibilities. In addition, you’re freer to add features and capabilities without implementing (and maintaining) them yourself—for example, adding a speech-to-text function in an application by relying on some cloud provider’s speech-to-text service. This is the tip of the iceberg, so to speak, in terms of the benefits of distributed archi‐ tectures. Of course, it’s not all sunshine and roses, and into every life, a little rain must fall… Deep Systems A distributed architecture is a prime example of what software architects often call a deep system.1 These systems are notable not because of their width, but because of their complexity. If you think about certain services or classes of services in a dis‐ tributed architecture, you should be able to identify the difference. A pool of cache nodes scales wide (as in, you simply add more instances to handle demand), but other services scale differently. Requests may route through three, four, fourteen, or forty different layers of services, and each of those layers may have other dependen‐ cies that you aren’t aware of. Even if you have a comparatively simple service, your software probably has dozens of dependencies on code that you didn’t write, or on managed services through a cloud provider, or even on the underlying orchestration software that manages its state. The problem with deep systems is ultimately a human one. It quickly becomes unre‐ alistic for a single human, or even a group of them, to understand enough of the Introduction: What Is Distributed Tracing? | xv
services that are in the critical path of even a single request and continue maintaining it. The scope of what you as a service owner can control versus what you’re implicitly responsible for is illustrated in Figure P-1. This calculus becomes a recipe for stress and burnout, as you’re forced into a reactive state against other service owners, con‐ stantly fighting fires, and trying to figure out how your services interact with each other. Figure P-1. The service that you can control has dependencies that you’re responsible for but have no direct control over. Distributed architectures require a reimagined approach to understanding the health and performance of software. It’s not enough to simply look at a single stack trace or watch graphs of CPU and memory utilization. As software scales—in depth, but also in breadth—telemetry data like logs and metrics alone don’t provide the clarity you require to quickly identify problems in production. The Difficulties of Understanding Distributed Architectures Distributing your software presents new and exciting challenges. Suddenly, failures and crashes become harder to pin down. The service that you’re responsible for may be receiving data that’s malformed or unexpected from a source that you don’t control because that service is managed by a team halfway across the globe (or a remote team). Failures in services that you thought were rock-solid suddenly cause cascading failures and errors across all of your services. To borrow a phrase from Twitter, you’ve got a microservices murder mystery (see Figure P-2) on your hands. xvi | Introduction: What Is Distributed Tracing?
Figure P-2. It’s funny because it’s true. To extend the metaphor, monitoring helps determine where the body is, but it doesn’t reveal why the murder occurred. Distributed tracing fills in those gaps by allowing you to easily comprehend your entire system by providing solutions to three major pain points: Obfuscation As your application becomes more distributed, the coherence of failures begins to decrease. That is to say, the distance between cause and effect increases. An outage at your cloud provider’s blob storage could fan out to cause huge cascad‐ ing latency for everyone, or a single difficult-to-diagnose failure at a particular service many hops away that prevents you from uncovering the proximate cause. Inconsistency Distributed applications might be reliable overall, but the state of individual com‐ ponents can be much less consistent than they would be in monolithic or non- distributed applications. In addition, since each component of a distributed application is designed to be highly independent, the state of those components will be inconsistent—what happens when someone does a deployment, for exam‐ ple? Do all of the other components understand what to do? How does that impact the overall application? Decentralized Critical data about the performance of your services will be, by definition, decen‐ tralized. How do you go looking for failures in a service when there may be a thousand copies of that service running, on hundreds of hosts? How do you cor‐ relate those failures? The greatest strength of distributing your application is also the greatest impediment to understanding how it actually functions! You may be wondering, “How do we address these difficulties?” Spoiler: distributed tracing. Introduction: What Is Distributed Tracing? | xvii
How Does Distributed Tracing Help? Distributed tracing emerges as a critical tool in managing the explosion of complexity that our deep systems bring. It provides context that spans the life of a request and can be used to understand the interactions and shape of your architecture. However, these individual traces are just the beginning—in aggregate, traces can give you important insights about what’s actually going on in your distributed system, allowing you not only to correlate interesting data about your services (for example, that most of your errors are happening on a specific host or in a specific database cluster), but also to filter and rank the importance of other types of telemetry. Effectively, dis‐ tributed traces provide context that helps you filter problem-solving down to only things that are relevant to your investigation, so you don’t have to guess and check multiple logs and dashboards. In this way, distributed tracing is actually at the center of a modern observability platform, and it becomes a critical component of your dis‐ tributed architecture rather than an isolated tool. So, what is a trace? The easiest way to understand is to think about your software in terms of requests. Each of your components is in the business of doing some sort of work in response to a request (aka RPC, from remote procedure call) from another service. This could be as prosaic as a web page requesting some structured data from a service endpoint to present to a user, or as complex as a highly parallelized search process. The actual nature of the work doesn’t matter too much, although there are certain patterns that we’ll discuss later on that lend themselves to certain styles of tracing. While distributed tracing can function in most distributed systems, as we’ll discuss in Chapter 4, its strengths are best demonstrated in modeling the RPC rela‐ tionships between your services. In addition to the RPC relationships, think about the work that each of those services does. Maybe they’re authenticating and authorizing user roles, performing mathemat‐ ical calculations, or simply transforming data from one format to another. These services are communicating with each other through RPCs, sending requests and receiving responses. Regardless of what they’re doing, one thing that all of these serv‐ ices have in common is that the work they’re performing takes some length of time. The basic pattern of services and RPCs is illustrated in Figure P-3. Figure P-3. A request from a client process to a service process. We call the work that each service is doing a span, as in the span of time that it takes for the work to occur. These spans can be annotated with metadata (known as xviii | Introduction: What Is Distributed Tracing?
Loading comments...
Reply to Comment
Edit Comment