Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Julien Pivotto, Brian Brazil

Rating No ratings yet

Get up to speed with Prometheus, the metrics-based monitoring system used in production by tens of thousands of organizations. This updated second edition provides site reliability engineers, Kubernetes administrators, and software developers with a hands-on introduction to the most important aspects of Prometheus, including dashboarding and alerting, direct code instrumentation, and metric collection from third-party systems with exporters. Prometheus server maintainer Julien Pivotto and core developer Brian Brazil demonstrate how you can use Prometheus for application and infrastructure monitoring. This book guides you through Prometheus setup, the Node Exporter, and the Alertmanager, and then shows you how to use these tools for application and infrastructure monitoring. You'll understand why this open source system has continued to gain popularity in recent years. You will: • Know where and how much instrumentation to apply to your application code • Monitor your infrastructure with Node Exporter and use new collectors for network system pressure metrics • Get an introduction to Grafana, a popular tool for building dashboards • Use service discovery and the new HTTP SD monitoring system to provide different views of your machines and services • Use Prometheus with Kubernetes and examine exporters you can use with containers • Discover Prom's new improvements and features, including trigonometry functions • Learn how Prometheus supports important security features including TLS and basic authentication

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# Prometheus: Up & Running — Infrastructure and Application Performance Monitoring, 2nd Edition ## 【One-Line Pitch】 A practical, hands-on guide for SREs, Kubernetes administrators, and developers who want to master Prometheus for monitoring both applications and infrastructure — covering everything from setup and instrumentation to alerting, exporters, and Kubernetes integration. ## 【Book Arc】 - **Opening (~0%–9%)**: Introduces Prometheus's role in modern monitoring, contrasts it with legacy tools like MRTG and Graphite, and explains why metrics-based monitoring fits cloud-native environments where services are treated as "cattle" rather than "pets." - **Early (~9%–25%)**: Walks through getting Prometheus running, configuring scrape jobs, using the expression browser, and setting up the Node Exporter for machine-level metrics — including a first alerting rule and Alertmanager configuration. - **Early (~25%–34%)**: Dives into instrumentation fundamentals: client libraries, the /metrics endpoint, basic metric types (counters, gauges), and how to expose application metrics to Prometheus. - **Middle (~34%–47%)**: Covers advanced instrumentation patterns — histograms and summaries for latency, unit testing instrumentation, handling multiprocess applications, and using the Pushgateway for batch jobs. - **Middle (~47%–end)**: Explores the OpenMetrics exposition format, additional metric types (StateSet, GaugeHistograms, Info), and then moves into common exporters (Consul, MySQLd, Blackbox), Kubernetes integration with cAdvisor and kube-state-metrics, and working with other monitoring systems. ## 【Key Takeaways】 - **Metrics-based monitoring is the right fit for cloud-native scale** (Early): Unlike event-based logging, metrics have memory usage tied to the number of metrics, not event volume — making them scalable for dynamic, containerized environments. - **Client libraries handle the hard parts of instrumentation** (Early): Thread safety, bookkeeping, and exposition format generation are abstracted away; instrumenting a shared library (like an RPC client) automatically instruments all dependent applications. - **Exporters bridge the gap for systems you can't instrument directly** (Early): For kernels, databases, and third-party software, exporters translate existing interfaces (SNMP, ad hoc formats) into Prometheus metrics — the Node Exporter being the canonical example. - **The `rate()` function is essential for working with counters** (Early): It automatically handles counter resets and sample misalignment, making it the correct way to derive per-second rates from monotonically increasing counters. - **Histograms beat averages for latency tracking** (Middle): Averages hide outliers; histograms let you compute quantiles and switch between percentiles and averages as needed, using the `_sum` and `_count` series. - **Unit test only critical instrumentation** (Middle): Testing every metric adds friction and reduces breadth; focus tests on key metrics (requests, latency, errors) in core libraries, accepting that ~5% of debug metrics may break. - **The Pushgateway is for batch jobs, not regular services** (Middle): Use `push`, `pushadd`, and `delete` semantics carefully — it's designed for short-lived jobs that must report metrics before they exit. - **OpenMetrics extends the exposition format** (Middle): New types like StateSet, GaugeHistograms, and Info provide richer semantics, while UNIT metadata clarifies metric meaning — all served with a versioned content type. ## 【Reading Tips】 - **Skim the historical background in Chapter 1** (Early): The MRTG/Graphite history is interesting context but not essential — jump ahead to the practical setup if you're already familiar with monitoring concepts. - **Deep-read the instrumentation chapters (3–4)** (Early–Middle): This is where the book earns its keep — understanding metric types, naming conventions, and multiprocess handling will save you hours of debugging later. - **Use the example configurations as templates** (Early): The `prometheus.yml` and `rules.yml` examples are directly usable — adapt them rather than starting from scratch. - **Pay special attention to the Pushgateway section** (Middle): It's a common source of misuse; understanding the three write methods (`push`, `pushadd`, `delete`) and when to use each is critical. - **The Kubernetes and exporter chapters are reference material** (Middle–Late): Skim these to know what's available, then return when you need to integrate a specific system — they're organized by exporter for easy lookup. ## 【Coverage Limits】 The excerpts cover roughly the first half of the book (through the exposition format chapter). Later chapters on Kubernetes integration, common exporters, and working with other monitoring systems are only visible through the table of contents — details on those topics are not covered in this guide. ##
Excerpt 1
ion Performance Monitoring Julien Pivotto & Brian Brazil Using the Textfile Collector 135 Timestamps 137 8. Service Discovery. . . . . . . . . . . . . . . ....
View in text
Excerpt 2
the early 2000s, writing scripts to report temperature and network usage on my home computers. 6 | Chapter 1: What Is Prometheus? Client libraries take care...
View in text
Excerpt 3
e equiva‐ lents make usage simpler and clearer. Example 3-7. The same example as Example 3-6 but using the gauge utilities from prometheus_client import Gaug...
View in text
Excerpt 4
minated by # EOF: my_counter_total 14 a_small_gauge 8.3e-96 # EOF Metric Types The metric types supported by the Prometheus text exposition format are also s...
View in text
Excerpt 5
e Exporter as a service on the node, without Docker. Docker attempts to isolate a container from the inner workings of the machine, which doesn’t work well w...
View in text
Excerpt 6
histicated relabeling rules you may find yourself needing a temporary label to put a value in. The __tmp prefix is reserved for this purpose. How to Scrape Y...
View in text
Excerpt 7
e_duration_seconds gauge probe_duration_seconds 0.019061888 # HELP probe_icmp_duration_seconds Duration of icmp request by phase # TYPE probe_icmp_duration_s...
View in text
Excerpt 8
partition across a metric, and if you take a sum or average across a metric it should be meaningful, as discussed in “When to Use Labels” on page 98. In part...
View in text
Tags
AI categories
Cloud NativeDevOpsProgramming
ISBN: 1098131142
Publisher: O'Reilly Media
Publish Year: 2023
Language: English
Pages: 418
File Format: PDF
File Size: 6.5 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…