Get up to speed with Prometheus, the metrics-based monitoring system used in production by tens of thousands of organizations. This updated second edition provides site reliability engineers, Kubernetes administrators, and software developers with a hands-on introduction to the most important aspects of Prometheus, including dashboarding and alerting, direct code instrumentation, and metric collection from third-party systems with exporters.
Prometheus server maintainer Julien Pivotto and core developer Brian Brazil demonstrate how you can use Prometheus for application and infrastructure monitoring. This book guides you through Prometheus setup, the Node Exporter, and the Alertmanager, and then shows you how to use these tools for application and infrastructure monitoring. You'll understand why this open source system has continued to gain popularity in recent years.
You will:
• Know where and how much instrumentation to apply to your application code
• Monitor your infrastructure with Node Exporter and use new collectors for network system pressure metrics
• Get an introduction to Grafana, a popular tool for building dashboards
• Use service discovery and the new HTTP SD monitoring system to provide different views of your machines and services
• Use Prometheus with Kubernetes and examine exporters you can use with containers
• Discover Prom's new improvements and features, including trigonometry functions
• Learn how Prometheus supports important security features including TLS and basic authentication
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
# Prometheus: Up & Running — Infrastructure and Application Performance Monitoring, 2nd Edition
## 【One-Line Pitch】
A practical, hands-on guide for SREs, Kubernetes administrators, and developers who want to master Prometheus for monitoring both applications and infrastructure — covering everything from setup and instrumentation to alerting, exporters, and Kubernetes integration.
## 【Book Arc】
- **Opening (~0%–9%)**: Introduces Prometheus's role in modern monitoring, contrasts it with legacy tools like MRTG and Graphite, and explains why metrics-based monitoring fits cloud-native environments where services are treated as "cattle" rather than "pets."
- **Early (~9%–25%)**: Walks through getting Prometheus running, configuring scrape jobs, using the expression browser, and setting up the Node Exporter for machine-level metrics — including a first alerting rule and Alertmanager configuration.
- **Early (~25%–34%)**: Dives into instrumentation fundamentals: client libraries, the /metrics endpoint, basic metric types (counters, gauges), and how to expose application metrics to Prometheus.
- **Middle (~34%–47%)**: Covers advanced instrumentation patterns — histograms and summaries for latency, unit testing instrumentation, handling multiprocess applications, and using the Pushgateway for batch jobs.
- **Middle (~47%–end)**: Explores the OpenMetrics exposition format, additional metric types (StateSet, GaugeHistograms, Info), and then moves into common exporters (Consul, MySQLd, Blackbox), Kubernetes integration with cAdvisor and kube-state-metrics, and working with other monitoring systems.
## 【Key Takeaways】
- **Metrics-based monitoring is the right fit for cloud-native scale** (Early): Unlike event-based logging, metrics have memory usage tied to the number of metrics, not event volume — making them scalable for dynamic, containerized environments.
- **Client libraries handle the hard parts of instrumentation** (Early): Thread safety, bookkeeping, and exposition format generation are abstracted away; instrumenting a shared library (like an RPC client) automatically instruments all dependent applications.
- **Exporters bridge the gap for systems you can't instrument directly** (Early): For kernels, databases, and third-party software, exporters translate existing interfaces (SNMP, ad hoc formats) into Prometheus metrics — the Node Exporter being the canonical example.
- **The `rate()` function is essential for working with counters** (Early): It automatically handles counter resets and sample misalignment, making it the correct way to derive per-second rates from monotonically increasing counters.
- **Histograms beat averages for latency tracking** (Middle): Averages hide outliers; histograms let you compute quantiles and switch between percentiles and averages as needed, using the `_sum` and `_count` series.
- **Unit test only critical instrumentation** (Middle): Testing every metric adds friction and reduces breadth; focus tests on key metrics (requests, latency, errors) in core libraries, accepting that ~5% of debug metrics may break.
- **The Pushgateway is for batch jobs, not regular services** (Middle): Use `push`, `pushadd`, and `delete` semantics carefully — it's designed for short-lived jobs that must report metrics before they exit.
- **OpenMetrics extends the exposition format** (Middle): New types like StateSet, GaugeHistograms, and Info provide richer semantics, while UNIT metadata clarifies metric meaning — all served with a versioned content type.
## 【Reading Tips】
- **Skim the historical background in Chapter 1** (Early): The MRTG/Graphite history is interesting context but not essential — jump ahead to the practical setup if you're already familiar with monitoring concepts.
- **Deep-read the instrumentation chapters (3–4)** (Early–Middle): This is where the book earns its keep — understanding metric types, naming conventions, and multiprocess handling will save you hours of debugging later.
- **Use the example configurations as templates** (Early): The `prometheus.yml` and `rules.yml` examples are directly usable — adapt them rather than starting from scratch.
- **Pay special attention to the Pushgateway section** (Middle): It's a common source of misuse; understanding the three write methods (`push`, `pushadd`, `delete`) and when to use each is critical.
- **The Kubernetes and exporter chapters are reference material** (Middle–Late): Skim these to know what's available, then return when you need to integrate a specific system — they're organized by exporter for easy lookup.
## 【Coverage Limits】
The excerpts cover roughly the first half of the book (through the exposition format chapter). Later chapters on Kubernetes integration, common exporters, and working with other monitoring systems are only visible through the table of contents — details on those topics are not covered in this guide.
##
Excerpt 1
ion Performance Monitoring Julien Pivotto & Brian Brazil Using the Textfile Collector 135 Timestamps 137 8. Service Discovery. . . . . . . . . . . . . . . ....
the early 2000s, writing scripts to report temperature and network usage on my home computers. 6 | Chapter 1: What Is Prometheus? Client libraries take care...
e equiva‐ lents make usage simpler and clearer. Example 3-7. The same example as Example 3-6 but using the gauge utilities from prometheus_client import Gaug...
minated by # EOF: my_counter_total 14 a_small_gauge 8.3e-96 # EOF Metric Types The metric types supported by the Prometheus text exposition format are also s...
e Exporter as a service on the node, without Docker. Docker attempts to isolate a container from the inner workings of the machine, which doesn’t work well w...
histicated relabeling rules you may find yourself needing a temporary label to put a value in. The __tmp prefix is reserved for this purpose. How to Scrape Y...
e_duration_seconds gauge probe_duration_seconds 0.019061888 # HELP probe_icmp_duration_seconds Duration of icmp request by phase # TYPE probe_icmp_duration_s...
partition across a metric, and if you take a sum or average across a metric it should be meaningful, as discussed in “When to Use Labels” on page 98. In part...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
Prometheus Up Running Infrastructure and Application Performance Monitoring, 2nd Edition (Julien Pivotto, Brian Brazil)(Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
Prometheus Up Running Infrastructure and Application Performance Monitoring, 2nd Edition (Julien Pivotto, Brian Brazil)(Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment