Become an expert in implementing observability methods for legacy technologies and discover how to use AIOps and OpenTelemetry to analyze root causes and solve problems in banking and telecommunications. Through this book, you will engage with issues that occur in kernels, networks, CPU, and IO by developing skills to handle traces and logs, as well as Profiles (eBPF) and debugging. The real-world examples in the book will enable you to analyze and aggregate observability data, helping you gain competence in automating systems and resolving business-critical issues rapidly and efficiently.
The book will introduce you to new observability approaches, describe different types of errors, and explain how observability addresses them. It will provide training on how to develop dashboards and charts and design a root cause analysis process. Emphasizing trace-centric observability, you will gain expertise in using EAI servers to integrate legacy tech and using extensions to complement the OpenTelemetry Agent. You will also understand the varied practical uses of OpenTelemetry through examples from multiple industries, as well as an OpenTelemetry demo application.
The book then takes you through infrastructure observability and infrastructure anomaly detection, enabling you to visualize and trace problems, and helping you identify and proactively respond to anomalies in system resources. In the final chapters, you will learn how to aggregate and analyze observability data using Presto and Druid. Finally, you will familiarize yourself with AIOps and learn how to implement it with Langchain and RAGs.
By the end of this book, you will be fully trained in the practical implementation of observability and using observability data to identify, analyze, and solve problems for large industries like finance and telecommunications.
For
System developers, data engineers, SREs, infrastructure engineers, system architects, Java developers, and DevOps engineers who…
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
【One-Line Pitch】
A practical field manual for engineers wrestling with observability in aging, mission-critical systems, showing how to apply OpenTelemetry, AIOps, and root-cause analysis to banking, telecom, and other legacy-heavy industries.
【Book Arc】
- **Opening (~0%–9%)**: Establishes the core problem—legacy systems lack modern observability—and introduces the Root Cause Analysis (RCA) process, from identifying problem areas to analyzing individual requests and understanding low-level methods. It also lays out the fundamental observability signals: logs, RUM, profiles, debugging, and events, plus a data model for RCA.
- **Early (~16%–25%)**: Moves into practical tooling and correlation. Covers Grafana's correlation features (metrics↔traces↔logs), introduces The New Stack (TNS) configuration, and walks through setting up Grafana and running the "o11y Shop" demo application, including dashboard development.
- **Early (~25%–34%)**: Dives deep into trace-centric RCA. Explains how traces work—context, propagators, baggage—and demonstrates propagation across managed services (AWS, GCP, Azure), message services (Solace, TIBCO, MQTT, Kafka, Spring Cloud Stream), EAI servers, and black-box systems.
- **Middle (~34%–44%)**: Applies observability to real industries: banking (from RUM to API server, Kafka, microservices, EAI, and Jaeger), telecom (order orchestration, network provisioning), online games, and ultra-low-latency trading systems. Also covers advanced OpenTelemetry topics like baggage context, span attributes, annotations, and Promscale with Kubernetes and SQL.
- **Middle (~44%–47%)**: Shifts to infrastructure-level RCA. Introduces system traces using KUtrace, ftrace, and Kubeshark, explains kernel internals and development, and covers eBPF with BCC and bpftrace, plus PCP. Ends with chaos engineering and a demo.
- **Late (~47%–end)**: The excerpts do not cover the final chapters in detail, but based on the book's stated scope, the remaining content covers infrastructure anomaly detection, aggregating observability data with Presto and Druid, and implementing AIOps using LangChain and RAGs.
【Key Takeaways】
- **RCA is a structured process, not a guessing game** (Opening): The book frames root cause analysis as a repeatable method—identify the problem area, analyze individual requests, then drill into low-level methods—rather than ad-hoc debugging. This gives teams a systematic starting point for legacy systems.
- **Observability is more than logs** (Opening): Logs, RUM, profiles, debugging, and events are presented as complementary signals, each with a role in the RCA data model. Understanding their interplay is essential before choosing tools.
- **Correlation between signals is the real value** (Early): Grafana's correlation features (metrics to traces, logs to traces, traces to metrics) are demonstrated as a way to navigate between data types. The o11y Shop demo provides a hands-on sandbox to practice these workflows.
- **Trace propagation is the hardest part of legacy integration** (Early): The book dedicates significant space to propagating context across managed services (AWS, GCP, Azure), message queues (Kafka, MQTT, JMS), and EAI servers. This is where most real-world implementations fail, and the demos are the core value.
- **Industry patterns differ significantly** (Middle): Banking, telecom, games, and trading each have distinct observability needs—from RUM-to-legacy flows in banking to ultra-low-latency design in trading. The book provides architecture-specific walkthroughs rather than generic advice.
- **Infrastructure observability requires kernel-level tools** (Middle): KUtrace, ftrace, eBPF (BCC, bpftrace), and Kubeshark are introduced for tracing system-level behavior. This is essential for diagnosing issues that span application and infrastructure boundaries.
- **AIOps is the endgame** (Late, inferred): The book's final chapters aim to close the loop by aggregating observability data with Presto and Druid, then applying AIOps via LangChain and RAGs to automate anomaly detection and response—though the excerpts do not detail these chapters.
【Reading Tips】
- **Skim the early signal definitions** (~0–9%) if you already know logs/metrics/traces; the RCA process framework is worth reading once, but the real meat is in the correlation and propagation chapters.
- **Deep-read the trace propagation chapters** (~25–34%): These demos (AWS, Kafka, EAI, black-box) are the most transferable to real legacy systems. If you only have time for one section, make it this one.
- **Use the industry chapters as reference** (~34–44%): Pick the chapter closest to your domain (banking, telecom, games, trading) and read that deeply; skim the others for patterns.
- **Expect a steep learning curve in the infrastructure chapters** (~44–47%): Kernel internals and eBPF require prior Linux knowledge. If you're not an infrastructure engineer, skim the concepts and focus on the tools' use cases.
- **The final AIOps chapters are not covered in the excerpts**: If your goal is AIOps implementation, you may need to supplement with external resources on LangChain and RAGs, as the book's coverage here is not verifiable from the available material.
【Coverage Limits】
This guide is based on excerpts covering roughly the first half of the book (through chaos engineering). The final chapters on Presto/Druid aggregation and AIOps with LangChain/RAGs are not covered in the source material and are only inferred from the book's stated scope.
Excerpt 1
large industries like finance and telecommunications. For System developers, data engineers, SREs, infrastructure engineers, system architects, Java develope...
................................................. 116 2.4.1. Introducing The New Stack (TNS) ...................................................................
................................................................... 611 8.1.1. Time Windows ....................................................................
stand and solve the problem, traces are discussed in detail. Traces consist of two types: distributed traces and system traces. My approach is to start with...
is book focuses more on infrastructure than on applications. Part 4, “Leveraging Observability Data,” includes the following: • Chapter 8, Analyze RCA Data:...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
Observability For Legacy Systems Methods and Solutions with OpenTelemetry and AIOps (Hyen Seuk Jeong)(Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
Observability For Legacy Systems Methods and Solutions with OpenTelemetry and AIOps (Hyen Seuk Jeong)(Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment