Share E-Book

Agentic DevOps with Claude Code Build governed AI platforms on Kubernetes with GitOps, observability, and self-service… (Michael Forrester) (z-library.sk, 1lib.sk, z-lib.sk)

Author

Rating No ratings yet

Log in to rate

DevOps
Language English

Build and validate a governed AI-native developer platform with Claude, Kubernetes, GitOps, policy controls, observability, and reproducible workflows Key Features • Control Claude with specifications, permissions, tests, and auditable workflows • Build governed agent and model infrastructure on Kubernetes with GitOps • Create a Backstage self-service path from developer request to traced agent Build agentic DevOps workflows without bypassing the controls your Kubernetes platform already depends on. This book shows you how to use Claude as a controlled platform-engineering worker while introducing agents, model serving, and developer self-service through reproducible GitOps workflows, explicit trust boundaries, policy checks, and testable completion gates. You’ll establish a cloud-native foundation with Argo CD, cert-manager, OpenBao, External Secrets Operator, Kyverno, Prometheus, Grafana, Loki, Tempo, and OpenTelemetry. You’ll then add governed AI traffic using Gateway API, kgateway, agentgateway, kagent, MCP tools, and LLM Guard before serving an OpenAI-compatible model with KServe and vLLM. The hands-on approach shows you how to constrain Claude with specifications, permissions, audit hooks, tests, and Git checkpoints. You’ll trace agent and model activity, diagnose failures from evidence, and turn operational fixes into reusable tests. You’ll also build a Backstage template and Argo CD ApplicationSet that provide a governed path from developer request to running agent. By the end, you’ll be able to build and validate an AI-native internal developer platform in phases, route agent and model activity through existing platform controls, and prepare the architecture for production use. For: Platform engineers, DevOps engineers, SREs, cloud engineers, and platform architects who want to use Claude to build and operate governed AI capabilities on Kubernetes. Engineering and technical leads can also use the architecture and production guidance to assess scope and

Format PDF
Size 2.6 MB
6
Views
(First 20 pages)

Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Page 1
Michael Rishi Forrester Build governed AI platforms on Kubernetes with GitOps, observability, and self-service workfl ows Agentic DevOps with Claude Code Agentic D evO ps w ith Claude Code MICHAEL RISHI FORRESTER • Control Claude with specifications, permissions, and test gates • Design an AI-native IDP with clear ownership and trust boundaries • Build a reproducible Kubernetes foundation with Argo CD • Trace platform, agent, and model activity with OpenTelemetry • Govern LLM, MCP, and agent traffic through shared gateways • Run Kubernetes-native agents with guardrails and policy controls • Serve OpenAI-compatible models using KServe and vLLM • Build a governed Backstage self-service path for agent services WHAT YOU WILL LEARN Build agentic DevOps workfl ows without bypassing the controls your Kubernetes platform already depends on. This book shows you how to use Claude as a controlled platform-engineering worker while introducing agents, model serving, and developer self-service through reproducible GitOps workfl ows, explicit trust boundaries, policy checks, and testable completion gates. You'll establish a cloud-native foundation with Argo CD, cert-manager, OpenBao, External Secrets Operator, Kyverno, Prometheus, Grafana, Loki, Tempo, and OpenTelemetry. You'll then add governed AI traff ic using Gateway API, kgateway, agentgateway, kagent, MCP tools, and LLM Guard before serving an OpenAI-compatible model with KServe and vLLM. The hands-on approach shows you how to constrain Claude with specifi cations, permissions, audit hooks, tests, and Git checkpoints. You'll trace agent and model activity, diagnose failures from evidence, and turn operational fi xes into reusable tests. You'll also build a Backstage template and Argo CD ApplicationSet that provide a governed path from developer request to running agent. By the end of the book, you'll be able to build and validate an AI-native internal developer platform in phases, route agent and model activity through existing platform controls, and prepare the architecture for production use. 1 S T E D I T I O N Agentic DevOps with Claude Code www.packtpub.com Get a free PDF copy of this book packtpub.com/unlock/9781808344190
Page 2
Agentic DevOps with Claude Code Build governed AI platforms on Kubernetes with GitOps, observability, and self-service workflows Michael Rishi Forrester
Page 3
Agentic DevOps with Claude Code Copyright © 2026 Packt Publishing All rights reserved. No part of this book may be reproduced, stored in a retrieval system, or transmitted in any form or by any means, without the prior written permission of the publisher, except in the case of brief quotations embedded in critical articles or reviews. Every effort has been made in the preparation of this book to ensure the accuracy of the information presented. However, the information contained in this book is sold without warranty, either express or implied. Neither the author, nor Packt Publishing or its dealers and distributors, will be held liable for any damages caused or alleged to have been caused directly or indirectly by this book. Packt Publishing has endeavored to provide trademark information about all of the companies and products mentioned in this book by the appropriate use of capitals. However, Packt Publishing cannot guarantee the accuracy of this information. Portfolio Director: Kartikey Pandey Relationship Lead: Preet Ahuja Project Manager: Sonam Pandey Content Engineer: Apramit Bhattacharya Technical Editor: Simran Ali Indexer: Manju Arasan Production Designer: Shantanu Zagade Growth Lead: Vikramaditya Vishwanath First published: September 2026 Production reference: 1220926 Published by Packt Publishing Ltd. Grosvenor House 11 St Paul's Square Birmingham B3 1RB, UK. ISBN 978-1-80834-419-0 www.packtpub.com
Page 4
To everyone who has worked for me, everyone I have worked for, every customer, every student, and every conversation that went somewhere I did not expect. This book is made of those. And to my family, who make the difference: my sister Lori, my son Liam, and my nephews James and Revan. – Michael Rishi Forrester
Page 5
Contributors About the author Michael Rishi Forrester is AI Workforce Transformation Lead at Accenture LearnVantage and founder of The Performant Professionals. He has 25+ years in operations and DevOps, holds CKA, CKAD, and CKS certifications along with ~10 AWS certifications, and delivered the 27- component predecessor of this workshop at KCD Texas, May 15, 2026, where it was the third- most-requested submission of the conference. He spoke at KubeCon EU Amsterdam in March 2026 (two sessions, one co-delivered with Mumshad Mannambeth), devopsdays Atlanta, SREday Austin, LLMday Austin, and KCD Texas, and is a confirmed speaker at Agentic Engineering Days Zurich in November 2026. He is an independent practitioner with no Anthropic affiliation; the content is vendor-neutral by design. Whitney Lee is the reason this workshop exists in the shape it does. She was the inspiration for building it and a hand in crafting the whole thing, and the book carries her fingerprints throughout. Rich Morrow, Ryan Dymek and George Sawyer got me into training in the first place, which is the decision most of my working life now rests on. Chris Kelly taught me a great deal about managing and leading, including the parts you only learn by watching what not to do. Gordon O'Reilly showed me there are still genuinely good leaders out there, which is not a small thing to be reminded of. Louis Puster is an extraordinary business partner and, more than that, an extraordinary human engineer. Bobby Wilson is the solid rock in every conversation about keeping a business sane. Greg Wagner is an excellent salesman and strategist, and the coolest music aficionado I know when it comes to records. Chris McNabb has been an inspiration on DevOps and on staying sane while doing it. Jeremy Morgan is a great fellow trainer and a better sounding board. Ari Waller and Vincent Mayers are doing good work at Neo4j and across their industry, and have done more than anyone to teach me how community and collaboration are built. Justin Zakrzeski, along with everyone I have ever worked with or who has worked for me, has been an inspiration.
Page 6
At Packt, thank you to Rachita Shukla, who opened the door; Preet Ahuja, who shaped the book and carried it through; Kartikey Pandey, Portfolio Director; Sonam Pandey, Project Manager; Apramit Bhattacharya, the Content Engineer who built this manuscript from the workshop and took every correction on the chin; Simran Ali, Technical Editor; Manju Arasan, Indexer; Shantanu Zagade, Production Designer; and Vikramaditya Vishwanath, Growth Lead. And to the many people I will inevitably have left out. The omission is mine. It says nothing about what you were worth to the work.
Page 7
(This page has no text content)
Page 8
Table of Contents Preface xxv Free benefits with your book ................................................................................................. xxx Part 1: Building a Controlled Cloud-Native Foundation 1 Chapter 1: Designing an AI-Native Internal Developer Platform 3 Defining the AI-native platform outcome ................................................................................. 4 From components to capabilities • 4 A practical outcome statement • 5 Why this matters to operations • 6 Defining done before selecting tools • 7 Hands-on exercise: write the platform outcome contract • 7 Troubleshooting the outcome definition • 8 Mapping the foundation, AI, and self-service planes ................................................................ 9 The cloud-native foundation • 10 What happens when the foundation fails • 10 The AI plane • 10 What happens when the AI plane fails • 11 The self-service plane • 11 Platform-owned defaults versus developer choices • 11 Dependency mapping exercise • 12 Troubleshooting the plane model • 13 Separating the builder agent from the served model ................................................................ 13 The builder agent • 14 The runtime model • 14 The runtime agent • 14 Comparing the roles • 14 A request path example • 15 Hands-on exercise: create the identity matrix • 16 Troubleshooting identity design • 16
Page 9
Drawing trust boundaries and production boundaries ............................................................ 17 Reasoning-layer controls • 18 Infrastructure-layer controls • 18 Classifying the workshop decisions • 18 Threat-model exercise • 20 Troubleshooting the trust model • 21 Production readiness checklist • 21 Summary ................................................................................................................................ 22 Chapter 2: Governing Claude Code with Specifications, Tests, and Permissions 23 Establishing ground truth with a no-change preflight ............................................................ 24 What bare means • 24 The preflight evidence set • 24 Using an explicit kubeconfig • 25 Hands-on lab: build the Phase 0 gate • 26 Prerequisites • 26 Files used • 26 Step 1: Read tests/test_phase_0_preflight.py • 26 Step 2: Run the test • 27 Troubleshooting tips • 28 Writing phased build specifications and completion gates ..................................................... 29 Why one giant prompt fails • 29 Anatomy of a phase specification • 29 Writing a testable goal • 30 Fixing decisions before execution • 30 The stop condition • 31 Hands-on lab: write a phase contract • 31 Create spec/phases/example-inventory.md • 31 Run a specification review • 32 Troubleshooting specification problems • 32 Pinning versions and controlling source provenance .............................................................. 33 Why chart and application versions both matter • 33 Validating components.yaml • 33 Run the component validator • 34 Table of Contents viii
Page 10
Approved sources and vendoring • 34 Source provenance checklist • 35 When a newer version is worse • 35 Hands-on lab: make an unpinned component fail CI • 36 Troubleshooting tips • 37 Constraining Claude Code with permissions and audit hooks ................................................ 37 Allow, ask, and deny • 37 The defense layers • 38 Audit hooks • 39 Protecting the audit trail • 40 Hands-on lab: test the permission boundary • 40 1. Set the audit identity • 40 2. Run three requests • 41 Expected behavior • 41 3. Inspect the audit record • 41 What just happened • 41 Troubleshooting tips • 41 Running the test-build-commit-stop control loop .................................................................. 42 Observe • 43 Prove the expected failure • 43 Build through Git • 43 Validate live behavior • 43 Review the diff • 43 Commit and checkpoint • 43 Stop and approve • 43 Hands-on lab: execute a controlled dry run • 44 1. Run the failing test • 44 2. Build and validate • 44 What just happened • 44 Troubleshooting the control loop • 45 Production reality check • 45 Summary ................................................................................................................................ 46 Chapter 3: Bootstrapping the GitOps Foundation 47 Bootstrapping Argo CD and defining the GitOps exception ..................................................... 48 ix Table of Contents
Page 11
The bootstrap boundary • 48 Why server-side apply is required • 48 Argo CD 3.x considerations • 49 Using in-cluster Gitea • 50 Hands-on lab: bootstrap Argo CD • 50 Prerequisites • 50 Files used • 50 Render the pinned chart • 50 Apply the bootstrap • 51 Apply the root Application • 51 Expected output • 52 Troubleshooting tips • 52 Structuring the App-of-Apps and sync-wave order ................................................................. 52 Reading the root Application • 53 Why sync waves exist • 53 Annotating a child Application • 54 Health and synchronization are different • 55 App-of-Apps refresh behavior • 55 Hands-on exercise: inspect the dependency graph • 56 Expected output • 56 Troubleshooting tips • 56 Managing certificates, secrets, storage, and workload identity ............................................... 57 cert-manager as an early dependency • 57 The secret flow • 57 The workshop OpenBao configuration • 58 Production OpenBao design • 59 gp3 storage • 59 EKS Pod Identity • 60 Hands-on lab: prove secret materialization and storage readiness • 60 1. Run the Phase 1 tests • 60 2. Inspect without disclosing the secret • 61 Troubleshooting tips • 61 Introducing platform policy in audit mode ............................................................................. 62 Why audit first • 62 Baseline policies • 62 Table of Contents x
Page 12
Controller default drift • 63 Policy review workflow • 63 Production reality check • 64 Hands-on exercise: inspect audit results • 64 1. List policy reports • 64 2. Inspect one result • 64 Troubleshooting tips • 65 Repairing Git-sourced faults at the source .............................................................................. 65 The diagnostic ladder • 65 1. Inspect Argo CD first • 66 2. Fix one source field • 67 3. Refresh the cascade • 67 Expected output • 68 Recurring fault classes • 68 Troubleshooting tips • 68 Production reality check • 69 Summary ................................................................................................................................ 69 Chapter 4: Building the Platform Observability Plane 71 Designing a shared telemetry pipeline .................................................................................... 72 OpenTelemetry as the seam • 72 Signal contracts • 73 Collector roles • 74 Processing order • 74 Resource identity and cardinality budgets • 75 Why this matters operationally • 76 Hands-on exercise: define the telemetry contract • 76 Expected result • 77 Troubleshooting the design • 77 Collecting metrics, logs, and traces on EKS ............................................................................. 77 Metrics collection • 77 Log collection • 78 Trace collection • 79 Collector deployment details • 79 The OpenTelemetry Operator • 80 xi Table of Contents
Page 13
Hands-on lab: verify collectors and scrape targets • 81 Prerequisites • 81 Files used • 81 Check the daemonset • 81 Inspect Prometheus targets • 81 Troubleshooting tips • 82 Persisting and querying telemetry with the observability stack .............................................. 82 Prometheus and Alertmanager persistence • 82 Grafana data sources • 83 Querying by API • 83 Loki label discipline • 84 Tempo search and trace identity • 84 Hands-on lab: query Grafana data sources • 85 Retrieve the password safely • 85 Query through a port-forward • 85 Troubleshooting tips • 86 Validating the telemetry path and controlling drift ................................................................ 86 The probe trace • 86 1. Create a compact trace probe • 87 2. Run the phase test • 88 3. Diagnosing a missing trace • 89 Controller-default drift • 89 Production reality check • 90 Troubleshooting tips • 90 Summary ................................................................................................................................ 91 Chapter 5: Delivering the Developer Portal and Platform Extensions 93 Treating Backstage as the platform interface .......................................................................... 94 Interface versus mechanism • 94 What a developer should see • 95 What the platform team should own • 95 Why Backstage arrives last • 96 The pre-built image • 96 Backstage data in the corrected workshop build • 97 Hands-on exercise: review the portal contract • 97 Table of Contents xii
Page 14
Expected result • 98 Troubleshooting the interface boundary • 98 Connecting the catalog, TechDocs, and live Argo CD state ...................................................... 98 Registering a component • 99 TechDocs as docs-as-code • 99 The Argo CD integration • 100 Querying the catalog API • 101 Expected output • 101 Querying the Argo CD proxy • 102 Credential boundaries • 102 Catalog eventual consistency • 102 Hands-on lab: verify catalog and live delivery state • 103 Prerequisites • 103 Files used • 103 Run the Phase 3 integration tests • 103 Troubleshooting tips • 104 Adding workflows, events, progressive delivery, and autoscaling ......................................... 104 Argo Workflows • 105 Argo Events • 105 Argo Rollouts • 106 KEDA • 106 Hands-on lab: verify extension APIs • 107 1. Check the CRDs • 107 2. Run the full phase gate • 108 Troubleshooting tips • 108 Diagnosing portal image, configuration, and credential failures .......................................... 109 The diagnostic layers • 109 Failure 1: doubled registry path • 110 Failure 2: the file exists but the process cannot reach it • 111 Failure 3: production configuration never loaded • 111 Failure 4: secret and credential drift • 112 Failure 5: database secret and storage assumptions • 112 Failure 6: eventually consistent catalog state • 112 Hands-on lab: diagnose a failed Backstage deployment • 112 1. Inspect Argo CD and the rendered workload • 112 xiii Table of Contents
Page 15
2. Inspect events and logs • 113 Expected evidence • 113 3. Test the outcome after the repair • 113 Troubleshooting checklist • 114 Production reality check ........................................................................................................ 114 Summary ............................................................................................................................... 115 Part 2: Adding the AI Plane 117 Chapter 6: Creating a Governed Gateway for Agent Traffic 119 Using Gateway API as the platform traffic contract ............................................................... 120 The three-object contract • 120 Why status conditions matter • 121 Hands-on lab: establish the gateway contract • 122 1. Create gateway.yaml • 122 2. Create application.yaml • 123 3. Create test_gateway_contract.py • 124 4. Run the lab • 125 What just happened • 126 Troubleshooting tips • 126 Separating kgateway and agentgateway responsibilities ....................................................... 127 The practical division • 127 Install order and ownership • 128 Inspect the pinned Applications • 129 Expected output • 129 The OCI detail that bites • 130 The chart-values correction • 130 Expected result • 130 What this command is doing • 130 Troubleshooting controller overlap • 131 Production reality check • 131 Mediating LLM, MCP, and agent-to-agent traffic ................................................................... 131 Three traffic types, three risk shapes • 131 Lab: route the local vLLM service • 132 1. Create vllm-backend.yaml • 133 Table of Contents xiv
Page 16
2. Create vllm-route.yaml • 133 3. Verify route attachment • 134 The MCP path • 134 The A2A path • 135 Test the mediated path from the agent namespace • 135 Expected result • 136 What this command is doing • 136 Production reality check • 136 Landing AI policies in audit mode .......................................................................................... 136 Audit first does not mean audit forever • 136 Review the guardrail-reference policy • 137 Test an audited violation • 138 Apply and inspect the fixture • 139 Expected output • 139 What just happened • 139 The registry policy caveat • 139 The OTel annotation caveat • 140 The bypass-policy caveat • 140 Troubleshooting policy drift • 140 Verifying routes, mTLS, and audit logging ............................................................................ 140 Build an evidence chain • 140 Add a mutual-TLS listener • 141 Test both sides of mTLS • 143 Expected behavior • 143 What just happened • 143 Verify route parents programmatically • 144 Define the audit record • 144 1. Create tests/check_audit_shape.py • 145 2. Run the audit check • 146 Expected output • 146 What just happened • 146 Troubleshooting the complete path • 146 Production reality check • 147 Summary ............................................................................................................................... 147 xv Table of Contents
Page 17
Chapter 7: Running Agents as Kubernetes Resources 149 Declaring and reconciling agents with kagent CRDs ............................................................. 150 An Agent is a resource, not a personality • 150 CRDs before controllers • 151 Expected output • 151 What this command is doing • 151 Hands-on lab: declare the platform helper • 151 1. Create agent.yaml • 152 2. Inspect the resource before commit • 153 3. Create test_agent_ready.py • 153 4. Run the reconciliation test • 154 Troubleshooting tips • 154 Production reality check • 155 Routing a ModelConfig to an in-cluster endpoint .................................................................. 155 OpenAI-compatible describes a protocol • 155 Create the placeholder key • 156 1. Create modelconfig.yaml • 157 2. Check the endpoint from the agent namespace • 157 Use a deterministic inference probe • 158 1. Run the probe • 159 Troubleshooting the ModelConfig • 159 Connecting MCP tools through agentgateway ...................................................................... 160 MCP resources in the pinned stack • 160 Repository truth before extension • 161 Create remote-mcp.yaml • 162 Bind named tools to the Agent • 162 1. Inspect discovered tool status • 163 2. Create test_mcp_contract.py • 164 Test one tool call • 164 Expected output • 165 Troubleshooting MCP • 165 Production reality check • 165 Blocking prompt injection with LLM Guard .......................................................................... 166 LLM Guard in this platform • 166 Table of Contents xvi
Page 18
The runtime integration gap • 166 Create llm-guard-policy.yaml • 167 Create the deterministic fixture • 168 1. Run the paired test • 169 2. Inspect the audit decision • 170 Troubleshooting guardrails • 170 Production reality check • 171 Instrumenting agent calls with OpenLLMetry ....................................................................... 171 Trace the useful boundaries • 171 Create trace_probe.py • 172 Pin the dependencies • 173 Set the semantic convention • 174 Expected output • 174 What just happened • 174 Query Tempo for GenAI evidence • 174 Expected output • 175 What this command is doing • 175 Create a polling test • 175 What this code is doing • 176 Troubleshooting telemetry • 176 Production reality check • 176 Summary ............................................................................................................................... 177 Chapter 8: Serving and Observing Models on Kubernetes 179 Deploying KServe and an OpenAI-compatible vLLM endpoint ............................................. 180 The serving topology • 180 Install APIs before resources • 181 Expected output • 181 What this command is doing • 181 Hands-on lab: create the predictor • 181 1. Create base/inferenceservice.yaml • 182 2. Create overlays/cpu/resources.yaml • 183 3. Create overlays/cpu/runtime.yaml • 184 4. Create base/kustomization.yaml • 185 Create the Kustomization • 186 xvii Table of Contents
Page 19
Render before reconciliation • 186 Expected result • 187 What just happened • 187 Handle KServe's annotation normalization • 187 Production reality check • 188 Tuning CPU inference for Kubernetes constraints ................................................................. 188 The live failure map • 188 Measure instead of guessing • 190 Expected evidence • 191 What these commands are doing • 191 Watch model startup • 191 Expected milestones • 191 What this command is doing • 191 Inspect shared memory • 192 Expected output • 192 NUMA policy and security • 192 Tune context length deliberately • 192 Troubleshooting CPU performance • 192 Running and tracing an end-to-end inference ....................................................................... 193 Start with the cheapest request • 193 Expected output • 193 What this command is doing • 194 Create inference_probe.py • 194 Run the probe twice • 195 Expected output • 196 What just happened • 196 Trace model and token attributes • 196 Expected evidence • 197 What this command is doing • 197 Validate tool-calling support • 197 Expected behavior • 198 Troubleshooting the round trip • 198 Production reality check • 198 Positioning llm-d, GPUs, and managed model routes ........................................................... 198 The serving decision • 199 Table of Contents xviii
Page 20
Moving to GPU • 200 Positioning llm-d • 201 Managed model routes • 201 Route by policy, not model fashion • 202 Production reality check • 202 Testing readiness, performance, and regression-sensitive settings ....................................... 202 Test at four layers • 202 The readiness complication • 203 Create test_phase_6_serving.py • 204 Add volume and memory assertions • 205 Test the API model name • 205 Create benchmark_inference.py • 206 Expected output • 207 Set a generous regression gate • 207 Customize Argo CD health carefully • 207 Run the phase gate • 207 Expected result • 207 Troubleshooting regression tests • 208 Production reality check • 208 Capture a phase evidence record • 208 Expected result • 209 What just happened • 209 Summary .............................................................................................................................. 209 Part 3: Productizing and Operating the Platform 211 Chapter 9: Shipping a Governed Self-Service Golden Path 213 Defining the agent-service template contract ....................................................................... 214 The contract before the YAML • 214 Designing inputs that fail early • 216 Define the generated repository as an API • 217 Acceptance criteria for the contract • 218 Generating governed manifests from a Backstage form ........................................................ 219 Lab goal and prerequisites • 219 Verify the scaffolder action inventory • 220 xix Table of Contents
The above is a preview of the first 20 pages. Register to read the complete e-book.

Recommended for You

Loading recommended books...
Failed to load, please try again later

Tip the Site

Scan the WeChat Pay or Alipay code to tip. No login required.

WeChat Pay
Alipay
← Back to List