Every day, companies struggle to scale critical applications. As traffic volume and data demands increase, these applications become more complicated and brittle, exposing risks and compromising availability. With the popularity of software as a service, scaling has never been more important.
Updated with an expanded focus on modern architecture paradigms such as microservices and cloud computing, this practical guide provides techniques for building systems that can handle huge quantities of traffic, data, and demand—without affecting the quality your customers expect. Architects, managers, and directors in engineering and operations organizations will learn how to build applications at scale that run more smoothly and reliably to meet the needs of customers.
• Learn how scaling affects the availability of your services, why that matters, and how to improve it
• Dive into a modern service-based application architecture that ensures high availability and reduces the effects of service failures
• Explore the Single Team Owned Service Architecture paradigm (STOSA)—a model for scaling your development organization in tandem with your application
• Understand, measure, and mitigate risk in your systems
• Use the cloud to build highly scalable applications
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
【One-Line Pitch】
A practical field guide for architects, engineering managers, and operations leaders who need to keep large cloud applications available and reliable as traffic, data, and teams grow. It reframes scaling as a risk-management and organizational problem, not just a technical one.
【Book Arc】
- **Opening (~0%–10%)**: Frames why scaling breaks systems and introduces the book's five tenets—availability, modern service-based architecture, organization, risk, and cloud—as the backbone for everything that follows.
- **Early (~10%–30%)**: Defines reliability versus availability, surveys what causes poor availability (resource exhaustion, load-based changes, more moving parts), and introduces tools like service tiers and risk matrices for tracking and improving it.
- **Middle (~30%–55%)**: Moves into architecture: monolith versus service-based design, guidelines for splitting into services, stateless versus stateful services, data partitioning, and how to respond predictably to cascading service failures.
- **Late (~55%–80%)**: Covers organizational scaling through the Single Team Owned Service Architecture (STOSA) paradigm, plus risk measurement, mitigation planning, and service-level agreement management.
- **Ending (~80%–100%)**: Ties the tenets together around cloud-based scaling, revisiting availability, architecture, organization, risk, and cloud as a unified practice for architecting for scale.
【Key Takeaways】
- **Availability is the reason to scale at all** (Early): the book separates reliability (doing operations correctly) from availability (being operational when needed), arguing availability is harder to fix and deserves architectural focus.
- **Growth itself causes outages** (Early): resource exhaustion, rushed load-driven changes, and multiplying moving parts are named as recurring causes of declining availability—useful as a diagnostic checklist.
- **"Two mistakes high" is a design discipline** (Middle): systems should retain enough redundancy to survive a failure *while already recovering from another*, since correlated failures can erase apparent redundancy.
- **Redundancy math must account for maintenance and correlation** (Middle): worked examples show that capacity headroom for node failures, upgrades, and data center outages requires more nodes than raw traffic calculations suggest.
- **Services should be split along business, ownership, and data lines** (Middle): the book offers concrete guidelines—specific business requirements, distinct team ownership, naturally separable data—while warning against going too far.
- **Risk mitigation means planning the degraded experience** (Middle): the "No-Search Web Store" example shows how a failed dependency can be handled with a fallback page and incentive rather than an error, turning a failure into a tolerable customer experience.
- **Architecture and organization scale together** (Late): STOSA ties service ownership to team ownership, so the development organization grows in tandem with the application rather than lagging behind it.
- **Risk must be measured and reviewed continuously** (Late): risk matrices, service tiers, and SLAs are presented as living tools to be revisited as the system and its risks change.
【Reading Tips】
- Read the availability and risk chapters (Early) closely—they define vocabulary used throughout; skim the table-of-contents-heavy opening chunks.
- Treat the "two mistakes high" and data center resiliency scenarios as worked problems; redo the node math yourself to internalize the reasoning.
- If you lead teams, prioritize the STOSA and organization material (Late); if you build systems, prioritize the service-splitting and failure-handling chapters (Middle).
- Keep the risk matrix and service-tier concepts as actionable artifacts—draft one for your own system as you read.
- The book is explicitly written to age well; focus on the durable principles rather than any specific cloud vendor details.
【Coverage Limits】
This guide is synthesized from stratified excerpts and the book's front matter and table of contents; specific chapter contents, examples, and later chapters are only partially represented, so some detail is inferred from headings rather than full text.
Page 6
. For more information, contact our corporate/institutional sales department: 800-998-9938 or corporate@oreilly.com. Acquisitions Editor: Kathleen Carr Index...
lutions look simplistic and minimalistic. Our industry will demand more and more complex systems and architectures to handle the scale of tomorrow. Naturally...
nt to review your risk management plans on a regular basis. Additionally, you should create and implement mitigation plans to reduce your appli‐ cation risks...
handle 300 req/sec each, we will need at least: Four nodes Which can handle the traffic but will not handle a node failure. Five nodes Which handle a single...
m to manage all of them. Separate team for security reasons Sometimes you want to restrict the number and scope of individuals who have access to the code an...
garbage or incomprehensible results. Summary | 69 CHAPTER 6 Service Ownership—STOSA In Chapter 3, we discussed what a service was and how it could be utilize...
rcentage is above the SLA (we say you are meeting your SLA). However, one time in late summer it dropped below your 80% SLA for a short period of time (we sa...
udes things like System, Owner, Date Identified, and Status. Make sure to assign a risk ID to each item (a sim‐ ple numbering from 1...n is reasonable). Are...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
Architecting for Scale How to Maintain High Availability and Manage Risk in the Cloud (Atchison, Lee)(Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
Architecting for Scale How to Maintain High Availability and Manage Risk in the Cloud (Atchison, Lee)(Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment