Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: it-ebooks

Rating No ratings yet

No description

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
【One-Line Pitch】 A practitioner's field report on how Alibaba rebuilt its Double 11 core systems into a fully cloud-native engine in 2020, useful for engineers and architects who want concrete patterns for running Kubernetes, Serverless, and middleware at extreme scale. 【Book Arc】 - **Opening (~0%–10%)**: Sets the strategic frame — why cloud-native is the shortest path to digital transformation, and how Alibaba moved from "Double 11 on the cloud" to a "cloud-native Double 11," with headline results across elasticity, compute, middleware, and Serverless. - **Early (~10%–35%)**: Drills into the technical foundation: Kubernetes/containers as the new cloud interface, the Sigma-to-ACK migration, Serverless landing challenges (cold start, dev workflow, middleware connectivity), RocketMQ's zero-fault messaging, and Sentinel Go for flow control and degradation. - **Middle (~35%–55%)**: Moves to operating large clusters — multi-cluster scale, change velocity, risk control, OpenKruise as the deployment substrate, and the ASI scheduler's mixed scheduling of online, batch, and best-effort workloads. - **Late (~55% onward)**: The excerpts do not cover the later chapters in detail; based on the table of contents and author list, the book continues into further cloud-native product and practice topics, but the provided material stops around the scheduling discussion. 【Key Takeaways】 - **Cloud-native is framed as an evolution, not a rewrite** (Opening): the book positions the journey as moving from Cloud Hosting to Cloud Native, with Alibaba's own path from 2008 distributed middleware through 2011 containerization to 2020 full core-system cloud-nativization. - **Kubernetes is the new cloud interface** (Early): containers standardize application delivery and decouple apps from infrastructure, while Kubernetes standardizes scheduling and hides underlying differences — this is the conceptual anchor for everything that follows. - **Scale demands specific engineering responses** (Middle): the book details concrete techniques for large clusters — using application names as namespaces, optimizing ETCD storage and read paths, and protecting the APIServer with inflight request limits. - **Serverless adoption hinges on solving cold start and workflow gaps** (Early): the book candidly lists three pain points — cold start stretching to seconds in core chains, disconnect from testing/gray-release/disaster-recovery workflows, and middleware connection overhead — then describes reserved-plus-on-demand modes and tooling as the response. - **Middleware is treated as a first-class cloud-native concern** (Early): RocketMQ's zero-fault Double 11 record and Sentinel's flow control, degradation, and adaptive protection illustrate how messaging and resilience are engineered for peak traffic. - **Deployment automation is the substrate for scale** (Middle): OpenKruise extends Kubernetes workloads to handle deployment, upgrade, scaling, QoS, health checks, and migration for millions of containers, and its "three-in-one" open-source/commercial/internal model is a recurring theme. - **Scheduling is a multi-layer problem, not just a central scheduler** (Middle): the book describes a "broad scheduler" combining central, single-machine, kernel, rescheduling, and orchestration layers, with differentiated SLO tiers (Product, Batch, Best Effort) to improve utilization. - **Stability is designed in, not bolted on** (Middle): default stability configurations, standardized probes, centralized self-healing, and resource isolation for sidecars are presented as practices that make stability "like water, electricity, and coal." 【Reading Tips】 - **Read the Opening for context, then jump to chapters matching your stack.** The strategic framing is useful once; the real value is in the Kubernetes, Serverless, middleware, and scheduling chapters. - **Treat the Serverless chapter as a decision aid.** The three pain points and the reserved/on-demand solution are the most transferable part for teams evaluating FaaS for latency-sensitive workloads. - **Skim the product-marketing passages.** Some sections read as Alibaba Cloud product promotion; extract the architectural patterns and skip the feature lists. - **Pay attention to the "three-in-one" theme.** The recurring idea of internal practice feeding open source and commercial products is a useful lens for evaluating any vendor's cloud-native claims. - **Use the scheduling chapter for utilization thinking.** The Product/Batch/Best Effort tiering is a concrete model for mixing workloads to raise cluster efficiency. 【Coverage Limits】 This guide is based on stratified excerpts covering roughly the first half of the book (through the scheduling discussion). Later chapters and specific implementation details beyond the provided material are not covered.
Excerpt 1
拥抱云原生架构、 用技术加速创新,将成为企业数字化转型升级成功的关键。 阿里云研究员、阿里云云原生应用平台负责人 丁宇 4985 亿交易额的背后,全面揭秘阿里巴巴双 11的云原生支撑力 < 12 当大促场景成为企业业务日常,这个问题的答案非常值得借鉴。 从“双 11 上云”到“云上双 11 云原生重构双 11“技...
View in text
Excerpt 2
的业务逻辑,还要依赖中间件、数据库、 存储等后端服务,这些服务的连接都要在实例启动的时候进行建连,这无形中加大了冷启动 的时间,进而把冷启动的时间加长到秒级别。对于核心在线业务场景来说,秒级别的冷启动 是不可接受的。 挑战二:与研发流程割裂 Serverless 主打的场景是像写业务函数一样去写业务代码,简单快速...
View in text
Excerpt 3
实现。Sentinel 通过淘 汰机制(如 LRU、LFU、ARC 策略等)来识别热点参数,通过令牌桶机制来控制每个热 点参数的访问量。目前的 Sentinel Go 版本采用 LRU 策略统计热点参数,社区也已有 贡献者提交了优化淘汰机制的 PR,在后续的版本中社区会引入更多的缓存淘汰机制来适 配不同的场景。...
View in text
Excerpt 4
主要包括斗鱼 TV、申通、有赞等,而开源社区中携程、Ly f t 等公司也都是 OpenKruise 的用户和贡献者。 OpenKruise 将基于阿里巴巴超大规模场景锤炼出的云原生应用负载能力开放出来, 不仅在云原生社区中补充了扩展应用负载的重要板块,还为云上客户提供了阿里巴巴多年应 用部署的管理经验和云原生化...
View in text
Excerpt 5
巴复杂任务资源混合调度技术面纱 科普:以一台 96 核(实际上我们说的都是 96 个逻辑核)的 X86 架构物理机或神 龙为例,它有 2 个 Socket,每个 Socket 有 48 个物理核,每个物理核下有 2 个逻辑 核。【当然,ARM 的架构又与X86 不同】。 由于 CPU 架构的 L1 L2 L3 C...
View in text
Excerpt 6
binlog 复制的方式,实现了数据库的在线迁移。 云原生趋势下的迁移与容灾思考 < 106 3)共享文件系统的容灾 图中采用了 Gluster 的文件系统,由于分布式系统的一致性通常由内部维护,单纯使 用块级别很难保证节点的一致性,所以这里面使用文件级别容灾更为精确。 4)数据库的容灾 单纯依靠存储层面是无法根...
View in text
Excerpt 7
、资源不足、OOM 等,比如 Pod 实例容器重启、 驱逐、健康检查失败、启动失败等。PaaS 平台上云应用实践过程。 用户为集群一键安装 NPD 组件,为集群和应用分别配置告警规则,设置关注的事件 类型和通知方式即可。 集群上所有事件自动采集到 SLS 日志服务,日志服务上的告警规则由我们根据事件 类型和用途自...
View in text
Excerpt 8
量,多个门店支付时好时 坏,短时间也无法维护,导致用户体验差,这让世纪联华的技术人决心改进这套使用了十多 年的老系统。 2014~2018 年:中央机房部署架构的演进 在 2014 年经历了 双 12 大促活动的问题后,联华技术人决心改进各项系统,于是将 交易系统和会员系统陆续迁移到自建的中央物理机房,商品系统也...
View in text
Tags
AI categories
Cloud NativeDevOpsKubernetes
Publisher: it-ebooks
Publish Year: 2021
Language: Chinese
File Format: PDF
File Size: 4.0 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…