本书聚焦学术前沿,围绕人工智能的两大核心要素,即数据和模型,对人工智能领域安全问题以及相关攻防算法展开系统全面、详细深入的介绍。
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
AI guide
【One-Line Pitch】
A systematic, research-frontier survey of AI security organized around the two things every model depends on—data and models—covering the attack and defense algorithms that define the field. Best for graduate students, researchers, and engineers who want a structured map of adversarial ML rather than a hands-on coding tutorial.
【Book Arc】
- **Opening (~0%–10%)**: Frames the whole field—a short history of AI (Dartmouth, Deep Blue, ImageNet-era deep learning) and a three-way taxonomy of AI safety: endogenous, derivative, and AI-assisted security. Then lays the ML groundwork (loss functions, ERM/SRM, overfitting, SGD, reinforcement-learning basics) that later attack chapters assume.
- **Early (~10%–30%)**: Moves into data-side threats. Data poisoning is formalized as a bi-level optimization problem; privacy attacks (membership inference, attribute inference) are tied to the train/test generalization gap; data stealing is split into black-box and white-box settings, with large language models and diffusion models as headline cases of unintended memorization.
- **Early–Middle (~30%–45%)**: Deepfakes and forgery. Face swapping (SimSwap, FSGAN), Transformer-based editing (TransEditor), and video generation are described, followed by detection methods (color-space statistics, frequency analysis) and the emerging problem of detecting AI-generated text (GPTZero, watermarking).
- **Middle (~45%–55%)**: Adversarial examples. The classic progression from L-BFGS to FGSM, BIM, and PGD, then query-based and transfer attacks, ending with physical-world attacks (RP2) that survive printing, distance, and camera angle.
- **Late (~55%–100%)**: Defense and privacy-preserving training. Differential privacy (DP-SGD, gradient clipping, noise addition, moment accounting), federated learning variants and frameworks, secure boosted trees (XGBoost under local-data constraints), and defenses against poisoning, backdoors, and model stealing. *Excerpts do not cover the exact chapter boundaries in this final stretch.*
【Key Takeaways】
- **AI safety splits into three layers** (Opening): endogenous (data/model flaws), derivative (misuse like deepfakes), and AI-assisted security. This taxonomy is the book's organizing spine—keep it in mind as a filing system for every later technique.
- **Poisoning is best understood as bi-level optimization** (Early): the attacker's outer loop picks poisoned data, the inner loop retrains the victim. Framing it this way makes nearly all poisoning variants special cases, and explains why gradient-based poisoning struggles with deep nets (vanishing/exploding gradients, memory cost).
- **Privacy leakage is a generalization-gap problem** (Early): membership inference works because models behave differently on train vs. test data. Metrics like prediction loss, confidence, and entropy are practical attack signals—and overfitting is a direct privacy risk.
- **Generative models memorize** (Early): LLMs and diffusion models can regurgitate training data, including private text and near-identical images. This is not a bug in one model but a structural property of large-scale training.
- **Adversarial attacks form a complexity ladder** (Middle): FGSM (one-step, cheap, weak) → BIM (iterative) → PGD (strongest white-box baseline) → query and transfer attacks (black-box). Transferability is what makes black-box attacks practical.
- **Physical-world attacks are real** (Middle): RP2 shows perturbations can be printed and remain effective across distance, angle, and lighting—so adversarial robustness is not just a digital concern.
- **Defense is layered, not singular** (Late): data augmentation and filtering for poisoning, norm thresholding in federated settings, DP-SGD with clipping and moment accounting for privacy. No single defense covers all threat models.
- **Detection of AI-generated content is an open problem** (Middle): watermarking and perplexity-based detectors exist, but false positives (e.g., the US Constitution flagged as AI-written) show reliability is unresolved.
【Reading Tips】
- **Deep-read Chapters 1–2** even if you know ML—the loss-function and optimization framing is reused constantly in attack derivations.
- **Skim the deepfake generation chapter** if your interest is defense; the detection half is more actionable and less model-specific.
- **Treat the attack chapters as a ladder**: read FGSM→BIM→PGD in sequence; skipping ahead makes the black-box methods feel unmotivated.
- **Watch the math notation**: the book uses a symbol table and reuses θ, f, L across chapters—keep it handy.
- **For practitioners**, jump to the defense and federated-learning sections first, then return to attacks to understand what you're defending against.
【Coverage Limits】
This guide is based on stratified excerpts covering roughly the first half of the book plus scattered late material; the exact structure and content of the final defense chapters are only partially represented.
Excerpt 1
上“学习”来逐步逼近。 在实际机器学习求解过程中,人们通常使用经验风险 (empirical risk)来近似期望损失。给定训练数据集D[式(2.3)], 经验损失可定义如下: 此处近似的合理性与前文提到的独立同分布采样假设紧密相关。实际 优化通常遵循经验风险最小化(empirical risk minimiza...
View in text
Excerpt 2
露其“拼图”行为,阻止其通过多次联合查询来窃取数据或隐私信 息。篡改和伪造数据检测则是假设修改数据与天然数据具有不同分 布,并通过寻找这种不同来检测修改或伪造的样本。测试阶段的对抗 和后门样本检测主要是通过这两种异常样本对模型产生的异常激活或 输出分布来检测。而对模型窃取来说,攻击者往往需要对目标模型发 出大量的...
View in text
Excerpt 3
4.17 TransEditor 及所采用的StyleGAN2 的架构图 YU S,TACK J,MO S,et al. Generating videos with dynamics- aware implicit generative adversarial networks[EB/OL]. 2022. ht...
View in text
Excerpt 4
然是不可信的。 解决这些问题需要长期的实践探索,而“什么样的方式才是人类与AI和谐 共处的正确方式”是值得思考的问题。 GOODFELLOW I J,SHLENS J,SZEGEDY C. Explaining and harnessing adversarial examples[C]//Internation...
View in text
Excerpt 5
容易因对抗扰动而发生预测错误,所 以有用但不鲁棒。 基于上述定义,Ilyas等人基于分类任务进行实验,以“对抗训练得到 的模型倾向于使用鲁棒特征”为前提,通过解耦鲁棒特征和非鲁棒特征的 方式来证明真实数据集中广泛存在有用不鲁棒特征,是导致对抗样本出现 的原因。具体来说,先通过对抗训练获得一个(在一定程度上)鲁棒的...
View in text
Excerpt 6
练 ALAYRAC J B,UESATO J,HUANG P S,et al. Are labels required for improving adversarial robustness?[C]//Advances in Neural Information Processing Systems,2019....
View in text
Excerpt 7
e-trained NLP foundation models[EB/OL]. 2021. https://arxiv.org/abs/2110.02467. 上述传统形式的触发模式虽然能够取得较强的攻击性能,但是往往容 易被相关防御方法检测或者移除。此外,当原始训练文本规模较大时,可 能会导致上述攻击难以收敛。...
View in text
Excerpt 8
inst DNN model stealing attacks[C]//IEEE European Symposium on Security and Privacy,2019. 一般来说,窃取模型需要对目标模型发起大量访问,并且窃取查询样 本应该与正常查询样本具有不同的分布。基于此假设,Juuti等人 在 20...
View in text
Tags
AI categories
Artificial IntelligenceCybersecurityData
Text Preview (First 20 pages)
Registered users can read the full content for free
Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.
Generating text preview…
Loading comments...
Reply to Comment
Edit Comment