Deep Reinforcement Learning in Action teaches you how to program AI agents that adapt and improve based on direct feedback from their environment. In this example-rich tutorial, you’ll master foundational and advanced DRL techniques by taking on interesting challenges like navigating a maze and playing video games. Along the way, you’ll work with core algorithms, including deep Q-networks and policy gradients, along with industry-standard tools like PyTorch and OpenAI Gym.
For readers with intermediate skills in Python and deep learning.
Alexander Zai is a machine learning engineer at Amazon AI.
Brandon Brown is a machine learning and data analysis blogger.
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
【One-Line Pitch】
A hands-on course that turns reinforcement learning theory into working PyTorch agents, walking you from bandits and Q-learning to policy gradients and multi-agent systems. Best for Python programmers with some deep learning background who learn by building.
【Book Arc】
- **Opening (~0%–10%)**: Frames what reinforcement learning is and why it differs from supervised learning, using a data-center cooling example to define objectives, environments, states, and the agent loop.
- **Early (~10%–30%)**: Builds the modeling vocabulary—Markov decision processes, rewards, transition probabilities—and starts with the multi-armed bandit, including a contextual bandit applied to ad placement.
- **Middle (~30%–50%)**: Moves into value-based methods: the Q-learning update rule, a from-scratch Gridworld game engine, and a neural network that learns to navigate it, with attention to PyTorch mechanics like detaching nodes.
- **Late (~50%–80%)**: Expands beyond basic DQN into policy gradient methods, actor-critic approaches, and alternative optimization such as evolutionary algorithms, plus distributional DQN.
- **Ending (~80%–100%)**: Samples frontier topics—curiosity-driven exploration, multi-agent reinforcement learning, and interpretable RL via attention and relational models—before a closing review and roadmap.
【Key Takeaways】
- **RL is framed as objective-driven control, not just game playing** (Opening): the data-center cooling example shows how to bundle costs and constraints into a single objective function, a reusable way to think about any control task.
- **The book teaches problem-first, then formalizes** (Early): terms and math are introduced only when a concrete problem demands them, so definitions stick better than in theory-first treatments.
- **Bandits are the gentlest entry point** (Early): the multi-armed bandit and its contextual variant (ad placement) demonstrate exploration, reward prediction, and softmax action selection before full RL enters.
- **Q-learning is presented as a general update rule, not a single algorithm** (Middle): the book connects the ad-placement network to the broader Q-learning family, then applies it to Gridworld with a small three-layer network.
- **Practical PyTorch discipline matters** (Middle): detaching target nodes and controlling what backpropagation touches are called out as common bug sources, not incidental details.
- **Policy gradients and actor-critic extend beyond value prediction** (Late): these methods shift from predicting values to directly learning policies, preparing readers for more complex decision problems.
- **The book deliberately samples, not exhausts** (Late): evolutionary algorithms, distributional DQN, curiosity-driven exploration, multi-agent RL, and interpretability are presented as exciting directions rather than complete treatments.
- **Code is the course** (Throughout): the authors recommend reading alongside the GitHub Jupyter Notebooks, which include fuller code and figure-generation scripts beyond the pared-down in-text listings.
【Reading Tips】
- **Read with the GitHub repository open**: the in-text code is minimized for space, and the notebooks contain the complete, maintained versions plus extra comments.
- **Deep-read the bandit and Gridworld chapters**: they establish the reward, state, and update-rule intuitions that every later algorithm reuses.
- **Skim the frontier chapters on a first pass**: curiosity, multi-agent, and interpretability are samples meant to orient you, not to be mastered immediately.
- **Watch the PyTorch graph mechanics closely**: the detach/no_grad discussion is a recurring source of training bugs and worth slowing down for.
- **Treat the closing roadmap as a reading list**: use it to decide which advanced area to pursue next rather than expecting full coverage inside the book.
【Coverage Limits】
The excerpts cover the book's structure, early foundations, bandit and Q-learning material, and chapter titles through the conclusion, but do not include detailed content from the later chapters on policy gradients, actor-critic, distributional DQN, curiosity, multi-agent, or interpretability. Specific algorithms, figures, and results from those chapters are therefore not summarized here.
Page 20
m- bered code listings that represented larger code blocks. At press time we are confident all the in-text code is working, but we cannot guar- antee that th...
nce a deep learning algorithm has a finite number of param- eters, we can use it to compress any possible state into something we can efficiently process, an...
om the resulting probability distribution over the actions. The chosen action will return a reward and updates the state of the environment. θ1 and θ2 repres...
the player is initialized at a random position on the board. Last, you can ini- tialize it so that all the objects are placed randomly (which is harder for t...
include it as a subscript rather than as an explicit input like π(x,θ) where x is some input data (i.e., the state of the game). Notations like π(x,θ) sugges...
the actions are generated, and “critic” refers to the value function, because that’s what (in part) tells the actor how good its actions are. Since we’re usi...
utionary strategies can scale better than other algorithms Neural networks were loosely inspired by real biological brains, and convolutional neural networks...
done: action by sampling from a s not lost params = unpack_params(agent['params']) categorical distribution probs = model(state,params) action = torch.distri...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
Deep Reinforcement Learning in Action (Alexander Zai, Brandon Brown)(Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
Deep Reinforcement Learning in Action (Alexander Zai, Brandon Brown)(Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment