Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: Stefan Jansen

Rating No ratings yet

No description

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# Machine Learning for Trading: A Disciplined Workflow from Research to Live Execution ## 【One-Line Pitch】 A comprehensive, practice-oriented guide for quantitative traders and data scientists who want to build robust ML-driven trading systems—emphasizing disciplined research workflows, data integrity, and adaptive strategies that survive real-world market regime changes. ## 【Book Arc】 - **Opening (~0%–10%)**: Establishes the core thesis that process discipline matters more than model sophistication, using 2020–2025 market shocks (pandemic liquidity crisis, meme-stock sentiment waves, inflation surge, equity concentration) to demonstrate why adaptive workflows are essential. - **Early (~10%–23%)**: Builds the data foundation—covering the financial data universe (FX conventions, fixed-income quirks, price aggregation rules), storage and querying strategies (Parquet + Polars/DuckDB as defaults), and the critical distinction between single-venue and consolidated market data. - **Early (~23%–32%)**: Dives into market microstructure—order book mechanics, top-of-book imbalance versus order flow imbalance, and rigorous data quality protocols including event sequencing, timestamp semantics, and book reconstruction invariants. - **Middle (~39%–48%)**: Explores advanced data topics—fundamental and alternative data with release-timestamp discipline, synthetic financial data generation with fidelity/utility/privacy trade-offs, and diffusion models for regime-specific scenario generation. - **Middle (~48%–end of sample)**: Introduces the strategy research framework, beginning with price-based strategies (trend-following, momentum, mean reversion) and their specific risk profiles, setting up the case studies that follow. ## 【Key Takeaways】 - **Process discipline beats model sophistication** (Opening): Durable trading performance depends more on a disciplined, adaptive workflow than on choosing sophisticated models—markets change, assumptions fail, and tooling amplifies both good and bad practices. - **Market regimes are real and recurring** (Early): The 2020–2025 shocks (liquidity, sentiment, macro, crowding) are not special cases but recurring patterns; workflows must include explicit checks for instability and decay rather than assuming stable conditions. - **Price data is a convention, not a fact** (Early): In FX and fixed income, "the price" depends on venue choice, aggregation rules, and close conventions—documenting these choices in dataset metadata is essential for reproducible research. - **Storage and query defaults matter** (Early): Start with partitioned Parquet plus Polars or DuckDB for research velocity; add server databases only when concurrency and governance requirements justify operational overhead. - **Microstructure quality is an accounting problem** (Early): Treat order book reconstruction as a stateful system with explicit invariants—no negative depth, best bid ≤ best ask, and order lifecycle consistency—rather than relying on statistical filters alone. - **Synthetic data validation is three-dimensional** (Middle): Evaluate generated financial data on fidelity (marginal, dependence, time-series dynamics), utility (downstream task performance), and privacy—these objectives trade off against each other. - **Release timestamps are first-class data** (Middle): For scheduled economic releases (EIA, USDA), the publication timestamp must be part of the dataset; avoid workflows that treat pre-release expectations as published facts. - **Regime-conditioned generation enables stress testing** (Middle): Diffusion models with classifier guidance can generate regime-specific scenarios (low/high volatility), but require careful tuning to avoid exaggerating rare regimes. ## 【Reading Tips】 - **Skim the opening chapters** (~0–10%) if you're already convinced about workflow discipline; the market shock examples are illustrative but the core argument is stated early. - **Deep-read the microstructure and data quality sections** (~23–32%): These contain the most actionable, hard-won engineering details—event sequencing, timestamp semantics, and book reconstruction invariants will save you from subtle bugs. - **Pay special attention to the storage/query decision matrix** (Early, ~23%): The table comparing Parquet+Polars, DuckDB, PostgreSQL, and specialized databases is a practical reference you'll return to. - **The synthetic data chapter** (Middle, ~39–48%) is valuable if you work with generative models; the fidelity/utility/privacy framework is transferable beyond trading. - **Treat the strategy research framework** (Middle, ~48%) as a preview of the case studies—note the emphasis on turnover and execution realism for short-cadence strategies. ## 【Coverage Limits】 This guide covers the first half of the book (through the strategy research framework introduction). The nine case studies, agentic frameworks, and knowledge graphs mentioned in the table of contents are not covered in the sampled excerpts. ##
Page 15
............................................................................................ 191 Fundamentals and characteristics • 191 Time, calendar, and e...
View in text
Excerpt 2
elation of 0.710, indicating that the dendrogram preserves the pairwise distance structure reasonably well—Ward, GMM, and K-Means give comparable regime stru...
View in text
Excerpt 3
double-check the two primary observables—trades and quotes: • Validate basic fields (positive price and size, valid symbols, expected tick size where availab...
View in text
Excerpt 4
a improves or preserves performance on the downstream task. Fidelity is usually necessary but not sufficient: synthetic data can match broad distributional p...
View in text
Excerpt 5
even though the original labels may overlap substantially. Weighting retains all data; subsampling is simpler but discards information. Most setups combine a...
View in text
Excerpt 6
builds from day -3 to day +10 without a pre-event reversal. This pattern supports the hypothesis that the breakout captures continuation rather than a mechan...
View in text
Excerpt 7
. Representation quality, however, is only half the problem. A FinBERT em- bedding of an earnings call is not yet a tradable signal. It becomes one only afte...
View in text
Excerpt 8
xpected prediction when only features in 𝑆𝑆 are known. Shapley (1953) proved that this is the unique allocation satisfying four axioms: • Efficiency: attri...
View in text
Tags
AI categories
algorithmDataBackend
ISBN: 1803246979
Publisher: Packt Publishing
Publish Year: 2027
Language: English
Pages: 865
File Format: PDF
File Size: 7.9 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…