Ian Goodfellow
Yoshua Bengio
Aaron Courville
,Contents
Website viii
Acknowledgments ix
Notation xiii
1 Introduction 1
1.1 Who Should Read This Book? . . . . . . . . . . . . . . . . . . . . 8
1.2 Historical Trends in Deep Learning . . . . . . . . . . . . . . . . . 12
I Applied Math and Machine Learning Basics 27
2 Linear Algebra 29
2.1 Scalars, Vectors, Matrices and Tensors . . . . . . . . . . . . . . . 29
2.2 Multiplying Matrices and Vectors . . . . . . . . . . . . . . . . . . 32
2.3 Identity and Inverse Matrices . . . . . . . . . . . . . . . . . . . . 34
2.4 Linear Dependence and Span . . . . . . . . . . . . . . . . . . . . 35
2.5 Norms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37
2.6 Special Kinds of Matrices and Vectors . . . . . . . . . . . . . . . 38
2.7 Eigendecomposition . . . . . . . . . . . . . . . . . . . . . . . . . . 40
2.8 Singular Value Decomposition . . . . . . . . . . . . . . . . . . . . 42
2.9 The Moore-Penrose Pseudoinverse . . . . . . . . . . . . . . . . . . 43
2.10 The Trace Operator . . . . . . . . . . . . . . . . . . . . . . . . . 44
2.11 The Determinant . . . . . . . . . . . . . . . . . . . . . . . . . . . 45
2.12 Example: Principal Components Analysis . . . . . . . . . . . . . 45
3 Probability and Information Theory 51
3.1 Why Probability? . . . . . . . . . . . . . . . . . . . . . . . . . . . 52
i
, CONTENTS
3.2 Random Variables . . . . . . . . . . . . . . . . . . . . . . . . . . 54
3.3 Probability Distributions . . . . . . . . . . . . . . . . . . . . . . . 54
3.4 Marginal Probability . . . . . . . . . . . . . . . . . . . . . . . . . 56
3.5 Conditional Probability . . . . . . . . . . . . . . . . . . . . . . . 57
3.6 The Chain Rule of Conditional Probabilities . . . . . . . . . . . . 57
3.7 Independence and Conditional Independence . . . . . . . . . . . . 58
3.8 Expectation, Variance and Covariance . . . . . . . . . . . . . . . 58
3.9 Common Probability Distributions . . . . . . . . . . . . . . . . . 60
3.10 Useful Properties of Common Functions . . . . . . . . . . . . . . 65
3.11 Bayes’ Rule . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 68
3.12 Technical Details of Continuous Variables . . . . . . . . . . . . . 69
3.13 Information Theory . . . . . . . . . . . . . . . . . . . . . . . . . . 71
3.14 Structured Probabilistic Models . . . . . . . . . . . . . . . . . . . 73
4 Numerical Computation 78
4.1 Overflow and Underflow . . . . . . . . . . . . . . . . . . . . . . . 78
4.2 Poor Conditioning . . . . . . . . . . . . . . . . . . . . . . . . . . 80
4.3 Gradient-Based Optimization . . . . . . . . . . . . . . . . . . . . 80
4.4 Constrained Optimization . . . . . . . . . . . . . . . . . . . . . . 91
4.5 Example: Linear Least Squares . . . . . . . . . . . . . . . . . . . 94
5 Machine Learning Basics 96
5.1 Learning Algorithms . . . . . . . . . . . . . . . . . . . . . . . . . 97
5.2 Capacity, Overfitting and Underfitting . . . . . . . . . . . . . . . 108
5.3 Hyperparameters and Validation Sets . . . . . . . . . . . . . . . . 118
5.4 Estimators, Bias and Variance . . . . . . . . . . . . . . . . . . . . 120
5.5 Maximum Likelihood Estimation . . . . . . . . . . . . . . . . . . 129
5.6 Bayesian Statistics . . . . . . . . . . . . . . . . . . . . . . . . . . 133
5.7 Supervised Learning Algorithms . . . . . . . . . . . . . . . . . . . 137
5.8 Unsupervised Learning Algorithms . . . . . . . . . . . . . . . . . 142
5.9 Stochastic Gradient Descent . . . . . . . . . . . . . . . . . . . . . 149
5.10 Building a Machine Learning Algorithm . . . . . . . . . . . . . . 151
5.11 Challenges Motivating Deep Learning . . . . . . . . . . . . . . . . 152
II Deep Networks: Modern Practices 162
6 Deep Feedforward Networks 164
6.1 Example: Learning XOR . . . . . . . . . . . . . . . . . . . . . . . 167
6.2 Gradient-Based Learning . . . . . . . . . . . . . . . . . . . . . . . 172
ii