CS 7643 Quiz 2 | Actual Questions and Answers
Latest Updated 2025/2026 (Graded A+) Georgia
Institute of Technology
1. Which of the following are common issues while optimizing
the weights of a deep neural network? (Select all that apply)
A. Existence of local minima
B. Ill-conditioned loss surface
C. Noisy gradient estimates
D. Saddle points
Correct Answer: B, C, D
2. Which of the following is the advantage of Leaky ReLU
compared to ReLU?
A. There’s no saturation on the positive end
B. Its output is always positive
C. It is cheap to compute
D. There’s no “dead” neuron when computing gradients
Correct Answer: D
3. The vanishing gradient problem occurs primarily with:
, A. ReLU activations
B. Tanh and Sigmoid activations
C. Linear transformations
D. Max pooling layers
Correct Answer: B
4. Which of the following best describes batch normalization?
A. Adds noise to gradients during backpropagation
B. Normalizes inputs within a mini-batch to stabilize training
C. Increases the learning rate dynamically
D. Removes neurons to prevent overfitting
Correct Answer: B
5. Why does stochastic gradient descent (SGD) often converge
better than full batch gradient descent?
A. It guarantees global minima
B. It uses second-order derivatives
C. The noise helps escape saddle points and local minima
D. It requires fewer epochs
Correct Answer: C
6. What is the role of momentum in gradient descent?
Latest Updated 2025/2026 (Graded A+) Georgia
Institute of Technology
1. Which of the following are common issues while optimizing
the weights of a deep neural network? (Select all that apply)
A. Existence of local minima
B. Ill-conditioned loss surface
C. Noisy gradient estimates
D. Saddle points
Correct Answer: B, C, D
2. Which of the following is the advantage of Leaky ReLU
compared to ReLU?
A. There’s no saturation on the positive end
B. Its output is always positive
C. It is cheap to compute
D. There’s no “dead” neuron when computing gradients
Correct Answer: D
3. The vanishing gradient problem occurs primarily with:
, A. ReLU activations
B. Tanh and Sigmoid activations
C. Linear transformations
D. Max pooling layers
Correct Answer: B
4. Which of the following best describes batch normalization?
A. Adds noise to gradients during backpropagation
B. Normalizes inputs within a mini-batch to stabilize training
C. Increases the learning rate dynamically
D. Removes neurons to prevent overfitting
Correct Answer: B
5. Why does stochastic gradient descent (SGD) often converge
better than full batch gradient descent?
A. It guarantees global minima
B. It uses second-order derivatives
C. The noise helps escape saddle points and local minima
D. It requires fewer epochs
Correct Answer: C
6. What is the role of momentum in gradient descent?