CS 559 Quiz 5: Neural Networks exam
with correct answers in option and
rationale graded A+
Section 1: Basics of Neural Networks (10 Questions)
1. What is the primary function of a neuron in a neural network?
A) To store data
B) To perform a weighted sum of inputs and apply an activation function
C) To optimize the learning rate
D) To initialize weights randomly
Correct Answer: B
Rationale:
A neuron computes a weighted sum of its inputs and applies an activation function (e.g., ReLU,
sigmoid) to introduce non-linearity. Other options are incorrect because:
A) Neurons process data, not store it.
C) Learning rate is a hyperparameter, not a neuron’s function.
D) Weight initialization is done before training, not by a neuron.
2. Which activation function is most prone to the vanishing gradient problem?
A) ReLU
B) Sigmoid
C) Tanh
D) Leaky ReLU
Correct Answer: B
Rationale:
The sigmoid function squashes inputs to (0,1), causing gradients to become extremely small during
backpropagation, leading to vanishing gradients. ReLU avoids this by having a constant gradient for
positive inputs, while Leaky ReLU and Tanh mitigate it to some extent.
,3. What is the purpose of backpropagation in neural networks?
A) To initialize weights
B) To compute gradients of the loss function w.r.t. weights
C) To select the best activation function
D) To increase the learning rate
Correct Answer: B
Rationale:
Backpropagation uses the chain rule to compute gradients of the loss function with respect to each
weight, enabling gradient descent optimization. Other options are incorrect because:
A) Weight initialization is done before training.
C) Activation functions are chosen based on the problem, not backpropagation.
D) Learning rate is a hyperparameter, not adjusted by backpropagation.
4. Which of the following is NOT a common loss function for classification?
A) Mean Squared Error (MSE)
B) Cross-Entropy Loss
C) Hinge Loss
D) Binary Cross-Entropy
Correct Answer: A
Rationale:
MSE is typically used for regression, not classification. The others are standard for classification:
B) Cross-Entropy: Multi-class classification.
C) Hinge Loss: Used in SVMs and some neural networks.
D) Binary Cross-Entropy: Binary classification.
5. What is the role of the bias term in a neuron?
A) To shift the activation function left or right
B) To increase the learning rate
,C) To prevent overfitting
D) To normalize inputs
Correct Answer: A
Rationale:
The bias term allows the activation function to shift left or right, enabling the model to fit data
better. It does not affect learning rate (B), overfitting (C), or input normalization (D).
6. Which optimization algorithm adapts the learning rate for each parameter individually?
A) Stochastic Gradient Descent (SGD)
B) Adam
C) Momentum
D) Batch Gradient Descent
Correct Answer: B
Rationale:
Adam (Adaptive Moment Estimation) adjusts learning rates per parameter using moving averages of
gradients. SGD (A) and Batch Gradient Descent (D) use a fixed learning rate, while Momentum (C)
accelerates SGD but does not adapt learning rates per parameter.
7. What is the main advantage of using ReLU over sigmoid?
A) ReLU is bounded between 0 and 1
B) ReLU avoids the vanishing gradient problem for positive inputs
C) ReLU is differentiable everywhere
D) ReLU is used only in output layers
Correct Answer: B
Rationale:
ReLU’s gradient is 1 for positive inputs, preventing vanishing gradients (unlike sigmoid). Other
options are incorrect:
A) ReLU is unbounded (output ≥ 0).
, C) ReLU is not differentiable at 0.
D) ReLU is used in hidden layers, not just output layers.
8. What is the purpose of dropout in neural networks?
A) To increase the learning rate
B) To prevent overfitting by randomly deactivating neurons
C) To initialize weights
D) To reduce the number of layers
Correct Answer: B
Rationale:
Dropout randomly deactivates neurons during training, forcing the network to learn robust features
and reducing overfitting. It does not affect learning rate (A), weight initialization (C), or network
depth (D).
9. Which of the following is a disadvantage of using a very deep neural network?
A) Increased computational cost
B) Higher risk of underfitting
C) Reduced feature extraction capability
D) Faster convergence
Correct Answer: A
Rationale:
Deep networks require more computations and are prone to vanishing/exploding gradients, not
underfitting (B). They improve feature extraction (C) and may slow convergence (D).
10. What is the output of a neuron with weights [0.5, -0.3], inputs [2, 1], bias 0.1, and ReLU
activation?
A) 0.8
B) 0.0
C) 1.0
D) 0.5
with correct answers in option and
rationale graded A+
Section 1: Basics of Neural Networks (10 Questions)
1. What is the primary function of a neuron in a neural network?
A) To store data
B) To perform a weighted sum of inputs and apply an activation function
C) To optimize the learning rate
D) To initialize weights randomly
Correct Answer: B
Rationale:
A neuron computes a weighted sum of its inputs and applies an activation function (e.g., ReLU,
sigmoid) to introduce non-linearity. Other options are incorrect because:
A) Neurons process data, not store it.
C) Learning rate is a hyperparameter, not a neuron’s function.
D) Weight initialization is done before training, not by a neuron.
2. Which activation function is most prone to the vanishing gradient problem?
A) ReLU
B) Sigmoid
C) Tanh
D) Leaky ReLU
Correct Answer: B
Rationale:
The sigmoid function squashes inputs to (0,1), causing gradients to become extremely small during
backpropagation, leading to vanishing gradients. ReLU avoids this by having a constant gradient for
positive inputs, while Leaky ReLU and Tanh mitigate it to some extent.
,3. What is the purpose of backpropagation in neural networks?
A) To initialize weights
B) To compute gradients of the loss function w.r.t. weights
C) To select the best activation function
D) To increase the learning rate
Correct Answer: B
Rationale:
Backpropagation uses the chain rule to compute gradients of the loss function with respect to each
weight, enabling gradient descent optimization. Other options are incorrect because:
A) Weight initialization is done before training.
C) Activation functions are chosen based on the problem, not backpropagation.
D) Learning rate is a hyperparameter, not adjusted by backpropagation.
4. Which of the following is NOT a common loss function for classification?
A) Mean Squared Error (MSE)
B) Cross-Entropy Loss
C) Hinge Loss
D) Binary Cross-Entropy
Correct Answer: A
Rationale:
MSE is typically used for regression, not classification. The others are standard for classification:
B) Cross-Entropy: Multi-class classification.
C) Hinge Loss: Used in SVMs and some neural networks.
D) Binary Cross-Entropy: Binary classification.
5. What is the role of the bias term in a neuron?
A) To shift the activation function left or right
B) To increase the learning rate
,C) To prevent overfitting
D) To normalize inputs
Correct Answer: A
Rationale:
The bias term allows the activation function to shift left or right, enabling the model to fit data
better. It does not affect learning rate (B), overfitting (C), or input normalization (D).
6. Which optimization algorithm adapts the learning rate for each parameter individually?
A) Stochastic Gradient Descent (SGD)
B) Adam
C) Momentum
D) Batch Gradient Descent
Correct Answer: B
Rationale:
Adam (Adaptive Moment Estimation) adjusts learning rates per parameter using moving averages of
gradients. SGD (A) and Batch Gradient Descent (D) use a fixed learning rate, while Momentum (C)
accelerates SGD but does not adapt learning rates per parameter.
7. What is the main advantage of using ReLU over sigmoid?
A) ReLU is bounded between 0 and 1
B) ReLU avoids the vanishing gradient problem for positive inputs
C) ReLU is differentiable everywhere
D) ReLU is used only in output layers
Correct Answer: B
Rationale:
ReLU’s gradient is 1 for positive inputs, preventing vanishing gradients (unlike sigmoid). Other
options are incorrect:
A) ReLU is unbounded (output ≥ 0).
, C) ReLU is not differentiable at 0.
D) ReLU is used in hidden layers, not just output layers.
8. What is the purpose of dropout in neural networks?
A) To increase the learning rate
B) To prevent overfitting by randomly deactivating neurons
C) To initialize weights
D) To reduce the number of layers
Correct Answer: B
Rationale:
Dropout randomly deactivates neurons during training, forcing the network to learn robust features
and reducing overfitting. It does not affect learning rate (A), weight initialization (C), or network
depth (D).
9. Which of the following is a disadvantage of using a very deep neural network?
A) Increased computational cost
B) Higher risk of underfitting
C) Reduced feature extraction capability
D) Faster convergence
Correct Answer: A
Rationale:
Deep networks require more computations and are prone to vanishing/exploding gradients, not
underfitting (B). They improve feature extraction (C) and may slow convergence (D).
10. What is the output of a neuron with weights [0.5, -0.3], inputs [2, 1], bias 0.1, and ReLU
activation?
A) 0.8
B) 0.0
C) 1.0
D) 0.5