Actual Questions and Answers | 2026 Updates | 100% Correct.
SECTION 1: NEURAL NETWORK FUNDAMENTALS
Question 1
What is the primary difference between parametric and non-parametric
models?
A) Parametric models have a fixed number of parameters; non-
parametric models grow with data
B) Non-parametric models have a fixed number of parameters;
parametric models grow with data
C) Parametric models cannot be trained; non-parametric models can
D) Both have identical parameter counts
Correct Answer: A
Rationale: Parametric models (e.g., logistic regression) have a fixed
number of parameters regardless of training data size, while non-
parametric models (e.g., K-NN) grow in complexity with more data.
Question 2
What is a key drawback of K-Nearest Neighbors (K-NN)?
A) It requires storing all training data and computing distances at
inference time
,B) It cannot handle classification tasks
C) It has too many trainable parameters
D) It requires gradient descent
Correct Answer: A
Rationale: K-NN is a non-parametric, lazy learner that stores all
training data and computes distances at inference, making it
computationally expensive for large datasets.
Question 3
Which type of model is K-NN?
A) Parametric
B) Non-parametric
C) Linear
D) Generative
Correct Answer: B
Rationale: K-NN is non-parametric because it does not learn a fixed
set of parameters; instead, it memorizes the training data.
Question 4
What decision boundary does logistic regression produce?
A) A linear combination passed through a sigmoid function
B) A circular boundary
,C) A polynomial of degree 3
D) A non-linear kernel boundary
Correct Answer: A
Rationale: Logistic regression computes a linear combination of
inputs (w·x + b) and passes it through a sigmoid to produce
probabilities.
Question 5
What is the difference between MLE and MAP estimation?
A) MAP incorporates a prior distribution; MLE does not
B) MLE incorporates a prior; MAP does not
C) Both incorporate priors
D) Neither incorporates priors
Correct Answer: A
Rationale: Maximum Likelihood Estimation (MLE) maximizes the
likelihood P(D|θ). Maximum A Posteriori (MAP) maximizes
P(D|θ)P(θ), incorporating a prior on parameters.
Question 6
Which statement about Naïve Bayes is correct?
A) It assumes conditional independence of features given the class
B) It assumes features are dependent on each other
, C) It requires gradient descent
D) It cannot handle text data
Correct Answer: A
Rationale: Naïve Bayes assumes features are conditionally
independent given the class label, simplifying computation.
Question 7
What causes the vanishing gradient problem?
A) Sigmoid and tanh activations saturate, producing gradients near zero
B) ReLU activations produce large gradients
C) Too high a learning rate
D) Too few layers in the network
Correct Answer: A
Rationale: Sigmoid and tanh saturate for large positive/negative
inputs, causing gradients to approach zero and preventing deep
layers from learning.
Question 8
Which activation function has a range of (0, 1)?
A) Sigmoid
B) Tanh