1. What is the difference between variance and bias in machine learning models?
A. Variance refers to the error introduced by the model's assumptions, while bias
refers to the variability in predictions across different datasets.
B. Variance refers to the error introduced by the model’s assumptions, while bias
refers to the variability of the model’s predictions on the same dataset.
C. Variance refers to the variability of the model's predictions, while bias refers to the
errors caused by oversimplifying the model.
D. There is no difference between variance and bias in machine learning.
Answer: C) Variance refers to the variability of the model's predictions, while bias
refers to the errors caused by oversimplifying the model.
Rationale: High variance indicates that the model is sensitive to fluctuations in the
training data (overfitting), while high bias indicates that the model is too simple and
does not capture the underlying patterns (underfitting).
2. Which of the following algorithms can be used for both regression and
classification tasks?
A. K-Nearest Neighbors (KNN)
,B. Support Vector Machine (SVM)
C. Naive Bayes
D. K-Means Clustering
Answer: B) Support Vector Machine (SVM)
Rationale: SVM can be used for both classification (e.g., SVC - Support Vector
Classification) and regression (e.g., SVR - Support Vector Regression).
3. Which of the following is the best way to handle large datasets in machine
learning?
A. Train the model on a small sample of the data
B. Use a simpler model with fewer features
C. Apply dimensionality reduction techniques
D. Increase the number of iterations in training
Answer: C) Apply dimensionality reduction techniques
Rationale: Dimensionality reduction techniques like PCA help to reduce the size and
complexity of the data, making it easier to handle large datasets.
4. In the context of regression, what does R-squared measure?
A. The accuracy of the predictions
, B. The proportion of variance in the dependent variable that is predictable from the
independent variables
C. The number of outliers in the dataset
D. The mean absolute error of the model
Answer: B) The proportion of variance in the dependent variable that is predictable
from the independent variables
Rationale: R-squared measures how well the independent variables explain the
variability of the dependent variable.
5. Which of the following is not a type of machine learning?
A. Supervised learning
B. Unsupervised learning
C. Reinforcement learning
D. Structured learning
Answer: D) Structured learning
Rationale: Structured learning is not a recognized category of machine learning. The
main types are supervised, unsupervised, and reinforcement learning.
6. What does the term “overfitting” mean in machine learning?
A. Variance refers to the error introduced by the model's assumptions, while bias
refers to the variability in predictions across different datasets.
B. Variance refers to the error introduced by the model’s assumptions, while bias
refers to the variability of the model’s predictions on the same dataset.
C. Variance refers to the variability of the model's predictions, while bias refers to the
errors caused by oversimplifying the model.
D. There is no difference between variance and bias in machine learning.
Answer: C) Variance refers to the variability of the model's predictions, while bias
refers to the errors caused by oversimplifying the model.
Rationale: High variance indicates that the model is sensitive to fluctuations in the
training data (overfitting), while high bias indicates that the model is too simple and
does not capture the underlying patterns (underfitting).
2. Which of the following algorithms can be used for both regression and
classification tasks?
A. K-Nearest Neighbors (KNN)
,B. Support Vector Machine (SVM)
C. Naive Bayes
D. K-Means Clustering
Answer: B) Support Vector Machine (SVM)
Rationale: SVM can be used for both classification (e.g., SVC - Support Vector
Classification) and regression (e.g., SVR - Support Vector Regression).
3. Which of the following is the best way to handle large datasets in machine
learning?
A. Train the model on a small sample of the data
B. Use a simpler model with fewer features
C. Apply dimensionality reduction techniques
D. Increase the number of iterations in training
Answer: C) Apply dimensionality reduction techniques
Rationale: Dimensionality reduction techniques like PCA help to reduce the size and
complexity of the data, making it easier to handle large datasets.
4. In the context of regression, what does R-squared measure?
A. The accuracy of the predictions
, B. The proportion of variance in the dependent variable that is predictable from the
independent variables
C. The number of outliers in the dataset
D. The mean absolute error of the model
Answer: B) The proportion of variance in the dependent variable that is predictable
from the independent variables
Rationale: R-squared measures how well the independent variables explain the
variability of the dependent variable.
5. Which of the following is not a type of machine learning?
A. Supervised learning
B. Unsupervised learning
C. Reinforcement learning
D. Structured learning
Answer: D) Structured learning
Rationale: Structured learning is not a recognized category of machine learning. The
main types are supervised, unsupervised, and reinforcement learning.
6. What does the term “overfitting” mean in machine learning?