SEC595 Exam 2026 | SANS Applied
Data Science & AI/ML
Prepare for the SEC595 exam with focused review of
Applied Data Science and AI/Machine Learning for
Cybersecurity Professionals. Topics include Python, data
acquisition, SQL, MongoDB, statistics, probability,
Bayesian inference, clustering, K-Means, SVMs, decision
trees, random forests, deep learning, neural networks,
autoencoders, anomaly detection, and model
deployment. Ideal for SEC595 and GMLE preparation.
Question 1:
Which of the following is a measure of central tendency?
A) Standard deviation
B) Mean
C) Range
D) Variance
Answer: B
,2|Page
Rationale: The mean is a measure of central tendency that represents the
average of a data set. In contrast, standard deviation, range, and variance
are all measures of dispersion or spread.
Question 2:
In machine learning, what is the primary purpose of a training dataset?
A) To evaluate the final model's performance
B) To tune hyperparameters
C) To teach the model patterns and relationships
D) To deploy the model into production
Answer: C
Rationale: The training dataset is used to fit the model by allowing it to
learn patterns and relationships within the data. Evaluation and
hyperparameter tuning typically use validation or test sets, while
deployment is a separate stage.
Question 3:
Which Python library is most commonly used for data manipulation and
analysis in cybersecurity data science workflows?
A) Matplotlib
B) Pandas
C) Flask
D) PyGame
Answer: B
Rationale: Pandas is the standard Python library for data manipulation
and analysis, offering DataFrames and Series structures ideal for handling
structured cybersecurity data. Matplotlib is for visualization, Flask is a web
framework, and PyGame is for game development.
,3|Page
Question 4:
What type of machine learning is used when the data has no labeled
outcomes?
A) Supervised learning
B) Reinforcement learning
C) Unsupervised learning
D) Semi-supervised learning
Answer: C
Rationale: Unsupervised learning works with unlabeled data to discover
hidden patterns or groupings, such as clustering network traffic. Supervised
learning requires labels, reinforcement learning uses rewards, and semi-
supervised learning uses a mix of labeled and unlabeled data.
Question 5:
Which of the following is an example of a supervised learning algorithm?
A) K-means clustering
B) Principal Component Analysis
C) Random Forest
D) DBSCAN
Answer: C
Rationale: Random Forest is a supervised learning algorithm used for
classification and regression tasks. K-means, PCA, and DBSCAN are all
unsupervised techniques used for clustering or dimensionality reduction.
Question 6:
In cybersecurity, anomaly detection is most closely associated with which
type of machine learning?
, 4|Page
A) Supervised classification
B) Unsupervised learning
C) Reinforcement learning
D) Transfer learning
Answer: B
Rationale: Anomaly detection typically uses unsupervised learning because
it seeks to identify unusual patterns without relying on labeled examples of
attacks. This makes it effective for detecting novel threats.
Question 7:
What does the term "feature" refer to in machine learning?
A) The output variable being predicted
B) An individual measurable property of the data
C) The loss function used during training
D) The number of layers in a neural network
Answer: B
Rationale: A feature is an individual measurable property or characteristic
of the data used as input to a model. The output variable is called the label
or target, the loss function measures error, and layers relate to network
architecture.
Question 8:
Which metric is most appropriate for evaluating a model on an
imbalanced dataset?
A) Accuracy
B) Mean Squared Error
C) F1 Score