AWS CERTIFIED MACHINE LEARNING
ENGINEER – ASSOCIATE (MLA-C01) STUDY
GUIDE | LATEST UPDATE 2026/2027 | PRACTICE
QUESTIONS AND ANSWERS | EXAM REVIEW |
100% CORRECT ANSWERS | VERIFIED
SOLUTIONS
This comprehensive practice examination is designed for candidates preparing for
the AWS Certified Machine Learning Engineer – Associate (MLA-C01)
certification. It serves as a complete study guide and exam review, reflecting the
latest update for 2026/2027. The questions are meticulously crafted to mirror the
difficulty, scenario-based style, and scope of the assessment, covering all four
domains: Data Preparation for Machine Learning (28%), ML Model Development
(26%), Deployment and Orchestration of ML Workflows (22%), and ML Solution
Monitoring, Maintenance, and Security (24%). By engaging with these practice
questions and answers, candidates can assess readiness, identify knowledge gaps,
and reinforce critical machine learning engineering concepts on AWS. The
detailed rationales provide verified solutions, ensuring deep comprehension and
enhancing preparation for this professional certification.
Table of Contents
1. Data Preparation for Machine Learning
2. ML Model Development
3. Deployment and Orchestration of ML Workflows
4. ML Solution Monitoring, Maintenance, and Security
, 2
SECTION 1: DATA PREPARATION FOR MACHINE LEARNING
Question 1: A machine learning engineer needs to store large volumes of training
data that will be accessed frequently during model training. The data is semi-
structured and requires high throughput. Which AWS storage service is MOST
appropriate?
A) Amazon S3
B) Amazon EBS
C) Amazon RDS
D) Amazon Redshift
Correct Answer: A
Amazon S3 is the most appropriate storage service for large volumes of training
data that require high throughput and frequent access. S3 provides virtually
unlimited scalability, high durability, and is the primary data lake service on AWS
for ML workloads. Option B (EBS) is block storage for EC2 instances and is not
designed for shared access. Option C (RDS) is a relational database service not
optimized for large-scale ML training data. Option D (Redshift) is a data
warehouse optimized for analytics, not primary ML training data storage.
Question 2: A data engineer is using AWS Glue to transform raw data into a
format suitable for machine learning. The transformation requires complex custom
logic that cannot be expressed with standard Glue transforms. What is the MOST
appropriate solution?
A) Use AWS Lambda functions within the Glue ETL job
B) Use Amazon EMR with Apache Spark
C) Use AWS Glue Studio with custom transforms
D) Use Amazon Athena for data transformation
Correct Answer: B
Amazon EMR with Apache Spark provides the flexibility to write complex custom
transformation logic using Spark's APIs. Option A (Lambda) has time and memory
limitations that may not suit large-scale ETL. Option C (Glue Studio) uses visual
transforms and may not support highly complex custom logic. Option D (Athena) is
a query service, not designed for complex ETL transformations.
, 3
Question 3: A machine learning engineer needs to label a large dataset of images
for a computer vision project. The labeling requires multiple annotators to ensure
accuracy. Which AWS service should be used?
A) Amazon SageMaker Ground Truth
B) Amazon Rekognition
C) Amazon Mechanical Turk
D) AWS Glue DataBrew
Correct Answer: A
Amazon SageMaker Ground Truth is specifically designed for data labeling with
support for multiple annotators, quality control mechanisms, and integration with
SageMaker. Option B (Rekognition) is a pre-trained computer vision service, not a
labeling tool. Option C (Mechanical Turk) is a crowdsourcing marketplace but
lacks the ML-specific workflow integration of Ground Truth. Option D (Glue
DataBrew) is a data preparation tool, not a labeling service.
Question 4: A data scientist is working with a dataset that contains missing values
in several columns. The data scientist wants to impute these missing values using
the median of each column. Which SageMaker Data Wrangler transform should be
used?
A) Fill missing values with mean
B) Fill missing values with median
C) Impute missing values with custom value
D) Handle missing values by dropping rows
Correct Answer: B
SageMaker Data Wrangler provides a "Fill missing values with median" transform
specifically for this purpose. Option A (mean) would not match the requirement.
Option C (custom value) is for specific values, not statistical measures. Option D
(dropping rows) would lose data.
Question 5: A machine learning engineer needs to ingest streaming data from IoT
devices for real-time inference. The data must be processed and delivered to
SageMaker endpoints with minimal latency. Which AWS service should be used?
, 4
A) Amazon Kinesis Data Streams
B) AWS Glue
C) Amazon S3
D) Amazon RDS
Correct Answer: A
Amazon Kinesis Data Streams is designed for real-time streaming data ingestion
and processing with low latency. Option B (Glue) is for batch ETL. Option C (S3)
is object storage and not optimized for real-time streaming. Option D (RDS) is a
relational database.
Question 6: A data engineer is preparing a dataset for training a binary
classification model. The dataset is highly imbalanced, with only 2% of samples
belonging to the positive class. Which technique should be used to address this
imbalance?
A) Oversample the minority class
B) Undersample the majority class
C) Use SMOTE (Synthetic Minority Over-sampling Technique)
D) All of the above
Correct Answer: D
All of the listed techniques can address class imbalance: oversampling the
minority class, undersampling the majority class, and using SMOTE to generate
synthetic samples. The choice depends on the specific dataset and requirements.
SageMaker Data Wrangler and custom processing jobs can implement these
techniques.
Question 7: A machine learning engineer needs to split a dataset into training,
validation, and test sets. The dataset contains a time series where chronological
order must be preserved. Which splitting strategy is MOST appropriate?
A) Random split
B) Stratified split
C) Time-based split
D) K-fold cross-validation