Online Learning at Chamberlain
MEDICINE 207/2028 EDITION
NURSE WISEMAN LATEST EDITION NURSING
,AWS - Analytics and Ingestion Exam
Review| Grade A + | Questions and
Verified Answers| 100% Correct|
A company stores structured logs in S3 and needs occasional ad hoc
SQL queries without loading the data into a database. Which service
should it use? - Answers ✔️✔️Amazon Athena.
A company runs Athena queries over unpartitioned multi-terabyte
data and costs are high. What optimization should it apply? - Answers
✔️✔️Partition data, use columnar formats such as Parquet or ORC,
compress files, and select only required columns.
An Athena query uses SELECT * when only two columns are needed.
What cost problem does this create? - Answers ✔️✔️Athena pricing is
based largely on data scanned, so reading unnecessary columns
increases cost.
A company has thousands of tiny S3 files and Athena performance is
poor. What improvement should it make? - Answers ✔️✔️Compact
small files into appropriately sized columnar objects and use sensible
partitioning.
A company needs shared metadata describing S3 datasets for Athena,
EMR, and Redshift Spectrum. Which service should it use? - Answers
✔️✔️The AWS Glue Data Catalog.
, A company adds new partitions to S3 every day and wants metadata
discovered automatically. Which Glue component can help? - Answers
✔️✔️An AWS Glue crawler, although partition projection or explicit
catalog updates may be more efficient for predictable layouts.
A company needs a serverless Spark-based ETL job to transform data
in S3. Which service should it use? - Answers ✔️✔️AWS Glue.
A company needs full control of Hadoop and Spark cluster
configuration and long-running big-data frameworks. Which service
should it use? - Answers ✔️✔️Amazon EMR.
A company needs transient Spark clusters that terminate after each
job. Which EMR deployment can it use? - Answers ✔️✔️An ephemeral
EMR cluster, EMR Serverless, or EMR on EKS depending on operational
and platform requirements.
A company wants to avoid managing EMR cluster instances for
intermittent Spark jobs. Which option should it evaluate? - Answers
✔️✔️Amazon EMR Serverless.
A company needs a petabyte-scale relational data warehouse for
repeated complex BI queries. Which service should it use? - Answers
✔️✔️Amazon Redshift.
MEDICINE 207/2028 EDITION
NURSE WISEMAN LATEST EDITION NURSING
,AWS - Analytics and Ingestion Exam
Review| Grade A + | Questions and
Verified Answers| 100% Correct|
A company stores structured logs in S3 and needs occasional ad hoc
SQL queries without loading the data into a database. Which service
should it use? - Answers ✔️✔️Amazon Athena.
A company runs Athena queries over unpartitioned multi-terabyte
data and costs are high. What optimization should it apply? - Answers
✔️✔️Partition data, use columnar formats such as Parquet or ORC,
compress files, and select only required columns.
An Athena query uses SELECT * when only two columns are needed.
What cost problem does this create? - Answers ✔️✔️Athena pricing is
based largely on data scanned, so reading unnecessary columns
increases cost.
A company has thousands of tiny S3 files and Athena performance is
poor. What improvement should it make? - Answers ✔️✔️Compact
small files into appropriately sized columnar objects and use sensible
partitioning.
A company needs shared metadata describing S3 datasets for Athena,
EMR, and Redshift Spectrum. Which service should it use? - Answers
✔️✔️The AWS Glue Data Catalog.
, A company adds new partitions to S3 every day and wants metadata
discovered automatically. Which Glue component can help? - Answers
✔️✔️An AWS Glue crawler, although partition projection or explicit
catalog updates may be more efficient for predictable layouts.
A company needs a serverless Spark-based ETL job to transform data
in S3. Which service should it use? - Answers ✔️✔️AWS Glue.
A company needs full control of Hadoop and Spark cluster
configuration and long-running big-data frameworks. Which service
should it use? - Answers ✔️✔️Amazon EMR.
A company needs transient Spark clusters that terminate after each
job. Which EMR deployment can it use? - Answers ✔️✔️An ephemeral
EMR cluster, EMR Serverless, or EMR on EKS depending on operational
and platform requirements.
A company wants to avoid managing EMR cluster instances for
intermittent Spark jobs. Which option should it evaluate? - Answers
✔️✔️Amazon EMR Serverless.
A company needs a petabyte-scale relational data warehouse for
repeated complex BI queries. Which service should it use? - Answers
✔️✔️Amazon Redshift.