CCA Spark and Hadoop Developer Exam
Independent Revision Guide and Practice
CCA SPARK AND HADOOP DEVELOPER EXAM – COMPLETE Q&A BANK
(200+ QUESTIONS)
EXAM CODE: CCA-175 | HANDS-ON PERFORMANCE-BASED | SPARK 2.4 |
CDH 6.1.1
UPDATED 2025/2026 | SCALA & PYTHON | VERIFIED ANSWERS WITH
RATIONALES
SECTION 1: DATA INGESTION – SQOOP & FLUME (Questions 1-25)
1. Which Sqoop command is used to import data from a relational database into
HDFS?
A) sqoop export
B) sqoop import
C) sqoop eval
D) sqoop list-tables
Answer: B
1
,Rationale: Sqoop import transfers data from RDBMS tables to HDFS. The export
command moves data from HDFS to RDBMS .
2. Which Sqoop command lists all tables in a given database?
A) sqoop list-databases
B) sqoop list-tables
C) sqoop show-tables
D) sqoop eval
Answer: B
Rationale: The sqoop list-tables command displays all tables in a specified
database .
3. The --target-dir parameter in Sqoop import specifies:
A) The database name
B) The HDFS output directory
C) The table name
D) The number of mappers
Answer: B
Rationale: --target-dir defines the HDFS directory where imported data will be
stored .
2
,4. The --as-parquetfile option in Sqoop import:
A) Imports data as text files
B) Imports data as Parquet files
C) Imports data as Avro files
D) Imports data as Sequence files
Answer: B
Rationale: --as-parquetfile creates Parquet-format output, which is more efficient
for analytics .
5. What does the --split-by parameter do in Sqoop import?
A) Specifies the column used for splitting the data
B) Specifies the number of mappers
C) Specifies the compression codec
D) Specifies the output format
Answer: A
Rationale: --split-by defines the column used to divide the data among parallel
mappers .
6. Which Sqoop command lists all databases in a MySQL server?
A) sqoop list-tables
3
, B) sqoop list-databases
C) sqoop show-databases
D) sqoop eval
Answer: B
Rationale: sqoop list-databases displays all databases available on the connected
RDBMS server .
7. Flume is primarily designed for:
A) Batch processing
B) Streaming data ingestion
C) Data storage
D) Data visualization
Answer: B
Rationale: Flume is a distributed service for efficiently collecting, aggregating, and
moving large amounts of streaming data .
8. The memory channel in Flume stores events in:
A) HDFS
B) The local file system
C) A JVM heap
4
Independent Revision Guide and Practice
CCA SPARK AND HADOOP DEVELOPER EXAM – COMPLETE Q&A BANK
(200+ QUESTIONS)
EXAM CODE: CCA-175 | HANDS-ON PERFORMANCE-BASED | SPARK 2.4 |
CDH 6.1.1
UPDATED 2025/2026 | SCALA & PYTHON | VERIFIED ANSWERS WITH
RATIONALES
SECTION 1: DATA INGESTION – SQOOP & FLUME (Questions 1-25)
1. Which Sqoop command is used to import data from a relational database into
HDFS?
A) sqoop export
B) sqoop import
C) sqoop eval
D) sqoop list-tables
Answer: B
1
,Rationale: Sqoop import transfers data from RDBMS tables to HDFS. The export
command moves data from HDFS to RDBMS .
2. Which Sqoop command lists all tables in a given database?
A) sqoop list-databases
B) sqoop list-tables
C) sqoop show-tables
D) sqoop eval
Answer: B
Rationale: The sqoop list-tables command displays all tables in a specified
database .
3. The --target-dir parameter in Sqoop import specifies:
A) The database name
B) The HDFS output directory
C) The table name
D) The number of mappers
Answer: B
Rationale: --target-dir defines the HDFS directory where imported data will be
stored .
2
,4. The --as-parquetfile option in Sqoop import:
A) Imports data as text files
B) Imports data as Parquet files
C) Imports data as Avro files
D) Imports data as Sequence files
Answer: B
Rationale: --as-parquetfile creates Parquet-format output, which is more efficient
for analytics .
5. What does the --split-by parameter do in Sqoop import?
A) Specifies the column used for splitting the data
B) Specifies the number of mappers
C) Specifies the compression codec
D) Specifies the output format
Answer: A
Rationale: --split-by defines the column used to divide the data among parallel
mappers .
6. Which Sqoop command lists all databases in a MySQL server?
A) sqoop list-tables
3
, B) sqoop list-databases
C) sqoop show-databases
D) sqoop eval
Answer: B
Rationale: sqoop list-databases displays all databases available on the connected
RDBMS server .
7. Flume is primarily designed for:
A) Batch processing
B) Streaming data ingestion
C) Data storage
D) Data visualization
Answer: B
Rationale: Flume is a distributed service for efficiently collecting, aggregating, and
moving large amounts of streaming data .
8. The memory channel in Flume stores events in:
A) HDFS
B) The local file system
C) A JVM heap
4