Apache Spark Developer Certification Exam Verified
Questions, Correct Answers, and Detailed
Explanations for Computer Science Students||Already
Graded A+
1. What is the default storage level of an RDD in Spark?
A) MEMORY_ONLY
B) MEMORY_AND_DISK
C) DISK_ONLY
D) MEMORY_ONLY_SER
By default, RDDs are not persisted, but if persisted without specifying,
Spark stores them in memory only.
2. Which of the following is a transformation in Spark?
A) map()
B) collect()
C) count()
D) saveAsTextFile()
Transformations return a new RDD and are lazy; actions return a
value or write data.
3. What is the main difference between a DataFrame and an RDD?
A) DataFrames are immutable, RDDs are mutable
B) RDDs have a schema, DataFrames do not
C) DataFrames have a schema and are optimized, RDDs are not
D) RDDs are faster than DataFrames
DataFrames provide schema information and use Catalyst optimizer
for query planning.
,4. Which Spark component is used for graph processing?
A) Spark SQL
B) GraphX
C) MLlib
D) Spark Streaming
GraphX is the Spark API for graphs and graph-parallel computation.
5. What does the persist() method do in Spark?
A) Deletes an RDD
B) Converts an RDD to a DataFrame
C) Stores an RDD in memory/disk for reuse
D) Triggers computation immediately
persist() allows caching an RDD in memory/disk to avoid
recomputation.
6. Which language(s) is/are supported by Spark?
A) Java only
B) Scala only
C) Python only
D) Scala, Java, Python, R, and SQL
Spark has APIs for multiple languages, making it versatile.
7. What type of operation is filter() in Spark?
A) Action
B) Transformation
C) Output
D) Persisting
filter() creates a new RDD based on a condition, so it is a
transformation.
, 8. Which Spark module provides machine learning functionality?
A) Spark SQL
B) GraphX
C) MLlib
D) Spark Streaming
MLlib provides scalable machine learning algorithms.
9. What is the role of the Spark driver program?
A) Stores data in HDFS
B) Executes tasks on the cluster
C) Coordinates the execution of tasks on executors
D) Reads data from Kafka
The driver schedules tasks and maintains metadata about the Spark
application.
10. In Spark, what is a narrow transformation?
A) Transformation requiring a shuffle
B) Transformation where each output partition depends on a single
input partition
C) Transformation that writes to disk
D) Transformation that triggers an action
Narrow transformations do not require data movement across
partitions.
11. What Spark feature allows iterative computations to be faster?
A) DAG
B) RDD caching/persistence
Questions, Correct Answers, and Detailed
Explanations for Computer Science Students||Already
Graded A+
1. What is the default storage level of an RDD in Spark?
A) MEMORY_ONLY
B) MEMORY_AND_DISK
C) DISK_ONLY
D) MEMORY_ONLY_SER
By default, RDDs are not persisted, but if persisted without specifying,
Spark stores them in memory only.
2. Which of the following is a transformation in Spark?
A) map()
B) collect()
C) count()
D) saveAsTextFile()
Transformations return a new RDD and are lazy; actions return a
value or write data.
3. What is the main difference between a DataFrame and an RDD?
A) DataFrames are immutable, RDDs are mutable
B) RDDs have a schema, DataFrames do not
C) DataFrames have a schema and are optimized, RDDs are not
D) RDDs are faster than DataFrames
DataFrames provide schema information and use Catalyst optimizer
for query planning.
,4. Which Spark component is used for graph processing?
A) Spark SQL
B) GraphX
C) MLlib
D) Spark Streaming
GraphX is the Spark API for graphs and graph-parallel computation.
5. What does the persist() method do in Spark?
A) Deletes an RDD
B) Converts an RDD to a DataFrame
C) Stores an RDD in memory/disk for reuse
D) Triggers computation immediately
persist() allows caching an RDD in memory/disk to avoid
recomputation.
6. Which language(s) is/are supported by Spark?
A) Java only
B) Scala only
C) Python only
D) Scala, Java, Python, R, and SQL
Spark has APIs for multiple languages, making it versatile.
7. What type of operation is filter() in Spark?
A) Action
B) Transformation
C) Output
D) Persisting
filter() creates a new RDD based on a condition, so it is a
transformation.
, 8. Which Spark module provides machine learning functionality?
A) Spark SQL
B) GraphX
C) MLlib
D) Spark Streaming
MLlib provides scalable machine learning algorithms.
9. What is the role of the Spark driver program?
A) Stores data in HDFS
B) Executes tasks on the cluster
C) Coordinates the execution of tasks on executors
D) Reads data from Kafka
The driver schedules tasks and maintains metadata about the Spark
application.
10. In Spark, what is a narrow transformation?
A) Transformation requiring a shuffle
B) Transformation where each output partition depends on a single
input partition
C) Transformation that writes to disk
D) Transformation that triggers an action
Narrow transformations do not require data movement across
partitions.
11. What Spark feature allows iterative computations to be faster?
A) DAG
B) RDD caching/persistence