Google Professional Data Engineer Practice Test
Exam
1. A retail company needs to store 50 TB of structured sales
transaction data for ad-hoc analytical queries by business
analysts. The data requires complex joins and must support SQL
queries with sub-second response times for dashboards. Data
ingestion occurs in daily batches. Which storage solution is most
appropriate?
A) Cloud Bigtable with Apache HBase API for fast random access
B) Cloud Storage with Avro format and BigQuery external tables
C) BigQuery with partitioned and clustered tables for analytical
workloads
D) Cloud SQL PostgreSQL with read replicas for query
distribution
Correct Answer: C
Rationale: BigQuery is designed for analytical workloads with
structured data, supporting complex SQL joins and fast query
performance. Partitioning and clustering optimize query
performance and cost. Cloud Bigtable is for high-throughput
transactional workloads, not complex analytical joins. Cloud
,Storage external tables have higher latency for frequent
queries.
2. A gaming company needs to store player session data with
high write throughput of 1 million events per second and low-
latency reads under 10 milliseconds for real-time leaderboards.
The data schema evolves frequently, and the system must
handle petabyte-scale storage. Which Google Cloud service best
meets these requirements?
A) BigQuery with streaming inserts
B) Cloud Bigtable with column-family design for time-series data
C) Cloud SQL with sharding
D) Firestore in Native mode with automatic scaling
Correct Answer: B
Rationale: Cloud Bigtable is a high-performance NoSQL
database designed for massive write throughput and low-
latency reads at petabyte scale. It supports flexible schema
evolution. BigQuery has higher query latency, Cloud SQL cannot
handle this scale, and Firestore is optimized for mobile and web
applications rather than high-throughput time-series data.
,3. A financial services firm must retain transaction logs for 7
years to comply with regulations. The data is rarely accessed
after 90 days but must remain immediately accessible. The
solution must minimize storage costs while ensuring data
integrity and encryption. Which approach is most cost-
effective?
A) Store all data in BigQuery with long-term storage pricing
B) Use Cloud Storage with lifecycle policies: Standard for 90
days, then Nearline for 2 years, then Coldline for remaining
period
C) Maintain everything in Cloud SQL with automated backups
D) Use Cloud Bigtable with time-to-live policies set to 7 years
Correct Answer: B
Rationale: Cloud Storage lifecycle policies automatically
transition data to lower-cost storage classes based on access
patterns. Nearline and Coldline provide cost savings while
maintaining immediate accessibility. BigQuery long-term
storage is not the most cost-effective for rarely accessed raw
logs, and Cloud SQL and Bigtable are not designed for archival
log storage.
, 4. A healthcare organization needs to design a data lake for
genomic research data. The system must support batch
processing of large files exceeding 100 GB each, accommodate
unstructured and semi-structured data, and integrate with
existing Apache Spark workloads. Data residency requirements
mandate storage in specific regions. Which architecture is
optimal?
A) Cloud Storage regional buckets with Dataproc for Spark
processing
B) BigQuery with clustered tables for genomic sequences
C) Cloud Bigtable for variant call format storage
D) Cloud SQL with JSON columns for flexible schema
Correct Answer: A
Rationale: Cloud Storage regional buckets provide cost-effective
storage for large unstructured and semi-structured files with
regional residency control. Dataproc provides managed Apache
Spark for batch processing. BigQuery is not ideal for raw
genomic sequence files, Cloud Bigtable is not designed for large
file storage, and Cloud SQL has size limitations for this use case.
Exam
1. A retail company needs to store 50 TB of structured sales
transaction data for ad-hoc analytical queries by business
analysts. The data requires complex joins and must support SQL
queries with sub-second response times for dashboards. Data
ingestion occurs in daily batches. Which storage solution is most
appropriate?
A) Cloud Bigtable with Apache HBase API for fast random access
B) Cloud Storage with Avro format and BigQuery external tables
C) BigQuery with partitioned and clustered tables for analytical
workloads
D) Cloud SQL PostgreSQL with read replicas for query
distribution
Correct Answer: C
Rationale: BigQuery is designed for analytical workloads with
structured data, supporting complex SQL joins and fast query
performance. Partitioning and clustering optimize query
performance and cost. Cloud Bigtable is for high-throughput
transactional workloads, not complex analytical joins. Cloud
,Storage external tables have higher latency for frequent
queries.
2. A gaming company needs to store player session data with
high write throughput of 1 million events per second and low-
latency reads under 10 milliseconds for real-time leaderboards.
The data schema evolves frequently, and the system must
handle petabyte-scale storage. Which Google Cloud service best
meets these requirements?
A) BigQuery with streaming inserts
B) Cloud Bigtable with column-family design for time-series data
C) Cloud SQL with sharding
D) Firestore in Native mode with automatic scaling
Correct Answer: B
Rationale: Cloud Bigtable is a high-performance NoSQL
database designed for massive write throughput and low-
latency reads at petabyte scale. It supports flexible schema
evolution. BigQuery has higher query latency, Cloud SQL cannot
handle this scale, and Firestore is optimized for mobile and web
applications rather than high-throughput time-series data.
,3. A financial services firm must retain transaction logs for 7
years to comply with regulations. The data is rarely accessed
after 90 days but must remain immediately accessible. The
solution must minimize storage costs while ensuring data
integrity and encryption. Which approach is most cost-
effective?
A) Store all data in BigQuery with long-term storage pricing
B) Use Cloud Storage with lifecycle policies: Standard for 90
days, then Nearline for 2 years, then Coldline for remaining
period
C) Maintain everything in Cloud SQL with automated backups
D) Use Cloud Bigtable with time-to-live policies set to 7 years
Correct Answer: B
Rationale: Cloud Storage lifecycle policies automatically
transition data to lower-cost storage classes based on access
patterns. Nearline and Coldline provide cost savings while
maintaining immediate accessibility. BigQuery long-term
storage is not the most cost-effective for rarely accessed raw
logs, and Cloud SQL and Bigtable are not designed for archival
log storage.
, 4. A healthcare organization needs to design a data lake for
genomic research data. The system must support batch
processing of large files exceeding 100 GB each, accommodate
unstructured and semi-structured data, and integrate with
existing Apache Spark workloads. Data residency requirements
mandate storage in specific regions. Which architecture is
optimal?
A) Cloud Storage regional buckets with Dataproc for Spark
processing
B) BigQuery with clustered tables for genomic sequences
C) Cloud Bigtable for variant call format storage
D) Cloud SQL with JSON columns for flexible schema
Correct Answer: A
Rationale: Cloud Storage regional buckets provide cost-effective
storage for large unstructured and semi-structured files with
regional residency control. Dataproc provides managed Apache
Spark for batch processing. BigQuery is not ideal for raw
genomic sequence files, Cloud Bigtable is not designed for large
file storage, and Cloud SQL has size limitations for this use case.