WGU D467 Exploring Data | Western Governors
University | Academic Year 2026/2027
Section 1: Data Structures, Types, Formats, and Integrity (Questions 1–30)
1. A data analyst is working with a dataset where each row represents a single
customer transaction. The columns include transaction ID, date, amount, and
product category. What type of data structure is this?
A. Unstructured data
B. Structured data
C. Semi-structured data
D. Metadata
B. Structured data
Structured data is organized into a tabular format with rows and columns,
making it easily searchable and analyzable. Transactional records with defined
fields are a classic example of structured data. Unstructured data lacks a
predefined format (e.g., free-text emails), and semi-structured data has some
organizational properties but not a rigid tabular schema (e.g., JSON).
2. Which of the following are examples of unstructured data? (Select all that
apply.)
A. A spreadsheet of monthly sales figures
B. A collection of customer service call recordings
C. A relational database table of employee records
D. A folder of scanned handwritten notes
E. A CSV file of website clickstream data
B. A collection of customer service call recordings
D. A folder of scanned handwritten notes
, Unstructured data does not conform to a predefined data model and is
difficult to organize in relational tables. Audio recordings and handwritten notes
are unstructured because their content cannot be easily parsed into rows and
columns. Spreadsheets, relational tables, and CSV files are structured formats.
3. An analyst discovers that a column labeled “Age” in a dataset contains values
such as “twenty-five” and “unknown” alongside numeric entries. This is
primarily an issue of:
A. Data redundancy
B. Data integrity
C. Data privacy
D. Data visualization
B. Data integrity
Data integrity refers to the accuracy, consistency, and reliability of data over
its lifecycle. Mixing text and numeric values in a field expected to be numeric
violates integrity constraints and will cause errors in analysis. Redundancy refers
to duplicate storage, privacy to access control, and visualization to presentation.
4. Which data type is most appropriate for storing a postal ZIP code in a
database when leading zeros must be preserved?
A. INTEGER
B. FLOAT
C. VARCHAR
D. BOOLEAN
C. VARCHAR
ZIP codes can have leading zeros (e.g., “01234”), which would be stripped if
stored as an integer or float. VARCHAR (variable-length character) preserves
leading zeros and allows alphanumeric characters if needed. BOOLEAN is for
true/false values.
,5. A dataset contains a column “Order_Date” with entries like “2026-01-15”,
“01/15/2026”, and “15-Jan-2026”. This is best described as an issue with:
A. Data format consistency
B. Data volume
C. Data governance
D. Data modeling
A. Data format consistency
When the same logical field is represented in multiple date formats, the data
lacks format consistency. This makes sorting, filtering, and joining unreliable.
Standardizing to a single format (e.g., ISO 8601: YYYY-MM-DD) is a standard
cleaning step.
6. What is the primary purpose of a data dictionary?
A. To store the actual data values
B. To describe the structure, meaning, and relationships of data elements
C. To visualize data trends
D. To encrypt sensitive data
B. To describe the structure, meaning, and relationships of data elements
A data dictionary (or metadata repository) documents field names, data
types, allowed values, and relationships. It does not store the data itself, perform
visualization, or handle encryption.
7. Which of the following are common data integrity constraints in a relational
database? (Select all that apply.)
A. Primary key uniqueness
B. Foreign key referential integrity
C. NOT NULL constraints
, D. Data visualization rules
E. Encryption algorithms
A. Primary key uniqueness
B. Foreign key referential integrity
C. NOT NULL constraints
Primary keys enforce uniqueness, foreign keys ensure referential integrity
between tables, and NOT NULL ensures a column cannot contain missing values.
Visualization rules and encryption are not integrity constraints in the database
schema sense.
8. A company stores customer feedback in a JSON file where each record may
have different fields. This is an example of:
A. Structured data
B. Semi-structured data
C. Unstructured data
D. Transactional data
B. Semi-structured data
Semi-structured data (e.g., JSON, XML) contains tags or markers that separate
semantic elements but does not enforce a rigid tabular schema. JSON records can
have varying fields, unlike structured data, but they are more organized than
unstructured data such as free text.
9. An analyst notices that the same customer appears three times in a dataset
with slight variations in name spelling (“Jon Smith”, “John Smith”, “J. Smith”).
This is an example of:
A. Data duplication
B. Data inconsistency
C. Data redundancy
D. Data loss
University | Academic Year 2026/2027
Section 1: Data Structures, Types, Formats, and Integrity (Questions 1–30)
1. A data analyst is working with a dataset where each row represents a single
customer transaction. The columns include transaction ID, date, amount, and
product category. What type of data structure is this?
A. Unstructured data
B. Structured data
C. Semi-structured data
D. Metadata
B. Structured data
Structured data is organized into a tabular format with rows and columns,
making it easily searchable and analyzable. Transactional records with defined
fields are a classic example of structured data. Unstructured data lacks a
predefined format (e.g., free-text emails), and semi-structured data has some
organizational properties but not a rigid tabular schema (e.g., JSON).
2. Which of the following are examples of unstructured data? (Select all that
apply.)
A. A spreadsheet of monthly sales figures
B. A collection of customer service call recordings
C. A relational database table of employee records
D. A folder of scanned handwritten notes
E. A CSV file of website clickstream data
B. A collection of customer service call recordings
D. A folder of scanned handwritten notes
, Unstructured data does not conform to a predefined data model and is
difficult to organize in relational tables. Audio recordings and handwritten notes
are unstructured because their content cannot be easily parsed into rows and
columns. Spreadsheets, relational tables, and CSV files are structured formats.
3. An analyst discovers that a column labeled “Age” in a dataset contains values
such as “twenty-five” and “unknown” alongside numeric entries. This is
primarily an issue of:
A. Data redundancy
B. Data integrity
C. Data privacy
D. Data visualization
B. Data integrity
Data integrity refers to the accuracy, consistency, and reliability of data over
its lifecycle. Mixing text and numeric values in a field expected to be numeric
violates integrity constraints and will cause errors in analysis. Redundancy refers
to duplicate storage, privacy to access control, and visualization to presentation.
4. Which data type is most appropriate for storing a postal ZIP code in a
database when leading zeros must be preserved?
A. INTEGER
B. FLOAT
C. VARCHAR
D. BOOLEAN
C. VARCHAR
ZIP codes can have leading zeros (e.g., “01234”), which would be stripped if
stored as an integer or float. VARCHAR (variable-length character) preserves
leading zeros and allows alphanumeric characters if needed. BOOLEAN is for
true/false values.
,5. A dataset contains a column “Order_Date” with entries like “2026-01-15”,
“01/15/2026”, and “15-Jan-2026”. This is best described as an issue with:
A. Data format consistency
B. Data volume
C. Data governance
D. Data modeling
A. Data format consistency
When the same logical field is represented in multiple date formats, the data
lacks format consistency. This makes sorting, filtering, and joining unreliable.
Standardizing to a single format (e.g., ISO 8601: YYYY-MM-DD) is a standard
cleaning step.
6. What is the primary purpose of a data dictionary?
A. To store the actual data values
B. To describe the structure, meaning, and relationships of data elements
C. To visualize data trends
D. To encrypt sensitive data
B. To describe the structure, meaning, and relationships of data elements
A data dictionary (or metadata repository) documents field names, data
types, allowed values, and relationships. It does not store the data itself, perform
visualization, or handle encryption.
7. Which of the following are common data integrity constraints in a relational
database? (Select all that apply.)
A. Primary key uniqueness
B. Foreign key referential integrity
C. NOT NULL constraints
, D. Data visualization rules
E. Encryption algorithms
A. Primary key uniqueness
B. Foreign key referential integrity
C. NOT NULL constraints
Primary keys enforce uniqueness, foreign keys ensure referential integrity
between tables, and NOT NULL ensures a column cannot contain missing values.
Visualization rules and encryption are not integrity constraints in the database
schema sense.
8. A company stores customer feedback in a JSON file where each record may
have different fields. This is an example of:
A. Structured data
B. Semi-structured data
C. Unstructured data
D. Transactional data
B. Semi-structured data
Semi-structured data (e.g., JSON, XML) contains tags or markers that separate
semantic elements but does not enforce a rigid tabular schema. JSON records can
have varying fields, unlike structured data, but they are more organized than
unstructured data such as free text.
9. An analyst notices that the same customer appears three times in a dataset
with slight variations in name spelling (“Jon Smith”, “John Smith”, “J. Smith”).
This is an example of:
A. Data duplication
B. Data inconsistency
C. Data redundancy
D. Data loss