What feature in AWS Glue must be used if you have a lot Catalog Partition Predicates
of partitions for a table and need to implement server-side
partition pruning?
Helps you create point-to-point integrations between event Amazon EventBridge Pipes
producers and consumers with optional transform, filter
and enrich steps.
It refers to a temporary cluster that is created on-demand Transient EMR cluster
to perform specific tasks or jobs and is terminated once the
tasks are completed.
Helps you monitor and track database activities, including Amazon Redshift audit logging
queries executed, connections established, and schema
changes, providing visibility and enhancing security and
compliance.
It gathers container-related metrics at the cluster, pod, and Amazon CloudWatch Container Insights
container levels. Also it collect performance and
application logs.
It's a feature of Amazon OpenSearch Service that stores UltraWarm data nodes
large amounts of read-only data cost-effectively through
standard data nodes that use "hot" storage for faster
performance.
An Amazon Redshift feature that helps you query and Amazon Redshift Spectrum
analyze data directly from files stored in Amazon S3,
extending Redshift's querying capabilities to exabytes of
data without loading it into the data warehouse.
An AWS Service tool for visual data preparation that allows AWS Glue DataBrew
users to clean and normalize data without writing code.
Users can easily modify their data structure with schema Apache Iceberg table format
evolution, allowing them to add, rename, or remove
columns from a data table without affecting the underlying
data.
A performance optimization technique used in databases Pushdown Predicates
and data processing frameworks where filtering is applied
as close to the data source as possible.
An integrated development environment (IDE) for creating, AWS Glue Studio
running, and monitoring ETL (Extract, Transform, Load)
jobs in AWS Glue.
Helps organizations manage the complexity of data AWS Glue Schema Registry
schemas in modern data architectures, ensuring data
consistency, quality, and reliability across various data
sources and applications.
A technique used in data querying and storage systems to Partition Projection
dynamically determine partitioning information at query
time rather than storing it statically.
It a sets of SQL statements that are stored and executed Stored Procedures
on the Redshift cluster, primarily used for encapsulating
complex data processing logic.
, AWS Certified Data Engineer Associate DEA-C01 Test
An Amazon Redshift feature that allows masking sensitive Amazon Redshift Dynamic Data Masking (DDM)
data in real-time during query execution without altering
the original data.
Its a process of distributing incoming data streams to Fan-out for streaming data distribution
multiple downstream consumers or processing systems
simultaneously. It is often facilitated by services like
Amazon Kinesis Data Streams or Amazon Managed
Streaming for Apache Kafka (Amazon MSK).
An open-source tool is available for creating, scheduling, Apache Airflow
and monitoring workflows.
A cloud-based big data platform that simplifies and Amazon EMR
accelerates big data processing tasks at scale using open-
source tools like Apache Hadoop, Spark, Hive, and HBase.
It is part of AWS Glue ETL library and are particularly well- AWS Glue DynamicFrame
suited for semi-structured data or data with varying
schemas, common in data lakes and other big data
scenarios.
Designed to provide an improved performance and Amazon Redshift RA3 nodes
flexibility for Redshift clusters, particularly for workloads
with varying compute and storage requirements.
It enables EMR clusters to directly access data stored in EMR File System (EMRFS)
S3 as if it were a native file system, without requiring data
transfer or replication.
It simplifies the management and processing of large-scale S3 Batch Operations
data operations in Amazon S3, making it easier for users
to perform bulk actions on their S3 objects efficiently and
cost-effectively.
Gives you the ability to orchestrate a complex ETL pipeline AWS Step Functions
with multiple steps and dependencies.
Enables you to capture ongoing changes after a full-load Change Data Capture (CDC)
migration to a supported target datastore.
It is an AWS serverless scheduler that allows you to Amazon EventBridge Scheduler
create, run, and manage tasks from one central, managed
service.
An AWS Glue tool that scans and analyzes data stored in AWS Glue Crawler
various data sources, extracts schema information, and
creates metadata tables in the AWS Glue Data Catalog.
This storage format is highly efficient for analytical queries, Apache Parquet
as it allows for selective column reads and compression
techniques optimized for columnar data.
It enables secure sharing of AWS resources across AWS Resource Access Manager (RAM)
different AWS accounts or within an organization,
simplifying access control and reducing overhead.