Study guide · Flashcards & memory notes

AWS Data Engineer Flashcards & Memory Notes

40 recall cards for the distinctions DEA-C01 keeps testing: how the data arrives, which store can answer, how you notice a bad run, and who may read a column.

Exam guide: DEA-C014 decks · 40 cardsNo sign-up required

How to use these cards: Say the answer out loud before you flip the card, and mark it honestly. Revisit the cards you missed tomorrow rather than rereading them right away. Progress is kept only while this page is open. For the full explanations behind each card, see the core notes.

Study mode

Choose a deck, flip each card, and mark what you already know.

All flashcards

Select a question to reveal its answer.

Deck 1 · Ingestion and transformation

Streams, batch jobs, formats, and the workflow that starts them.

10 cards
1.1What is the difference between Kinesis Data Streams and Kinesis Data Firehose?

Data Streams keeps records for consumers to read and replay. Firehose buffers a stream and delivers it to a destination such as S3.

1.2What does replay mean for an ingestion pipeline?

You can read the same records again from a stream retention window or a Kafka offset. A copy in S3 is the replay path when the stream itself is not retained.

1.3What is Kinesis enhanced fan-out for?

A consumer that needs its own read throughput on a shard when several consumers must keep up. Shared consumers divide the shard's standard read throughput.

1.4When do you choose AWS Glue instead of Amazon EMR?

When you want a managed Spark or Python job more than a cluster to tune. EMR fits a cluster you size, including Spot task nodes.

1.5Why convert CSV to Parquet before Athena queries it?

Parquet is columnar and compresses well, so Athena scans less data. CSV stores whole rows and makes every query read more bytes.

1.6What starts work when a new object lands in Amazon S3?

An S3 event notification to Lambda, SQS, SNS, or EventBridge. A crawler registers schema and partitions. It is not the arrival signal itself.

1.7When do you choose Amazon AppFlow?

When a supported SaaS application must send records to AWS on a schedule or an event. A custom high-volume stream is Kinesis or Amazon MSK.

1.8How does Lambda read from Kinesis Data Streams?

Through an event source mapping. The function does not poll a public URL you attach to the stream.

1.9What is the difference between stateless and stateful stream processing?

Stateless handling treats each record on its own and can scale out. Stateful handling, such as a windowed aggregate, needs a service that keeps state, such as Managed Service for Apache Flink.

1.10What do volume, velocity, and variety describe?

How much data arrives, how fast it arrives, and whether it is structured, semi-structured, or unstructured.

Deck 2 · Stores, catalog, and lifecycle

Where the data lives, how the schema is registered, and when it expires.

10 cards
2.1When do you choose Amazon Athena instead of loading Amazon Redshift?

Athena runs occasional SQL on data that already lives in S3. Load Redshift when many concurrent analytic queries need a warehouse.

2.2What is the difference between Redshift COPY, Spectrum, and a federated query?

COPY loads data into the warehouse. Spectrum queries S3 through the catalog. A federated query reads a remote database such as RDS without that load.

2.3What does DynamoDB TTL do?

It expires items after a timestamp you store on the item. Deletion is eventual, so TTL is not a precise clock and it is not a backup.

2.4What does Apache Iceberg add on a data lake?

ACID behavior, schema evolution, and partition evolution on files in S3. Amazon S3 Tables is the managed Iceberg table feature on the DEA-C01 in-scope list.

2.5What is the difference between the Glue Data Catalog and SageMaker Catalog?

The Glue Data Catalog is the technical metastore engines share. Exam guide 1.1 names SageMaker Catalog as the business catalog.

2.6Why can an Athena query miss today's partition?

The folder exists, and the catalog does not know about it yet. A crawler, MSCK REPAIR, an explicit add, or partition projection registers it.

2.7What is the difference between HNSW and IVF vector indexes?

HNSW is a graph of neighboring vectors. IVF groups vectors into clusters. The exam guide places both with Aurora PostgreSQL.

2.8When do you choose Amazon MemoryDB?

When the workload needs fast key-value access and the store should be durable. The DEA-C01 in-scope list names MemoryDB for that pattern.

2.9What does AWS Lake Formation control?

Permissions on cataloged data, including column and row filters. The files still live in the data store. IAM still controls the API calls.

2.10What does an S3 Lifecycle rule do?

It transitions objects to a colder storage class or expires them, including noncurrent versions when versioning is on. A hot prefix belongs in S3 Standard, because infrequent-access classes charge for constant reads.

Deck 3 · Operations and quality

Orchestration failures, SQL, logs, and checks that catch a bad run.

10 cards
3.1When do you choose AWS Step Functions instead of Amazon MWAA?

Step Functions is a serverless state machine with retries when you want to own the workflow as states. MWAA is managed Apache Airflow for DAGs, backfills, and teams that already run Airflow.

3.2What is data skew, and what do you change?

One key or partition is much larger than the others, so one worker runs long after the rest finish. Salt the hot key or repartition. Adding a worker without changing the key often leaves the hot partition on one task.

3.3What is the difference between CloudTrail and CloudWatch Logs?

CloudTrail records AWS API calls. CloudWatch Logs stores what the job printed. CloudTrail Lake queries API records. Logs Insights searches application logs.

3.4What is Amazon QuickSight for, compared with Amazon Athena?

QuickSight visualizes data for people. Athena runs the SQL. A view in Athena does not copy the table.

3.5What does an AWS Glue job bookmark remember?

Which data the job already processed, so the next run picks up new data. It is not a data catalog, and a skipped day needs a rerun of that interval.

3.6When is serverless capacity a better fit than a provisioned cluster?

When the workload is spiky or unattended and you would otherwise pay for idle time. A steady concurrency every business hour can justify a provisioned warehouse.

3.7Where do you put a check for empty fields?

In the pipeline, with a DataBrew rule or a Glue data quality check, before the dashboard is the only place someone notices. A sample that reads only the first file of a day can miss a bad partition.

3.8How do partitions change an Athena bill?

A filter on the partition column skips files, so Athena scans less data. Selecting every column from CSV keeps the scan large even when a partition exists.

3.9Which service notifies people that a pipeline failed?

Amazon SNS, often from a CloudWatch alarm or an EventBridge rule. A queue holds work for a consumer. It does not fan the alert out by itself.

3.10What is the difference between a random sample and a stratified sample?

A random sample gives each row a chance of being chosen. A stratified sample keeps a share from each group so a rare partition is still represented.

Deck 4 · Security and governance

Roles, column grants, encryption, audit, and where data is allowed to go.

10 cards
4.1Where should a pipeline rotate a database password?

AWS Secrets Manager, which stores the secret and rotates it. Parameter Store can hold a value. The exam guide points rotation at Secrets Manager.

4.2How do you hide one column from an analyst on a data lake?

Grant a column filter in AWS Lake Formation. A bucket policy that blocks the whole prefix is broader than the column the stem named.

4.3Which service finds sensitive data already stored in Amazon S3?

Amazon Macie. AWS WAF inspects HTTP requests. It does not scan objects at rest.

4.4What has to allow a cross-account decrypt with AWS KMS?

The key policy must allow the other account, and that account's IAM policy must allow the decrypt. An IAM allow in only one account leaves the other side closed.

4.5What does least privilege mean on a data pipeline role?

The role can run its job and nothing more. Write a custom policy when a managed policy is wider than that task.

4.6When do you encrypt in transit, and when do you encrypt before transit?

TLS encrypts the connection. Encrypt before transit when the requirement is that AWS stores ciphertext the client produced.

4.7How do you keep a copy out of a disallowed Region?

Block the replication or backup destination that names that Region. Cross-Region replication improves resilience and breaks data sovereignty when the destination is forbidden.

4.8Which CloudTrail events show object-level Amazon S3 reads?

Data events. Management events record control-plane calls such as creating a bucket. They do not include every object read.

4.9What does Amazon Redshift data sharing do?

It lets another Redshift warehouse query live data without a second COPY. UNLOAD writes files to S3. Sharing is not a file export.

4.10What is the difference between masking a column and deleting it?

Masking hides the value in a view, a Lake Formation filter, or a transform, and keeps the source. Deletion removes the data and is the requirement when the rule says the value must not be retained.

Memory notes

Groupings that make the highest-yield DEA-C01 distinctions easier to recall.

Replay versus deliver

  • Data Streams or MSK read again
  • Firehose buffer and land
  • S3 copy the replay when the stream is gone

Query in place or load

  • Athena occasional SQL on S3
  • Spectrum Redshift SQL on S3
  • COPY data inside the warehouse
  • Federated query rows left in RDS

Catalog versus permission

Glue Data Catalog names the schema. Lake Formation decides who can read a column. SageMaker Catalog is the business catalog in exam guide 1.1.

Hold, notify, order

  • SQS one consumer later
  • SNS tell many people
  • Step Functions order the steps
  • MWAA Airflow DAGs

Two logs

CloudTrail is who called the API. CloudWatch Logs is what the job printed. Data events are the object-level S3 trail.

Hide versus remove

Mask in a view or a Lake Formation filter when the source must remain. Expire with a lifecycle rule or TTL when the value must go away. Block replication into a Region the data must not enter.

Continue the AWS Data Engineer study path

Test what you have learned with a timed practice exam.

AWS Data Engineer hub →
Available

Overview

Exam format, the 720 passing score, domain weights, and a study plan.

Available

Core Notes

Reference notes for ingestion, stores, operations, and governance.

Available · You are here

Flashcards & Memory Notes

40 recall cards in four decks, plus memory notes for streams, stores, and access.

Available

Practice Exams

200 original questions with custom exams, explanations, and a score report by domain.

Related Tools

Useful companions while you study.

All study topics →