Replay versus deliver
- Data Streams or MSK read again
- Firehose buffer and land
- S3 copy the replay when the stream is gone
Study guide · Flashcards & memory notes
40 recall cards for the distinctions DEA-C01 keeps testing: how the data arrives, which store can answer, how you notice a bad run, and who may read a column.
How to use these cards: Say the answer out loud before you flip the card, and mark it honestly. Revisit the cards you missed tomorrow rather than rereading them right away. Progress is kept only while this page is open. For the full explanations behind each card, see the core notes.
Choose a deck, flip each card, and mark what you already know.
You marked every card in this deck as known. Shuffle and run through it again, or choose another deck.
Select a question to reveal its answer.
Streams, batch jobs, formats, and the workflow that starts them.
Data Streams keeps records for consumers to read and replay. Firehose buffers a stream and delivers it to a destination such as S3.
You can read the same records again from a stream retention window or a Kafka offset. A copy in S3 is the replay path when the stream itself is not retained.
A consumer that needs its own read throughput on a shard when several consumers must keep up. Shared consumers divide the shard's standard read throughput.
When you want a managed Spark or Python job more than a cluster to tune. EMR fits a cluster you size, including Spot task nodes.
Parquet is columnar and compresses well, so Athena scans less data. CSV stores whole rows and makes every query read more bytes.
An S3 event notification to Lambda, SQS, SNS, or EventBridge. A crawler registers schema and partitions. It is not the arrival signal itself.
When a supported SaaS application must send records to AWS on a schedule or an event. A custom high-volume stream is Kinesis or Amazon MSK.
Through an event source mapping. The function does not poll a public URL you attach to the stream.
Stateless handling treats each record on its own and can scale out. Stateful handling, such as a windowed aggregate, needs a service that keeps state, such as Managed Service for Apache Flink.
How much data arrives, how fast it arrives, and whether it is structured, semi-structured, or unstructured.
Where the data lives, how the schema is registered, and when it expires.
Athena runs occasional SQL on data that already lives in S3. Load Redshift when many concurrent analytic queries need a warehouse.
COPY loads data into the warehouse. Spectrum queries S3 through the catalog. A federated query reads a remote database such as RDS without that load.
It expires items after a timestamp you store on the item. Deletion is eventual, so TTL is not a precise clock and it is not a backup.
ACID behavior, schema evolution, and partition evolution on files in S3. Amazon S3 Tables is the managed Iceberg table feature on the DEA-C01 in-scope list.
The Glue Data Catalog is the technical metastore engines share. Exam guide 1.1 names SageMaker Catalog as the business catalog.
The folder exists, and the catalog does not know about it yet. A crawler, MSCK REPAIR, an explicit add, or partition projection registers it.
HNSW is a graph of neighboring vectors. IVF groups vectors into clusters. The exam guide places both with Aurora PostgreSQL.
When the workload needs fast key-value access and the store should be durable. The DEA-C01 in-scope list names MemoryDB for that pattern.
Permissions on cataloged data, including column and row filters. The files still live in the data store. IAM still controls the API calls.
It transitions objects to a colder storage class or expires them, including noncurrent versions when versioning is on. A hot prefix belongs in S3 Standard, because infrequent-access classes charge for constant reads.
Orchestration failures, SQL, logs, and checks that catch a bad run.
Step Functions is a serverless state machine with retries when you want to own the workflow as states. MWAA is managed Apache Airflow for DAGs, backfills, and teams that already run Airflow.
One key or partition is much larger than the others, so one worker runs long after the rest finish. Salt the hot key or repartition. Adding a worker without changing the key often leaves the hot partition on one task.
CloudTrail records AWS API calls. CloudWatch Logs stores what the job printed. CloudTrail Lake queries API records. Logs Insights searches application logs.
QuickSight visualizes data for people. Athena runs the SQL. A view in Athena does not copy the table.
Which data the job already processed, so the next run picks up new data. It is not a data catalog, and a skipped day needs a rerun of that interval.
When the workload is spiky or unattended and you would otherwise pay for idle time. A steady concurrency every business hour can justify a provisioned warehouse.
In the pipeline, with a DataBrew rule or a Glue data quality check, before the dashboard is the only place someone notices. A sample that reads only the first file of a day can miss a bad partition.
A filter on the partition column skips files, so Athena scans less data. Selecting every column from CSV keeps the scan large even when a partition exists.
Amazon SNS, often from a CloudWatch alarm or an EventBridge rule. A queue holds work for a consumer. It does not fan the alert out by itself.
A random sample gives each row a chance of being chosen. A stratified sample keeps a share from each group so a rare partition is still represented.
Roles, column grants, encryption, audit, and where data is allowed to go.
AWS Secrets Manager, which stores the secret and rotates it. Parameter Store can hold a value. The exam guide points rotation at Secrets Manager.
Grant a column filter in AWS Lake Formation. A bucket policy that blocks the whole prefix is broader than the column the stem named.
Amazon Macie. AWS WAF inspects HTTP requests. It does not scan objects at rest.
The key policy must allow the other account, and that account's IAM policy must allow the decrypt. An IAM allow in only one account leaves the other side closed.
The role can run its job and nothing more. Write a custom policy when a managed policy is wider than that task.
TLS encrypts the connection. Encrypt before transit when the requirement is that AWS stores ciphertext the client produced.
Block the replication or backup destination that names that Region. Cross-Region replication improves resilience and breaks data sovereignty when the destination is forbidden.
Data events. Management events record control-plane calls such as creating a bucket. They do not include every object read.
It lets another Redshift warehouse query live data without a second COPY. UNLOAD writes files to S3. Sharing is not a file export.
Masking hides the value in a view, a Lake Formation filter, or a transform, and keeps the source. Deletion removes the data and is the requirement when the rule says the value must not be retained.
Groupings that make the highest-yield DEA-C01 distinctions easier to recall.
Glue Data Catalog names the schema. Lake Formation decides who can read a column. SageMaker Catalog is the business catalog in exam guide 1.1.
CloudTrail is who called the API. CloudWatch Logs is what the job printed. Data events are the object-level S3 trail.
Mask in a view or a Lake Formation filter when the source must remain. Expire with a lifecycle rule or TTL when the value must go away. Block replication into a Region the data must not enter.
Test what you have learned with a timed practice exam.
Exam format, the 720 passing score, domain weights, and a study plan.
Reference notes for ingestion, stores, operations, and governance.
40 recall cards in four decks, plus memory notes for streams, stores, and access.
200 original questions with custom exams, explanations, and a score report by domain.
Useful companions while you study.
Convert structured data between JSON and YAML with formatting and validation.
Convert CSV spreadsheet data to JSON or turn JSON arrays into downloadable CSV files.
Encode text and files to Base64 or decode Base64 data with UTF-8 support.
Compare two text blocks and highlight added, removed, and changed content.
NodnWebTools provides general informational, educational, and convenience resources. Calculations, conversions, estimates, and learning materials may contain errors or become outdated. Financial, tax, medical, legal, and travel information is not professional advice. Verify important results and current requirements with qualified professionals or authoritative sources. Protect sensitive files and personal information, review each tool’s privacy limitations, and use only content you are authorized to process. You are responsible for how you use and share results. Study resources are independent and do not guarantee exam success or imply certification-provider endorsement. Amazon Web Services, AWS, and related marks are trademarks of Amazon.com, Inc. or its affiliates. NodnWebTools is not affiliated with, endorsed by, or sponsored by Amazon.