Study guide · Core notes

Data Engineer Notes: move the data that fits

Reference notes for DEA-C01: how records arrive, which store can answer the query, how you notice a bad run, and who is allowed to read a column.

Exam guide: DEA-C01 v1.18 modulesReviewed October 2026

How to read these notes: Items tagged Extra sit outside the DEA-C01 version 1.1 in-scope list, including Amazon FinSpace and AWS SCT. Learn the unmarked items first. New to the exam? Start with the overview for the format and domain weights. Foundational service names are in the Cloud Practitioner notes.

Module 1: Ingestion

Domain 1 starts with how the bytes arrive. Volume is how much, velocity is how fast and how steady, and variety is structured, semi-structured, or unstructured. A file that lands once a night is a batch problem. A clickstream that must be read again tomorrow is a stream problem. Name that before you name the service.

Ingestion tools on DEA-C01. A delivery pipe is not a replay log.
ToolUse it whenIt does not
Kinesis Data StreamsProducers append records and one or more consumers read them, including a replay inside the retention window. Default retention is 24 hours and can be extended up to 365 days.Load S3 for you. That delivery step is Firehose, or your own consumer.
Kinesis Data FirehoseYou want a stream buffered and written to S3, Redshift, OpenSearch, or another supported destination, with an optional format conversion.Keep a consumer-readable log you can rewind. Once it delivers, that buffer is not your replay history.
Amazon MSKThe pipeline is already Kafka: topics, consumer groups, and offsets.Remove the need to understand partitions and consumer lag. It manages the cluster, not the topic design.
DynamoDB StreamsItem inserts, updates, and deletes on a table must trigger downstream work, often through a Lambda event source mapping.Carry an unrelated clickstream. The stream is the table’s change feed.
AWS DMSA database must be migrated or kept in sync with change data capture.Replace a Spark transform. DMS moves rows. Glue and EMR reshape them.
Amazon AppFlowA supported SaaS application must land in S3 or another AWS destination on a schedule or an event.Ingest a custom high-volume stream. That is Kinesis or MSK.
Amazon S3 event notificationsA new object should start a Lambda function, a queue, a topic, or an EventBridge rule.Infer a schema. A crawler or an explicit partition update does that.

Fan-in is many producers writing one stream. Fan-out is many consumers reading it. Standard Kinesis consumers share a shard’s read throughput. Enhanced fan-out gives a consumer its own read throughput, which matters when several applications must keep up. A Lambda function reads Kinesis through an event source mapping, not through a public URL you invent.

Throttling shows up as Kinesis shard limits, DynamoDB capacity, or a database that runs out of connections. The durable fix is more shards, a capacity mode that matches the spike, or a pool such as RDS Proxy, plus backoff. Immediate retries make the limit worse. An IP allow list for a database is a security group or a source allow list on that engine. A public endpoint with a wide open rule is the pattern to retire.

Replay means you can read the same records again from the stream’s retention or from a Kafka offset. If the requirement is “rebuild yesterday after a bad job,” land a copy in S3 or keep a replayable stream. Firehose alone does not give you that rewind. Stateless processing handles each record on its own and can scale out. Stateful processing, such as a windowed aggregate, needs a service that keeps state, such as Amazon Managed Service for Apache Flink, or an external store. A Lambda function that forgets everything between invocations is the stateless side of that line.

Transform and programming concepts

Change the shape of the data with the smallest engine that meets the deadline.

AWS Glue runs managed Spark or Python shell jobs, crawlers, and the Data Catalog. Use it when you want the job more than the cluster. Amazon EMR runs the Hadoop and Spark stack when you need to tune nodes, install libraries, or place task nodes on Spot. AWS Lambda fits a short function, including one record at a time from a stream. Its timeout is 15 minutes, so a two-hour Spark job does not belong there. AWS Batch and containers on Amazon ECS or Amazon EKS fit custom batch workers when Glue’s job model is the wrong shape.

Columnar formats such as Apache Parquet and Apache ORC store columns together, compress well, and let Athena or Spark skip columns. CSV is the usual landing format and a poor analytics format. Convert on the way in with Glue, with Firehose record conversion, or with an Athena CTAS statement that writes Parquet. Partition by a column you filter on, often a date, and avoid a huge number of tiny files. Compaction is part of the job, not a later surprise.

Glue job bookmarks record what a job already processed so the next run picks up new data. They are not a data catalog. JDBC and ODBC connections are how Glue, Athena federated queries, and Redshift reach a database. Cost levers on this domain are format and partition pruning, Glue’s flexible execution for jobs that can wait, EMR Spot for interruptible task nodes, and Lambda memory, because more memory also allocates more CPU.

The exam guide asks for programming concepts without language-specific syntax. Distributed computing means the work is split across partitions and a hot key can stall the whole stage. Infrastructure as code on this exam includes AWS CloudFormation and the AWS CDK. AWS SAM packages serverless pieces such as Lambda, Step Functions, and DynamoDB. CI/CD for a pipeline uses the in-scope developer tools, including CodeBuild, CodePipeline, and CodeDeploy. A Lambda function that needs more than its writable temporary space can mount Amazon EFS.

Version 1.1 adds large language models as a processing step. Amazon Bedrock can classify or extract fields inside a pipeline. The same guide says performing machine learning training and inference is out of scope for the target candidate, so the exam-shaped answer is a call to a managed model, not a training job.

Module 2: Orchestration

A pipeline is a sequence with retries, not a single job that hopes the next one starts. Amazon EventBridge runs schedules and routes events. AWS Step Functions is a serverless state machine with retries and catches, and it fits a workflow you want to own as states. Amazon MWAA is managed Apache Airflow for DAGs, backfills, and teams that already speak Airflow. AWS Glue workflows tie Glue triggers, crawlers, and jobs together when the whole graph is Glue.

Amazon SQS holds work for a consumer. Amazon SNS notifies many subscribers. A failed job should publish to SNS or raise a CloudWatch alarm. A queue in front of workers absorbs a burst. Step Functions does not replace that buffer. It orders the steps.

Build the pipeline so one bad task can retry without reprocessing the whole day. Bookmarks, checkpoints, and idempotent writes are what make a rerun safe. A schedule that overlaps itself needs a concurrency limit, or two copies of the job will fight over the same partition.

Module 3: Choose a data store

The access pattern picks the engine. The copy strategy picks whether the bytes move.

Store choices on DEA-C01. Querying in place and loading a warehouse answer different workloads.
NeedStoreWatch for
Occasional SQL on S3Amazon AthenaCost follows data scanned. Partitions, Parquet, and compression are the levers. Athena Spark notebooks explore with Spark.
Busy analytic SQLAmazon RedshiftCOPY loads from S3. UNLOAD writes S3. Spectrum queries S3 through the catalog. A federated query reads RDS or Aurora in place. Materialized views store a repeated aggregate. Distribution and sort keys, and locks, still matter.
Transactions for an applicationAmazon RDS or Amazon AuroraThis is the system of record, not the warehouse. Aurora PostgreSQL can hold vector indexes such as HNSW and IVF.
Key lookup at any scaleAmazon DynamoDBDesign the partition key for the access pattern. TTL expires items eventually. It is not a precise clock and it is not a backup.
Fast key-value with durabilityAmazon MemoryDBThe exam guide names MemoryDB for fast key-value access. Extra Amazon ElastiCache is not on the in-scope list.
Documents, wide columns, or a graphAmazon DocumentDB, Amazon Keyspaces, or Amazon NeptuneMatch the model the application already uses. Do not force a graph into a warehouse question.
Files, archive, and shared filesAmazon S3, S3 Glacier, Amazon S3 Tables, Amazon EFS, Amazon EBSS3 is the lake. Glacier classes are colder tiers. S3 Tables is the managed Iceberg table feature added in guide 1.1. EFS is a shared file system. EBS is a volume for one instance.

Redshift Serverless fits intermittent warehouse work. A provisioned cluster fits a steady, known concurrency. The same tradeoff shows up with Athena, which is serverless, versus a cluster you keep warm. Transfer into the lake with AWS DataSync when the network can finish on time, with the AWS Snow Family when the data set is too large for the link, and with AWS Transfer Family when partners already speak SFTP.

Open table formats are in skill 2.1.7, and Apache Iceberg is the example. Iceberg adds ACID behavior, schema evolution, and partition evolution on files in S3. S3 Tables is the in-scope managed place to keep those tables. Vector search uses an index type: HNSW is a graph of neighbors, and IVF groups vectors into clusters. The guide places those indexes with Aurora PostgreSQL and places vectorization concepts with Amazon Bedrock knowledge bases. Amazon Kendra is the in-scope enterprise search service. It is not a warehouse.

Module 4: Catalog, lifecycle, and schema

The AWS Glue Data Catalog is the technical metastore Athena, EMR, Glue, and Redshift Spectrum share. An Apache Hive metastore is the same idea running with a cluster. Point EMR at the Glue catalog when several engines must see one schema. A crawler samples objects and writes tables. New partitions appear only after a crawler, an MSCK REPAIR, an explicit add, or partition projection. A query that misses today’s folder is often a partition that was never registered.

Exam guide 1.1 separates that technical catalog from a business catalog. Amazon SageMaker Catalog is the example for the business catalog and for data lineage, alongside Amazon SageMaker ML Lineage Tracking. CloudTrail records API calls. It does not draw lineage between datasets.

Lifecycle is a policy, not a hope. An S3 Lifecycle rule transitions objects to a colder class or expires them, including noncurrent versions when versioning is on. Put hot prefixes in S3 Standard. Standard-IA and Glacier classes charge you back when something reads them constantly. DynamoDB TTL deletes expired items in the background. A legal deletion requirement is an expiration or a delete job with a date, plus a check that replicas and backups follow the same rule.

Protect what you must keep. S3 versioning keeps overwritten objects. Redshift snapshots and AWS Backup cover the warehouse and other supported stores. Cross-Region replication improves resilience and can violate data sovereignty if the destination Region is disallowed. Block that replication when the stem names a Region the data must not enter.

Schema evolution means new columns and changed types without rewriting history. Iceberg is built for that. A Parquet dataset can add columns if readers tolerate them. Changing a partition key is a new layout, not a metadata tweak. AWS DMS Schema Conversion converts a schema between engines. Extra Version 1.1 removed AWS SCT from the in-scope service list, so prefer DMS Schema Conversion when the question is schema conversion.

Indexing, partitioning, and compression are the optimization trio. A Redshift sort key matches the filter. A DynamoDB partition key matches the lookup. An Athena partition matches the WHERE clause. Snappy-compressed Parquet is the usual lake default because it splits and scans less data than gzip-compressed CSV.

Module 5: Operations and analysis

After the pipeline exists, the exam asks how you run it, query it, and notice that it stalled.

Automation reuses the orchestrators from Module 2 and adds the failure mode. An MWAA DAG that will not import is a code or dependency problem in the environment, not a Redshift sort key. A Step Functions execution that times out needs a longer state timeout or a smaller unit of work, plus a catch that records the input. EventBridge is how a schedule or a service event starts the next step. Lambda is how a small automated action runs. An SDK call from that function uses an IAM role on the function, not a long-lived key pasted into the environment.

Analysis tools stay close to the data. Amazon QuickSight is the visualization service. AWS Glue DataBrew profiles and cleans with visual rules. Amazon SageMaker Data Wrangler is named for verifying and cleaning data. Athena and Redshift are where SQL lives: filters, joins, GROUP BY, windowed rolling averages, pivots, and views. Creating a view does not copy the data. A materialized view in Redshift does.

Provisioned capacity is cheaper and more predictable when the workload is steady. Serverless capacity, including Athena, Redshift Serverless, and Lambda, removes idle cluster management and bills the work. A stem that says “spiky and unattended” leans serverless. A stem that says “the same concurrency every business hour” can justify a provisioned warehouse.

Module 6: Quality, monitoring, and logs

Data quality is a check in the pipeline, not a meeting after the dashboard looks wrong. Look for empty fields, failed row counts, unexpected nulls, and schema drift. DataBrew can hold those rules for datasets it prepares. AWS Glue Data Quality can enforce rules in a Glue job because Glue itself is in scope. A sample can be random or stratified. A sample that only reads the first file of a partitioned day will miss a bad partition.

Skew means one key or one partition is much larger than the others, so one worker runs long after the rest finish. Salt the hot key, repartition, or broadcast the small side of a join. Adding a worker without changing the key often leaves the hot partition on one task.

Glue and EMR failures are usually memory, skew, or bad input. Raise workers or memory when the error is an out-of-memory condition. Fix the key when one task never finishes. A bookmark that skipped a day needs a reset or a rerun of that interval, not a new cluster size.

Logs a data pipeline actually uses. API history and application output are different stores.
QuestionLog
Who called which API?AWS CloudTrail. Turn on data events when you need object-level S3 activity. Management events do not include every object read.
Query those API records with SQLAWS CloudTrail Lake.
What did the job print?Amazon CloudWatch Logs, then CloudWatch Logs Insights for search.
Large log files already in S3Amazon Athena, Amazon EMR, or Amazon OpenSearch Service, depending on whether you want SQL, a big processing job, or search.
Did a setting drift?AWS Config records configuration changes. It is not the application log.
Tell a personAmazon SNS, often from a CloudWatch alarm or an EventBridge rule. Amazon Managed Grafana is the in-scope dashboard service.

Module 7: Security and governance

Domain 4 is access, encryption, audit, and the promise that data stays where it is allowed to stay.

Least privilege is a custom IAM policy when a managed policy is wider than the task. Applications, Lambda, API Gateway, the CLI, and CloudFormation should assume roles. A security group is the stateful allow list on a data store’s network interface. Update it when a pipeline host must reach Redshift or RDS, and keep the rule as narrow as the client.

AWS Secrets Manager stores a database password and rotates it. AWS Systems Manager Parameter Store stores configuration and can store a secret. The exam guide’s rotation example is Secrets Manager. Inside Redshift, database users, groups, and roles still exist on top of IAM. Redshift data sharing grants another warehouse access to live data without a second COPY.

AWS Lake Formation is the permission layer for cataloged lake data used by Athena, EMR, Redshift, and S3. Use it for column-level and row-level grants. IAM still has to allow the person to call the service. Lake Formation does not replace the bucket as storage. Tag-based and attribute-based rules belong here when the stem says permissions should follow a tag or a trait instead of a named person. S3 Access Points and AWS PrivateLink are the private paths into data, not a second copy of the files.

Encrypt at rest with AWS KMS when you must control the key. A cross-account read needs both a key policy that allows the other account and an IAM policy on that account. Encrypt in transit with TLS. Encrypt before transit when the requirement is that AWS stores ciphertext the client produced. Masking and anonymization hide a column in a view, a Lake Formation filter, or a transform. Deleting the column is a different requirement. Amazon Macie finds sensitive data already stored in S3. Pair that finding with a Lake Formation grant or a remediation job. Macie does not block an HTTP request. AWS WAF does, and AWS Shield is for DDoS.

Privacy and governance are about sharing and place. Grant data sharing on purpose. Identify PII with Macie. Stop backups and replication into a disallowed Region. AWS Config shows the configuration change that turned a risky setting on. Data sovereignty means the data stays in the Regions and accounts the rule allows. Exam guide 1.1 also names SageMaker Unified Studio: a domain is the boundary, domain units group ownership, and projects are where a team uses cataloged data under those permissions. Amazon SageMaker Catalog projects are how that access is managed in the guide’s wording.

AWS Budgets and AWS Cost Explorer are in scope when the stem is the bill for a pipeline. They do not fix a hot partition. They show the spend that a bad scan or an idle cluster created.

Frequently asked questions

When do you choose Kinesis Data Streams instead of Kinesis Data Firehose?

Choose Data Streams when consumers must read the stream and replay records inside the retention window. Choose Firehose when the job is to buffer a stream and deliver it to a destination such as Amazon S3.

When do you query with Athena instead of loading Amazon Redshift?

Use Athena for occasional SQL on data that already lives in Amazon S3. Load Redshift with COPY when many concurrent analytic queries need a warehouse. Redshift Spectrum queries S3 without that load. A federated query reads a remote database.

What does AWS Lake Formation grant?

Lake Formation grants permissions on cataloged data, including column and row filters, for services such as Amazon Athena, Amazon EMR, Amazon Redshift, and Amazon S3. The files still live in the data store. IAM still controls the API calls.

Service names and prices change. Use the current AWS exam guide and the AWS certification site as the authority when they differ from these notes.

Continue the AWS Data Engineer study path

Use the notes, then practice recall and timed questions.

AWS Data Engineer hub →
Available

Overview

Exam format, the 720 passing score, domain weights, and a study plan.

Available · You are here

Core Notes

Reference notes for ingestion, stores, operations, and governance.

Available

Practice Exams

200 original questions with custom exams, explanations, and a score report by domain.

Related Tools

Useful companions while you study.

All study topics →