AWS Solutions Architect Associate Diaries: SAA-C03 Study Notes

37 minute read

I started pulling these notes together during SAA-C03 prep because the official documentation kept sending me in circles. What I actually needed was a way to pattern-match quickly — given a scenario, which service, and why. These are those notes. I am adding modules as I work through them, ten total.


Module 1: Architecting Fundamentals

The exam keeps coming back to one question: given a scenario, what is the right design decision and why. Module 1 is about the mental model behind those decisions — the Well-Architected Framework and the principles AWS expects you to apply.

Well-Architected Framework

Six pillars. The exam does not ask you to recite them, but it tests whether you can apply them.

Pillar Core Question
Operational Excellence Can we operate and improve the system easily?
Security Are data, systems, and users protected?
Reliability Does the system keep working when things fail?
Performance Efficiency Are we using resources efficiently?
Cost Optimization Are we wasting money?
Sustainability Are we minimising environmental impact?

Sustainability shows up occasionally. The others come up constantly.

Core Design Principles

These four come up in almost every scenario question.

AWS assumes failures will happen. A single EC2 with a single database is a bad design. The expected answer is multiple AZs, Auto Scaling, and a load balancer.

When two answers both work, pick the one with less operational effort. RDS over EC2 with MySQL. Lambda over EC2 for event processing. AWS wants you to use managed services.

If one application directly calls another, a failure in one cascades to the other. SQS between them breaks that dependency.

Adding more servers scales better than making one server bigger. Horizontal scaling requires a load balancer and is what the exam expects when availability is mentioned.

Key Concepts

Scalability is the ability to grow. Elasticity is automatic growth and shrink. The exam uses both words but means different things by them.

High availability minimises downtime. Fault tolerance means the system keeps running through a failure. Multi-AZ gives you both.

The shared responsibility model trips people up. AWS is responsible for the physical infrastructure. You are responsible for everything you put on top of it.

AWS  = Security OF the cloud (hardware, data centres, networking)
You  = Security IN the cloud (IAM, encryption, OS patches, security groups)

Common Exam Architectures

Highly available web application:

Users
  |
 ALB
  |
Auto Scaling Group
  |
EC2 (multiple AZs)

Decoupled architecture:

App --> SQS --> Workers

Event-driven serverless:

S3 Upload --> Lambda --> Process

Exam Keyword Map

Keyword Answer
Highly available Multi-AZ
Scalable / Elastic Auto Scaling
Decouple / Buffer SQS
Managed service Let AWS manage it
Fault tolerant Multiple AZs
Single point of failure Eliminate it
Monitoring CloudWatch
Audit CloudTrail

Cheat Sheet

Pillars
  Operational Excellence, Security, Reliability,
  Performance Efficiency, Cost Optimization, Sustainability

Principles
  Design for Failure
  Use Managed Services
  Decouple Components
  Scale Horizontally

When stuck on a scenario question, ask:
  What is the bottleneck?
  What is the single point of failure?
  Can AWS manage this instead?

Module 2: Account Security

Every security question comes down to three things: who is making the request, what are they allowed to do, and how do we enforce that at scale. IAM is the answer to all three.

Identities

A principal is anything that can make a request to AWS — a person, an application, or an AWS service.

Identity Represents Use when
IAM User One person Permanent credentials for a human
IAM Group Collection of users Multiple people need the same permissions
IAM Role Temporary access A service or app needs to call another AWS service

The one that trips people up is the role. A role has no username and no password. It issues temporary credentials. When EC2 needs to read from S3, you attach a role to the EC2 — you do not put access keys on the server.

Human        -> User
Many humans  -> Group
Service/App  -> Role

Policies

Policies are JSON documents that say what is allowed or denied. You attach them to users, groups, or roles.

{
  "Effect": "Allow",
  "Action": "s3:GetObject",
  "Resource": "*"
}

Two types come up on the exam:

Identity-based policies attach to a user, group, or role. Resource-based policies attach directly to the resource — an S3 bucket policy is the most common example. If you need to grant another AWS account access to your bucket, you use a bucket policy.

One rule that never changes: an explicit Deny always wins over an Allow, regardless of what other policies say.

Least Privilege

Grant only what is required. Nothing more. This is the answer whenever the question asks about security best practice for permissions. AdministratorAccess for everyone is the wrong answer even if it works.

MFA and Root

The root account is created automatically and has unlimited access. Best practice is to enable MFA on it and then not use it for day-to-day work.

MFA adds a second factor on top of a password. Any question about improving account security without changing permissions points to MFA.

Access keys are for programmatic access — CLI, SDK, applications. They are not for console login.

Managing Multiple Accounts

Large organisations rarely run everything in one AWS account. Production, development, and security typically live in separate accounts.

AWS Organizations manages all of them from a single place.

Organization
     |
    OU (Production)
   /    \
Acct A  Acct B

    OU (Development)
        |
      Acct C

Organizational Units (OUs) are logical containers. You apply policies to an OU and every account inside it inherits them.

Service Control Policies (SCPs) set the maximum permissions any account in that OU can have. An SCP does not grant permissions — it only limits them. Even an account admin cannot exceed what the SCP allows.

IAM Policy  -> grants permissions
SCP         -> limits permissions

IAM Identity Center (formerly Single Sign-On) gives users one login that works across all accounts and applications. If the question mentions one login for many AWS accounts, this is the answer.

Cross-account access works through roles. Account A assumes a role in Account B. No credential sharing, no permanent access.

Cheat Sheet

Human needs access?          -> IAM User
Many humans, same access?    -> IAM Group
Service needs AWS access?    -> IAM Role
Secure root account?         -> Enable MFA
Programmatic access?         -> Access Keys
EC2 accessing S3?            -> IAM Role, not access keys

Grant access to S3 bucket?   -> Bucket Policy
Restrict entire account?     -> SCP
Manage many accounts?        -> AWS Organizations
One login across accounts?   -> IAM Identity Center
Cross-account access?        -> Assume Role

Explicit Deny always wins.
SCP limits, IAM grants.

Module 3: Networking 1

Networking questions are about understanding how traffic moves — into AWS, through it, and between services. Most of the confusion comes from mixing up which component does what. Once the mental model is clear, the questions get straightforward.

IP Addressing

Every resource in AWS has an IP address. Public IPs are reachable from the internet. Private IPs only work inside your VPC.

Public IP   54.22.18.7    anyone on the internet can reach it
Private IP  10.0.1.10     only reachable inside your network

CIDR notation defines how large a network is. The number after the slash tells you how many addresses are in the block. Smaller number means more addresses.

10.0.0.0/16   65,536 addresses   large
10.0.1.0/24      256 addresses   small

VPC

A VPC (Virtual Private Cloud) is your isolated private network inside AWS. Nothing else runs in it unless you put it there. Every AWS account gets a default VPC that works immediately. Most production environments use a custom VPC for more control.

A VPC lives in one region. It cannot span regions. It can span multiple Availability Zones, which is how you get high availability.

Inside a VPC you create subnets. A subnet lives in one AZ.

VPC (us-east-1)
  |
  +-- Public Subnet  (AZ-A)
  +-- Private Subnet (AZ-A)
  +-- Public Subnet  (AZ-B)
  +-- Private Subnet (AZ-B)

Public subnets have a route to the internet. They hold load balancers, bastion hosts, and NAT Gateways. Private subnets have no direct internet access. Databases and application servers belong here.

Internet Gateway and Route Tables

An Internet Gateway is what connects a VPC to the internet. Without one, nothing in the VPC can reach the internet and nothing from the internet can reach the VPC.

Route tables tell traffic where to go. A public subnet has a route table entry that sends outbound traffic to the Internet Gateway.

Destination   Target
0.0.0.0/0     Internet Gateway

That single entry is what makes a subnet public.

Elastic IP and NAT Gateway

A regular public IP on an EC2 instance changes when the instance stops and starts. An Elastic IP is a static public IP that stays the same. Useful when external systems need to connect to a fixed address.

NAT Gateway solves a specific problem: a private EC2 needs to reach the internet (to download updates, call an API) but you do not want the internet to be able to initiate connections back to it.

Private EC2
     |
Private Subnet
     |
NAT Gateway  (lives in public subnet)
     |
Internet Gateway
     |
Internet

The private EC2 can initiate outbound connections. The internet cannot initiate inbound ones. NAT Gateway needs to live in a public subnet and needs an Elastic IP.

Security Groups and Network ACLs

These are the two layers of traffic control. They are often confused because they both filter traffic, but they operate at different levels and behave differently.

Security Groups attach to individual EC2 instances. They are stateful — if you allow inbound traffic, the response is automatically allowed out without a separate rule. They only support allow rules. Everything not explicitly allowed is blocked.

Network ACLs (NACLs) attach to subnets. They are stateless — you need explicit rules for both inbound and outbound traffic. They support both allow and deny rules.

Feature Security Group Network ACL
Applies to EC2 instance Subnet
Stateful Yes No
Deny rules No Yes
Default behaviour Deny all inbound Allow all (default NACL)

The practical difference: if you need to block a specific IP address from reaching your subnet, you need a NACL because Security Groups cannot deny. If you just need to control what ports an EC2 accepts, a Security Group is enough.

Full Architecture

How all the pieces fit together in a standard two-tier setup:

Internet
    |
Internet Gateway
    |
Public Subnet
  Load Balancer
  NAT Gateway
    |
Private Subnet
  EC2 App Servers
  RDS Database

Security Groups protect each instance. NACLs protect each subnet. Route tables direct traffic between them.

Cheat Sheet

Private network in AWS?          -> VPC
Section of a VPC in one AZ?      -> Subnet
Connect VPC to internet?         -> Internet Gateway
Private EC2 outbound internet?   -> NAT Gateway
Static public IP?                -> Elastic IP
Instance-level firewall?         -> Security Group
Subnet-level firewall?           -> Network ACL
Traffic routing rules?           -> Route Table

Security Group  = stateful, allow only, per instance
Network ACL     = stateless, allow + deny, per subnet
NAT Gateway     = outbound only, lives in public subnet
VPC             = one region, can span multiple AZs

Module 4: Compute

Compute questions ask one thing: who runs your code, and under what conditions. The answer depends on how much control you need, how long the code runs, and how predictable the traffic is.

The Control Spectrum

Most control    EC2          you manage the OS
                Containers   ECS / EKS / Fargate
Least control   Lambda       AWS manages everything

More control means more maintenance. Less control means more automation. The exam usually rewards choosing the least control that still meets the requirements.

EC2

A virtual server in AWS. You choose the CPU, RAM, OS, and disk. AWS manages the physical hardware. You manage everything above it — the OS, patches, applications, and firewall rules.

When you launch an EC2 instance you pick two things:

An AMI (Amazon Machine Image) is the blueprint — the OS and any pre-installed software. An instance type defines the hardware size.

Instance type families tell you what the hardware is optimised for:

Family Optimised for Use when
T Burstable, cheap Dev, testing, low-traffic sites
M General purpose Most standard workloads
C CPU Gaming, encoding, scientific compute
R Memory Databases, in-memory caching
P / G GPU Machine learning, AI, video rendering

EC2 Storage

Two options. They behave very differently.

EBS (Elastic Block Store) is a persistent volume attached to the instance. Data survives reboots and stop/start cycles. You can take snapshots of it, which are stored in S3 and used for backup or cloning volumes.

Instance Store is temporary storage physically attached to the host. It is fast, but if the instance stops or terminates, the data is gone. Use it for caches and scratch space, never for anything you need to keep.

EBS            -> persists after stop/start
Instance Store -> gone when instance stops

EC2 Pricing

This is tested heavily. The question usually gives you a workload description and asks which pricing model fits.

Model Commitment Cost Use when
On-Demand None Highest Unpredictable or short-term workloads
Reserved 1 or 3 years Up to 72% off Always-on, predictable workloads
Savings Plans 1 or 3 years Similar to Reserved More flexibility across compute services
Spot None Up to 90% off Interruptible jobs — AWS can reclaim anytime
Dedicated Host On-Demand or Reserved Highest Licensing or compliance requirements

Spot is the cheapest but AWS can terminate the instance with two minutes notice. Never use it for anything that cannot tolerate interruption.

Savings Plans are generally preferred over Reserved Instances now because they apply more broadly across compute services.

Lambda

You upload code. AWS runs it when an event triggers it, then stops. No server to provision, patch, or manage.

Event (S3 upload, API call, schedule)
    |
  Lambda runs your code
    |
  Stops

Key constraints: maximum execution time is 15 minutes. No OS access. Scales automatically.

Good for event-driven processing, APIs, automation, and scheduled jobs. Not suitable for long-running processes or anything that needs OS-level control.

  EC2 Lambda
Server management You AWS
Runs Continuously Only when invoked
Billing While running Per invocation and duration
OS access Yes No
Max runtime Unlimited 15 minutes

Cheat Sheet

Virtual server, full OS control?     -> EC2
Serverless, event-driven?            -> Lambda
Persistent EC2 storage?              -> EBS
Temporary fast storage?              -> Instance Store
Backup an EBS volume?                -> Snapshot
Distribute traffic across EC2s?      -> ELB
Auto add/remove EC2s?                -> Auto Scaling

Pricing
  Unpredictable workload?            -> On-Demand
  Always-on, long-term?              -> Savings Plans
  Interruptible batch jobs?          -> Spot
  Compliance / licensing?            -> Dedicated Host

Lambda max runtime = 15 minutes
Instance Store = lost on stop/terminate
EBS = persists, snapshotable

Module 5: Storage

Storage questions come down to matching the data access pattern to the right service. Get that mapping right and most storage questions answer themselves.

Service Overview

Service Think of it as Best for
S3 Infinite file cabinet Files, backups, images, logs
EBS Hard drive attached to a server EC2 storage
EFS Shared network drive Multiple EC2s accessing same files
FSx Managed enterprise file systems Windows/HPC workloads
Glacier Deep archive vault Long-term backups

Amazon S3

Objects live in buckets. S3 is not a file system — photos/2025/image.jpg is just the object name, not a real folder path. There are no actual directories.

11 nines durability, regional, scales automatically, no capacity planning needed.

Storage classes — the exam tests whether you can match access frequency to the right class:

Class Use when Cost
Standard Frequently accessed Highest
Standard-IA Occasional access, immediate retrieval Lower storage, retrieval fee
One Zone-IA Re-creatable data, one AZ is acceptable Cheaper, riskier
Glacier Instant Rare access, still immediate Low
Glacier Flexible Archive, minutes-to-hours retrieval Very low
Glacier Deep Archive Almost never accessed Cheapest

Lifecycle policies move objects between classes automatically. No manual work:

Day 0   -> Standard
Day 30  -> Standard-IA
Day 90  -> Glacier
Day 365 -> Delete

Versioning keeps old copies when you overwrite or delete. If the question mentions recovering a deleted file, versioning is the answer.

Cross-Region Replication copies objects to another region automatically. Requires versioning on both buckets. Used for disaster recovery and reducing latency for global users.

Encryption options:

Type Who manages keys When to use
SSE-S3 AWS Simplest option
SSE-KMS AWS KMS Need auditing and key control
SSE-C Customer Customer-managed keys

SSE-KMS comes up most when the question mentions auditing or key control alongside encryption.

S3 can also host static websites — HTML, CSS, JavaScript, no servers needed. If the question says cheap static website, S3 is the answer.

Amazon EBS

A block storage volume attached to one EC2 instance. Data persists across reboots.

EC2
 |
EBS

Snapshots back up to S3. Use them for backup, restore, and copying volumes across regions.

Amazon EFS

A shared file system that multiple Linux EC2 instances can mount at the same time.

EC2-A \
EC2-B -- EFS
EC2-C /

EBS vs EFS:

Feature EBS EFS
Storage type Block File
Shared across EC2s No Yes
OS Any Linux only

FSx

FSx for Windows supports SMB and Active Directory. If the question mentions a Windows file share, this is the answer.

FSx for Lustre is for high-performance computing workloads — machine learning, analytics, scientific computing.

Storage Gateway

Connects on-premises storage to AWS. Used for hybrid storage and migrations.

Mode Protocol AWS backend Use when
File Gateway NFS / SMB S3 File shares stored in S3
Volume Gateway (Cached) iSCSI AWS primary Primary data in AWS, cache locally
Volume Gateway (Stored) iSCSI AWS snapshots Primary data on-prem, backup to AWS
Tape Gateway Virtual tape S3 + Glacier Replace physical tape backups

The mode matters. Cached Volumes means primary data lives in AWS. Stored Volumes means primary data stays on-prem with AWS as the backup.

Snow Family

Physical devices for moving data when the internet is too slow.

Device Scale
Snowcone Small
Snowball Medium (up to hundreds of TB)
Snowmobile Petabytes / exabytes

The rule of thumb: if uploading over the internet would take weeks, use a Snow device instead.

Cheat Sheet

Files?                  -> S3
One EC2 disk?           -> EBS
Many EC2 shared disk?   -> EFS
Windows file share?     -> FSx for Windows
HPC / ML workloads?     -> FSx for Lustre
Archive?                -> Glacier
Hybrid storage?         -> Storage Gateway
Massive data move?      -> Snowball
Recover deleted S3?     -> Versioning
Automatic tiering?      -> Lifecycle Policy
Cross-region copy?      -> CRR
Encryption + auditing?  -> SSE-KMS
Primary in AWS?         -> Cached Volumes
Primary on-prem?        -> Stored Volumes
Replace tape backups?   -> Tape Gateway

Module 6: Database Services

Database questions are about matching the data model to the right service. The exam does not expect you to know SQL syntax or DynamoDB internals. It expects you to read a scenario and pick the right tool.

Service Overview

Service Type Use when
RDS Relational SQL Structured data, joins, transactions
Aurora Relational SQL Need MySQL/PostgreSQL with better performance
DynamoDB NoSQL key-value Massive scale, millisecond latency, serverless
ElastiCache In-memory cache Reduce database load, sub-millisecond reads
Redshift Data warehouse Analytics on large datasets
DocumentDB Document (MongoDB) Migrating MongoDB workloads
Neptune Graph Relationship-heavy data
Timestream Time-series IoT sensors, metrics over time
QLDB Immutable ledger Cryptographically verifiable audit trail

RDS

RDS manages relational databases — MySQL, PostgreSQL, MariaDB, Oracle, SQL Server. AWS handles backups, patching, and failover. You handle the data and queries.

Use RDS when the question mentions SQL, ACID transactions, joins, or structured data.

Multi-AZ keeps a standby replica in a second AZ. If the primary fails, AWS fails over automatically. This is for availability, not performance.

Read Replicas offload read traffic from the primary. If the question says the database is overloaded by reads, add a read replica.

  Multi-AZ Read Replica
Purpose High availability Read scaling
Failover Automatic Manual promotion
Helps with DR Yes No

The exam loves this comparison. Multi-AZ is the availability answer. Read Replica is the performance answer. They are not interchangeable.

Aurora

Aurora is AWS’s own database engine, compatible with MySQL and PostgreSQL. It stores 6 copies of data across 3 AZs automatically. It supports up to 15 read replicas. It is faster and more available than standard RDS.

Aurora Serverless scales capacity up and down automatically. Use it for unpredictable or infrequent workloads where you do not want to provision a fixed instance size.

DynamoDB

A serverless NoSQL key-value database. No SQL, no joins, no schema. You look up items by partition key. It scales to millions of requests per second with millisecond latency.

DAX (DynamoDB Accelerator) is an in-memory cache in front of DynamoDB. If the question asks for microsecond latency on DynamoDB reads, DAX is the answer.

DynamoDB Global Tables replicate data across multiple regions, all writable. Use it when the question asks for a global active-active database with low latency everywhere.

ElastiCache

An in-memory cache that sits in front of a database. Frequently read data is served from memory instead of hitting the database on every request.

Two engines: Redis and Memcached. Redis supports replication and persistence. Memcached is simpler. When in doubt, Redis is the answer.

Without cache:  App -> Database (every request)
With cache:     App -> Cache   (most requests)
                App -> Database (cache miss only)

Redshift

A data warehouse for analytics. Not for transactional workloads. Use it when the question involves querying petabytes of historical data, business intelligence, or reporting.

RDS       = run transactions
Redshift  = run reports

Specialist Databases

DocumentDB is MongoDB-compatible. If the question mentions migrating a MongoDB application, DocumentDB is the answer.

Neptune is a graph database. Use it when the data is about relationships — social networks, fraud detection, recommendation engines.

Timestream is for time-series data — IoT sensors, CPU metrics, anything where every record has a timestamp and you query across time ranges.

QLDB is an immutable ledger. Records cannot be altered or deleted without leaving a trace. Use it when the question mentions cryptographically verifiable history or financial audit trails.

Cheat Sheet

SQL / relational?                -> RDS
SQL with better performance?     -> Aurora
Unpredictable DB workload?       -> Aurora Serverless
Availability / failover?         -> Multi-AZ
Read scaling?                    -> Read Replica
NoSQL, massive scale?            -> DynamoDB
Microsecond DynamoDB reads?      -> DAX
Global active-active NoSQL?      -> DynamoDB Global Tables
Reduce DB load with cache?       -> ElastiCache
Analytics on large datasets?     -> Redshift
MongoDB migration?               -> DocumentDB
Relationship data?               -> Neptune
Sensor / time-series data?       -> Timestream
Immutable audit trail?           -> QLDB

Multi-AZ   = availability, automatic failover
Read Replica = performance, read scaling
These are not the same thing.

Module 7: Monitoring and Scaling

Two problems this module solves: knowing when something is wrong, and handling more traffic without manual intervention. The services split cleanly along those two lines.

Monitoring

CloudWatch collects metrics from AWS resources — CPU, network, disk, memory (with a custom metric agent). You set alarms on those metrics to trigger notifications, Auto Scaling actions, or Lambda functions. CloudWatch Logs stores log output from applications, Lambda, and EC2.

CloudTrail records every API call made in your account — who did what, when, and from where. It is not for performance monitoring. It is for auditing.

AWS Config tracks configuration changes to resources over time. If a security group was modified or an S3 bucket became public, Config has the history.

Trusted Advisor scans your account and flags cost, security, performance, and fault tolerance issues. Unused EBS volumes, open security group ports, underutilised instances — that kind of thing.

The confusion between these four comes up constantly:

Service Answers the question
CloudWatch Is something performing badly right now?
CloudTrail Who made that change?
AWS Config What did this resource look like before?
Trusted Advisor What should I fix to save money or improve security?

Load Balancing

A load balancer distributes incoming traffic across multiple EC2 instances. It also stops sending traffic to unhealthy instances automatically.

Three types, each for a different layer:

Type Layer Use when
ALB 7 (HTTP/HTTPS) Route by URL path, hostname, or headers
NLB 4 (TCP/UDP) Highest throughput, lowest latency
GWLB 3 Routing traffic through third-party firewalls

ALB is the most common exam answer. If the question mentions routing /api to one target group and /images to another, that is ALB. If the question asks for the highest performance TCP load balancer, that is NLB.

Auto Scaling

Auto Scaling adds and removes EC2 instances based on demand. You define three numbers: minimum (never go below), desired (target), and maximum (never exceed).

Three scaling policies:

  • Dynamic: reacts to a metric, typically CPU. When CPU exceeds a threshold, add instances.
  • Scheduled: fires at a known time. Use when traffic patterns are predictable — scale up every Monday morning.
  • Predictive: AWS analyses historical patterns and scales ahead of expected demand.

ELB and Auto Scaling Together

This architecture appears in almost every exam section:

Users
  |
 ALB
  |
Auto Scaling Group
 / | \
EC2 EC2 EC2

When traffic rises, Auto Scaling adds instances and ALB starts routing to them. When an instance fails a health check, ALB stops sending it traffic and Auto Scaling replaces it.

Cheat Sheet

Resource metrics and alarms?        -> CloudWatch
Who made an API call?               -> CloudTrail
Configuration change history?       -> AWS Config
Cost and security recommendations?  -> Trusted Advisor

Route by URL path?                  -> ALB
Highest performance TCP?            -> NLB
Third-party firewall appliances?    -> GWLB

Auto add/remove EC2s?               -> Auto Scaling
CPU-based scaling?                  -> Dynamic scaling
Known traffic pattern?              -> Scheduled scaling
Forecast-based scaling?             -> Predictive scaling

CloudWatch  = what is happening
CloudTrail  = who did it
AWS Config  = what changed

Module 8: Automation

The core idea here is infrastructure as code. Instead of clicking through the console to build an environment, you write a template that describes what you want and AWS builds it. The same template can recreate the same environment identically every time.

CloudFormation

CloudFormation is AWS’s infrastructure as code service. You write a YAML or JSON template describing your resources — VPC, EC2, RDS, security groups, load balancers — and CloudFormation creates them as a single unit called a stack.

Resources:
  MyBucket:
    Type: AWS::S3::Bucket
  MyEC2:
    Type: AWS::EC2::Instance

When you update the template and apply it, CloudFormation changes only what is different. When you delete the stack, it deletes all the resources inside it.

The exam tests four specific CloudFormation concepts:

Drift happens when someone manually changes a resource that CloudFormation manages. The deployed infrastructure no longer matches the template. CloudFormation can detect this. If the question describes a manual change to a CloudFormation-managed resource, the answer involves drift detection.

Change Sets let you preview what will change before applying an update — which resources will be added, modified, or deleted. Think of it as a diff before you commit.

Rollback happens automatically if a stack update fails partway through. CloudFormation returns the infrastructure to its last known good state.

Stack is the running collection of resources created from a template. Template defines what you want. Stack is what exists.

Template -> CloudFormation -> Stack -> Resources

Why It Matters for the Exam

The exam uses CloudFormation as the answer whenever the question involves:

  • identical environments across dev, test, and production
  • recreating infrastructure after an outage
  • version-controlled infrastructure
  • consistent, repeatable deployments

Amazon Q Developer

Amazon Q Developer is an AI assistant for software development and AWS tasks. It can write code, explain errors, generate CloudFormation templates, and answer AWS architecture questions. It shows up occasionally in the exam as the answer when the question asks about an AI coding assistant or automated code generation within AWS.

Cheat Sheet

Infrastructure as code?              -> CloudFormation
Identical dev/test/prod environments?-> CloudFormation
Manual change to managed resource?   -> Drift
Preview changes before applying?     -> Change Set
Failed update, restore previous?     -> Rollback
AI coding assistant for AWS?         -> Amazon Q Developer

Template  = what you want
Stack     = what exists
Drift     = reality no longer matches template

Module 9: Containers

Containers sit between EC2 and Lambda on the control spectrum. More portable than EC2, more flexible than Lambda. The exam does not test container internals — it tests when to use ECS vs EKS vs Fargate.

What a Container Is

A container packages an application together with everything it needs to run — runtime, libraries, dependencies, config. It runs the same way regardless of where it is deployed. The image is the blueprint. The running container is the instance of it.

Docker images are stored in Amazon ECR (Elastic Container Registry). Services pull images from ECR to run them.

Developer builds image
        |
     Amazon ECR  (stores images)
        |
   ECS or EKS   (runs containers)
        |
     Customers

Microservices

A monolithic application bundles everything — login, payments, orders, search — into one deployable unit. If one part fails or needs scaling, the whole thing is affected.

Microservices split those into independent services. Each runs separately, scales separately, and fails independently. Containers are the natural fit for microservices because each service becomes its own container.

Container Services

ECS (Elastic Container Service) is AWS’s own container orchestration service. It handles scheduling, scaling, and placement. Use it when you want AWS-native container management and are not already invested in Kubernetes.

EKS (Elastic Kubernetes Service) runs managed Kubernetes on AWS. Use it when the company already uses Kubernetes or needs portability across clouds. If the question mentions Kubernetes, the answer is EKS.

Fargate is the serverless compute layer for containers. It removes the need to provision or manage EC2 instances underneath your containers. Fargate is not a standalone service — it is a launch type used with ECS or EKS.

No Kubernetes requirement?   -> ECS
Already using Kubernetes?    -> EKS
Don't want to manage EC2?    -> add Fargate to either
Service Manages orchestration Manages servers
ECS Yes (AWS-native) You (unless using Fargate)
EKS Yes (Kubernetes) You (unless using Fargate)
Fargate No AWS

The most common exam scenario: containers, no Kubernetes requirement, no server management. Answer is ECS with Fargate.

Cheat Sheet

Store Docker images?             -> Amazon ECR
AWS-native container management? -> ECS
Already using Kubernetes?        -> EKS
No server management?            -> Fargate (with ECS or EKS)
App split into independent parts?-> Microservices

ECR   = stores images, does not run them
Fargate = execution engine, not a standalone service
ECS   = AWS proprietary
EKS   = Kubernetes

Module 10: Networking 2

Module 3 covered how traffic moves inside a VPC. This module is about how VPCs talk to each other, how AWS services are accessed privately, and how on-premises networks connect to AWS.

VPC Endpoints

By default, when an EC2 instance in a private subnet calls an AWS service like S3, that traffic routes out through a NAT Gateway and over the public internet — even though both are AWS. A VPC Endpoint keeps that traffic on the private AWS network.

Two types:

Gateway Endpoint works only for S3 and DynamoDB. It is free and added to a route table. If the question says EC2 needs to access S3 or DynamoDB without going through the internet, this is the answer.

Interface Endpoint (AWS PrivateLink) works for almost every other AWS service — Secrets Manager, SQS, SNS, CloudWatch, KMS, and more. It creates an Elastic Network Interface inside your subnet. More flexible, but has a cost.

Gateway Endpoint   -> S3 and DynamoDB only, free
Interface Endpoint -> everything else, uses ENI

VPC Peering

VPC Peering connects two VPCs so their resources can communicate privately over the AWS backbone, not the internet.

Two rules the exam tests repeatedly:

CIDR blocks cannot overlap. If VPC A is 10.0.0.0/16 and VPC B is 10.0.1.0/24, they overlap and cannot be peered.

Peering is not transitive. If A peers with B, and B peers with C, A cannot reach C through B. You need a direct peering connection between A and C.

A -- B -- C    A cannot reach C
A -- B         direct peering required
A -- C

Hybrid Networking

Hybrid means some infrastructure stays on-premises and some runs in AWS. Two ways to connect them:

Site-to-Site VPN creates an encrypted tunnel over the public internet. It has two components: a Virtual Private Gateway on the AWS side (attached to the VPC) and a Customer Gateway on the on-premises side (your router or firewall). AWS always creates two redundant IPsec tunnels for high availability.

On-premises
     |
Customer Gateway
     || (two tunnels)
Virtual Private Gateway
     |
    VPC

Direct Connect is a dedicated private fibre connection between your data centre and AWS. No internet. Predictable latency, consistent bandwidth. More expensive and slower to set up than VPN, but the right answer when the question asks for reliable, high-bandwidth, low-latency hybrid connectivity.

  VPN Direct Connect
Path Public internet Dedicated fibre
Setup time Fast Weeks to months
Cost Low High
Latency Variable Consistent
Use when Quick or temporary Enterprise, high bandwidth

A common architecture combines both: Direct Connect as the primary path, VPN as the failover. The exam calls this resilient hybrid connectivity.

Transit Gateway

When you have many VPCs, peering every pair becomes unmanageable. Transit Gateway acts as a central hub. Each VPC, VPN, or Direct Connect connection attaches to it once, and Transit Gateway routes traffic between them.

The connection from any network to a Transit Gateway is called an attachment — a VPC attachment, a VPN attachment, a Direct Connect attachment.

        Transit Gateway
       /    |    |    \
    VPC   VPC   VPN   Direct Connect

Transit Gateway has its own route tables that control which attachments can communicate with each other.

VPC Peering is still the right answer for connecting exactly two VPCs simply. Transit Gateway is the answer when you have many VPCs or need to combine VPCs with on-premises connectivity.

Cheat Sheet

Private access to S3 or DynamoDB?    -> Gateway Endpoint
Private access to other AWS services?-> Interface Endpoint (PrivateLink)
Connect two VPCs?                    -> VPC Peering
Connect many VPCs?                   -> Transit Gateway
Quick encrypted on-prem connection?  -> Site-to-Site VPN
Dedicated high-bandwidth connection? -> Direct Connect
Resilient hybrid (primary + backup)? -> Direct Connect + VPN

Site-to-Site VPN components:
  AWS side      -> Virtual Private Gateway
  On-prem side  -> Customer Gateway
  Tunnels       -> Two (redundant)

Connection to Transit Gateway        -> Attachment
VPC Peering: no overlapping CIDRs, not transitive

Module 11: Serverless

Serverless does not mean no servers. It means AWS manages the servers and you only think about the code and the architecture. You pay for what runs, not for what sits idle.

Services Overview

Service Role
Lambda Run code on events
API Gateway Front door for HTTP APIs
SQS Queue work between services
SNS Broadcast one message to many subscribers
Kinesis Process continuous real-time data streams
Step Functions Orchestrate multi-step workflows
EventBridge Route events based on rules

Lambda

Lambda runs code in response to events. An event can come from API Gateway, S3, SQS, SNS, Kinesis, EventBridge, or a schedule. It scales automatically and you pay per invocation and duration. Maximum execution time is 15 minutes.

Use Lambda for APIs, file processing, automation, and event-driven tasks. Do not use it for long-running processes or anything that needs OS-level access — use EC2 or containers for those.

Synchronous invocation means the caller waits for a response (API Gateway calling Lambda for a login request). Asynchronous means the caller does not wait (SNS triggering Lambda to send an email).

API Gateway

API Gateway sits in front of Lambda and exposes HTTP endpoints. It handles authentication, authorisation, rate limiting (throttling), caching, and monitoring. Clients call the API Gateway URL, not Lambda directly.

Use API Gateway for serverless REST APIs. Use ALB when you are load balancing EC2 or containers.

SQS

SQS is a message queue. A producer puts messages in. A consumer picks them up and processes them. This decouples the two sides — the producer does not need to wait for the consumer, and a spike in messages does not crash the consumer.

Two queue types:

Standard Queue — at-least-once delivery, very high throughput, order not guaranteed. Messages can occasionally be delivered more than once.

FIFO Queue — exactly-once processing, strict ordering, lower throughput. Use when order matters — banking transactions, inventory updates.

Four SQS behaviours the exam tests:

Visibility Timeout — when a consumer picks up a message, SQS hides it from other consumers for a set period. If the consumer finishes, it deletes the message. If it crashes, the timeout expires and the message becomes visible again for another consumer to retry.

Dead-Letter Queue (DLQ) — if a message fails processing too many times, SQS moves it to a DLQ instead of retrying forever. Use this to isolate and investigate problem messages.

Long Polling — instead of asking the queue repeatedly when it is empty, the consumer waits until a message arrives. Fewer API calls, lower cost.

Delay Queue — messages are hidden for a configured period after being sent before becoming available to consumers.

SNS

SNS broadcasts one message to many subscribers simultaneously. A single publish to an SNS topic delivers a copy to every subscriber — Lambda functions, SQS queues, email addresses, HTTP endpoints, SMS.

SQS  = one consumer processes each message
SNS  = every subscriber gets a copy

A common architecture combines both: SNS fans out to multiple SQS queues, each consumed by a different service independently at its own pace.

Kinesis

Kinesis processes continuous high-volume data streams in real time — IoT sensors, website clickstreams, financial trades, GPS telemetry. Unlike SQS which handles discrete tasks, Kinesis handles an ongoing flow of events where multiple consumers can read the same stream.

Tasks to process?          -> SQS
Continuous stream of data? -> Kinesis

Step Functions

Step Functions orchestrates sequences of Lambda functions and other AWS services into a workflow. It handles retries, error handling, branching, parallel execution, and state tracking. Use it when you have multiple steps that need to run in order, with conditions or retries between them — loan approvals, order fulfilment, image processing pipelines.

EventBridge

EventBridge routes events based on rules. An EC2 state change, a scheduled time, or a custom application event can trigger a rule that routes to Lambda, SNS, SQS, or other targets. Unlike SNS which broadcasts to all subscribers, EventBridge filters and routes based on event patterns.

Cheat Sheet

Run code on events?                  -> Lambda
HTTP API without servers?            -> API Gateway + Lambda
Queue work, decouple services?       -> SQS
Broadcast to many subscribers?       -> SNS
Real-time data streams?              -> Kinesis
Multi-step workflow with retries?    -> Step Functions
Rule-based event routing?            -> EventBridge

SQS Standard  = at-least-once, high throughput, no order
SQS FIFO      = exactly-once, ordered, lower throughput
Visibility Timeout = prevents duplicate processing
DLQ           = captures repeatedly failing messages
Long Polling  = reduces empty responses and API calls

Lambda max runtime = 15 minutes
API Gateway vs ALB: serverless API -> API Gateway,
                    servers/containers -> ALB

Module 12: Edge Services

Edge services reduce latency by moving content or processing closer to users, and protect applications from attacks before traffic reaches your infrastructure.

Route 53

Route 53 is AWS’s DNS service. It translates domain names to IP addresses, registers domains, and monitors endpoint health.

Routing policies — the exam tests which one fits the scenario:

Policy Use when
Simple One server, no logic needed
Weighted Split traffic by percentage (e.g. canary deployments)
Latency Route users to the AWS region with lowest latency
Geolocation Route based on the user’s country or continent
Failover Send traffic to a backup endpoint when primary is unhealthy
Multivalue Return multiple healthy IPs for basic load distribution

Health checks integrate with Failover routing. If the primary endpoint fails a health check, Route 53 automatically routes to the secondary.

CloudFront

CloudFront is a CDN (Content Delivery Network). It caches content at edge locations around the world so users are served from a nearby location instead of the origin region.

Origins can be S3, EC2, ALB, or API Gateway. The first request for a file goes to the origin. Subsequent requests are served from the edge cache.

Use CloudFront when the question mentions slow global performance, reducing load on an origin, or caching static content. CloudFront also supports signed URLs and signed cookies to restrict access to private content.

Edge Locations are not full AWS Regions. They run CloudFront, Route 53, and Shield — not EC2 or RDS.

DDoS Protection

AWS Shield Standard is free and enabled automatically on CloudFront, Route 53, and Elastic Load Balancing. It protects against common network and transport layer DDoS attacks.

AWS Shield Advanced is a paid subscription. It adds enhanced protection, 24/7 access to the AWS DDoS Response Team, and cost protection against scaling charges caused by an attack.

AWS WAF (Web Application Firewall) operates at the application layer. It blocks malicious HTTP requests — SQL injection, cross-site scripting, bot traffic, IP-based rules, rate limiting. WAF attaches to CloudFront, ALB, or API Gateway.

DDoS flood attack?          -> Shield
SQL injection / XSS?        -> WAF
Manage WAF/Shield at scale? -> Firewall Manager

Do not confuse these with GuardDuty (threat detection from logs), Inspector (vulnerability scanning), or Macie (sensitive data in S3) — those are security monitoring tools, not attack prevention.

AWS Outposts

Outposts installs AWS hardware inside your own data centre. You run EC2, EBS, ECS, EKS, and RDS locally, but they are managed through the AWS console and connected to an AWS region. Use it when regulations require data to stay on-premises, or when applications need very low latency to local systems.

Outposts is different from Local Zones. Local Zones are AWS-owned infrastructure placed near a city. Outposts is AWS hardware inside your building.

Cheat Sheet

DNS?                              -> Route 53
Domain registration?              -> Route 53
Route to lowest latency region?   -> Latency routing
Route by country?                 -> Geolocation routing
Automatic DNS failover?           -> Failover routing
Canary / A-B traffic split?       -> Weighted routing

Global content caching?           -> CloudFront
Reduce origin load?               -> CloudFront
Restrict access to content?       -> Signed URLs / Signed Cookies

DDoS protection (free)?           -> Shield Standard
DDoS protection (enterprise)?     -> Shield Advanced
Block SQL injection / XSS?        -> WAF
Centralise WAF policies?          -> Firewall Manager

AWS hardware on-premises?         -> Outposts
AWS near a city (AWS-owned)?      -> Local Zone

Module 13: Backup and Recovery

This module is about designing for failure at the data level. High availability keeps applications running. Backup and recovery gets your data back when something goes wrong — accidental deletion, corruption, ransomware, or a regional outage.

RPO and RTO

Two terms the exam uses constantly:

RPO (Recovery Point Objective) — how much data loss is acceptable. If you back up every hour and a failure happens at 2:45, you lose 45 minutes of data. That is your RPO.

RTO (Recovery Time Objective) — how long the application can be down. If it takes 3 hours to restore, your RTO is 3 hours.

RPO = data loss tolerance
RTO = downtime tolerance
Lower is better for both.

Disaster Recovery Strategies

Four strategies, ordered from cheapest/slowest to most expensive/fastest:

Strategy What runs in DR Recovery time Cost
Backup and Restore Nothing — restore from backup when needed Hours Lowest
Pilot Light Core data tier only (e.g. database) Minutes to hours Low
Warm Standby Full environment at reduced scale Minutes Medium
Multi-Site / Active-Active Full environment at full scale, serving traffic Near zero Highest

The exam gives you a scenario with cost and recovery time constraints and asks which strategy fits. Backup and Restore is the answer when cost is the priority. Active-Active is the answer when downtime must be near zero.

High Availability vs Backup

These solve different problems. Multi-AZ RDS keeps the database running if an AZ fails. But if someone deletes a table, Multi-AZ replicates that deletion to the standby. You need a backup to recover the data. The exam tests this distinction — availability and backup are complementary, not interchangeable.

AWS Backup

AWS Backup is a centralised service for managing backups across EC2, EBS, RDS, Aurora, DynamoDB, EFS, FSx, and Storage Gateway. Instead of configuring backups separately per service, you define backup plans in one place.

Key features the exam tests:

Backup Plans define what to back up, how often, and how long to keep it. Lifecycle rules can move backups to cold storage after a set period to reduce cost.

Backup Vault is where recovery points are stored. Vaults support encryption via KMS and access control policies.

Tag-based backup automatically includes any resource with a matching tag in a backup plan. Useful for large environments where resources are created frequently.

Cross-Region copy copies backups to another region, protecting against a regional disaster destroying both the production data and its backups.

Cross-Account copy copies backups to a separate AWS account. If the production account is compromised by ransomware or an insider, the backups in the separate account remain intact.

Backup Vault Lock makes backups immutable for a defined retention period. No one — including account administrators — can delete them before the period expires. This is the WORM (Write Once Read Many) answer for backups.

Point-in-Time Recovery

RDS and DynamoDB both support point-in-time recovery. You can restore to any second within the retention window, not just the last backup. Use this when data was corrupted or deleted at a known time and you need to recover to just before that moment.

Cheat Sheet

How much data loss is acceptable?    -> RPO
How long can the app be down?        -> RTO

Cheapest DR strategy?                -> Backup and Restore
Fastest DR strategy?                 -> Multi-Site Active-Active
Core data always running, rest off?  -> Pilot Light
Scaled-down full environment?        -> Warm Standby

Centralised backup management?       -> AWS Backup
Automatic backup by resource tag?    -> Tag-based backup
Backups stored in?                   -> Backup Vault
Protect backups from deletion?       -> Backup Vault Lock
Survive regional disaster?           -> Cross-Region copy
Survive account compromise?          -> Cross-Account copy
Restore to specific moment?          -> Point-in-Time Recovery

Multi-AZ = availability (keeps app running)
Backup   = recovery (gets data back)
These are not the same thing.

Comments