AWS offers a very large service catalog, but most production workloads are built from a consistent, well-understood set of core services. Understanding how they fit together is more valuable than knowing every service that exists.
A VPC defines an isolated network within AWS. Production designs typically span at least two Availability Zones, with public subnets for internet-facing load balancers and NAT gateways, and private subnets for application instances and databases. Route tables and security groups then control exactly what traffic can move where — the same pattern discussed in general cloud network design, applied to AWS's specific constructs.
AWS IAM controls who and what can call which APIs. Production accounts should use IAM roles for EC2 instances and Lambda functions (rather than long-lived access keys embedded in application code), scoped policies limited to specific resources and actions, and separate roles for humans versus automated pipelines. For organizations running multiple accounts, AWS Organizations and Service Control Policies add guardrails at the account level.
EC2 instances are typically placed in an Auto Scaling Group spanning multiple Availability Zones, sized by a scaling policy tied to actual load (CPU utilization, request count, or a custom metric) rather than a fixed instance count. An Application Load Balancer distributes incoming traffic across healthy instances and removes unhealthy ones from rotation automatically based on health checks.
S3 handles object storage — static assets, backups, logs — with lifecycle rules to transition older objects to cheaper storage classes over time. RDS provides managed relational databases with Multi-AZ deployment for automatic failover and read replicas for scaling read-heavy workloads, removing much of the operational burden of running a database directly on EC2.
CloudWatch collects metrics and logs across most AWS services by default, and supports custom metrics from applications. CloudWatch Alarms can trigger notifications or automated actions (like scaling policies) when a metric crosses a defined threshold, forming the backbone of both monitoring and automated response on AWS.
Security groups act as a stateful firewall at the instance level; network ACLs add a stateless layer at the subnet level. Data should be encrypted at rest (using KMS-managed keys for S3, RDS and EBS) and in transit (TLS for load balancer listeners and internal service calls). CloudTrail logs API activity across the account, which is essential for audit and incident investigation.
High availability on AWS generally means spreading resources across multiple Availability Zones within a region, so a single zone's failure doesn't take down the whole workload. Backup strategy typically combines RDS automated backups and snapshots, EBS snapshots, and S3 versioning, with retention aligned to the RPO defined for that workload. Disaster recovery across regions is a separate, more involved decision — usually reserved for workloads where a full regional outage is a risk the business has explicitly decided to plan for.
Whether using Terraform or AWS's native CloudFormation, defining AWS infrastructure as code gives the same benefits discussed generally for IaC: reviewable changes, repeatable environments, and a clear history of how the account's infrastructure evolved over time.
AWS
A practical look at the core architecture decisions involved in building secure and scalable workloads on AWS.
More on this and related topics.
A complete guide to Amazon CloudWatch Omni, AWS's AI-powered observability platform for applications and AI agents, built on OpenTelemetry.
A practical guide to designing cloud infrastructure with the right balance of reliability, security, scalability and operational control.
How modern engineering teams can automate build, test, security and deployment workflows while keeping releases consistent and recoverable.