Designing Resilient AWS Architectures
Resilient architecture starts with the failure a system must withstand and its recovery objectives.
- Availability across failure zones, replication, backup, and disaster recovery address different needs.
- Map dependencies, define acceptable downtime and data loss, choose controls that fit, and test recovery rather than assuming a redundant service is recoverable.
On this page8 sections
- Start with business recovery needs
- Availability zones and regions
- Replication is not the same as backup
- Worked scenario: an order service
- Scale and queue for controlled recovery
- Common resilience mistakes
- A decision sequence for exam scenarios
- Map failure scope before choosing a pattern A component failure, Availability Zone disruption, regional outage, data corruption, and account compromise are different events. A design should state which it handles. Multi-zone deployment can reduce exposure to a zone-level event when compute, data, network paths and dependencies are configured accordingly. A regional recovery plan addresses a broader event and has additional replication, routing, and operational requirements. Recovery time objective and recovery point objective turn resilience into measurable targets. RTO is the target time to restore service. RPO is the acceptable data loss measured in time. A design with frequent replication may support a smaller RPO, while a tested restore process determines how quickly service can return. Define both before comparing architecture cost. Do not use replication as the only defense against data loss. If an application writes an incorrect value or a user deletes a record, replication can propagate that change. A backup or point-in-time restore capability gives another recovery path. Protect the recovery mechanism with separate access controls so a compromised production identity cannot trivially erase every copy. Worked recovery review Suppose a small retailer can tolerate 30 minutes of service interruption and five minutes of order data loss after a localized infrastructure failure. The team should compare a design's tested failover time and replication behavior with those targets. A diagram showing two zones is not enough. Test that the application reconnects, that capacity is available, and that order data meets the acceptable loss window. Now suppose the concern is a developer accidentally deleting a table. The same availability design may preserve service through a hardware issue but reproduce the deletion to its standby. A point-in-time recovery option can address the logical error, but the team must test restoration and understand how writes made after the restore point will be reconciled. A recovery exercise should involve the people who will act during a real incident. Record who authorizes failover, who validates data, how users are notified, and which dependency is restored first. Measure elapsed time rather than assuming a runbook can be completed quickly. After a test, update the architecture and runbook when a step depends on an undocumented account, credential, DNS change, or vendor action. Review resilience as the workload evolves. More users, new regions, changed data types, or a new downstream service can change failure behavior. Revisit recovery objectives when business impact changes and repeat restore tests after major modifications. This operational discipline is part of a resilient design because redundancy without a working recovery process can create false confidence.
Start with business recovery needs
Resilience is the ability of a workload to continue operating or recover when components fail. Before selecting services, identify the business process, acceptable interruption, acceptable data loss, and likely failure modes. Recovery time objective describes the target time to restore service; recovery point objective describes the acceptable amount of data loss measured in time.
These objectives shape architecture. A system that can tolerate several hours of downtime and a limited amount of data loss may use a different recovery design from a payment service that must remain available with minimal interruption. More aggressive objectives usually require more engineering, testing, replication, and cost.
List dependencies from the user request to the data store: DNS, network paths, load balancers, compute, identity, databases, queues, and external providers. Redundancy at the web tier will not help if every instance depends on one database or one external service. Identify the failure domains the design is expected to survive.
Availability zones and regions
An Availability Zone design can reduce the impact of a localized infrastructure failure when the application and its dependencies are configured to use more than one zone. Load balancing and health checks can route requests away from an unhealthy component. A managed database with a multi-zone standby can support failover for certain failure cases.
Multi-region designs address a broader failure scope and can support geographic recovery or residency requirements. They introduce complexity: data replication, conflict handling, failover decisions, global routing, and consistent operations. Do not choose multi-region simply because it sounds more resilient. Use it when the business requirement justifies the additional moving parts.
Availability and backup are different. A second instance or database standby can keep service running, but may also reproduce corrupted writes or accidental deletion. A backup provides a recovery point, but restoration takes time and requires a tested procedure. A robust design can need both.
Replication is not the same as backup
Replication copies data or changes to another location, often to support availability, read scaling, or disaster recovery. If an authorized user deletes data or an application writes incorrect records, replication may copy the same mistake. A backup preserves a point-in-time version that can be restored, subject to retention and backup design.
A recovery plan should define backup frequency, retention, isolation, access, encryption, and restore testing. Protect backup administration separately from the production workload so that one compromised identity cannot easily delete both. Confirm that backups include configuration and dependencies needed to make the data usable.
A snapshot that has never been restored is an assumption, not proof of recovery. Test restoration in an isolated environment, measure time, validate data integrity, and document application steps. Re-test after major architecture or database changes.
Worked scenario: an order service
An online retailer runs a stateless application tier behind a load balancer. Application instances are spread across two Availability Zones, but a single database instance and its storage reside in one zone. The business wants the site to continue through a zone outage and recover orders after accidental deletion.
For the zone-outage requirement, address the database single point of failure with an appropriate multi-zone database design and verify that the application can reconnect after failover. Ensure the load balancer health checks and compute capacity operate across zones. Review whether shared session state, queues, or another dependency remains tied to one zone.
For accidental deletion, configure a separate backup or point-in-time recovery capability with a retention period that meets the business need. Test restoration and define who can perform it. A standby database alone may receive the deletion, so it does not satisfy the data recovery requirement by itself.
Measure the resulting recovery time and data loss against the objectives. If the service must recover faster than the selected pattern allows, change the design or the business objective through an authorized decision. Do not claim that the architecture meets a target until tests demonstrate it.
Scale and queue for controlled recovery
Resilience also involves handling load spikes and downstream slowdown. Auto scaling can add capacity as demand rises, but it cannot compensate for a database bottleneck or an external service limit. Queues can absorb bursts and decouple producers from consumers, but the application must handle retries, duplicate messages, ordering, and poison messages appropriately.
Use health checks and monitoring that reflect actual user experience, not only whether a process responds. An instance can return a healthy response while it cannot reach its database. Alert thresholds should connect to action, such as scaling, failover, throttling, or incident escalation.
Common resilience mistakes
- Assuming a load balancer removes every single point of failure.
- Treating a read replica as a complete backup strategy.
- Choosing a multi-region design without a recovery requirement or operational plan.
- Failing to test restore procedures and application dependencies.
- Ignoring identity, DNS, certificates, quotas, or third-party services in the recovery path.
- Measuring infrastructure status without testing a real transaction.
A resilience review should cover the full service path and its operational procedures. Well-Architected guidance emphasizes reliability through design, recovery, and ongoing testing. The exam may ask for the best service pattern, but real architecture work also needs ownership, runbooks, monitoring, and exercises.
A decision sequence for exam scenarios
- Name the failure or disruption the requirement describes.
- Identify acceptable downtime and data loss, if supplied.
- Map single points of failure and dependencies.
- Distinguish continuous availability, replication, and point-in-time recovery needs.
- Choose the simplest architecture that meets the objective.
- Confirm that the plan can be monitored and tested.
Map failure scope before choosing a pattern A component failure, Availability Zone disruption, regional outage, data corruption, and account compromise are different events. A design should state which it handles. Multi-zone deployment can reduce exposure to a zone-level event when compute, data, network paths and dependencies are configured accordingly. A regional recovery plan addresses a broader event and has additional replication, routing, and operational requirements. Recovery time objective and recovery point objective turn resilience into measurable targets. RTO is the target time to restore service. RPO is the acceptable data loss measured in time. A design with frequent replication may support a smaller RPO, while a tested restore process determines how quickly service can return. Define both before comparing architecture cost. Do not use replication as the only defense against data loss. If an application writes an incorrect value or a user deletes a record, replication can propagate that change. A backup or point-in-time restore capability gives another recovery path. Protect the recovery mechanism with separate access controls so a compromised production identity cannot trivially erase every copy. Worked recovery review Suppose a small retailer can tolerate 30 minutes of service interruption and five minutes of order data loss after a localized infrastructure failure. The team should compare a design's tested failover time and replication behavior with those targets. A diagram showing two zones is not enough. Test that the application reconnects, that capacity is available, and that order data meets the acceptable loss window. Now suppose the concern is a developer accidentally deleting a table. The same availability design may preserve service through a hardware issue but reproduce the deletion to its standby. A point-in-time recovery option can address the logical error, but the team must test restoration and understand how writes made after the restore point will be reconciled. A recovery exercise should involve the people who will act during a real incident. Record who authorizes failover, who validates data, how users are notified, and which dependency is restored first. Measure elapsed time rather than assuming a runbook can be completed quickly. After a test, update the architecture and runbook when a step depends on an undocumented account, credential, DNS change, or vendor action. Review resilience as the workload evolves. More users, new regions, changed data types, or a new downstream service can change failure behavior. Revisit recovery objectives when business impact changes and repeat restore tests after major modifications. This operational discipline is part of a resilient design because redundancy without a working recovery process can create false confidence.
Map failure scope before choosing a pattern A component failure, Availability Zone disruption, regional outage, data corruption, and account compromise are different events. A design should state which it handles. Multi-zone deployment can reduce exposure to a zone-level event when compute, data, network paths and dependencies are configured accordingly. A regional recovery plan addresses a broader event and has additional replication, routing, and operational requirements. Recovery time objective and recovery point objective turn resilience into measurable targets. RTO is the target time to restore service. RPO is the acceptable data loss measured in time. A design with frequent replication may support a smaller RPO, while a tested restore process determines how quickly service can return. Define both before comparing architecture cost. Do not use replication as the only defense against data loss. If an application writes an incorrect value or a user deletes a record, replication can propagate that change. A backup or point-in-time restore capability gives another recovery path. Protect the recovery mechanism with separate access controls so a compromised production identity cannot trivially erase every copy. Worked recovery review Suppose a small retailer can tolerate 30 minutes of service interruption and five minutes of order data loss after a localized infrastructure failure. The team should compare a design's tested failover time and replication behavior with those targets. A diagram showing two zones is not enough. Test that the application reconnects, that capacity is available, and that order data meets the acceptable loss window. Now suppose the concern is a developer accidentally deleting a table. The same availability design may preserve service through a hardware issue but reproduce the deletion to its standby. A point-in-time recovery option can address the logical error, but the team must test restoration and understand how writes made after the restore point will be reconciled. A recovery exercise should involve the people who will act during a real incident. Record who authorizes failover, who validates data, how users are notified, and which dependency is restored first. Measure elapsed time rather than assuming a runbook can be completed quickly. After a test, update the architecture and runbook when a step depends on an undocumented account, credential, DNS change, or vendor action. Review resilience as the workload evolves. More users, new regions, changed data types, or a new downstream service can change failure behavior. Revisit recovery objectives when business impact changes and repeat restore tests after major modifications. This operational discipline is part of a resilient design because redundancy without a working recovery process can create false confidence.
Common questions
Does multi-AZ replace backup?
No. It can improve availability, while backups support recovery from logical mistakes or unwanted changes.
Is replication a backup?
Not by itself. Replication can copy corruption or deletion with valid changes.
When is multi-region appropriate?
When broader recovery, geographic, or business requirements justify the added complexity.
Why test restores?
To prove data and services can be recovered within the required objectives.
What do RTO and RPO measure?
Target recovery time and acceptable data loss, respectively.