“Recovery after disaster” = Preparing for and recovering from hardware failures, network outages, fires, or human error.
The 4 DR Strategies
(From Cheapest/Slowest to Most Expensive/Fastest)
Backup and Restore: You back up data to S3 or take EBS snapshots. If disaster strikes, you manually rebuild the environment and restore data. (Longest recovery time, lowest cost).
Pilot Light: The core, essential components (like a database) are always running and updated in AWS. The rest of the servers are turned off but have AMIs ready to launch quickly. (Faster recovery, moderate cost).
Warm Standby: A scaled-down, fully functional version of your entire environment is always running in AWS. During a disaster, you simply “scale up” the instance sizes to handle full production traffic. (Fast recovery, higher cost).
Multi-Site (Active/Active): Your full production environment runs simultaneously in your data center and in AWS. Traffic is load-balanced between both. If one fails, the other handles 100% of the load instantly. (Near-zero recovery time, highest cost).
Key DR Metrics
RTO (Recovery Time Objective): How long can you afford to be down? (e.g., “We must be back up within 4 hours”).
RPO (Recovery Point Objective): How much data can you afford to lose? (e.g., “We can only lose the last 15 minutes of transactions”).
Scenario Spotlight
You need a DR strategy that keeps a few key servers running with smaller instance types than production to save money, but allows quick scaling.
➔ Pilot Light (or Warm Standby, depending on how much is running).