devHọc Dev
Bài học

Tuần 7 - Ngày 1: High Availability Architectures

Tuần 7 – Ngày 1

Mục tiêu học tập

  • Nắm 4 nguyên tắc thiết kế HA và pattern Multi-AZ chuẩn cho web app
  • Chọn đúng scaling policy theo traffic pattern
  • Hiểu HA pattern của từng loại database (RDS, Aurora, DynamoDB)

1. HA Design Principles

HAPRINCIPLES1.EliminateSinglePointsofFailure-Multi-AZdeployments-Redundantcomponents2.DesignforFailure-Assumecomponentswillfail-Implementhealthchecks-Auto-healingmechanisms3.LooseCoupling-Usequeuesbetweencomponents-Asyncwherepossible4.StatelessApplications-Storestateexternally-Easyhorizontalscaling

2. Multi-AZ Architecture

MULTI-AZWEBAPPRoute53(DNS)ALB(Multi-AZ)AZ-aAZ-bAZ-cEC2EC2EC2RDSMulti-AZorAurora

3. Auto Scaling Strategies

SCALING POLICIES:

1. Target Tracking (Recommended)
   - Maintain CPU at 50%
   - Simple configuration

2. Step Scaling
   - Multiple thresholds
   - Different actions per step

3. Scheduled Scaling
   - Known traffic patterns
   - Time-based

4. Predictive Scaling
   - ML-based prediction
   - Proactive scaling

COOLDOWN PERIODS:
- Default: 300 seconds
- Prevents rapid scale in/out
- Can customize per policy

4. Session Management

STATELESSPATTERN:UserALBAnyEC2instanceElastiCache(Redis)SessionStoreBenefits:-Instancecanfail-Scalein/outfreely-Nostickysessionsneeded

5. Database HA Patterns

RDS:Multi-AZ(syncreplication,autofailover)ReadReplicas(async,manualpromotion)CombinationofbothAurora:6copiesacross3AZsWriter+upto15readersAutofailover<30secondsGlobalDatabaseforcross-regionDynamoDB:Multi-AZbydefaultGlobalTablesformulti-region

6. Câu hỏi ôn tập

  1. Vì sao stateless application là điều kiện tiên quyết cho HA?

    Xem đáp án

    Khi state (session, upload tạm, cache cục bộ) nằm trên instance, mất instance = mất state và ALB không thể route tự do — phải sticky sessions, làm scale-in gây lỗi người dùng. Đưa state ra ngoài (ElastiCache/DynamoDB cho session, S3/EFS cho file) thì mọi instance thay thế được nhau: ASG tự heal, scale in/out không ảnh hưởng user, Multi-AZ failover trong suốt. Đây cũng là điều kiện để đi tiếp lên Multi-Region active-active.

  2. Target tracking vs step scaling vs predictive — chọn thế nào?

    Xem đáp án

    Target tracking (recommended, mặc định): giữ metric quanh target (CPU 50%) — đơn giản, đủ cho đa số. Step scaling: cần phản ứng khác nhau theo mức vượt ngưỡng (vượt 20% thêm 2 instance, vượt 50% thêm 5). Scheduled: pattern biết trước theo giờ (batch 8AM). Predictive: ML dự báo — chỉ hiệu quả khi traffic có chu kỳ lặp lại (daily/weekly); spike ngẫu nhiên thì không dự báo được, phải dựa vào target/step với threshold thấp và warm pool.

  3. Aurora failover khác RDS Multi-AZ failover ra sao?

    Xem đáp án

    RDS Multi-AZ: failover sang standby (sync replica) mất 60–120s — chủ yếu do DNS propagation + crash recovery. Aurora: reader được promote trong < 30s (storage layer chia sẻ 6 copies/3 AZs nên không cần replay), chọn reader theo priority tier; nếu không có reader, Aurora tạo writer mới lâu hơn. Muốn failover nhanh hơn nữa: RDS Proxy (giữ connection) hoặc Aurora Global Database cho region-level failure.

  4. Khi nào Multi-AZ là đủ và khi nào phải Multi-Region?

    Xem đáp án

    Multi-AZ đủ cho hầu hết yêu cầu HA (99.99%) — AZ failure là kịch bản chính, chi phí thấp hơn nhiều. Multi-Region chỉ khi: (1) DR yêu cầu RTO/RPO rất thấp kể cả khi cả region mất, (2) compliance bắt buộc geographic redundancy, (3) user toàn cầu cần low-latency writes (active-active). Bẫy exam kinh điển: chọn Multi-Region khi đề chỉ yêu cầu "highly available" là over-engineering — sai theo tiêu chí "simplest solution that meets requirements".

7. Bài tập thực hành

  1. Dựng Multi-AZ web tier: ALB + ASG (min 2, across 2 AZ) + RDS Multi-AZ; terminate 1 instance bằng tay và quan sát ASG tự thay thế, ALB health check loại instance chết khỏi rotation.

  2. Session store ngoài instance: chuyển demo app từ session in-memory sang ElastiCache Redis; tắt sticky sessions trên ALB và verify user không bị logout khi instance bị terminate.

  3. So sánh scaling policies: cấu hình target tracking CPU 50%, bơm tải bằng stress, đo thời gian scale-out; đổi sang step scaling threshold thấp và so sánh tốc độ phản ứng.


Tài liệu tham khảo chính thức


Ngày tiếp theo: Disaster Recovery Strategies