Mục tiêu học tập
- Phân biệt CloudWatch (metrics/logs/alarms) và CloudTrail (API audit)
- Thuộc các metric quan trọng của Glue, Kinesis, Lambda, Redshift
- Xây alerting cho pipeline: alarms, EventBridge, SNS
- Biết Logs Insights và subscription filters cho log pipeline
1. CloudWatch vs CloudTrail — phân vai
| CloudWatch | CloudTrail | |
|---|---|---|
| Trả lời câu hỏi | "Hệ thống hoạt động thế nào?" (metrics, logs, alarms) | "Ai đã làm gì, khi nào?" (API calls) |
| Dữ liệu | Metrics theo thời gian, log events | Event ghi lại mọi API call (console/CLI/SDK) |
| Use case | Hiệu năng, alert, dashboard | Audit, security investigation, compliance |
| Lưu trữ | Metrics 15 tháng; Logs theo retention | 90 ngày event history; lâu hơn → trail ra S3 |
Câu "xác định ai đã xóa bảng Glue / thay đổi bucket policy" → CloudTrail. Câu "job chạy chậm/fail, alert" → CloudWatch.
2. Metrics phải thuộc theo service
| Service | Metric | Ý nghĩa khi bất thường |
|---|---|---|
| Kinesis | GetRecords.IteratorAgeMilliseconds | Consumer tụt hậu |
| Kinesis | Write/ReadProvisionedThroughputExceeded | Thiếu shard / hot shard / thiếu EFO |
| Lambda | Errors, Throttles, Duration, IteratorAge | Lỗi code, thiếu concurrency, gần timeout |
| Glue | glue.driver.aggregate.numFailedTasks, memory/CPU per executor | Task fail, thiếu memory |
| Firehose | DeliveryToS3.Success, ThrottledRecords | Đích lỗi, quá tải |
| Redshift | QueryDuration, WLMQueueLength, CPU, PercentageDiskSpaceUsed | Queue nghẽn, hết disk |
| SQS | ApproximateAgeOfOldestMessage | Consumer không theo kịp |
| DMS | CDCLatencySource/Target | Replication lag |
3. Alerting pattern chuẩn cho pipeline
- Composite alarm: gộp nhiều alarm giảm noise
- Alarm trên math expression: vd error rate = Errors/Invocations
- Step Functions: state
Catch→ SNS ngay trong workflow (chủ động) + EventBridge ruleExecution Status Change(lưới an toàn)
4. CloudWatch Logs cho data pipeline
- Glue continuous logging, Lambda logs, Redshift audit logs... đều đổ về Logs
- Logs Insights: query log bằng ngôn ngữ riêng (
fields,filter,stats) — điều tra lỗi job nhanh - Subscription filter: stream log realtime → Kinesis/Firehose/Lambda/OpenSearch (log chính nó cũng là data pipeline!)
- Retention: mặc định không hết hạn — đặt retention và/hoặc export S3 để giảm chi phí (điểm cost hay thi)
5. CloudTrail chi tiết đáng nhớ
- Management events (mặc định): control plane — CreateTable, PutBucketPolicy...
- Data events (bật thêm, tốn phí): data plane —
s3:GetObject,lambda:Invokeper-resource → điều tra "ai đọc object nào" - Trail ghi liên tục ra S3 (+ CloudWatch Logs); organization trail cho multi-account
- CloudTrail Lake: query event bằng SQL (không cần tự dựng Athena trên trail)
- Athena query trail logs trên S3: pattern phân tích audit kinh điển
Câu hỏi ôn tập
-
Cần biết ai đã DROP một bảng trong Glue Data Catalog tuần trước. Dùng gì?
Xem đáp án
CloudTrail: tìm event
DeleteTable(management event) trong Event history (90 ngày) hoặc query trail trên S3/CloudTrail Lake — thấy identity, thời gian, source IP. CloudWatch không ghi "ai gọi API nào". -
Pipeline Kinesis → Lambda thỉnh thoảng "âm thầm" tụt hậu 30 phút. Đặt cảnh báo thế nào?
Xem đáp án
CloudWatch Alarm trên
GetRecords.IteratorAgeMillisecondscủa stream (vd > 5 phút trong 3 datapoint liên tiếp) → SNS. Kèm alarmErrors/Throttlescủa Lambda vì tụt hậu thường do lỗi lặp hoặc thiếu concurrency. Iterator age là "chỉ số vàng" cho độ trễ consumer. -
Muốn được thông báo mọi Glue job FAILED và Step Functions execution FAILED trong account. Cách gọn nhất?
Xem đáp án
EventBridge rules: (1) pattern
source: aws.glue, detail.state: FAILED→ SNS; (2) patternsource: aws.states, detail.status: FAILED→ SNS. Service tự phát event — không sửa job/workflow, phủ toàn account. Đây là pattern "notification on failure, no code change". -
Đếm số dòng log chứa "ERROR" trong log group của Glue và alert khi > 10 trong 5 phút. Làm sao?
Xem đáp án
Metric filter trên log group (pattern
ERROR) sinh custom metric → CloudWatch Alarm ngưỡng 10/5 phút → SNS. Logs Insights dùng để điều tra thủ công, còn metric filter là cách biến log thành alert tự động. -
Yêu cầu audit "ai đã đọc các object trong bucket PII" — cấu hình gì và lưu ý chi phí?
Xem đáp án
Bật CloudTrail data events cho S3 (chọn đúng bucket PII, event đọc
GetObject) — management events mặc định KHÔNG ghi data plane. Data events tính phí theo số event, bucket busy sẽ tốn — chỉ bật trên bucket cần audit. Query bằng Athena/CloudTrail Lake. (S3 server access logs là lựa chọn rẻ hơn nhưng ít chi tiết, delivery best-effort.)
Bài tập thực hành
- Tạo alarm IteratorAge cho stream lab tuần 3 → SNS email
- Tạo EventBridge rule Glue FAILED → SNS, chạy một job cố tình lỗi để test
- Dùng Logs Insights query log Lambda: đếm error theo 5 phút
- Bật CloudTrail data events cho 1 bucket và xem event GetObject xuất hiện
Tài liệu tham khảo chính thức
- CloudWatch concepts
- Monitoring Glue with CloudWatch metrics
- CloudTrail concepts
- CloudWatch Logs Insights
Tiếp theo: AWS Glue Data Quality