</>Học Dev
Bài học

Tuần 6 - Ngày 1: Monitoring pipeline với CloudWatch và CloudTrail

Tuần 6 – Ngày 1

Mục tiêu học tập

  • Phân biệt CloudWatch (metrics/logs/alarms) và CloudTrail (API audit)
  • Thuộc các metric quan trọng của Glue, Kinesis, Lambda, Redshift
  • Xây alerting cho pipeline: alarms, EventBridge, SNS
  • Biết Logs Insights và subscription filters cho log pipeline

1. CloudWatch vs CloudTrail — phân vai

CloudWatchCloudTrail
Trả lời câu hỏi"Hệ thống hoạt động thế nào?" (metrics, logs, alarms)"Ai đã làm gì, khi nào?" (API calls)
Dữ liệuMetrics theo thời gian, log eventsEvent ghi lại mọi API call (console/CLI/SDK)
Use caseHiệu năng, alert, dashboardAudit, security investigation, compliance
Lưu trữMetrics 15 tháng; Logs theo retention90 ngày event history; lâu hơn → trail ra S3

Câu "xác định ai đã xóa bảng Glue / thay đổi bucket policy" → CloudTrail. Câu "job chạy chậm/fail, alert" → CloudWatch.

2. Metrics phải thuộc theo service

ServiceMetricÝ nghĩa khi bất thường
KinesisGetRecords.IteratorAgeMillisecondsConsumer tụt hậu
KinesisWrite/ReadProvisionedThroughputExceededThiếu shard / hot shard / thiếu EFO
LambdaErrors, Throttles, Duration, IteratorAgeLỗi code, thiếu concurrency, gần timeout
Glueglue.driver.aggregate.numFailedTasks, memory/CPU per executorTask fail, thiếu memory
FirehoseDeliveryToS3.Success, ThrottledRecordsĐích lỗi, quá tải
RedshiftQueryDuration, WLMQueueLength, CPU, PercentageDiskSpaceUsedQueue nghẽn, hết disk
SQSApproximateAgeOfOldestMessageConsumer không theo kịp
DMSCDCLatencySource/TargetReplication lag

3. Alerting pattern chuẩn cho pipeline

MetricCloudWatchAlarmSNSemail/Slack/PagerDutyEventEventBridgerule(GlueFAILED,StepFunctionsFAILED,DMSstatechange)SNS/Lambda/ticketLogMetricfilter(đếm"ERROR")AlarmSNS
  • Composite alarm: gộp nhiều alarm giảm noise
  • Alarm trên math expression: vd error rate = Errors/Invocations
  • Step Functions: state Catch → SNS ngay trong workflow (chủ động) + EventBridge rule Execution Status Change (lưới an toàn)

4. CloudWatch Logs cho data pipeline

  • Glue continuous logging, Lambda logs, Redshift audit logs... đều đổ về Logs
  • Logs Insights: query log bằng ngôn ngữ riêng (fields, filter, stats) — điều tra lỗi job nhanh
  • Subscription filter: stream log realtime → Kinesis/Firehose/Lambda/OpenSearch (log chính nó cũng là data pipeline!)
  • Retention: mặc định không hết hạn — đặt retention và/hoặc export S3 để giảm chi phí (điểm cost hay thi)

5. CloudTrail chi tiết đáng nhớ

  • Management events (mặc định): control plane — CreateTable, PutBucketPolicy...
  • Data events (bật thêm, tốn phí): data plane — s3:GetObject, lambda:Invoke per-resource → điều tra "ai đọc object nào"
  • Trail ghi liên tục ra S3 (+ CloudWatch Logs); organization trail cho multi-account
  • CloudTrail Lake: query event bằng SQL (không cần tự dựng Athena trên trail)
  • Athena query trail logs trên S3: pattern phân tích audit kinh điển

Câu hỏi ôn tập

  1. Cần biết ai đã DROP một bảng trong Glue Data Catalog tuần trước. Dùng gì?

    Xem đáp án

    CloudTrail: tìm event DeleteTable (management event) trong Event history (90 ngày) hoặc query trail trên S3/CloudTrail Lake — thấy identity, thời gian, source IP. CloudWatch không ghi "ai gọi API nào".

  2. Pipeline Kinesis → Lambda thỉnh thoảng "âm thầm" tụt hậu 30 phút. Đặt cảnh báo thế nào?

    Xem đáp án

    CloudWatch Alarm trên GetRecords.IteratorAgeMilliseconds của stream (vd > 5 phút trong 3 datapoint liên tiếp) → SNS. Kèm alarm Errors/Throttles của Lambda vì tụt hậu thường do lỗi lặp hoặc thiếu concurrency. Iterator age là "chỉ số vàng" cho độ trễ consumer.

  3. Muốn được thông báo mọi Glue job FAILED và Step Functions execution FAILED trong account. Cách gọn nhất?

    Xem đáp án

    EventBridge rules: (1) pattern source: aws.glue, detail.state: FAILED → SNS; (2) pattern source: aws.states, detail.status: FAILED → SNS. Service tự phát event — không sửa job/workflow, phủ toàn account. Đây là pattern "notification on failure, no code change".

  4. Đếm số dòng log chứa "ERROR" trong log group của Glue và alert khi > 10 trong 5 phút. Làm sao?

    Xem đáp án

    Metric filter trên log group (pattern ERROR) sinh custom metric → CloudWatch Alarm ngưỡng 10/5 phút → SNS. Logs Insights dùng để điều tra thủ công, còn metric filter là cách biến log thành alert tự động.

  5. Yêu cầu audit "ai đã đọc các object trong bucket PII" — cấu hình gì và lưu ý chi phí?

    Xem đáp án

    Bật CloudTrail data events cho S3 (chọn đúng bucket PII, event đọc GetObject) — management events mặc định KHÔNG ghi data plane. Data events tính phí theo số event, bucket busy sẽ tốn — chỉ bật trên bucket cần audit. Query bằng Athena/CloudTrail Lake. (S3 server access logs là lựa chọn rẻ hơn nhưng ít chi tiết, delivery best-effort.)

Bài tập thực hành

  • Tạo alarm IteratorAge cho stream lab tuần 3 → SNS email
  • Tạo EventBridge rule Glue FAILED → SNS, chạy một job cố tình lỗi để test
  • Dùng Logs Insights query log Lambda: đếm error theo 5 phút
  • Bật CloudTrail data events cho 1 bucket và xem event GetObject xuất hiện

Tài liệu tham khảo chính thức


Tiếp theo: AWS Glue Data Quality