</>Học Dev
Bài học

Tuần 2 - Ngày 1: AWS Glue Data Catalog và Crawlers

Tuần 2 – Ngày 1

Mục tiêu học tập

  • Hiểu vai trò trung tâm của Glue Data Catalog trong hệ analytics AWS
  • Nắm cách crawler hoạt động: classifier, schema inference, partition detection
  • Biết cấu hình crawler đúng: schedule, schema change policy, incremental crawl
  • Hiểu Glue Schema Registry cho streaming

1. Glue Data Catalog — metadata trung tâm

Khái niệm

Glue Data Catalog là kho metadata (Hive Metastore-compatible) mô tả dữ liệu: database → table → (schema, location S3, định dạng, SerDe, partitions). Một catalog per region per account.

GlueDataCatalogDatabase:sales_dbTable:ordersColumns:idBIGINT,amountDOUBLE...Location:s3://lake/curated/orders/Format:Parquet(SerDe)Partitions:year=2026/month=01...cùngmtmetadatachoAthenaRedshiftEMRGlueETLSpectrum(Hive/Spark)

Điểm thi quan trọng: catalog chỉ chứa metadata — dữ liệu thật vẫn nằm trên S3 (hoặc JDBC source). Xóa table trong catalog không xóa dữ liệu.

Cách đưa metadata vào catalog

  1. Crawler — tự động dò schema (phổ biến nhất)
  2. Athena DDLCREATE EXTERNAL TABLE ...
  3. Glue API / CloudFormation / Terraform — infrastructure as code
  4. Glue ETL jobenableUpdateCatalog khi ghi output

2. Crawler hoạt động thế nào

Luồng

Data store (S3 / JDBC / DynamoDB / MongoDB...)
      │ (1) connect bằng IAM role / connection
      ▼
Classifiers: xác định format (built-in: Parquet, ORC, Avro,
      │       JSON, CSV, ...; custom: Grok, XML, JSON path, CSV)
      ▼
(2) Infer schema + phát hiện partition từ cấu trúc thư mục
      ▼
(3) Ghi/cập nhật table + partitions vào Data Catalog
  • Crawler so khớp similarity giữa các file/prefix: nếu schema tương tự → gộp thành 1 table với nhiều partition; nếu khác → nhiều table
  • Thư mục dạng key=value được nhận thành partition keys
  • Custom classifier chạy trước built-in, theo thứ tự khai báo

Schema change policy

Khi crawl lại và schema nguồn thay đổi:

Tuỳ chọnHành vi
Update the table definitionCập nhật schema trong catalog (mặc định)
Add new columns onlyChỉ thêm cột mới, không đổi/cột cũ
Ignore change, don't updateGiữ nguyên catalog
Object deletion: delete/deprecate/ignoreXử lý khi table biến mất khỏi nguồn

Incremental crawl và tối ưu chi phí crawler

  • Incremental crawls (crawl new folders only): chỉ quét thư mục mới — nhanh, rẻ cho dữ liệu append theo partition
  • S3 event-based crawler: crawler đọc S3 event notifications qua SQS để chỉ crawl object thay đổi — "least cost/latency" cho bucket lớn thay vì full re-crawl theo lịch
  • Thay thế crawler hoàn toàn: partition projection (Athena) hoặc Glue job tự ALTER TABLE ADD PARTITION

3. Connections — crawl/ETL nguồn ngoài S3

Glue Connection chứa thông tin kết nối (JDBC URL, VPC/subnet/security group, credentials từ Secrets Manager):

  • JDBC: RDS, Redshift, on-premises database (qua VPC/VPN/DX)
  • MongoDB, DynamoDB, Kafka (cho streaming job)
  • Khi Glue chạy trong VPC: cần security group self-referencing rule và VPC endpoint/NAT cho S3, Glue API — lỗi kết nối VPC là câu hỏi troubleshooting phổ biến

4. Glue Schema Registry (cho streaming)

  • Kho schema Avro/JSON Schema/Protobuf với versioning và compatibility mode (BACKWARD, FORWARD, FULL...)
  • Producer/consumer (MSK, Kinesis, Flink) serialize kèm schema ID — registry validate schema mới có tương thích không trước khi cho đăng ký
  • Khác Data Catalog: Schema Registry cho message streaming, Data Catalog cho dataset at-rest

5. Câu hỏi tình huống điển hình

Tình huốngĐáp án hướng tới
Partition mới mỗi giờ, Athena không thấy dữ liệu mớiCrawler schedule/event-based, MSCK REPAIR, hoặc partition projection
Crawler tạo hàng trăm table thay vì 1 tableCấu trúc S3 không đồng nhất; bật "Create a single schema for each S3 path" / chuẩn hóa layout
Crawl bucket 10TB mỗi ngày tốn kémIncremental crawl hoặc S3 event-based crawler
File log format lạ (custom text)Custom classifier (Grok pattern)
Cần schema validation cho Kafka messagesGlue Schema Registry với compatibility mode

Câu hỏi ôn tập

  1. Xóa một table trong Glue Data Catalog có xóa dữ liệu trên S3 không?

    Xem đáp án

    Không. Catalog chỉ lưu metadata (schema, location, partitions). Dữ liệu S3 độc lập hoàn toàn — có thể tạo lại table trỏ về cùng location. Ngược lại, xóa dữ liệu S3 cũng không tự cập nhật catalog (table "mồ côi" vẫn còn cho đến khi crawler/DDL dọn).

  2. Crawler chạy trên bucket có s3://lake/orders/year=2026/month=01/ sẽ tạo ra gì trong catalog?

    Xem đáp án

    Một table orders với partition keys year, month (nhận diện từ cấu trúc thư mục Hive-style key=value) và các partition tương ứng. Schema cột được infer từ nội dung file (Parquet tự mô tả; CSV/JSON qua classifier).

  3. Bucket rất lớn, dữ liệu mới rơi vào ngẫu nhiên nhiều prefix. Cách giữ catalog cập nhật với chi phí crawl thấp nhất?

    Xem đáp án

    S3 event-based crawler: bật S3 event notifications → SQS queue; crawler đọc queue và chỉ crawl các object/prefix thay đổi thay vì quét toàn bucket. Incremental crawl ("crawl new folders only") phù hợp khi dữ liệu mới luôn nằm trong folder mới; event-based tổng quát hơn và rẻ nhất cho bucket lớn.

  4. Khi nào cần custom classifier?

    Xem đáp án

    Khi built-in classifiers không nhận diện đúng format: log text tùy biến (dùng Grok pattern), XML với row tag cụ thể, JSON cần JSON path để lấy đúng phần tử, CSV có quy tắc đặc biệt (delimiter/heading lạ). Custom classifier được thử trước built-in theo thứ tự khai báo trong crawler.

  5. Glue Schema Registry khác Glue Data Catalog ở điểm nào?

    Xem đáp án

    Data Catalog: metadata của dataset at-rest (table trên S3/JDBC) phục vụ query engine. Schema Registry: quản lý version schema của message streaming (Avro/JSON Schema/Protobuf) cho Kafka/MSK/Kinesis/Flink, enforce compatibility mode (BACKWARD/FORWARD/FULL) khi schema tiến hóa, giúp producer/consumer không vỡ khi schema đổi.

  6. Glue job chạy trong VPC không kết nối được RDS. Những nguyên nhân phổ biến?

    Xem đáp án

    (1) Security group của connection thiếu self-referencing inbound rule (Glue ENI nói chuyện với nhau); (2) SG của RDS không allow từ SG của Glue; (3) subnet thiếu đường ra S3/Glue API — cần S3 gateway endpoint hoặc NAT; (4) credentials sai — nên dùng Secrets Manager. Kiểm tra bằng cách chạy connection test trong Glue console.

Bài tập thực hành

  • Tạo crawler trỏ vào bucket raw từ tuần 1, chạy và xem table sinh ra trong catalog
  • Thêm thư mục partition year=2026/month=02/ với file mới, chạy lại crawler ở chế độ incremental
  • Query bảng qua Athena, xác nhận partition được nhận
  • Đọc Scheduling a crawler to keep the Data Catalog in sync

Tài liệu tham khảo chính thức


Tiếp theo: Glue Jobs và Job Bookmarks