AWS msk: Replaced configuration table with MSK Data Delivery overview
Summary
Deleted supported configurations table and added feature descriptions for MSK Data Delivery capabilities.
Security assessment
Removed technical configuration details without introducing security-related content; new text focuses on feature capabilities.
Evidence
-Configuration | S3 Tables (Iceberg) | S3 bucket
Diff
diff --git a/msk/latest/developerguide/msk-data-delivery-supported-configs.md b/msk/latest/developerguide/msk-data-delivery-supported-configs.md index 7b17c4d23..61b71a0b5 100644 --- a//msk/latest/developerguide/msk-data-delivery-supported-configs.md +++ b//msk/latest/developerguide/msk-data-delivery-supported-configs.md @@ -1 +1 @@ -[View a markdown version of this page](msk-data-delivery-supported-configs.md) +[View a markdown version of this page](msk-data-delivery.md) @@ -3 +3 @@ -[](/pdfs/msk/latest/developerguide/MSKDevGuide.pdf#msk-data-delivery-supported-configs "Open PDF") +[](/pdfs/msk/latest/developerguide/MSKDevGuide.pdf#msk-data-delivery "Open PDF") @@ -7 +7,18 @@ -# Supported configurations +# Amazon MSK Data Delivery + +With Amazon MSK data delivery, you can deliver Apache Kafka data from Amazon MSK Express brokers directly to Amazon S3, without connectors or additional infrastructure to manage. Amazon MSK Express automatically handles scaling, retries, and backpressure, and manages routine operations such as capacity scaling and version upgrades without introducing delivery gaps. Because these are native broker capabilities, they add no broker egress throughput, so you avoid the incremental infrastructure costs that scaling connector-based pipelines typically incurs and match capacity to actual workload demand rather than provisioning for peak. Each capability supports throughput of up to 10 GBps. + +The two capabilities are: + + * **Data delivery to streaming tables for Apache Iceberg** — With Amazon MSK Data Delivery, you can continuously materialize Apache Kafka topics as Apache Iceberg tables on Amazon S3 Tables. Intelligent inline compaction eliminates the performance impact of small files and keeps query performance predictable without sacrificing data freshness. Built-in coordination resolves concurrent writer conflicts across high-throughput consumers. Amazon S3 Tables automatically handles ongoing table maintenance, including compaction, snapshot expiration, and unreferenced file cleanup. + + * **Data delivery to Amazon S3 general purpose buckets** — With Amazon MSK Data Delivery, you can deliver Apache Kafka data in the source format to Amazon S3 general purpose buckets for downstream processing, with end-to-end reliability for mission-critical workloads. Use it to land Kafka data in Amazon S3 for use cases such as log archival, compliance retention, Kafka replay, and training AI/ML models. This approach removes the need to build self-managed connector pipelines that grow costly and operationally complex as workloads scale. + + + + +###### Topics + + * [data delivery for streaming tables to Apache Iceberg](./msk-data-delivery-iceberg.html) + + * [data delivery to Amazon S3 general purpose buckets](./msk-data-delivery-s3.html) @@ -9,10 +25,0 @@ -Configuration | S3 Tables (Iceberg) | S3 bucket ----|---|--- -Cluster / broker type| Amazon MSK Provisioned with Express brokers only| Amazon MSK Provisioned with Express brokers only -Input format| JSON or JSON_SCHEMA_GSR| JSON, ByteArray, String -Output format| Apache Iceberg tables; Parquet files with ZSTD or Snappy compression| Objects (compression: NONE, GZIP, or ZSTD) -Schema source| AWS Glue Schema Registry (required)| Not required -Schema evolution| Not supported| Not applicable -Partitioning| Time-based (TIME_HOUR)| Object key template -Data freshness| 5–15 minutes (default 10)| 5–15 minutes (default 10) -Storage class| Managed by S3 Tables| STANDARD, INTELLIGENT_TIERING, GLACIER_IR @@ -20 +26,0 @@ Storage class| Managed by S3 Tables| STANDARD, INTELLIGENT_TIERING, GLACIER_IR -###### Note @@ -22 +27,0 @@ Storage class| Managed by S3 Tables| STANDARD, INTELLIGENT_TIERING, GLACIER_IR -**S3 Tables input formats:** `JSON` is plain JSON objects — you provide the Glue Schema Registry ARN that defines the schema. `JSON_SCHEMA_GSR` is GSR-serialized JSON, where the schema ID is embedded in each record. @@ -30 +35 @@ To use the Amazon Web Services Documentation, Javascript must be enabled. Please -Get started +Troubleshooting @@ -32 +37 @@ Get started -Prerequisites +data delivery for streaming tables to Apache Iceberg