AWS clean-rooms: Updated Spark configuration parameters in Clean Rooms analysis template
Summary
Corrected parameter names and added new Spark configuration options including cross-join controls and memory safeguards.
Security assessment
The evidence line describes a resource management improvement to prevent memory exhaustion, which is a hardening measure but not tied to a specific vulnerability.
Evidence
+spark.maxRemoteBlockSizeFetchToMem | Sets the size threshold above which Spark fetches remote blocks to disk instead of memory. This avoids a single large request consuming too much memory. | 200m
Diff
diff --git a/clean-rooms/latest/userguide/use-analysis-template.md b/clean-rooms/latest/userguide/use-analysis-template.md index 59d85657a..8ccdbe08b 100644 --- a//clean-rooms/latest/userguide/use-analysis-template.md +++ b//clean-rooms/latest/userguide/use-analysis-template.md @@ -129 +129 @@ spark.sql.legacy.json.allowEmptyString.enabled | Specifies whether to allow emp -spark.sql.legacy.parquet.int96RebaseModelRead | Specifies whether to use legacy INT96 timestamp rebase mode when reading Parquet files. This legacy setting maintains compatibility with older Spark versions' timestamp handling. | (none) +spark.sql.legacy.parquet.int96RebaseModeInRead | Specifies whether to use legacy INT96 timestamp rebase mode when reading Parquet files. This legacy setting maintains compatibility with older Spark versions' timestamp handling. | (none) @@ -155 +155 @@ spark.sql.optimizer.runtime.bloomFilter.numBits | Defines the default number of -spark.sql.optimizer.runtime.rowlevelOperationGroupFilter.enabled | Specifies whether to enable runtime group filtering for row-level operations. Allows data sources to: +spark.sql.optimizer.runtime.rowLevelOperationGroupFilter.enabled | Specifies whether to enable runtime group filtering for row-level operations. Allows data sources to: @@ -180,0 +181,7 @@ spark.sql.hive.metastorePartitionPruningFallbackOnException | Specifies whether +spark.sql.crossJoin.enabled | Specifies whether to allow queries that contain a cartesian product without explicit CROSS JOIN syntax. | true +spark.sql.analyzer.maxIterations | Sets the maximum number of iterations the query analyzer runs before giving up. Higher values allow the analyzer to process very large or deeply nested queries. | 100 +spark.sql.dataprefetch.filescan.maxParallelismPerTask | Sets the maximum number of file splits to pre-fetch concurrently for each task when scanning files. | 4 +spark.sql.iceberg.data-prefetch.enabled | Specifies whether to enable data pre-fetch optimization when reading Iceberg tables. | true +spark.sql.legacy.nullValueWrittenAsQuotedEmptyStringCsv | Specifies whether to restore the legacy behavior of writing nulls as quoted empty strings in CSV output. When false, Spark writes nulls as unquoted empty strings. | false +spark.maxRemoteBlockSizeFetchToMem | Sets the size threshold above which Spark fetches remote blocks to disk instead of memory. This avoids a single large request consuming too much memory. | 200m +spark.emr-serverless.allocation.batch.size | Sets the number of executors to request at once in each round of executor allocation. | 20 @@ -185,0 +193 @@ spark.dynamicAllocation.enabled | Specifies whether to use dynamic resource all +spark.files.fetchFailure.unRegisterOutputOnHost | Specifies whether to unregister all map outputs on a host when a fetch failure occurs. When false, Spark unregisters only the outputs of the specific executor that failed, which reduces unnecessary stage recomputation. | false