AWS general: Updated Bedrock service quotas for multiple models
Summary
Adjusted quotas for AgenticRetrieveStream and Retrieve APIs; removed quotas for Claude 3.5 Sonnet models; added quotas for GPT-5.6 Luna/Sol/Terra models.
Security assessment
Changes involve service quota adjustments and model updates without security implications.
Evidence
+(Managed Knowledge Bases) AgenticRetrieveStream requests per minute per account | Each supported Region: 60 | No | The maximum number of AgenticRetrieveStream API requests per minute per account for Managed KBs.
Diff
diff --git a/general/latest/gr/bedrock.md b/general/latest/gr/bedrock.md index 37dd45d7f..529e43d73 100644 --- a//general/latest/gr/bedrock.md +++ b//general/latest/gr/bedrock.md @@ -358 +358 @@ Name | Default | Adjustable | Description -(Managed Knowledge Bases) AgenticRetrieveStream requests per minute per account | us-east-1: 60 eu-central-1: 60 eu-west-1: 60 Each of the other supported Regions: 1 | No | The maximum number of AgenticRetrieveStream API requests per minute per account for Managed KBs. +(Managed Knowledge Bases) AgenticRetrieveStream requests per minute per account | Each supported Region: 60 | No | The maximum number of AgenticRetrieveStream API requests per minute per account for Managed KBs. @@ -373 +373 @@ Name | Default | Adjustable | Description -(Managed Knowledge Bases) Retrieve requests per minute per knowledge base | us-east-1: 600 eu-central-1: 600 eu-west-1: 600 Each of the other supported Regions: 5 | No | The maximum number of Retrieve API requests per minute per Managed KB. +(Managed Knowledge Bases) Retrieve requests per minute per knowledge base | Each supported Region: 600 | No | The maximum number of Retrieve API requests per minute per Managed KB. @@ -448 +447,0 @@ Batch inference input file size (in GB) for Claude 3 Opus | Each supported Regio -Batch inference input file size (in GB) for Claude 3 Sonnet | Each supported Region: 1 | No | The maximum size of a single file (in GB) submitted for batch inference for Claude 3 Sonnet. @@ -450,2 +448,0 @@ Batch inference input file size (in GB) for Claude 3.5 Haiku | Each supported Re -Batch inference input file size (in GB) for Claude 3.5 Sonnet | Each supported Region: 1 | No | The maximum size of a single file (in GB) submitted for batch inference for Claude 3.5 Sonnet. -Batch inference input file size (in GB) for Claude 3.5 Sonnet v2 | Each supported Region: 1 | No | The maximum size of a single file (in GB) submitted for batch inference for Claude 3.5 Sonnet v2. @@ -520 +516,0 @@ Batch inference job size (in GB) for Claude 3 Opus | Each supported Region: 5 | -Batch inference job size (in GB) for Claude 3 Sonnet | Each supported Region: 5 | No | The maximum cumulative size of all input files (in GB) included in the batch inference job for Claude 3 Sonnet. @@ -522,2 +517,0 @@ Batch inference job size (in GB) for Claude 3.5 Haiku | Each supported Region: 5 -Batch inference job size (in GB) for Claude 3.5 Sonnet | Each supported Region: 5 | No | The maximum cumulative size of all input files (in GB) included in the batch inference job for Claude 3.5 Sonnet. -Batch inference job size (in GB) for Claude 3.5 Sonnet v2 | Each supported Region: 5 | No | The maximum cumulative size of all input files (in GB) included in the batch inference job for Claude 3.5 Sonnet v2. @@ -591 +584,0 @@ Cross-Region model inference requests per minute for Anthropic Claude 3.5 Haiku -Cross-Region model inference requests per minute for Anthropic Claude 3.5 Sonnet V2 | us-west-2: 500 Each of the other supported Regions: 100 | No | The maximum number of times that you can call model inference in one minute for Anthropic Claude 3.5 Sonnet V2. The quota considers the combined sum of Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. @@ -593 +585,0 @@ Cross-Region model inference tokens per minute for Anthropic Claude 3.5 Haiku | -Cross-Region model inference tokens per minute for Anthropic Claude 3.5 Sonnet V2 | us-west-2: 4,000,000 Each of the other supported Regions: 800,000 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-FF8B4E28) | The maximum number of tokens that you can submit for model inference in one minute for Anthropic Claude 3.5 Sonnet V2. The quota considers the combined sum of Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. @@ -603,2 +594,0 @@ Cross-region model inference requests per minute for Anthropic Claude 3 Opus | E -Cross-region model inference requests per minute for Anthropic Claude 3 Sonnet | us-west-2: 1,000 Each of the other supported Regions: 200 | No | The maximum number of times that you can call model inference in one minute. The quota considers the combined sum of Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream for Anthropic Claude 3 Sonnet. -Cross-region model inference requests per minute for Anthropic Claude 3.5 Sonnet | us-west-2: 500 ap-northeast-1: 40 ap-southeast-1: 40 Each of the other supported Regions: 100 | No | The maximum number of times that you can call model inference in one minute for Anthropic Claude 3.5 Sonnet. The quota considers the combined sum of Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. @@ -654,2 +643,0 @@ Cross-region model inference tokens per minute for Anthropic Claude 3 Opus | Eac -Cross-region model inference tokens per minute for Anthropic Claude 3 Sonnet | us-west-2: 2,000,000 Each of the other supported Regions: 400,000 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-5DF13F64) | The maximum number of on-demand tokens that you can submit for model inference in one minute. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream for Anthropic Claude 3 Sonnet. -Cross-region model inference tokens per minute for Anthropic Claude 3.5 Sonnet | us-west-2: 4,000,000 ap-northeast-1: 400,000 ap-southeast-1: 400,000 Each of the other supported Regions: 800,000 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-479B647F) | The maximum number of tokens that you can submit for model inference in one minute for Anthropic Claude 3.5 Sonnet. The quota considers the combined sum of Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. @@ -673,0 +662,3 @@ Cross-region model inference tokens per minute for DeepSeek R1 V1 | Each support +Cross-region model inference tokens per minute for GPT-5.6 Luna | Each supported Region: 20,000,000 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-E3C70889) | The maximum number of cross-region tokens that you can submit for model inference in one minute for GPT-5.6 Luna. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. +Cross-region model inference tokens per minute for GPT-5.6 Sol | Each supported Region: 10,000,000 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-D94ECE3B) | The maximum number of cross-region tokens that you can submit for model inference in one minute for GPT-5.6 Sol. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. +Cross-region model inference tokens per minute for GPT-5.6 Terra | Each supported Region: 20,000,000 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-99936FBF) | The maximum number of cross-region tokens that you can submit for model inference in one minute for GPT-5.6 Terra. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. @@ -724,0 +716,3 @@ Global cross-region model inference tokens per day for Cohere Embed V4 | Each su +Global cross-region model inference tokens per day for GPT-5.6 Luna | Each supported Region: 28,800,000,000 | No | The maximum number of global cross-region tokens that you can submit for model inference in one day for GPT-5.6 Luna. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. +Global cross-region model inference tokens per day for GPT-5.6 Sol | Each supported Region: 14,400,000,000 | No | The maximum number of global cross-region tokens that you can submit for model inference in one day for GPT-5.6 Sol. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. +Global cross-region model inference tokens per day for GPT-5.6 Terra | Each supported Region: 28,800,000,000 | No | The maximum number of global cross-region tokens that you can submit for model inference in one day for GPT-5.6 Terra. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. @@ -740,0 +735,3 @@ Global cross-region model inference tokens per minute for Cohere Embed V4 | Each +Global cross-region model inference tokens per minute for GPT-5.6 Luna | Each supported Region: 20,000,000 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-D983F7C0) | The maximum number of global cross-region tokens that you can submit for model inference in one minute for GPT-5.6 Luna. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. +Global cross-region model inference tokens per minute for GPT-5.6 Sol | Each supported Region: 10,000,000 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-4D7537CF) | The maximum number of global cross-region tokens that you can submit for model inference in one minute for GPT-5.6 Sol. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. +Global cross-region model inference tokens per minute for GPT-5.6 Terra | Each supported Region: 20,000,000 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-8E752C84) | The maximum number of global cross-region tokens that you can submit for model inference in one minute for GPT-5.6 Terra. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. @@ -752 +748,0 @@ Minimum number of records per batch inference job for Claude 3 Opus | Each suppo -Minimum number of records per batch inference job for Claude 3 Sonnet | Each supported Region: 100 | No | The minimum number of records across all input files in a batch inference job for Claude 3 Sonnet. @@ -754,2 +749,0 @@ Minimum number of records per batch inference job for Claude 3.5 Haiku | Each su -Minimum number of records per batch inference job for Claude 3.5 Sonnet | Each supported Region: 100 | No | The minimum number of records across all input files in a batch inference job for Claude 3.5 Sonnet. -Minimum number of records per batch inference job for Claude 3.5 Sonnet v2 | Each supported Region: 100 | No | The minimum number of records across all input files in a batch inference job for Claude 3.5 Sonnet v2. @@ -831,2 +824,0 @@ Model invocation max tokens per day for Anthropic Claude 3.5 Haiku (doubled for -Model invocation max tokens per day for Anthropic Claude 3.5 Sonnet V1 (doubled for cross-region calls) | Each supported Region: 2,880,000,000 | No | Daily maximum tokens for model inference for Anthropic Claude 3.5 Sonnet V1. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase. -Model invocation max tokens per day for Anthropic Claude 3.5 Sonnet V2 (doubled for cross-region calls) | us-west-2: 2,880,000,000 Each of the other supported Regions: 576,000,000 | No | Daily maximum tokens for model inference for Anthropic Claude 3.5 Sonnet V2. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase. @@ -854,0 +847,3 @@ Model invocation max tokens per day for GPT OSS Safeguard 20B (doubled for cross +Model invocation max tokens per day for GPT-5.6 Luna (doubled for cross-region calls) | Each supported Region: 14,400,000,000 | No | Daily maximum tokens for model inference for GPT-5.6 Luna. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase. +Model invocation max tokens per day for GPT-5.6 Sol (doubled for cross-region calls) | Each supported Region: 7,200,000,000 | No | Daily maximum tokens for model inference for GPT-5.6 Sol. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase. +Model invocation max tokens per day for GPT-5.6 Terra (doubled for cross-region calls) | Each supported Region: 14,400,000,000 | No | Daily maximum tokens for model inference for GPT-5.6 Terra. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase. @@ -917,2 +911,0 @@ Model units per provisioned model for Anthropic Claude 3 Haiku 48K | Each suppor -Model units per provisioned model for Anthropic Claude 3 Sonnet 200K | Each supported Region: 0 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-1F7657F1) | The maximum number of model units that can be allotted to a provisioned model for Anthropic Claude 3 Sonnet 200K. -Model units per provisioned model for Anthropic Claude 3 Sonnet 28K | Each supported Region: 0 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-B3C19043) | The maximum number of model units that can be allotted to a provisioned model for Anthropic Claude 3 Sonnet 28K. @@ -922,6 +914,0 @@ Model units per provisioned model for Anthropic Claude 3.5 Haiku 64K | Each supp -Model units per provisioned model for Anthropic Claude 3.5 Sonnet 18K | Each supported Region: 0 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-259C746F) | The maximum number of model units that can be allotted to a provisioned model for Anthropic Claude 3.5 Sonnet 18K. -Model units per provisioned model for Anthropic Claude 3.5 Sonnet 200K | Each supported Region: 0 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-2590C31B) | The maximum number of model units that can be allotted to a provisioned model for Anthropic Claude 3.5 Sonnet 200K. -Model units per provisioned model for Anthropic Claude 3.5 Sonnet 51K | Each supported Region: 0 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-208A3F5C) | The maximum number of model units that can be allotted to a provisioned model for Anthropic Claude 3.5 Sonnet 51K. -Model units per provisioned model for Anthropic Claude 3.5 Sonnet V2 18K | Each supported Region: 0 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-02710C34) | The maximum number of model units that can be allotted to a provisioned model for Anthropic Claude 3.5 Sonnet V2 18K. -Model units per provisioned model for Anthropic Claude 3.5 Sonnet V2 200K | Each supported Region: 0 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-24060791) | The maximum number of model units that can be allotted to a provisioned model for Anthropic Claude 3.5 Sonnet V2 200K. -Model units per provisioned model for Anthropic Claude 3.5 Sonnet V2 51K | Each supported Region: 0 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-B2718619) | The maximum number of model units that can be allotted to a provisioned model for Anthropic Claude 3.5 Sonnet V2 51K. @@ -1005 +991,0 @@ On-demand model inference requests per minute for Anthropic Claude 3 Opus | Each -On-demand model inference requests per minute for Anthropic Claude 3 Sonnet | us-west-2: 500 Each of the other supported Regions: 100 | No | The maximum number of times that you can call model inference in one minute. The quota considers the combined sum of Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream for Anthropic Claude 3 Sonnet. @@ -1007,2 +992,0 @@ On-demand model inference requests per minute for Anthropic Claude 3.5 Haiku | -On-demand model inference requests per minute for Anthropic Claude 3.5 Sonnet | us-west-2: 250 ap-southeast-2: 50 eu-central-2: 50 Each of the other supported Regions: 20 | No | The maximum number of times that you can call model inference in one minute for Anthropic Claude 3.5 Sonnet. The quota considers the combined sum of Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. -On-demand model inference requests per minute for Anthropic Claude 3.5 Sonnet V2 | us-west-2: 250 Each of the other supported Regions: 50 | No | The maximum number of times that you can call model inference in one minute for Anthropic Claude 3.5 Sonnet V2. The quota considers the combined sum of Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. @@ -1104 +1087,0 @@ On-demand model inference tokens per minute for Anthropic Claude 3 Opus | Each s -On-demand model inference tokens per minute for Anthropic Claude 3 Sonnet | us-west-2: 1,000,000 Each of the other supported Regions: 200,000 | No | The maximum number of on-demand tokens that you can submit for model inference in one minute. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream for Anthropic Claude 3 Sonnet. @@ -1106,2 +1088,0 @@ On-demand model inference tokens per minute for Anthropic Claude 3.5 Haiku | us -On-demand model inference tokens per minute for Anthropic Claude 3.5 Sonnet | us-west-2: 2,000,000 ap-southeast-2: 400,000 eu-central-2: 400,000 Each of the other supported Regions: 200,000 | No | The maximum number of tokens that you can submit for model inference in one minute for Anthropic Claude 3.5 Sonnet. The quota considers the combined sum of Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. -On-demand model inference tokens per minute for Anthropic Claude 3.5 Sonnet V2 | us-west-2: 2,000,000 Each of the other supported Regions: 400,000 | No | The maximum number of tokens that you can submit for model inference in one minute for Anthropic Claude 3.5 Sonnet V2. The quota considers the combined sum of Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. @@ -1174 +1154,0 @@ Records per batch inference job for Claude 3 Opus | Each supported Region: 100,0 -Records per batch inference job for Claude 3 Sonnet | Each supported Region: 100,000 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-16E25672) | The maximum number of records across all input files in a batch inference job for Claude 3 Sonnet. @@ -1176,2 +1155,0 @@ Records per batch inference job for Claude 3.5 Haiku | Each supported Region: 10 -Records per batch inference job for Claude 3.5 Sonnet | Each supported Region: 100,000 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-1E2B9998) | The maximum number of records across all input files in a batch inference job for Claude 3.5 Sonnet. -Records per batch inference job for Claude 3.5 Sonnet v2 | Each supported Region: 100,000 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-6EBFEB27) | The maximum number of records across all input files in a batch inference job for Claude 3.5 Sonnet v2. @@ -1245 +1222,0 @@ Records per input file per batch inference job for Claude 3 Opus | Each supporte -Records per input file per batch inference job for Claude 3 Sonnet | Each supported Region: 100,000 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-E93C745B) | The maximum number of records in an input file in a batch inference job for Claude 3 Sonnet. @@ -1247,2 +1223,0 @@ Records per input file per batch inference job for Claude 3.5 Haiku | Each suppo -Records per input file per batch inference job for Claude 3.5 Sonnet | Each supported Region: 100,000 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-5AB0EE48) | The maximum number of records in an input file in a batch inference job for Claude 3.5 Sonnet. -Records per input file per batch inference job for Claude 3.5 Sonnet v2 | Each supported Region: 100,000 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-897F8151) | The maximum number of records in an input file in a batch inference job for Claude 3.5 Sonnet v2. @@ -1316 +1290,0 @@ Sum of in-progress and submitted batch inference jobs using a base model for Cla -Sum of in-progress and submitted batch inference jobs using a base model for Claude 3 Sonnet | Each supported Region: 100 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-67BD0D49) | The maximum number of in-progress and submitted batch inference jobs using a base model for Claude 3 Sonnet. @@ -1318,2 +1291,0 @@ Sum of in-progress and submitted batch inference jobs using a base model for Cla -Sum of in-progress and submitted batch inference jobs using a base model for Claude 3.5 Sonnet | Each supported Region: 100 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-4E7EE0B5) | The maximum number of in-progress and submitted batch inference jobs using a base model for Claude 3.5 Sonnet. -Sum of in-progress and submitted batch inference jobs using a base model for Claude 3.5 Sonnet v2 | Each supported Region: 100 | [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-C2FA9AEC) | The maximum number of in-progress and submitted batch inference jobs using a base model for Claude 3.5 Sonnet v2.