AWS Security ChangesHomeSearch

AWS general: Bedrock quotas table: model additions, removals, and adjustable flags

Service: general · 2026-09-27 · Documentation low

File: general/latest/gr/bedrock.md

Summary

Updates the Amazon Bedrock service quotas table: marks several Guardrails ApplyGuardrail burst-rate quotas as adjustable (adding Service Quotas links), adds new model quotas (GPT-6 Astra/Luna/Sol, Moonshot AI Kimi K3, Claude Opus 5.5, Cross-Model Account-Level Tokens Per Day, bedrock-mantle endpoint input/output tokens), removes many deprecated model entries (Claude 3 Opus, Sonnet 4, Fable, DeepSeek, Qwen, etc.), and adjusts per-region values for Claude 3 Haiku and Claude 3.7 Sonnet.

Security assessment

The diff only modifies service quota limits, region values, adjustable flags, and model listings in the Bedrock quotas reference table. No authentication, encryption, credential, IAM, or vulnerability content is present; quota changes are operational/capacity related, not security fixes.

Evidence

+Cross-Model Account-Level Tokens Per Day | Each supported Region: 700,000,000 | No | Maximum total threshold for estimated billed tokens across all models based on on-demand inference pricing; actual usage will depend on factors including model pricing, input/output token ratio, and cache hit ratio.

Diff

diff --git a/general/latest/gr/bedrock.md b/general/latest/gr/bedrock.md
index ccb0d6d15..557bd5f1b 100644
--- a//general/latest/gr/bedrock.md
+++ b//general/latest/gr/bedrock.md
@@ -293,2 +293,2 @@ Name | Default | Adjustable | Description
-(Guardrails) On-demand ApplyGuardrail Content filter policy text units burst rate (Classic tier) |  us-east-1: 1,000 us-east-2: 1,000 us-west-2: 1,000 ap-northeast-1: 1,000 ap-northeast-2: 1,000 ap-south-1: 1,000 ap-southeast-1: 1,000 ap-southeast-2: 1,000 eu-central-1: 1,000 Each of the other supported Regions: 25 | No | The maximum number of text units in one burst that can be processed for content filters. While this limit applies to the classic tier, we recommend migrating to standard tier due to its superior robustness, additional capabilities, and multi-lingual support.  
-(Guardrails) On-demand ApplyGuardrail Content filter policy text units burst rate (Standard tier - Recommended) |  us-east-1: 1,000 us-east-2: 1,000 us-west-2: 1,000 ap-northeast-1: 500 ap-northeast-2: 1,000 ap-south-1: 500 ap-southeast-1: 1,000 ap-southeast-2: 400 eu-central-1: 500 eu-west-1: 1,000 Each of the other supported Regions: 25 | No | The maximum number of text units in one burst that can be processed for content filters. This applies to the standard tier, which is recommended.  
+(Guardrails) On-demand ApplyGuardrail Content filter policy text units burst rate (Classic tier) |  us-east-1: 1,000 us-east-2: 1,000 us-west-2: 1,000 ap-northeast-1: 1,000 ap-northeast-2: 1,000 ap-south-1: 1,000 ap-southeast-1: 1,000 ap-southeast-2: 1,000 eu-central-1: 1,000 Each of the other supported Regions: 25 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-4A0EAD70) | The maximum number of text units in one burst that can be processed for content filters. While this limit applies to the classic tier, we recommend migrating to standard tier due to its superior robustness, additional capabilities, and multi-lingual support.  
+(Guardrails) On-demand ApplyGuardrail Content filter policy text units burst rate (Standard tier - Recommended) |  us-east-1: 1,000 us-east-2: 1,000 us-west-2: 1,000 ap-northeast-1: 500 ap-northeast-2: 1,000 ap-south-1: 500 ap-southeast-1: 1,000 ap-southeast-2: 400 eu-central-1: 500 eu-west-1: 1,000 Each of the other supported Regions: 25 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-9807BAA2) | The maximum number of text units in one burst that can be processed for content filters. This applies to the standard tier, which is recommended.  
@@ -297,2 +297,2 @@ Name | Default | Adjustable | Description
-(Guardrails) On-demand ApplyGuardrail Denied topic policy text units burst rate (Classic tier) |  us-east-1: 200 us-west-2: 200 Each of the other supported Regions: 25 | No | The maximum number of text units in one burst that can be processed for denied topics. While this limit applies to the classic tier, we recommend migrating to standard tier due to its superior robustness, additional capabilities, and multi-lingual support.  
-(Guardrails) On-demand ApplyGuardrail Denied topic policy text units burst rate (Standard tier - Recommended) |  us-east-1: 1,000 us-east-2: 1,000 us-west-2: 1,000 ap-northeast-1: 500 ap-northeast-2: 1,000 ap-south-1: 500 ap-southeast-1: 1,000 ap-southeast-2: 400 eu-central-1: 500 eu-west-1: 1,000 Each of the other supported Regions: 25 | No | The maximum number of text units in one burst that can be processed for denied topics. This applies to the standard tier, which is recommended.  
+(Guardrails) On-demand ApplyGuardrail Denied topic policy text units burst rate (Classic tier) |  us-east-1: 200 us-west-2: 200 Each of the other supported Regions: 25 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-AFD14A53) | The maximum number of text units in one burst that can be processed for denied topics. While this limit applies to the classic tier, we recommend migrating to standard tier due to its superior robustness, additional capabilities, and multi-lingual support.  
+(Guardrails) On-demand ApplyGuardrail Denied topic policy text units burst rate (Standard tier - Recommended) |  us-east-1: 1,000 us-east-2: 1,000 us-west-2: 1,000 ap-northeast-1: 500 ap-northeast-2: 1,000 ap-south-1: 500 ap-southeast-1: 1,000 ap-southeast-2: 400 eu-central-1: 500 eu-west-1: 1,000 Each of the other supported Regions: 25 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-7ABD27EC) | The maximum number of text units in one burst that can be processed for denied topics. This applies to the standard tier, which is recommended.  
@@ -301 +301 @@ Name | Default | Adjustable | Description
-(Guardrails) On-demand ApplyGuardrail Sensitive information filter policy text units burst rate |  us-east-1: 1,000 us-east-2: 1,000 us-west-2: 1,000 ap-northeast-1: 1,000 ap-northeast-2: 1,000 ap-south-1: 1,000 ap-southeast-1: 1,000 ap-southeast-2: 1,000 eu-central-1: 1,000 Each of the other supported Regions: 25 | No | The maximum number of text units in one burst that can be processed for sensitive information filters.  
+(Guardrails) On-demand ApplyGuardrail Sensitive information filter policy text units burst rate |  us-east-1: 1,000 us-east-2: 1,000 us-west-2: 1,000 ap-northeast-1: 1,000 ap-northeast-2: 1,000 ap-south-1: 1,000 ap-southeast-1: 1,000 ap-southeast-2: 1,000 eu-central-1: 1,000 Each of the other supported Regions: 25 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-CC3FC244) | The maximum number of text units in one burst that can be processed for sensitive information filters.  
@@ -303 +303 @@ Name | Default | Adjustable | Description
-(Guardrails) On-demand ApplyGuardrail Word filter policy text units burst rate |  us-east-1: 1,000 us-east-2: 1,000 us-west-2: 1,000 ap-northeast-1: 1,000 ap-northeast-2: 1,000 ap-south-1: 1,000 ap-southeast-1: 1,000 ap-southeast-2: 1,000 eu-central-1: 1,000 Each of the other supported Regions: 25 | No | The maximum number of text units in one burst that can be processed for word filters.  
+(Guardrails) On-demand ApplyGuardrail Word filter policy text units burst rate |  us-east-1: 1,000 us-east-2: 1,000 us-west-2: 1,000 ap-northeast-1: 1,000 ap-northeast-2: 1,000 ap-south-1: 1,000 ap-southeast-1: 1,000 ap-southeast-2: 1,000 eu-central-1: 1,000 Each of the other supported Regions: 25 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-3CD2AF6F) | The maximum number of text units in one burst that can be processed for word filters.  
@@ -305 +305 @@ Name | Default | Adjustable | Description
-(Guardrails) On-demand ApplyGuardrail contextual grounding policy text units burst rate | Each supported Region: 106 | No | The maximum number of text units in one burst that can be processed for contextual grounding.  
+(Guardrails) On-demand ApplyGuardrail contextual grounding policy text units burst rate | Each supported Region: 106 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-F86AE9DD) | The maximum number of text units in one burst that can be processed for contextual grounding.  
@@ -307 +307 @@ Name | Default | Adjustable | Description
-(Guardrails) On-demand ApplyGuardrail requests burst rate |  us-east-1: 100 us-east-2: 100 us-west-1: 100 us-west-2: 100 ap-northeast-1: 100 ap-northeast-2: 100 ap-south-1: 100 ap-southeast-1: 100 eu-central-1: 100 Each of the other supported Regions: 25 | No | The maximum number of ApplyGuardrail API calls that you can send in one burst.  
+(Guardrails) On-demand ApplyGuardrail requests burst rate |  us-east-1: 100 us-east-2: 100 us-west-1: 100 us-west-2: 100 ap-northeast-1: 100 ap-northeast-2: 100 ap-south-1: 100 ap-southeast-1: 100 eu-central-1: 100 Each of the other supported Regions: 25 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-C07B5EA2) | The maximum number of ApplyGuardrail API calls that you can send in one burst.  
@@ -309 +309 @@ Name | Default | Adjustable | Description
-(Guardrails) On-demand InvokeGuardrailChecks requests burst rate | Each supported Region: 1,500 | No | The maximum number of InvokeGuardrailChecks API calls that you can send in one burst  
+(Guardrails) On-demand InvokeGuardrailChecks requests burst rate | Each supported Region: 1,500 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-70AC14D9) | The maximum number of InvokeGuardrailChecks API calls that you can send in one burst  
@@ -385 +384,0 @@ Name | Default | Adjustable | Description
-(Model customization) Maximum student model fine tuning context length for Anthropic Claude 3 haiku 20240307 V1 distillation customization jobs | Each supported Region: 32,000 | No | The maximum student model fine tuning context length for Anthropic Claude 3 haiku 20240307 V1 distillation customization jobs.  
@@ -409 +407,0 @@ Name | Default | Adjustable | Description
-(Model customization) Sum of training and validation records for a Claude 3 Haiku v1 Fine-tuning job | Each supported Region: 10,000 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-679179D2) | The maximum combined number of training and validation records allowed for a Claude 3 Haiku Fine-tuning job.  
@@ -447 +444,0 @@ Batch inference input file size (in GB) for Claude 3 Haiku | Each supported Regi
-Batch inference input file size (in GB) for Claude 3 Opus | Each supported Region: 1 | No | The maximum size of a single file (in GB) submitted for batch inference for Claude 3 Opus.  
@@ -457 +453,0 @@ Batch inference input file size (in GB) for Claude Opus 5 | Each supported Regio
-Batch inference input file size (in GB) for Claude Sonnet 4 | Each supported Region: 1 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-F611997D) | The maximum size of a single file (in GB) submitted for batch inference for Claude Sonnet 4.  
@@ -519 +514,0 @@ Batch inference job size (in GB) for Claude 3 Haiku | Each supported Region: 5 |
-Batch inference job size (in GB) for Claude 3 Opus | Each supported Region: 5 | No | The maximum cumulative size of all input files (in GB) included in the batch inference job for Claude 3 Opus.  
@@ -529 +523,0 @@ Batch inference job size (in GB) for Claude Opus 5 | Each supported Region: 5 |
-Batch inference job size (in GB) for Claude Sonnet 4 | Each supported Region: 5 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-E31283B0) | The maximum cumulative size of all input files (in GB) included in the batch inference job for Claude Sonnet 4.  
@@ -589,0 +584 @@ CreateAgentAlias requests per second | Each supported Region: 2 | No | The maxim
+Cross-Model Account-Level Tokens Per Day | Each supported Region: 700,000,000 | No | Maximum total threshold for estimated billed tokens across all models based on on-demand inference pricing; actual usage will depend on factors including model pricing, input/output token ratio, and cache hit ratio.  
@@ -601,2 +596 @@ Cross-region model inference requests per minute for Amazon Nova Pro | Each supp
-Cross-region model inference requests per minute for Anthropic Claude 3 Haiku |  us-east-1: 2,000 us-west-2: 2,000 ap-northeast-1: 400 ap-southeast-1: 400 Each of the other supported Regions: 800 | No | The maximum number of times that you can call model inference in one minute. The quota considers the combined sum of Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream for Anthropic Claude 3 Haiku.  
-Cross-region model inference requests per minute for Anthropic Claude 3 Opus | Each supported Region: 100 | No | The maximum number of cross-region requests that you can submit for model inference in one minute for Anthropic Claude 3 Opus. This quota represents the total number of requests to the Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream actions.  
+Cross-region model inference requests per minute for Anthropic Claude 3 Haiku |  ap-southeast-1: 400 Each of the other supported Regions: 800 | No | The maximum number of times that you can call model inference in one minute. The quota considers the combined sum of Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream for Anthropic Claude 3 Haiku.  
@@ -605 +599 @@ Cross-region model inference requests per minute for Anthropic Claude 3.5 Sonnet
-Cross-region model inference requests per minute for Anthropic Claude 3.7 Sonnet V1 |  us-east-1: 250 us-east-2: 250 us-west-2: 250 eu-central-1: 100 eu-north-1: 100 eu-west-1: 100 eu-west-3: 100 Each of the other supported Regions: 50 | No | The maximum number of cross-region requests that you can submit for model inference in one minute for Anthropic Claude 3.7 Sonnet V1. This quota represents the total number of requests to the Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream actions.  
+Cross-region model inference requests per minute for Anthropic Claude 3.7 Sonnet V1 | Each supported Region: 50 | No | The maximum number of cross-region requests that you can submit for model inference in one minute for Anthropic Claude 3.7 Sonnet V1. This quota represents the total number of requests to the Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream actions.  
@@ -607,2 +600,0 @@ Cross-region model inference requests per minute for Anthropic Claude Haiku 4.5
-Cross-region model inference requests per minute for Anthropic Claude Opus 4 V1 | Each supported Region: 200 | No | The maximum number of cross-region requests that you can submit for model inference in one minute for Anthropic Claude Opus 4 V1. This quota represents the total number of requests to the Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream actions.  
-Cross-region model inference requests per minute for Anthropic Claude Opus 4.1 | Each supported Region: 50 | No | The maximum number of cross-region requests that you can submit for model inference in one minute for Anthropic Claude Opus 4.1. This quota represents the total number of requests to the Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream actions.  
@@ -611,2 +602,0 @@ Cross-region model inference requests per minute for Anthropic Claude Opus 4.6 V
-Cross-region model inference requests per minute for Anthropic Claude Sonnet 4 V1 | Each supported Region: 200 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-559DCC33) | The maximum number of cross-region requests that you can submit for model inference in one minute for Anthropic Claude Sonnet 4 V1. This quota represents the total number of requests to the Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream actions.  
-Cross-region model inference requests per minute for Anthropic Claude Sonnet 4 V1 1M Context Length | Each supported Region: 5 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-534E5E05) | The maximum number of cross-region requests that you can submit for model inference in one minute for Anthropic Claude Sonnet 4 V1 1M Context Length. This quota represents the total number of requests to the Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream actions.  
@@ -652,2 +642 @@ Cross-region model inference tokens per minute for Amazon Nova Pro | Each suppor
-Cross-region model inference tokens per minute for Anthropic Claude 3 Haiku |  us-east-1: 4,000,000 us-west-2: 4,000,000 ap-northeast-1: 400,000 ap-southeast-1: 400,000 Each of the other supported Regions: 600,000 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-DCADBC78) | The maximum number of on-demand tokens that you can submit for model inference in one minute. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream for Anthropic Claude 3 Haiku.  
-Cross-region model inference tokens per minute for Anthropic Claude 3 Opus | Each supported Region: 800,000 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-6C86825E) | The maximum number of cross-region tokens that you can submit for model inference in one minute for Anthropic Claude 3 Opus. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
+Cross-region model inference tokens per minute for Anthropic Claude 3 Haiku |  ap-southeast-1: 400,000 Each of the other supported Regions: 600,000 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-DCADBC78) | The maximum number of on-demand tokens that you can submit for model inference in one minute. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream for Anthropic Claude 3 Haiku.  
@@ -656 +645 @@ Cross-region model inference tokens per minute for Anthropic Claude 3.5 Sonnet |
-Cross-region model inference tokens per minute for Anthropic Claude 3.7 Sonnet V1 |  us-east-1: 1,000,000 us-east-2: 1,000,000 us-west-2: 1,000,000 eu-central-1: 100,000 eu-north-1: 100,000 eu-west-1: 100,000 eu-west-3: 100,000 Each of the other supported Regions: 50,000 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-6E888CC2) | The maximum number of cross-region tokens that you can submit for model inference in one minute for Anthropic Claude 3.7 Sonnet V1. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
+Cross-region model inference tokens per minute for Anthropic Claude 3.7 Sonnet V1 | Each supported Region: 50,000 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-6E888CC2) | The maximum number of cross-region tokens that you can submit for model inference in one minute for Anthropic Claude 3.7 Sonnet V1. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
@@ -660,2 +648,0 @@ Cross-region model inference tokens per minute for Anthropic Claude Haiku 4.5 |
-Cross-region model inference tokens per minute for Anthropic Claude Opus 4 V1 | Each supported Region: 200,000 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-29C2B0A3) | The maximum number of cross-region tokens that you can submit for model inference in one minute for Anthropic Claude Opus 4 V1. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
-Cross-region model inference tokens per minute for Anthropic Claude Opus 4.1 | Each supported Region: 500,000 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-BD85BFCD) | The maximum number of cross-region tokens that you can submit for model inference in one minute for Anthropic Claude Opus 4.1. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
@@ -667,2 +654 @@ Cross-region model inference tokens per minute for Anthropic Claude Opus 5 | Eac
-Cross-region model inference tokens per minute for Anthropic Claude Sonnet 4 V1 | Each supported Region: 200,000 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-59759B4A) | The maximum number of cross-region tokens that you can submit for model inference in one minute for Anthropic Claude Sonnet 4 V1. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
-Cross-region model inference tokens per minute for Anthropic Claude Sonnet 4 V1 1M Context Length | Each supported Region: 1,000,000 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-1FA095B8) | The maximum number of cross-region tokens that you can submit for model inference in one minute for Anthropic Claude Sonnet 4 V1 1M Context Length. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
+Cross-region model inference tokens per minute for Anthropic Claude Opus 5.5 | Each supported Region: 30,000,000 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-A4430697) | The maximum number of cross-region tokens that you can submit for model inference in one minute for Anthropic Claude Opus 5.5. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
@@ -677,0 +664,3 @@ Cross-region model inference tokens per minute for GPT-5.6 Terra | Each supporte
+Cross-region model inference tokens per minute for GPT-6 Astra | Each supported Region: 4,000,000 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-53A144CD) | The maximum number of cross-region tokens that you can submit for model inference in one minute for GPT-6 Astra. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
+Cross-region model inference tokens per minute for GPT-6 Luna | Each supported Region: 4,000,000 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-AB39D760) | The maximum number of cross-region tokens that you can submit for model inference in one minute for GPT-6 Luna. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
+Cross-region model inference tokens per minute for GPT-6 Sol | Each supported Region: 2,000,000 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-8FD128CF) | The maximum number of cross-region tokens that you can submit for model inference in one minute for GPT-6 Sol. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
@@ -687,0 +677 @@ Cross-region model inference tokens per minute for Mistral Pixtral Large 25.02 V
+Cross-region model inference tokens per minute for Moonshot AI Kimi K3 | Each supported Region: 10,000,000 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-51FF35D5) | The maximum number of cross-region tokens that you can submit for model inference in one minute for Moonshot AI Kimi K3. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
@@ -709 +698,0 @@ Global cross-region model inference requests per minute for Anthropic Claude Opu
-Global cross-region model inference requests per minute for Anthropic Claude Sonnet 4 V1 | Each supported Region: 200 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-C63AA5DA) | The maximum number of global cross-region requests that you can submit for model inference in one minute for Anthropic Claude Sonnet 4 V1. This quota represents the total number of requests to the Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream actions.  
@@ -717,3 +705,0 @@ Global cross-region model inference tokens per day for Amazon Nova 2 Pro Preview
-Global cross-region model inference tokens per day for Anthropic Claude Fable 5 | Each supported Region: 5,760,000,000 | No | The maximum number of global cross-region tokens that you can submit for model inference in one day for Anthropic Claude Fable 5. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
-Global cross-region model inference tokens per day for Anthropic Claude Fable 5.1 | Each supported Region: 14,400,000,000 | No | The maximum number of global cross-region tokens that you can submit for model inference in one day for Anthropic Claude Fable 5.1. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
-Global cross-region model inference tokens per day for Anthropic Claude Haiku 4.5 | Each supported Region: 7,200,000,000 | No | The maximum number of global cross-region tokens that you can submit for model inference in one day for Anthropic Claude Haiku 4.5. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
@@ -721,6 +706,0 @@ Global cross-region model inference tokens per day for Anthropic Claude Opus 4.5
-Global cross-region model inference tokens per day for Anthropic Claude Opus 4.6 V1 | Each supported Region: 4,320,000,000 | No | The maximum number of global cross-region tokens that you can submit for model inference in one day for Anthropic Claude Opus 4.6 V1. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
-Global cross-region model inference tokens per day for Anthropic Claude Opus 4.7 | Each supported Region: 14,400,000,000 | No | The maximum number of global cross-region tokens that you can submit for model inference in one day for Anthropic Claude Opus 4.7. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
-Global cross-region model inference tokens per day for Anthropic Claude Opus 4.8 | Each supported Region: 43,200,000,000 | No | The maximum number of global cross-region tokens that you can submit for model inference in one day for Anthropic Claude Opus 4.8. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
-Global cross-region model inference tokens per day for Anthropic Claude Opus 5 | Each supported Region: 43,200,000,000 | No | The maximum number of global cross-region tokens that you can submit for model inference in one day for Anthropic Claude Opus 5. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
-Global cross-region model inference tokens per day for Anthropic Claude Sonnet 4 V1 | Each supported Region: 288,000,000 | No | The maximum number of global cross-region tokens that you can submit for model inference in one day for Anthropic Claude Sonnet 4 V1. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
-Global cross-region model inference tokens per day for Anthropic Claude Sonnet 4.5 V1 | Each supported Region: 7,200,000,000 | No | The maximum number of global cross-region tokens that you can submit for model inference in one day for Anthropic Claude Sonnet 4.5 V1. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
@@ -728,2 +707,0 @@ Global cross-region model inference tokens per day for Anthropic Claude Sonnet 4
-Global cross-region model inference tokens per day for Anthropic Claude Sonnet 4.6 | Each supported Region: 8,640,000,000 | No | The maximum number of global cross-region tokens that you can submit for model inference in one day for Anthropic Claude Sonnet 4.6. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
-Global cross-region model inference tokens per day for Anthropic Claude Sonnet 5 | Each supported Region: 8,640,000,000 | No | The maximum number of global cross-region tokens that you can submit for model inference in one day for Anthropic Claude Sonnet 5. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
@@ -731,4 +709,3 @@ Global cross-region model inference tokens per day for Cohere Embed V4 | Each su
-Global cross-region model inference tokens per day for GPT-5.6 Luna | Each supported Region: 28,800,000,000 | No | The maximum number of global cross-region tokens that you can submit for model inference in one day for GPT-5.6 Luna. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
-Global cross-region model inference tokens per day for GPT-5.6 Sol | Each supported Region: 14,400,000,000 | No | The maximum number of global cross-region tokens that you can submit for model inference in one day for GPT-5.6 Sol. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
-Global cross-region model inference tokens per day for GPT-5.6 Terra | Each supported Region: 28,800,000,000 | No | The maximum number of global cross-region tokens that you can submit for model inference in one day for GPT-5.6 Terra. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
-Global cross-region model inference tokens per day for Grok 4.6 | Each supported Region: 14,400,000,000 | No | The maximum number of global cross-region tokens that you can submit for model inference in one day for Grok 4.6. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
+Global cross-region model inference tokens per day for GPT-6 Luna | Each supported Region: 5,760,000,000 | No | The maximum number of global cross-region tokens that you can submit for model inference in one day for GPT-6 Luna. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
+Global cross-region model inference tokens per day for GPT-6 Sol | Each supported Region: 2,880,000,000 | No | The maximum number of global cross-region tokens that you can submit for model inference in one day for GPT-6 Sol. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
+Global cross-region model inference tokens per day for Moonshot AI Kimi K3 | Each supported Region: 14,400,000,000 | No | The maximum number of global cross-region tokens that you can submit for model inference in one day for Moonshot AI Kimi K3. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
@@ -746 +723 @@ Global cross-region model inference tokens per minute for Anthropic Claude Opus
-Global cross-region model inference tokens per minute for Anthropic Claude Sonnet 4 V1 | Each supported Region: 200,000 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-97E41E39) | The maximum number of global cross-region tokens that you can submit for model inference in one minute for Anthropic Claude Sonnet 4 V1. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
+Global cross-region model inference tokens per minute for Anthropic Claude Opus 5.5 | Each supported Region: 30,000,000 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-A103A344) | The maximum number of global cross-region tokens that you can submit for model inference in one minute for Anthropic Claude Opus 5.5. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
@@ -754,0 +732,3 @@ Global cross-region model inference tokens per minute for GPT-5.6 Terra | Each s
+Global cross-region model inference tokens per minute for GPT-6 Astra | Each supported Region: 4,000,000 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-B38B530A) | The maximum number of global cross-region tokens that you can submit for model inference in one minute for GPT-6 Astra. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
+Global cross-region model inference tokens per minute for GPT-6 Luna | Each supported Region: 4,000,000 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-CCC92354) | The maximum number of global cross-region tokens that you can submit for model inference in one minute for GPT-6 Luna. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
+Global cross-region model inference tokens per minute for GPT-6 Sol | Each supported Region: 2,000,000 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-039AF68B) | The maximum number of global cross-region tokens that you can submit for model inference in one minute for GPT-6 Sol. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
@@ -755,0 +736 @@ Global cross-region model inference tokens per minute for Grok 4.6 | Each suppor
+Global cross-region model inference tokens per minute for Moonshot AI Kimi K3 | Each supported Region: 10,000,000 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-2B5A3FA2) | The maximum number of global cross-region tokens that you can submit for model inference in one minute for Moonshot AI Kimi K3. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
@@ -766 +746,0 @@ Minimum number of records per batch inference job for Claude 3 Haiku | Each supp
-Minimum number of records per batch inference job for Claude 3 Opus | Each supported Region: 100 | No | The minimum number of records across all input files in a batch inference job for Claude 3 Opus.  
@@ -776 +755,0 @@ Minimum number of records per batch inference job for Claude Opus 5 | Each suppo
-Minimum number of records per batch inference job for Claude Sonnet 4 | Each supported Region: 100 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-F72F26EE) | The minimum number of records across all input files in a batch inference job for Claude Sonnet 4.  
@@ -844 +823 @@ Model invocation max tokens per day for Amazon Nova Pro (doubled for cross-regio
-Model invocation max tokens per day for Anthropic Claude 3 Haiku (doubled for cross-region calls) |  us-east-1: 2,880,000,000 us-west-2: 2,880,000,000 ap-northeast-1: 288,000,000 ap-southeast-1: 288,000,000 Each of the other supported Regions: 432,000,000 | No | Daily maximum tokens for model inference for Anthropic Claude 3 Haiku. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
+Model invocation max tokens per day for Anthropic Claude 3 Haiku (doubled for cross-region calls) |  ap-southeast-1: 288,000,000 Each of the other supported Regions: 432,000,000 | No | Daily maximum tokens for model inference for Anthropic Claude 3 Haiku. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
@@ -849,5 +827,0 @@ Model invocation max tokens per day for Anthropic Claude 3.7 Sonnet V1 (doubled
-Model invocation max tokens per day for Anthropic Claude Fable 5 (doubled for cross-region calls) | Each supported Region: 2,880,000,000 | No | Daily maximum tokens for model inference for Anthropic Claude Fable 5. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Anthropic Claude Fable 5.1 (doubled for cross-region calls) | Each supported Region: 7,200,000,000 | No | Daily maximum tokens for model inference for Anthropic Claude Fable 5.1. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Anthropic Claude Haiku 4.5 (doubled for cross-region calls) | Each supported Region: 3,600,000,000 | No | Daily maximum tokens for model inference for Anthropic Claude Haiku 4.5. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Anthropic Claude Opus 4 V1 (doubled for cross-region calls) | Each supported Region: 144,000,000 | No | Daily maximum tokens for model inference for Anthropic Claude Opus 4 V1. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Anthropic Claude Opus 4.1 (doubled for cross-region calls) | Each supported Region: 360,000,000 | No | Daily maximum tokens for model inference for Anthropic Claude Opus 4.1. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
@@ -855,7 +828,0 @@ Model invocation max tokens per day for Anthropic Claude Opus 4.5 (doubled for c
-Model invocation max tokens per day for Anthropic Claude Opus 4.6 V1 (doubled for cross-region calls) | Each supported Region: 2,160,000,000 | No | Daily maximum tokens for model inference for Anthropic Claude Opus 4.6 V1. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Anthropic Claude Opus 4.7 (doubled for cross-region calls) | Each supported Region: 7,200,000,000 | No | Daily maximum tokens for model inference for Anthropic Claude Opus 4.7. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Anthropic Claude Opus 4.8 (doubled for cross-region calls) | Each supported Region: 21,600,000,000 | No | Daily maximum tokens for model inference for Anthropic Claude Opus 4.8. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Anthropic Claude Opus 5 (doubled for cross-region calls) | Each supported Region: 21,600,000,000 | No | Daily maximum tokens for model inference for Anthropic Claude Opus 5. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Anthropic Claude Sonnet 4 V1 (doubled for cross-region calls) | Each supported Region: 144,000,000 | No | Daily maximum tokens for model inference for Anthropic Claude Sonnet 4 V1. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Anthropic Claude Sonnet 4 V1 1M Context Length (doubled for cross-region calls) | Each supported Region: 720,000,000 | No | Daily maximum tokens for model inference for Anthropic Claude Sonnet 4 V1 1M Context Length. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Anthropic Claude Sonnet 4.5 V1 (doubled for cross-region calls) | Each supported Region: 3,600,000,000 | No | Daily maximum tokens for model inference for Anthropic Claude Sonnet 4.5 V1. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
@@ -863,2 +829,0 @@ Model invocation max tokens per day for Anthropic Claude Sonnet 4.5 V1 1M Contex
-Model invocation max tokens per day for Anthropic Claude Sonnet 4.6 (doubled for cross-region calls) | Each supported Region: 4,320,000,000 | No | Daily maximum tokens for model inference for Anthropic Claude Sonnet 4.6. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Anthropic Claude Sonnet 5 (doubled for cross-region calls) | Each supported Region: 4,320,000,000 | No | Daily maximum tokens for model inference for Anthropic Claude Sonnet 5. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
@@ -866,14 +831,2 @@ Model invocation max tokens per day for Cohere Embed V4 (doubled for cross-regio
-Model invocation max tokens per day for DeepSeek R1 V1 (doubled for cross-region calls) | Each supported Region: 144,000,000 | No | Daily maximum tokens for model inference for DeepSeek R1 V1. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for DeepSeek V3 V1 (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for DeepSeek V3 V1. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for DeepSeek V3.2 (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for DeepSeek V3.2. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for GPT OSS Safeguard 120B (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for GPT OSS Safeguard 120B. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for GPT OSS Safeguard 20B (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for GPT OSS Safeguard 20B. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for GPT-5.6 Luna (doubled for cross-region calls) | Each supported Region: 14,400,000,000 | No | Daily maximum tokens for model inference for GPT-5.6 Luna. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for GPT-5.6 Sol (doubled for cross-region calls) | Each supported Region: 7,200,000,000 | No | Daily maximum tokens for model inference for GPT-5.6 Sol. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for GPT-5.6 Terra (doubled for cross-region calls) | Each supported Region: 14,400,000,000 | No | Daily maximum tokens for model inference for GPT-5.6 Terra. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Gemma 3 12B (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Gemma 3 12B. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Gemma 3 27B (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Gemma 3 27B. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Gemma 3 4B (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Gemma 3 4B. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Grok 4.6 (doubled for cross-region calls) | Each supported Region: 7,200,000,000 | No | Daily maximum tokens for model inference for Grok 4.6. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Kimi K2 Thinking (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Kimi K2 Thinking. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Magistral Small 1.2 (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Magistral Small 1.2. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
+Model invocation max tokens per day for GPT-6 Luna (doubled for cross-region calls) | Each supported Region: 57,600,000,000 | No | Daily maximum tokens for model inference for GPT-6 Luna. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
+Model invocation max tokens per day for GPT-6 Sol (doubled for cross-region calls) | Each supported Region: 28,800,000,000 | No | Daily maximum tokens for model inference for GPT-6 Sol. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
@@ -884,8 +836,0 @@ Model invocation max tokens per day for Meta Llama 3.2 90B Instruct (doubled for
-Model invocation max tokens per day for Meta Llama 4 Maverick V1 (doubled for cross-region calls) | Each supported Region: 432,000,000 | No | Daily maximum tokens for model inference for Meta Llama 4 Maverick V1. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Meta Llama 4 Scout V1 (doubled for cross-region calls) | Each supported Region: 432,000,000 | No | Daily maximum tokens for model inference for Meta Llama 4 Scout V1. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for MiniMax M2.5 (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for MiniMax M2.5. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Minimax M2 (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Minimax M2. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Minimax M2.1 (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Minimax M2.1. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Ministral 14B 3.0 (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Ministral 14B 3.0. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Ministral 3B 3.0 (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Ministral 3B 3.0. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Ministral 8B 3.0 (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Ministral 8B 3.0. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
@@ -896,2 +840,0 @@ Model invocation max tokens per day for Mistral AI Mixtral 8X7B Instruct (double
-Model invocation max tokens per day for Mistral Devstral 2 123b (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Mistral Devstral 2 123b. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Mistral Large 3 (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Mistral Large 3. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
@@ -899,16 +842 @@ Model invocation max tokens per day for Mistral Pixtral Large 25.02 V1 (doubled
-Model invocation max tokens per day for Moonshot AI Kimi K2.5 (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Moonshot AI Kimi K2.5. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for NVIDIA Nemotron 3 Super 120B A12B (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for NVIDIA Nemotron 3 Super 120B A12B. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for NVIDIA Nemotron Nano 2 (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for NVIDIA Nemotron Nano 2. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for NVIDIA Nemotron Nano 2 VL (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for NVIDIA Nemotron Nano 2 VL. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Nemotron Nano 3 30B (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Nemotron Nano 3 30B. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for OpenAI GPT OSS 120B (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for OpenAI GPT OSS 120B. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for OpenAI GPT OSS 20B (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for OpenAI GPT OSS 20B. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Qwen3 235B a22b 2507 V1 (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Qwen3 235B a22b 2507 V1. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Qwen3 32B V1 (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Qwen3 32B V1. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Qwen3 Coder 30B a3b V1 (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Qwen3 Coder 30B a3b V1. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Qwen3 Coder 480B a35b V1 (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Qwen3 Coder 480B a35b V1. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Qwen3 Coder Next (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Qwen3 Coder Next. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Qwen3 Next 80B A3B (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Qwen3 Next 80B A3B. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Qwen3 VL 235B A22B (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Qwen3 VL 235B A22B. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Voxtral Mini 1.0 (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Voxtral Mini 1.0. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Voxtral Small 1.0 (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Voxtral Small 1.0. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
+Model invocation max tokens per day for Moonshot AI Kimi K3 (doubled for cross-region calls) | Each supported Region: 14,400,000,000 | No | Daily maximum tokens for model inference for Moonshot AI Kimi K3. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
@@ -917,4 +844,0 @@ Model invocation max tokens per day for Writer AI Palmyra X5 V1 (doubled for cro
-Model invocation max tokens per day for Writer Palmyra Vision 7B (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Writer Palmyra Vision 7B. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Z.ai GLM 5 (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Z.ai GLM 5. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Z.ai GLM-4.7 (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Z.ai GLM-4.7. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
-Model invocation max tokens per day for Z.ai GLM-4.7 Flash (doubled for cross-region calls) | Each supported Region: 144,000,000,000 | No | Daily maximum tokens for model inference for Z.ai GLM-4.7 Flash. Combines sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream. Doubled for cross-region calls; not applicable in case of approved TPM increase.  
@@ -945,3 +868,0 @@ Model units per provisioned model for Anthropic Claude 3.5 Sonnet 51K | Each sup
-Model units per provisioned model for Anthropic Claude 3.7 V1.0 Sonnet 18K | Each supported Region: 0 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-D09F1612) | The maximum number of model units that can be allotted to a provisioned model for Anthropic Claude 3.7 V1.0 Sonnet 18K.  
-Model units per provisioned model for Anthropic Claude 3.7 V1.0 Sonnet 200K | Each supported Region: 0 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-F4131C39) | The maximum number of model units that can be allotted to a provisioned model for Anthropic Claude 3.7 V1.0 Sonnet 200K.  
-Model units per provisioned model for Anthropic Claude 3.7 V1.0 Sonnet 51K | Each supported Region: 0 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-0B0CDE73) | The maximum number of model units that can be allotted to a provisioned model for Anthropic Claude 3.7 V1.0 Sonnet 51K.  
@@ -1020,2 +941 @@ On-demand model inference requests per minute for Amazon Titan Text Premier | Ea
-On-demand model inference requests per minute for Anthropic Claude 3 Haiku |  us-east-1: 1,000 us-west-2: 1,000 ap-northeast-1: 200 ap-southeast-1: 200 Each of the other supported Regions: 400 | No | The maximum number of times that you can call model inference in one minute. The quota considers the combined sum of Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream for Anthropic Claude 3 Haiku.  
-On-demand model inference requests per minute for Anthropic Claude 3 Opus | Each supported Region: 50 | No | The maximum number of on-demand requests that you can submit for model inference in one minute for Anthropic Claude 3 Opus. This quota represents the total number of requests to the Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream actions.  
+On-demand model inference requests per minute for Anthropic Claude 3 Haiku |  ap-southeast-1: 200 Each of the other supported Regions: 400 | No | The maximum number of times that you can call model inference in one minute. The quota considers the combined sum of Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream for Anthropic Claude 3 Haiku.  
@@ -1119,2 +1039 @@ On-demand model inference tokens per minute for Amazon Titan Text Premier | Each
-On-demand model inference tokens per minute for Anthropic Claude 3 Haiku |  us-east-1: 2,000,000 us-west-2: 2,000,000 ap-northeast-1: 200,000 ap-southeast-1: 200,000 Each of the other supported Regions: 300,000 | No | The maximum number of on-demand tokens that you can submit for model inference in one minute. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream for Anthropic Claude 3 Haiku.  
-On-demand model inference tokens per minute for Anthropic Claude 3 Opus | Each supported Region: 400,000 | No | The maximum number of on-demand tokens that you can submit for model inference in one minute for Anthropic Claude 3 Opus. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream.  
+On-demand model inference tokens per minute for Anthropic Claude 3 Haiku |  ap-southeast-1: 200,000 Each of the other supported Regions: 300,000 | No | The maximum number of on-demand tokens that you can submit for model inference in one minute. The quota considers the combined sum of input and output tokens across all requests to Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream for Anthropic Claude 3 Haiku.  
@@ -1190 +1108,0 @@ Records per batch inference job for Claude 3 Haiku | Each supported Region: 100,
-Records per batch inference job for Claude 3 Opus | Each supported Region: 100,000 |  [Yes](https://console.aws.amazon.com/servicequotas/home/services/bedrock/quotas/L-E8FA49DB) | The maximum number of records across all input files in a batch inference job for Claude 3 Opus.  
@@ -1200 +1117,0 @@ Records per batch inference job for Claude Opus 5 | Each supported Region: 100,0