Service eligibility and regional restrictions
PuppyIP serves only compliant overseas businesses and their authorized personnel. Proxy services are not available in mainland China. The service may only be used for lawful business activities outside mainland China. Use of this service within mainland China is prohibited.
Hosting a proxy IP or server overseas does not change these restrictions. The service must not be provided to end users in mainland China through relaying, forwarding, sharing or resale. Before use, read the Terms of Service.
Key Takeaways
- The AWS model card records Grok 4.6's Bedrock launch date as August 18, 2026. September 21 brought a more detailed practical article, not the first day of model availability.
- bedrock-mantle uses xai.grok-4.6 with OpenAI-compatible interfaces, and the current model card lists only In-Region support. bedrock-runtime supports Responses, Chat Completions, and Converse, but requires the cross-Region inference ID us.xai.grok-4.6 or global.xai.grok-4.6.
- The model offers a 500,000-token context and four reasoning effort levels: low, medium, high, and xhigh. Converse requires reasoning_effort through additionalModelRequestFields; do not copy its field location from another interface.
- Published Standard pricing for In-Region/US Geo is $2.20 per million input tokens, $6.60 for output, and $0.55 for cache reads. Global pricing is $2.00 for input, $6.00 for output, and $0.50 for cache reads.
- Priority is billed at 1.75 times Standard, and Flex at 0.5 times. Lower-priced Global broadens the geographic routing scope, so price comparisons must not overlook data residency and organizational compliance requirements.
- The OpenAI-compatible bearer token path requires bedrock:CallWithBearerToken. SDK/Converse uses ordinary AWS credentials and SigV4. Long-term Bedrock API keys are better suited to controlled exploration; prioritize short-term credentials or roles in production. PRODUCT_FIT=NONE: changing IP address cannot grant model or IAM permissions.
Correct the timeline first: the model launched in August; September brought a practical guide
AWS published “xAI’s Grok 4.6 is now available in Amazon Bedrock” on September 21, 2026, but both the article and model card mark the Bedrock launch date as August 18. AWS What's New had already confirmed cross-Region inference on August 19. The value of this update is therefore its consolidated explanation of dual endpoints, Converse, service tiers, IAM, and examples, not a claim that access first opened on September 21.
Grok 4.6 targets coding, long-running agents, and knowledge work. The model card lists a 500,000-token context and low, medium, high, and xhigh reasoning levels. Capability descriptions do not establish that your account, source Region, quota, or budget qualifies. Before rollout, check model access, Region, Service Quotas, and billing in your own AWS account.
Step 1: Choose Mantle or Runtime for your client, and do not mix model IDs
If an existing application uses the OpenAI SDK and you want to minimize request-shape changes, evaluate bedrock-mantle. The model ID is xai.grok-4.6, with a Base URL such as https://bedrock-mantle.us-west-2.api.aws/openai/v1. The current model card lists Mantle only as In-Region, without Geo or Global inference IDs. Use the latest model-card table to confirm the actual Region.
If you need the AWS SDK, unified Converse/ConverseStream, Bedrock Guardrails, or cross-Region capacity, choose bedrock-runtime. Runtime does not support invoking the bare xai.grok-4.6 as an In-Region model. Use us.xai.grok-4.6 or global.xai.grok-4.6 from a supported source Region. The former confines processing to the United States geographic area, while the latter can route through a global capacity pool. They are not freely interchangeable aliases.
Step 2: Put reasoning effort in the correct field and perform a minimal read-only check
For a Converse call, set modelId to the cross-Region inference ID, place messages in messages, and set the output limit in inferenceConfig. Pass reasoning_effort through additionalModelRequestFields, for example {"reasoning_effort":"low"}. If it is mistakenly placed in inferenceConfig, do not assume the service will use the intended level.
For the first check, use a short prompt with no sensitive data or tool writes. Record the request's source Region, model ID, service_tier, reasoning_effort, request-id, input/output tokens, cache reads, latency, and actual billing fields. Test low and high separately before deciding whether xhigh is needed. A 500,000-token context is a capability limit, not a target to fill by default.
Step 3: Calculate cost by inference scope and service tier
As checked on September 22, 2026, the model card lists Standard prices per million tokens as follows: In-Region and US Geo cost $2.20 for input, $6.60 for output, and $0.55 for cache reads; Global costs $2.00 for input, $6.00 for output, and $0.50 for cache reads. Estimates in this article apply only to the current public table. Formal budgets must use the latest AWS pricing page, request usage, and billing.
Priority costs 1.75 times Standard and suits online workloads that genuinely require preferential processing. Flex costs 0.5 times Standard and suits tasks that can wait. Global has a lower published price but broadens the geography of data processing. If residency requirements, customer contracts, or regulations apply, a $0.20 reduction in the input unit price is not a reason to skip compliance review.
Step 4: Check the authentication path and IAM resources together
OpenAI-compatible interfaces can use a Bedrock API key as a bearer token, and the relevant IAM path requires bedrock:CallWithBearerToken. boto3 and Converse use AWS credentials and SigV4; bearer tokens are not the same authentication mechanism. If an application uses both interfaces, verify credential sources separately. Never put a long-term key in a code repository, logs, or client bundle.
Authorization for bedrock:InvokeModel must also cover the account's default project, the inference profile actually invoked, and foundation-model resources that may span Regions. Authorizing only us.xai.grok-4.6 does not automatically cover global.xai.grok-4.6. Long-term Bedrock API keys suit controlled exploration. In production, prioritize roles or automatically expiring short-term credentials, with least privilege, rotation, and revocation procedures.
Step 5: Validate Guardrails, streaming events, and rollback before rollout
If you need content filtering, denied topics, PII handling, or word-list policies, verify the actual input and output behavior of Bedrock Guardrails on the Runtime path. With ConverseStream, handle messageStart, contentBlockDelta, contentBlockStop, messageStop, and metadata individually. Do not reuse a Mantle SSE parser unchanged for a Converse event stream.
Compare the old model and Grok 4.6 using representative tasks, measuring success rate, tool calls, human rework, latency, and total cost. Stop expansion and return to the verified model ID and configuration if the Region is noncompliant, policy prohibits use, permissions are too broad, streaming parsing drops events, costs exceed the threshold, or tool-write outcomes are unclear. For API configuration, see troubleshooting AI API keys, Base URLs, and model names; for network and TLS errors, see the proxy connection troubleshooting checklist.
What still needs confirmation in your own AWS account
Public documentation cannot prove that your account has access, quotas are sufficient, a particular source Region is currently available, a Guardrail policy fits your application, or Global routing satisfies internal data requirements. Examples in the AWS article are not your own performance, quality, or cost tests. Validate any benchmark or recommended tier against a fixed task set.
This article does not claim that a proxy or fixed IP can change Bedrock model access, IAM, Service Quotas, data residency, or billing. A fixed egress point can at most help stabilize an allowed network environment. Insufficient permissions should be addressed by the AWS account administrator under least privilege. When a model is unavailable or a Region is unsupported, do not try to bypass the restriction by rotating accounts, guessing model names, or broadening network scope.
Sources
Frequently Asked Questions
Did Grok 4.6 only arrive in Amazon Bedrock on September 21, 2026?
No. The AWS model card records a launch date of August 18, 2026, and the August 19 What's New announcement confirmed cross-Region inference. September 21 brought a more detailed practical guide.
Which Grok 4.6 model ID should I use with bedrock-runtime?
Use the cross-Region inference ID us.xai.grok-4.6 or global.xai.grok-4.6. The current model card explicitly states that Runtime does not support In-Region invocation using the bare xai.grok-4.6.
What is the main difference between Mantle and Runtime?
Mantle focuses on OpenAI-compatible interfaces and uses xai.grok-4.6. Runtime supports Responses, Chat Completions, Converse/ConverseStream, Guardrails, and cross-Region inference. Consult the latest model card for specific Regions.
What is the Standard price for Grok 4.6 in Bedrock?
Current Standard pricing per million tokens for In-Region/US Geo is $2.20 for input, $6.60 for output, and $0.55 for cache reads. Global prices are $2.00, $6.00, and $0.50 respectively. Priority is 1.75 times Standard, and Flex is 0.5 times.
Why is a compliance check still needed before choosing cheaper Global inference?
Global can route through a worldwide capacity pool, whereas US Geo confines processing to the United States geographic area. Price is not the only consideration: data residency, customer contracts, and organizational policy may require a narrower inference scope.
Can changing proxies or using a fixed IP grant access to Grok 4.6 in Bedrock?
No. Model access, IAM, quotas, Regions, and billing are governed by the AWS account and service rules. Network egress cannot grant bedrock:InvokeModel, bedrock:CallWithBearerToken, or eligibility for cross-Region inference.