Service eligibility and regional restrictions
PuppyIP serves only compliant overseas businesses and their authorized personnel. Proxy services are not available in mainland China. The service may only be used for lawful business activities outside mainland China. Use of this service within mainland China is prohibited.
Hosting a proxy IP or server overseas does not change these restrictions. The service must not be provided to end users in mainland China through relaying, forwarding, sharing or resale. Before use, read the Terms of Service.
Key Takeaways
- AWS announced GLM 5.3 support on October 5; its model card is Active. Eligible enterprise customers must confirm access with their AWS account team.
- The model supports one-million-token context and up to 128K output tokens. Reasoning stays enabled with adjustable effort, and explicit caching has its own conditions.
- Use us.zai.glm-5.3 or global.zai.glm-5.3 cross-region IDs. In-Region inference is not supported.
- Standard Global prices are $1.68 input and $5.28 output per million tokens; US prices are $1.848 and $5.808. Cache pricing and other tiers are separate.
- This is not an automatic replacement for models retiring on Alibaba Cloud. Provider IDs, credentials, regions and behavior require separate migration checks.
Support is announced, but account access is conditional
AWS announced Z.ai GLM 5.3 on Bedrock on October 5, 2026, and its model card lists Active status. Announcement, publication of this guide and activation for a particular account are different events. No account activation was verified here.
AWS limits availability to eligible enterprise customers and directs them to their account team. The sources do not provide a universal account-qualification table or establish access for every account. Confirm your own eligibility before planning a rollout.
Long context serves engineering tasks, not guaranteed results
GLM 5.3 targets agentic coding and long engineering workflows, including repository generation, refactoring and tool use. One-million-token context and 128K output are limits, not promises of accurate results or low cost for every long task.
Reasoning remains enabled; adjust reasoning effort to balance latency, tokens and results. The model supports function calling, structured JSON, streaming and reasoning. Production tool workflows still need permission boundaries, timeouts and explicit failure handling.
Use cross-region IDs with the correct API
The supported inference IDs are us.zai.glm-5.3 and global.zai.glm-5.3. The base ID zai.glm-5.3 is not a regional on-demand invocation route. Check the official API examples for Chat Completions, InvokeModel and Converse under bedrock-runtime.
Match the API, authentication and endpoint you actually use. Do not reuse another provider's model name, key or Responses endpoint just because it also offers GLM. No Bedrock request was executed for this article.
US and Global describe routing scope
US geographic inference and Global inference are cross-region modes. US does not mean one AWS region, and Global does not mean requests remain in the originating region. Decide residency and routing requirements before choosing the model ID.
Check the current compatibility table for source regions, mode and account quota, and confirm eligibility with the account team. An endpoint does not establish support for every routing mode; another customer's successful request does not establish yours.
Budget input, output and caching separately
On October 6, Standard Global inference was $1.68 input and $5.28 output per million tokens. Standard US inference was $1.848 input and $5.808 output. These are the documented table rates, not a claim about every region, pricing tier or contract.
Cache reads and 30-minute cache writes have separate pricing. Include input, output, cache writes, actual hits and the chosen tier in your budget rather than counting prompt words alone. Contract rates, taxes and actual usage were not verified; no savings were measured.
Explicit caching has minimums and retention rules
Explicit cache points apply to supported system and message content. The model card describes implicit caching defaults and explicit caching with at least 1,024 tokens and at least 30 minutes of retention. Check the current docs for the exact applicable conditions.
Cache reads and writes are billed differently from ordinary input. A miss, changed prompt or retention expiry can change the result. Enabling caching does not by itself prove lower latency or cost for your workload.
Treat access and configuration as one migration checklist
Confirm eligibility, source region, cross-region ID, API, tier, quota and pricing as one configuration set. Changing a model string alone neither grants access nor makes migration free or automatic.
After authorization, use isolated, redacted fixed samples to check code, JSON, tool completion, latency and charges. Define a budget stop and fallback before testing. Business acceptance and actual billing matter more than a successful model-list response.
A GLM family name does not join separate providers
Alibaba Cloud's retirement of 11 Model Studio models involves its own IDs, regions, protocols and deadlines. GLM branding does not make AWS authentication, prices, outputs or APIs equivalent. Keep that migration checklist separate.
Before switching, confirm any unknown AWS eligibility conditions with the official docs and account team. Preserve evidence of your current integration and select a supported, validated replacement rather than assuming this announcement completes the migration.
Sources
Frequently Asked Questions
Can every Bedrock account use GLM 5.3?
No. AWS specifies eligible enterprise customers and directs them to their account team. Active model status does not grant every account access.
Which inference IDs should I use?
The cross-region IDs are us.zai.glm-5.3 and global.zai.glm-5.3. In-Region inference using the base ID is not supported.
Does one-million-token context guarantee better coding?
No. It is a capacity limit. Compare actual task completion, output quality, latency and cost on representative work.
What are the Standard prices?
Per million tokens, Global input/output are $1.68/$5.28 and US input/output are $1.848/$5.808. Cache operations, other tiers and contract conditions must be checked separately.
Does US inference stay in one region?
No. It is geographic cross-region routing. Check supported source regions and your residency requirements; Global also does not promise origin-region processing.
Does this replace retiring Alibaba Cloud GLM models automatically?
No. Provider IDs, credentials, regions, API contracts and behavior differ. Validate that migration separately.