Service eligibility and regional restrictions
PuppyIP serves only compliant overseas businesses and their authorized personnel. Proxy services are not available in mainland China. The service may only be used for lawful business activities outside mainland China. Use of this service within mainland China is prohibited.
Hosting a proxy IP or server overseas does not change these restrictions. The service must not be provided to end users in mainland China through relaying, forwarding, sharing or resale. Before use, read the Terms of Service.
Key Takeaways
- Microsoft announced GA October 6, including Chat Completions, Responses, tool calls, and streaming. Actual projects still need permissions and quota.
- Global Standard short-context input, output, and cache rates are $2, $6, and $0.50 per million tokens; at or above 200,000 context tokens, they are $4, $12, and $1.
- US Data Zone Standard rates are $2.20, $6.60, and $0.55 for short context and $4.40, $13.20, and $1.10 for long context. Other channels' bills do not apply.
- Select grok-4.7 version 1 in the catalog, but put the actual deployment name in request model. Combined input and output must fit the current context limit.
Who can deploy the new Foundry model?
Microsoft product-blog author Mahesh Balachandran announced Grok 4.7 general availability in Foundry Models October 6, 2026. It provides an additional channel for Azure teams; xAI, Bedrock, or Copilot access does not establish a Foundry deployment.
Current guidance requires an Azure subscription with a valid payment method, a Foundry project in a supported region, and resource creation and management permissions. Deployment also requires Cognitive Services Contributor on the resource. GA does not guarantee quota for every project.
How should Global and US Data Zone be chosen?
The announcement lists Global Standard and US Data Zone Standard. Check your data-processing requirements, regional availability, and capacity before comparing prices. A Global or US name does not replace actual deployment conditions.
Choose grok-4.7 in the catalog; current guidance lists version 1. Region tables and quotas are separate checks. No subscription's guaranteed regions were confirmed here, and older Grok regional or quota limits are not inherited by 4.7.
What changes at 200,000 tokens?
Microsoft divides short and long context at 200,000 tokens. Below that threshold, Global Standard input, output, and cache cost $2, $6, and $0.50 per million tokens. At or above it, they cost $4, $12, and $1.
US Data Zone Standard costs $2.20, $6.60, and $0.55 for short context and $4.40, $13.20, and $1.10 for long context. These are dollar prices in Microsoft's announcement. Confirm each request's tier and billing definition against deployment pricing and usage.
Hypothetical calculation: a short-context request with 10,000 input and 1,000 output tokens and no cache costs about $0.026 on Global Standard or $0.0286 on US Data Zone Standard at the published rates. This is not a measured bill or full application cost.
Does model mean the model ID or deployment name?
Select grok-4.7 in the catalog, but send the deployment name you created in model; they may differ. Guidance supports Chat Completions and Responses through the Foundry resource endpoint. Use the chosen interface rather than copying xAI or Bedrock identifiers.
Authentication supports Microsoft Entra ID or the Foundry resource's API key, with Microsoft recommending Entra ID. Tool calls and streaming are supported. Begin with a simple request without sensitive data or external writes before adding business tools, so failures can be located by layer.
Is 500,000-token context still only planned?
The October 6 blog says planned, but current Microsoft Learn deployment guidance explicitly lists a 500000-token context window. The older announcement alone should no longer describe it as unsupported. Check current documentation and actual deployment details during integration.
Input and generated output, including reasoning output, share that window. Leave output space near the limit and check both context-price tier and per-minute quota. The maximum window is neither a recommendation to fill every request nor equal per-minute throughput.
How can you verify a first integration?
Retain deployment type, region, name, and authentication method. Check a simple result and usage, then record input, output, cache, and latency. For migration, compare the same authorized representative tasks for quality and cost; vendor benchmarks cannot replace your acceptance.
For 401 check authentication, 403 roles and model access, 404 endpoint and deployment name, and 429 capacity and quota. Do not switch channels to solve a permission or naming error without locating its stage. No actual Azure deployment or account operation was executed here.
Sources
Frequently Asked Questions
Can an existing xAI key call Foundry directly?
Microsoft requires Entra ID or the Foundry resource's own key. An xAI direct-API key does not establish Foundry access. Obtain endpoint and authentication from the selected deployment.
Are reasoning parameters identical in Chat Completions and Responses?
No. Current guidance uses reasoning_effort with low, medium, high, and xhigh for Chat Completions; Responses uses reasoning.effort with low, medium, and high. Follow the interface's parameter table rather than changing the endpoint alone.