Service eligibility and regional restrictions
PuppyIP serves only compliant overseas businesses and their authorized personnel. Proxy services are not available in mainland China. The service may only be used for lawful business activities outside mainland China. Use of this service within mainland China is prohibited.
Hosting a proxy IP or server overseas does not change these restrictions. The service must not be provided to end users in mainland China through relaying, forwarding, sharing or resale. Before use, read the Terms of Service.
Key Takeaways
- From September 15, 2026, account information must be completed before activating Model Studio. Official documentation does not specify that day's exact switch time or time zone.
- The new 90-day rule applies only to users first activating Singapore Model Studio from September 8, 2026 at 03:00 UTC. It does not automatically apply to users activated earlier.
- Validity starts on the latest of activation, model release or model-request approval. Remaining quota expires without pausing or replenishment, and registering again does not grant new quota.
- Quota covers only real-time inference for participating models in Singapore with International deployment scope. It excludes batch processing, fine-tuning, deployment, custom models, PAI-DSW and OSS.
- Each model is metered independently, and latest and dated snapshots may be metered separately. The official typical allowance of 1,000,000 tokens is not guaranteed for every model.
- Accounts activated before September 15 with incomplete information may use remaining quota, but must complete their information before PAYG after exhaustion. With completed information and Free Quota Only disabled, over-quota calls can continue to incur charges. Switches and bills have delays; network egress changes neither eligibility nor billing.
What changed: New-user free quota now lasts 90 days
Alibaba Cloud's New user free quota page, updated as of September 22, 2026, states that users first activating Model Studio in Singapore from September 8, 2026 at 03:00 UTC, or 11:00 Beijing time, receive new-user model quota valid for 90 days. A separate official page explicitly says users who activated earlier are unaffected by this validity change.
The 90 days do not always begin on account registration or announcement day. For a particular model, validity starts on the latest of Model Studio activation, that model's release or approval of its access request. Quota expires at the end of that period. Unused balances are not paused, replenished or extended, and registering another account does not provide additional new-user quota.
Eligibility and scope: Complete account information, then check Singapore and model coverage
The current official page states that from September 15, 2026, account information must be completed before Model Studio activation. Previously activated users with incomplete information can continue using remaining free quota, but must complete the information before entering pay-as-you-go after exhaustion. The page gives only a date, without an exact switch time or time zone. Do not apply the new activation requirement retroactively to previously activated accounts.
Quota covers only participating models in the Singapore region with International service deployment scope, and offsets only real-time inference or invocation. Batch invocation, fine-tuning, model deployment, custom models, PAI-DSW, OSS storage and request charges are excluded. First activation after eligibility is met grants quota automatically, usually taking up to 2 hours to appear. Some speech models additionally require access to be enabled per model.
Per-model accounting: A typical one million tokens is not a universal promise
Official guidance calculates free quota independently per model, with no pooling or transfer across models. The account and its RAM users share that account's quota. Input and output tokens are deducted together. An undated latest version and a dated snapshot may also count as separate models; exhausting one does not automatically switch to another.
The documented typical 1,000,000 tokens is only a common value. Model allowances in the official pricing table differ: examples include one million tokens, 500,000 tokens, quantities measured by voice, or no free quota at all. Before launch, check every model in Singapore Model Square or the current pricing table. Do not put an example figure into a uniform budget.
Pay-as-you-go deduction order: Identify the API Key and account state first
For a general pay-as-you-go API Key, the deduction order is free quota, resource plan, savings plan and finally pay-as-you-go. Dedicated Coding Plan or Token Plan Keys do not consume this API-key new-user quota. OAuth's daily free invocation allowance is a separate scheme and must not be mixed into this one.
Distinguish activation eligibility before and from September 15: accounts activated earlier with incomplete information may use remaining free quota, then must complete their information; new activations from September 15 require completed information first. Users with completed information and Free Quota Only disabled incur PAYG charges at input and output prices for over-quota calls. An overdue account may be unable to call even if other models still show balances. Record Key type, information-completion status, activation date, model, region and billing owner in one checklist.
Free Quota Only: Spending protection, not an instant switch
Free Quota Only is disabled by default. When enabled, exhausting a participating model's free quota returns HTTP 403 with AllocationQuota.FreeTierOnly and stops calls from entering PAYG. Disable it to restore pay-as-you-go only after paid usage is explicitly approved. On this error, stop retrying and check balance, expiry and switch state first.
Official guidance warns that enabling or disabling the switch is not immediate. Calls made during propagation may still incur charges after quota exhaustion; restoring PAYG after disabling can take around 30 minutes. It is therefore cost protection, not a real-time production-availability control. If zero unexpected charges are required, enable it early, wait for propagation and stop load testing.
Usage and billing delays: Absence on screen does not mean no charges
The free-quota page updates on a minute-level basis but needs manual refresh. Invocation records usually lag by several minutes, and documentation warns of hour-level delays during peaks. Model Studio large-model PAYG bills may lag by 1 to 3 hours; a final bill can still appear after an API Key is deleted or resources are stopped.
Perform only one small nonproduction call for verification. Record the request ID, full model name, input/output tokens, Key type, region, time and balances before and after. After stopping calls, wait through at least the official billing-delay window before comparing usage and bills. Do not expand traffic if balances, errors or charges cannot be explained.
Stop conditions, rollback and network boundaries
Stop immediately if any of these applies: preparing to activate after September 15 without completed information; a model outside Singapore/International scope or marked No free quota; a request for excluded batch, fine-tuning or deployment work; expired or exhausted quota; quota not yet visible in the console; AllocationQuota.FreeTierOnly; or a switch or bill still within its delay window. Retain the old model or a budgeted alternative workflow; do not switch models silently.
For 401/403, check the Key, account, quota and switch first. For 404/model not found, check region and model name. For 429, check model quota. Only with DNS, TLS, proxy 407 or connection-timeout evidence should you use the systematic proxy connection troubleshooting checklist. Fixed egress cannot extend the 90 days, create free quota, change deployment scope or prevent PAYG. PRODUCT_FIT=CONDITIONAL_NETWORK_ONLY.
Sources
Frequently Asked Questions
What is required before activating Model Studio from September 15?
Official guidance requires completing account information before activation. Users activated before September 15 with incomplete information may use remaining free quota, but must complete it before PAYG after exhaustion.
Which date starts the 90 days?
For each model, use the latest of Model Studio activation, model release and model-request approval. Do not always use registration or announcement day.
Does every model receive one million tokens?
There is no universal guarantee. Official documentation says one million is typical; pricing tables show some models with 500,000, different units or no free quota. Check each model.
Will charges start automatically after quota runs out?
With completed account information and Free Quota Only disabled, over-quota calls move to subsequent deductions and pay-as-you-go. Users activated before September 15 with incomplete information must complete it first to continue PAYG.
How should I handle AllocationQuota.FreeTierOnly?
It means Free Quota Only has blocked over-quota calls. Stop retrying and check the balance. Disable the switch only after PAYG is explicitly approved; restoration may take around 30 minutes.
Can a fixed IP restore quota or prevent charges?
No. Quota, expiry, region, API Key type and billing depend on platform state. Fixed egress has troubleshooting value only for network evidence such as DNS, TLS, 407 or connection timeouts.