Service eligibility and regional restrictions
PuppyIP serves only compliant overseas businesses and their authorized personnel. Proxy services are not available in mainland China. The service may only be used for lawful business activities outside mainland China. Use of this service within mainland China is prohibited.
Hosting a proxy IP or server overseas does not change these restrictions. The service must not be provided to end users in mainland China through relaying, forwarding, sharing or resale. Before use, read the Terms of Service.
Key Takeaways
- Within the same five-minute interval, multiple 5XX failures submitted by the same Alibaba Cloud account to the same model with the same request-parameter hash count as one failure for SLA calculations.
- If one request in that identical-request group succeeds, the announcement says the group does not count as an SLA failure. This does not erase client errors, latency or retry costs.
- Formally retired models no longer receive the availability commitment, and their service errors are excluded from the SLA error rate. Include a model-lifecycle inventory in monitoring.
- The current SLA defines five-minute intervals and calculates availability from the monthly average error rate. Intervals with fewer than five requests count as zero error rate. The full agreement still governs scope and exclusions.
- Preserve raw request times, models, regions, endpoints, response codes, retry associations and parameter fingerprints now. Do not retain only figures aggregated under the provider's rules.
Business failures and SLA failures are not the same number
On August 26, 2026, Alibaba Cloud announced a revision to the Model Studio model-inference SLA, effective September 28, with an expected time of 00:00 (UTC+8), subject to the actual change time. The announced date has now arrived, but this article has no account-level execution logs and cannot claim every instance switched at midnight. The change concerns how the provider calculates its SLA error rate, not a promise that the API will automatically reduce 5XX responses, retries or user-visible faults. The announcement treats continued use as acceptance of the revised agreement; account terms and the full SLA still govern applicability.
Internal monitoring should continue recording every failure. If the same request fails three times within five minutes and succeeds on the fourth attempt, the application still experienced three errors and additional latency. Under the new rules, the group may not count as an SLA failure. Treating provider SLA figures as the business SLO would understate retry storms, tail latency and user failure rates.
The four conditions for five-minute deduplication
The announcement combines four conditions: the same time interval, Alibaba Cloud account, model and hash calculated from request parameters. Multiple 5XX failures are combined into one for SLA calculations only when all conditions hold. If any request in the group succeeds, the group is not counted as a failure. The announcement does not disclose the hash algorithm, normalized fields or how streaming requests are grouped, nor establish how regions, workspaces or different endpoints are combined.
Do not recreate a hash yourself and claim it matches the provider's accounting. Instead, retain original requests or safely redacted reproducible parameters, client request IDs, model IDs, endpoints and regions, start/end times, response codes, retry numbers and final results. Generate your own team retry-group ID separately. This supports both real reliability measurement and original evidence in a dispute.
Why retired-model errors need separate handling
The second revision excludes service errors from models formally retired by the platform from the SLA error rate, because those models no longer have the availability commitment. This does not mean failures of old models have no impact; it means teams can no longer rely on this SLA for their availability. Model Studio's retirement policy lists notification periods, retirement effects and models. Account email, site messages and official announcements may also provide specific dates.
Compare exact production model IDs with the official retirement list now, distinguishing mainline models, dated snapshots and aliases. Only formally retired models fall under this exclusion. Do not exclude every older, preview or soon-to-retire model indiscriminately. If a model is retiring or retired, test replacement outputs, tool calls, costs, rate limits and rollback conditions on test traffic before migrating. Changing egress IP cannot restore lost model SLA eligibility.
Audit and acceptance checklist before the effective date
First, record the verification time and official announcement version, plus the team's accounts, regions, workspaces, endpoints and models. Second, export raw 5XX and successful-request samples covering at least a five-minute interval, and verify that sampling or aggregation has not lost retries. Third, calculate both the per-request business error rate and a simulation of the provider's new SLA error rate, marking which retry groups explain the difference. Fourth, check every model's lifecycle and assign migration owners and deadlines to retiring models.
Acceptance does not mean making the two rates equal. It means tracing either aggregate back to original requests and explaining why a success changes the provider's SLA numerator without changing failures users already experienced. If logs lack model IDs, time zones are confused, fingerprints are irreproducible or post-effective behavior differs from the announcement, stop using simulated figures for claim decisions, preserve evidence and verify through existing support channels.
Service credits and unknowns
The current public SLA says it applies to paid model-inference services and lists a 99.9% monthly availability commitment, service-credit tiers and claim evidence requirements. Do not automatically include free or trial services, other Model Studio products or contractual exclusions. Claims must be submitted within 30 days after the end of the billing month containing the incident. Whether credits are granted remains subject to the agreement's review process. This article has not accessed account bills, tickets or actual credit results.
On September 28, 2026, the official announcement and public agreement were rechecked. The effective date, deduplication conditions for the same account, model, interval and parameter hash, treatment of a success within a group, and formally retired-model exclusion can be verified directly. The agreement page was updated before the announcement, so its existence alone does not prove the platform switched at a particular minute. Actual switch time, hash implementation, account-level settlement, dispute handling and real credit outcomes remain UNKNOWN. This is an EXPLAINER with an audit checklist, not a guarantee of successful claims. PRODUCT_FIT=NONE: proxies or IP changes do not alter SLA accounting.
Sources
Frequently Asked Questions
If the same request fails three times within five minutes, does it count as three SLA failures?
If the account, model and request-parameter hash are identical, the announcement says multiple 5XX responses count as one failure in that five-minute interval. Actual grouping remains subject to platform records.
If a retry eventually succeeds, can earlier 5XX responses be omitted from logs?
No. The new SLA rules may exclude the group, but the business still experienced errors, latency and retry costs. Internal SLOs and troubleshooting require every event.
Does this rule also deduplicate 4XX, rate limits or client timeouts?
The announcement explicitly discusses 5XX. The current SLA separately defines total requests, authentication-validation failures, 4XX and exclusions. Do not extend the 5XX rule directly to other errors.
Can I claim service credits while continuing to call a retired model?
The announcement says formally retired models no longer have the availability commitment and their errors do not count toward the SLA error rate. Check lifecycle status and migrate; changing IP cannot restore eligibility.
Does this checklist guarantee service credits will be approved?
No. It preserves internal evidence and explains the two error rates. Account scope, materials, deadlines and final credits are determined by the full agreement and platform review.