Service eligibility and regional restrictions
PuppyIP serves only compliant overseas businesses and their authorized personnel. Proxy services are not available in mainland China. The service may only be used for lawful business activities outside mainland China. Use of this service within mainland China is prohibited.
Hosting a proxy IP or server overseas does not change these restrictions. The service must not be provided to end users in mainland China through relaying, forwarding, sharing or resale. Before use, read the Terms of Service.
Key Takeaways
- The DeepSeek scope includes deepseek-v3, v3.1, v3.2, v3.2-exp, r1, r1-0528 and three R1 distilled models; the GLM scope is glm-4.6 and glm-4.7, for 11 model IDs total.
- The DeepSeek page recommends qwen3.7-plus, qwen3.7-max and qwen3.6-flash; the GLM page recommends qwen3.7-plus, qwen3.8-max and qwen3.8-flash. Do not merge the lists into one-to-one mappings.
- The deadline is October 10, 2026. Neither official page specifies an execution time, time zone, automatic routing, compatibility guarantee, price equivalence or post-retirement error body. All remain UNKNOWN.
- The DeepSeek and GLM documents list different regions, endpoints and Responses API capabilities. Validate Workspace, API Key, base URL, protocol and target model as a complete set.
- Migration acceptance should cover output quality, thinking parameters, tool calls, structured output, stream termination, context, rate limits, cost, monitoring and failure handling.
- PRODUCT_FIT=CONDITIONAL_NETWORK_ONLY: Fixed egress cannot extend model availability. Investigate the network only for DNS, TLS, 407, timeouts or path differences with identical configuration.
What is happening: Nine DeepSeek and two GLM IDs retire on the same day
As directly checked on September 21, 2026, the official DeepSeek page lists deepseek-v3, deepseek-v3.1, deepseek-v3.2, deepseek-v3.2-exp, deepseek-r1, deepseek-r1-0528, deepseek-r1-distill-qwen-7b, deepseek-r1-distill-qwen-14b and deepseek-r1-distill-qwen-32b. The official GLM page lists glm-4.6 and glm-4.7. Both pages give October 10, 2026 as the retirement date.
Official documentation does not publish an exact execution time, time zone, per-tenant or per-region order, grace period, automatic fallback, or post-retirement error codes or bodies. First locate all 11 IDs in code, environment variables, configuration systems, gateways, scheduled tasks, batches, evaluation scripts and billing tags. Record the business owner, protocol, region and last successful call.
Keep replacement lists separate: Recommendation is neither equivalence nor automatic mapping
The DeepSeek page recommends qwen3.7-plus, qwen3.7-max and qwen3.6-flash; the GLM page recommends qwen3.7-plus, qwen3.8-max and qwen3.8-flash. The shared qwen3.7-plus recommendation does not make it suitable for every workload. Nor should the other four candidates be substituted across families with an assumption of identical behavior.
There is no official one-to-one mapping from old IDs to new IDs, or promise of equivalent context, thinking, tool calls, structured output, latency, quotas or prices. Build candidate matrices separately for coding, reasoning, low latency, batch processing and tool use. Mark prices and limits UNKNOWN until checked against account billing and real requests.
Fix the region and endpoint first: Do not transfer one family's region list to the other
The DeepSeek page lists OpenAI-compatible and DashScope endpoints in Beijing, Virginia, Singapore, Frankfurt and Tokyo. The GLM page lists Beijing, Virginia, Frankfurt, Hong Kong and Singapore. Models and rate limits can differ by region, so success in one region does not establish availability in another.
For every call path, retain the region, base URL, Workspace ID reference, Key owner, model, protocol and rate-limit configuration. Do not configure a Beijing Key against a Singapore endpoint or replace only the model string while retaining unverified region, data-boundary or quota assumptions. Never write the key itself into logs, test fixtures or resource pages.
Retest protocols and parameters: Separate Chat Completions, Responses and DashScope
Existing calls may use OpenAI-compatible Chat Completions, DashScope or the Anthropic-compatible Messages interface listed on the DeepSeek page. The DeepSeek page says Responses API currently supports only its listed DeepSeek V4 family, in Beijing and Singapore only. The GLM page says Responses API currently supports only glm-5.2 and glm-5.3, also only in Beijing and Singapore.
Neither current capability statement proves that old DeepSeek V3/R1 or GLM 4.6/4.7 requests automatically become Responses requests. enable_thinking is not a standard OpenAI parameter: the Python SDK example passes it through extra_body, while the Node.js example places it at the top level. After migration, check the target model's documentation instead of copying the parameter and assuming it works.
Select with fixed samples: One 200 response does not complete migration
Save redacted fixed samples and human evaluation criteria for every old call. Record input type, thinking settings, temperature, max tokens, tool schema, JSON constraints, streaming parser, retry behavior and concurrency. Run the same samples against each official candidate for that family, comparing actual model, request ID, time to first token, total duration, token usage, tool calls, structured results and human-assessed quality.
Check prices against the target region's official pricing page and actual account bills: models, regions, context tiers and caching rules may all alter costs. Without quality, cost, quota and failure-path verification, a model is only a candidate. Do not treat one HTTP 200 or demonstration prompt as production acceptance.
Staged rollout and fallback: Change one variable through an auditable switch
Complete fixed-sample testing outside production first, then expand gradually through low-risk internal tasks, read-only business workloads and a small amount of production traffic. Switch the target model per workload with an auditable configuration toggle. Do not simultaneously change region, SDK, prompts, tool schema and retry strategy, or you will not be able to attribute changes in results.
Set error-rate, quality, cost, latency and tool-use thresholds at each stage. Stop expanding if structured outputs break, tools exceed authorization, costs cross the threshold, quota is insufficient or results cannot be explained. Before October 10, you can return to an old model that is still available; after the deadline, use another validated model or pause the affected queues.
Investigate model, region, authentication, quota and network separately
For 404 or model unavailable, first check model ID, retirement date and regional availability. For 401/403, check Workspace, Key owner, permissions and account state. For 429, check regional quota, concurrency and rate limits. For 400, check protocol and parameters. For 200 with changed output, return to fixed samples, actual model and tool schema. Do not hide a retired model or an unknown outcome from a side-effecting request behind infinite retries.
Use the site's Systematic Proxy Connection Troubleshooting Checklist only for DNS-resolution failures, TLS-handshake errors, proxy 407, connection timeouts, or consistent differences between compliant egress paths with the same account, region and endpoint. Fixed egress cannot restore retired models, grant permissions, change quota or make a cross-region Key valid.
Completion and rechecks: Remove implicit dependencies on all 11 old IDs
Completion requires at least: zero remaining uses of the 11 old IDs or explicit deactivation records; every production path bound to the correct region, Workspace, Key and protocol; target-model acceptance for functionality, quality, cost, quotas and failure handling; monitoring of the actual model; rehearsed fallback or queue-pausing procedures; and removal of temporary permissions and test data.
Recheck the DeepSeek, GLM, model-list, region and pricing pages before switching, one week before the deadline and again on October 9. If the precise time remains unspecified, finish early using the most conservative boundary; do not invent 00:00. If a candidate is unavailable in the target region, account prices cannot be read or a critical task has no acceptable replacement, stop expansion and have the business owner choose degradation or suspension.
Sources
Frequently Asked Questions
Which Model Studio models retire on October 10?
There are 11: nine DeepSeek V3, R1 and distilled model IDs, plus glm-4.6 and glm-4.7. Inventory the exact IDs individually against both official pages.
Should DeepSeek and GLM migrate to the same Qwen models?
Do not generalize that way. Only qwen3.7-plus appears on both official recommendation lists. Other candidates differ, and no one-to-one mapping or behavioral equivalence is promised.
At what time on October 10 will the models retire?
Both official pages give only a date, without a time or time zone. Mark the exact execution time UNKNOWN and finish migration before the deadline.
Is replacing the model parameter enough?
No. Also check regional endpoints, Workspace, API Key, protocol, thinking parameters, tool calls, output quality, quotas, prices, monitoring and failure handling.
Can old Chat Completions calls be changed directly to Responses API?
Do not assume so. Both documents currently specify supported model and region restrictions for Responses. Verify protocol migration separately from model replacement.
Can I fall back to old model IDs after the deadline?
Do not treat old IDs as a reliable fallback. Use another validated model that remains available, or pause the affected queues and handle them manually.
Can switching to a fixed IP let me keep calling retired models?
No. Retirement is a platform availability rule. Fixed egress is useful for investigation only when evidence points to DNS, TLS, 407, timeouts or consistent path differences.