Service eligibility and regional restrictions
PuppyIP serves only compliant overseas businesses and their authorized personnel. Proxy services are not available in mainland China. The service may only be used for lawful business activities outside mainland China. Use of this service within mainland China is prohibited.
Hosting a proxy IP or server overseas does not change these restrictions. The service must not be provided to end users in mainland China through relaying, forwarding, sharing or resale. Before use, read the Terms of Service.
Key Takeaways
- An HTTP status code usually means the request reached some service layer. DNS failures, connection timeouts and TLS failures should first be treated as network-layer issues.
- 401, 403, 404, 429 and 5xx have different responsible layers; they cannot all be blamed on keys, balances or proxies.
- A relay may rewrite errors. Record HTTP status, error type, request ID, Retry-After and the redacted entry-point domain together.
- Use limited backoff only for transient failures such as 429, interrupted connections and some 5xx responses. Correct configuration, permission and model errors before sending again.
- Change only one variable per round and set maximum retries and stop conditions.
Save evidence before deciding whether to retry
At the first error, record the time and time zone, HTTP status, the error object’s type or code, request-id, Retry-After, model ID, and the redacted Base URL domain and path structure. For keys, record only the owning project and most recent rotation time; never include keys in screenshots, logs, tickets or chats.
If the relay dashboard provides request logs, check whether the request arrived, which upstream model it actually reached, whether tokens or charges were generated, and whether the final status came from the relay or upstream service. If these fields are unavailable, make one minimal request rather than repeatedly refreshing to guess the cause.
Distinguish network failures, relay responses and upstream responses
DNS resolution failure, connection timeout, TLS handshake failure or an interrupted connection may mean no valid HTTP response was received. Check DNS, ports, certificate chains, system time and proxy connectivity first. For the general layers, see the proxy connection troubleshooting checklist; for unclear address structure, see the proxy URI format guide.
If you received a 4xx or 5xx, do not rely solely on the client’s final line. Compare response headers, JSON error type, request ID and relay logs to determine whether the entry gateway, model router or upstream service returned it. When a relay wraps several upstream errors in one Chinese message, HTTP status and original error fields are more reliable than the message heading.
401 and 403: repair authentication and permissions before retrying
For 401, first check whether the API Key is complete, revoked or expired, whether the authentication header is correct, and whether the key belongs to the current endpoint and project. A 403 is more likely to concern project, organization, workspace, model or API permissions. Recognizing a key does not mean it can access the requested resource.
These errors generally should not be automatically retried unchanged. Check the current service’s official documentation and console for authentication methods, key status and permissions, then send one minimal request. For the basic relationship between API Key, Base URL and model fields, see the AI API configuration and error guide.
404 and model not found: check paths, versions and mappings
A 404 may mean the endpoint path or resource does not exist, the model ID is missing, the model is unavailable to the current project, or the relay has no route for it. First check whether the client appends a version path automatically, then verify the Base URL, API version, resource path and model ID.
Do not change the key, endpoint, model and proxy all at once after model not found. Copy a real ID from the current service’s model catalog or supported-models endpoint, then confirm the relay dashboard’s mapping name. Change only the model and retest once. If it still fails, retain the request ID and entry logs and contact the relevant service’s support.
429 does not always mean insufficient balance: read rate-limit and quota evidence
A 429 may come from requests per minute, input or output tokens, concurrency, daily quotas, spending limits, the relay’s own package, or upstream acceleration limits. Read the error body, Retry-After, remaining limits, reset time and dashboard usage first. Do not diagnose it from “too many requests” or “insufficient balance” alone.
Once a transient rate limit is confirmed, wait according to Retry-After. Without an explicit wait time, use exponential backoff with random jitter and a maximum attempt count. Do not have several tasks retry unchanged at once. For bills and whether failed requests are charged, see the AI relay billing and request-record guide.
5xx, timeouts and interrupted streams: retry only transient faults
For 500, 502, 503, 504 or service overload, first inspect the official status page, relay status and logs for the same request ID. Retry a limited number of times only after confirming a temporary failure. If a fixed model and endpoint keep returning the same error, stop and submit the request ID instead of switching exits indefinitely.
A streaming request may fail midway after returning 200. Save the last complete event, client timeout, received tokens and the point where the connection closed. Then determine whether the server failed, the client closed early or an intermediate network interrupted the stream. For context and streaming boundaries, see the AI API context and streaming guide.
Establish a comparison with a five-minute minimal request
Keep the machine, client version, endpoint and model fixed, and send only a minimal text request without business data. Preserve the original configuration for the first attempt. For the second, change just one variable already supported by evidence. Record both times, status codes, error types, request IDs, durations and whether charges occurred.
After the minimal request succeeds, gradually add real context, streaming or tool calls. Stop if the minimal request still fails with unchanged evidence; enlarging the request adds only cost and diagnostic noise. For the responsibilities of official and relay endpoints, see the official API versus relay entry-point guide.
At the stop condition, send redacted evidence to the right support team
Stop automatic retries if 401/403 remains unresolved, a 404 path or model is unconfirmed, a 429 wait window is still active, a bounded sequence of 5xx attempts has not recovered, or every request incurs charges without useful output. Support material should include at least the time, entry-point domain, model ID, status code, error type, request ID, Retry-After and minimal reproduction steps.
A proxy affects only the connection path; it cannot fix invalid keys, model permissions, quotas or upstream outages. For a stable fixed exit for lawful development and documentation access, visit the PuppyIP website and follow the PuppyIP tutorials after purchase. Continue to follow the service’s region, account, rate-limit and usage rules.
Sources
Frequently Asked Questions
Can an AI API 401 be retried automatically?
Usually not unchanged. Correct the key, authentication header, project or endpoint relationship first, then send one minimal request.
Does 429 always mean the relay balance is insufficient?
No. It may also reflect requests, tokens, concurrency, daily quotas, spending limits or upstream acceleration limits. Inspect the error body, rate-limit headers and dashboard usage.
Why does model not found sometimes return 404?
A missing model ID, API version, endpoint path or relay mapping can all produce 404. Confirm the models and paths actually supported by the current endpoint.
Should I switch proxies immediately after 502, 503 or 504?
Do not switch first. Check official or relay status, request IDs and logs. Investigate the proxy network when DNS, connection or TLS evidence points there.
Why can a streaming request fail after returning 200?
The server may send an error event after the stream begins, or a client timeout or intermediate network may close it early. Save the last event and the point of closure.
What information should I retain before contacting support?
Keep at least the time, entry-point domain, model ID, status code, error type, request ID, Retry-After, minimal reproduction steps and redacted logs. Never send real keys.