PuppyIP Resource Center
Developer Tool Updates 10 min Published 2026-09-05 Updated 2026-09-30

GitHub Contents API degradation: the capacity-change cause and a sync reconciliation checklist

GitHub's postmortem identifies an impact window of approximately 05:45–06:07 Beijing time on September 5. Investigation was announced at 06:02 and resolution at 06:23. Newly added capacity across availability zones still routed requests to a few servers in the old zone, overloading them while new servers sat idle; rollback restored balance. Publishers and synchronization jobs should reconcile uncertain writes before read-only checks and a reversible, single-file PUT.

GitHub API Contents API API degradation Content synchronization Idempotent retries Incident recovery

Service eligibility and regional restrictions

PuppyIP serves only compliant overseas businesses and their authorized personnel. Proxy services are not available in mainland China. The service may only be used for lawful business activities outside mainland China. Use of this service within mainland China is prohibited.

Hosting a proxy IP or server overseas does not change these restrictions. The service must not be provided to end users in mainland China through relaying, forwarding, sharing or resale. Before use, read the Terms of Service.

Key Takeaways

  • Separate the actual impact window, approximately 05:45–06:07, from the 06:02 investigation and 06:23 resolution announcements.
  • The cause was routing during capacity expansion across availability zones. It does not make every 404, 409, 422, or timeout an incident-related error.
  • Inventory GET, PUT, and DELETE calls to /repos/{owner}/{repo}/contents/{path}, including repository, path, ref, time, status, request ID, and redacted logs.
  • A timed-out write or delete has an uncertain outcome. Read the target branch and blob SHA before retrying.
  • Recover with a noncritical GET, one PUT using the current SHA, and readback. Expand gradually and serialize PUT and DELETE operations.
  • A fixed outbound IP does not repair GitHub's service. Investigate connectivity only when multiple endpoints show DNS, TLS, 407, or connection-timeout evidence.

Which times describe the impact rather than the announcement?

The official postmortem places the impact at approximately 05:45–06:07 Beijing time on September 5, 2026. GitHub announced investigation at 06:02 and resolution at 06:23. Match your requests to their actual timestamps and outcomes rather than using the public announcement time as the entire outage window.

Capacity was expanded across availability zones, but routing continued to send traffic to a small set of servers in the old zone. Those servers overloaded while the new capacity remained idle; rollback balanced traffic again. The report does not break down error rates by endpoint, region, or method. A 401, 403, 404, 409, or 422 outside this window still needs its own explanation.

Which jobs need reconciliation?

Search publishing jobs, documentation synchronization, configuration distribution, bots, and CI for actual Contents API calls. Record owner, repository, path, branch or ref, method, start time, HTTP status, request ID, and retry count. Prioritize uncertain operations within and immediately around the impact window.

A working repository webpage, successful git clone, or an Operational Pull Requests component does not prove a Contents API PUT succeeded. A 404 may reflect the path, ref, or permissions. Requests outside the affected window should first be checked for authentication, parameters, branch protection, and connectivity.

What should you do with a timed-out PUT or DELETE?

A timeout means the client did not receive a confirmed response; it does not prove that the server made no change. After a PUT, GET the same path and ref and compare the blob SHA, expected content hash, and commit. After a DELETE, check whether the file still exists and inspect the latest branch commit. Pause automatic retries until the outcome is reconciled.

Updating an existing file with PUT requires its current sha; deletion also requires the blob sha. An old SHA can cause a 409, while invalid parameters or abuse checks can cause a 422. Do not overwrite newly changed content blindly or bypass branch protection to clear the backlog.

How do you make retries safe?

Track a business idempotency key, branch, old SHA, new content hash, and commit message for each operation. Read the current SHA before a retry. If the intended content is already present, mark the job complete instead of creating another write. If another operation advanced the branch, stop and recalculate the change.

GitHub warns that parallel create, update, and delete operations can conflict. Serialize writes and keep concurrency and retry limits bounded. Recover one file first, then a small queue; avoid sending an accumulated backlog as a sudden burst.

What does a controlled recovery look like?

First, GET a noncritical repository file or a test-branch file and confirm a 200 response, content, ref, and SHA. Then make one reversible PUT with a fresh SHA, record the commit SHA, and GET the file again to confirm its blob, content, and branch head. Expand from a few tasks to one repository and then to the full queue.

Restore DELETE operations last and keep them serialized with PUT operations. At each stage, inspect error rates, latency, 409 and 422 responses, duplicates, and backlog growth. If anomalies return, stop expansion and retain a verified local read-only copy or a single task approved for manual execution.

Which errors belong to permissions, content, or connectivity?

Private-repository reads need Contents read permission; writes need Contents write permission. Editing .github/workflows also requires Workflows write permission. Check tokens, repository access, and organization policy for 401 or 403; owner, repository, path, ref, and visibility for 404; SHA and concurrency for 409; parameters, content, and request frequency for 422.

Only DNS failures, TLS errors, 407 responses, TCP timeouts, or connection failures across multiple APIs justify a network-path investigation. The GitHub enterprise proxy and CA guide covers that layer. A fixed outbound address can help reproduce a stable path; it cannot repair a status incident, token permissions, or a stale SHA.

When is recovery complete, and when should writes stop?

Recovery needs more than a Resolved status: all jobs in the impact window should have a known outcome, the noncritical GET and PUT readback should pass, and there should be no duplicate writes, accidental deletions, or growing backlog. Preserve rate limits, branch protection, and any required human approval. Resume large writes and deletes only after a stable observation period.

Stop writes if the incident reopens, several repositories fail, outcomes remain unknown, SHA conflicts persist, duplicates appear, or deletions cannot be explained. Only independently established network-path evidence makes it useful to visit PuppyIP for a fixed outbound path. Otherwise, focus on GitHub status, permissions, branches, and idempotency.

Sources

Frequently Asked Questions

Has GitHub recovered from this Contents API incident?

GitHub marked the incident resolved at 06:23 Beijing time on September 5. Its postmortem places the actual impact at approximately 05:45–06:07 and identifies availability-zone routing during capacity expansion as the cause. Rollback restored balanced traffic; your uncertain operations still need reconciliation.

Can I immediately retry a PUT that timed out?

Read the same path and ref first. Compare the blob SHA, content hash, and branch commit with the intended change. The write may already have succeeded. Retry only after you establish the outcome and recompute the operation with the current SHA.

Does a 409 or 422 prove this incident affected my job?

No. A 409 can come from an old SHA or a concurrent change, and a 422 can reflect invalid parameters, content, or abuse checks. Match the actual request time and evidence rather than assigning the incident as the cause automatically.

Can PUT and DELETE calls run in parallel during recovery?

GitHub's documentation warns that concurrent create, update, and delete operations can conflict. Serialize these writes and use bounded retries and concurrency while restoring the queue.

How can I tell whether batch synchronization is safe to resume?

Confirm a noncritical GET, one reversible PUT with a fresh SHA, and readback. Check for duplicates, conflicts, permission failures, and backlog growth, then expand in stages while continuing to observe results.

Will a fixed IP solve Contents API degradation?

It will not repair GitHub's server incident, token access, or stale SHAs. Consider network-path troubleshooting only when independent evidence shows DNS, TLS, proxy-authentication, or connection failures across multiple endpoints.