Summary
On July 23, 2026, between 7:44 a.m. and 2:26 p.m. PDT, some users experienced elevated errors while using ChatGPT, the API Platform, Codex, and file-related features. The issue was caused by a network maintenance failure at an infrastructure provider that disrupted traffic entering and leaving one region. Following initial recovery at 11:16 a.m. PDT, automated DDoS protections from the infrastructure provider caused a second wave of network disruptions. All impacted services recovered by approximately 2:26 p.m. PDT.
Impact
During this period:
Some users experienced elevated errors or latency while using ChatGPT, the API Platform, Codex, and file-related features.
File-related features include File Search, file-download, attachment, and image-moderation requests.
Root Cause
The infrastructure provider’s maintenance system unintentionally expanded the scope of routine network maintenance to additional network devices, including devices that carried traffic into and out of its regions. As part of that maintenance, network routes used to direct traffic were removed between part of the provider’s wider network. This disrupted regional, cross-region, and private network traffic, causing requests between our services and their dependencies to fail or time out. A second period of impact occurred when the provider’s automated DDoS protections activated across additional regions and dropped traffic.
Resolution
Engineers began mitigating the incident at approximately 8:49 a.m. PDT by shifting traffic away from the affected region. The infrastructure provider rolled back the maintenance change, restoring regional network connectivity by approximately 11:26 a.m. PDT. When automated DDoS protections caused additional errors, engineers coordinated with the provider to lift the DDoS mitigations affecting critical infrastructure and increase network limits. All impacted services recovered by approximately 2:26 p.m. PDT.
Prevention and Improvements
We have implemented or are working on several improvements to reduce the likelihood and impact of similar incidents:
Expand regional failover. Increase multi-region coverage for critical customer-facing and internal services so a regional network disruption has a smaller impact.
Automate traffic steering. Improve detection and automated traffic-shifting recommendations when regional networking or provider limits degrade.
Improve network capacity planning. Add proactive monitoring and capacity planning for high-volume private network links and shared infrastructure limits.
Strengthen provider safeguards. Work with our infrastructure provider on faster, more targeted DDoS mitigation and clearer alerts when limits are approached or reached.
Exercise regional-failure recovery. Expand testing for regional traffic loss, dependency failover, and recovery after broad network mitigations.
We apologize for the disruption and are continuing to strengthen the reliability of the systems that support ChatGPT, the API Platform, Codex, and File Search.