Incident Summary
On 29th October 2025, OrbusInfinity experienced a major service disruption affecting users across all regions. The incident was caused by a failure in the underlying cloud provider’s global content delivery network (CDN) and domain name system (DNS) services, which are responsible for routing user requests and ensuring reliable access to OrbusInfinity infrastructure worldwide. The outage resulted in connection timeouts, DNS resolution errors, and 502 Bad Gateway errors for users attempting to access the OrbusInfinity platform. The global incident began at 15:41 UTC. The first OrbusInfinity region was recovered at 20:19 UTC, with full-service restoration achieved by 22:14 UTC, when the last impacted region was restored.
Impact
Regions Affected
All OrbusInfinity regions were impacted, including:
- Australia East
- Canada
- IRAP (Australia)
- Qatar
- South Africa
- United Arab Emirates (UAE)
- United Kingdom South
- United States East
- United States West
- West Europe
User Experience
- Users experienced slow or failed logins, timeouts, DNS errors, 502 Bad Gateway and 504 Gateway Time-out errors and an inability to access OrbusInfinity services.
Regional Recovery Timeline
Region
Incident Start (UTC)
Partial Recovery (UTC)
Full Recovery (UTC)
Australia East
15:41
20:35
20:50
Canada
15:41
20:40
21:47
IRAP (Australia)
15:41
20:40
20:50
Qatar
15:41
20:45
20:45
South Africa
15:41
20:19
20:19
UAE
15:41
20:39
20:53
UK South
15:41
20:37
20:37
US East
15:41
20:44
22:14
US West
15:41
20:41
22:14
West Europe
15:41
20:38
20:38
Root Cause Analysis
What Happened:
The incident was triggered by a specific sequence of configuration changes within the cloud provider’s global CDN and DNS services. These changes generated incompatible configuration metadata, which exposed a latent bug in the cloud provider’s data processing layer.
Technical Details (Cloud Provider):
- The incompatible configuration was deployed globally, causing asynchronous crashes across the CDN and DNS infrastructure, impacting all OrbusInfinity regions.
- The protection system, designed to validate configurations, did not detect the issue in time due to the delayed manifestation of the bug.
- The resulting failures disrupted both traffic routing and DNS resolution, leading to widespread service degradation and interruption for OrbusInfinity users.
- The cloud provider’s engineering teams responded by blocking further configuration changes, manually updating the healthy configuration snapshot, and gradually reloading configurations across all edge sites to restore service.
Resolution:
Service was restored as the cloud provider stabilized the affected infrastructure and traffic management systems. OrbusInfinity operations returned to normal with no direct remediation required at the application layer.
As with the incident on 10th October, it was not possible to mitigate the impact of the outage as it was a global incident for the cloud provider, affecting static content delivery and routing. Failover to an alternative availability zone would have had no impact in this instance. Alternatives to the OrbusInfinity routing and content delivery infrastructure are being investigated as fall-backs for such scenarios, to reduce the impact of future global traffic management and content delivery service incidents (ongoing).
Preventive Measures (Cloud Provider):
- The cloud provider has fixed the underlying defects in both the control and data processing layers.
- Additional validation stages and extended “bake time” have been introduced in the configuration deployment process.
- Enhanced isolation and recovery procedures have been implemented to reduce the risk and impact of similar incidents in the future.