OrbusInfinity Outage - majority of regions

Incident Report for Orbus Software

Postmortem

Incident Summary

On 29th October 2025, OrbusInfinity experienced a major service disruption affecting users across all regions. The incident was caused by a failure in the underlying cloud provider’s global content delivery network (CDN) and domain name system (DNS) services, which are responsible for routing user requests and ensuring reliable access to OrbusInfinity infrastructure worldwide. The outage resulted in connection timeouts, DNS resolution errors, and 502 Bad Gateway errors for users attempting to access the OrbusInfinity platform. The global incident began at 15:41 UTC. The first OrbusInfinity region was recovered at 20:19 UTC, with full-service restoration achieved by 22:14 UTC, when the last impacted region was restored.

Impact

Regions Affected

All OrbusInfinity regions were impacted, including:

  • Australia East
  • Canada
  • IRAP (Australia)
  • Qatar
  • South Africa
  • United Arab Emirates (UAE)
  • United Kingdom South
  • United States East
  • United States West
  • West Europe

User Experience

  • Users experienced slow or failed logins, timeouts, DNS errors, 502 Bad Gateway and 504 Gateway Time-out errors and an inability to access OrbusInfinity services.

Regional Recovery Timeline

Region Incident Start (UTC) Partial Recovery (UTC) Full Recovery (UTC) Australia East 15:41 20:35 20:50 Canada 15:41 20:40 21:47 IRAP (Australia) 15:41 20:40 20:50 Qatar 15:41 20:45 20:45 South Africa 15:41 20:19 20:19 UAE 15:41 20:39 20:53 UK South 15:41 20:37 20:37 US East 15:41 20:44 22:14 US West 15:41 20:41 22:14 West Europe 15:41 20:38 20:38

Root Cause Analysis

What Happened:

The incident was triggered by a specific sequence of configuration changes within the cloud provider’s global CDN and DNS services. These changes generated incompatible configuration metadata, which exposed a latent bug in the cloud provider’s data processing layer.

Technical Details (Cloud Provider):

  • The incompatible configuration was deployed globally, causing asynchronous crashes across the CDN and DNS infrastructure, impacting all OrbusInfinity regions.
  • The protection system, designed to validate configurations, did not detect the issue in time due to the delayed manifestation of the bug.
  • The resulting failures disrupted both traffic routing and DNS resolution, leading to widespread service degradation and interruption for OrbusInfinity users.
  • The cloud provider’s engineering teams responded by blocking further configuration changes, manually updating the healthy configuration snapshot, and gradually reloading configurations across all edge sites to restore service.

Resolution:

Service was restored as the cloud provider stabilized the affected infrastructure and traffic management systems. OrbusInfinity operations returned to normal with no direct remediation required at the application layer.

As with the incident on 10th October, it was not possible to mitigate the impact of the outage as it was a global incident for the cloud provider, affecting static content delivery and routing. Failover to an alternative availability zone would have had no impact in this instance. Alternatives to the OrbusInfinity routing and content delivery infrastructure are being investigated as fall-backs for such scenarios, to reduce the impact of future global traffic management and content delivery service incidents (ongoing).

Preventive Measures (Cloud Provider):

  • The cloud provider has fixed the underlying defects in both the control and data processing layers.
  • Additional validation stages and extended “bake time” have been introduced in the configuration deployment process.
  • Enhanced isolation and recovery procedures have been implemented to reduce the risk and impact of similar incidents in the future.
Posted Nov 12, 2025 - 14:02 UTC

Resolved

All regions are now operational. The incident has been marked as resolved.

A full incident report will published in due course.
Posted Oct 30, 2025 - 00:11 UTC

Monitoring

While the fix has been implemented in all regions, we continue to monitor the situation. As of now, North America is still intermittently reporting failures, all other regions have been operational for the past 30mins.
Posted Oct 29, 2025 - 21:24 UTC

Update

While issue continues to affect majority of regions, we start to see the signs of recovery in more regions.
Posted Oct 29, 2025 - 20:35 UTC

Update

While issue continues to affect majority of regions, we start to see the signs of recovery.
Posted Oct 29, 2025 - 20:20 UTC

Identified

OrbusInfinity is currently impacted by a global incident with our cloud provider. The provider is working on a fix, and we’ll keep you informed with updates as progress is made.
Posted Oct 29, 2025 - 17:07 UTC

Investigating

We are investigating an outage across all OrbusInfinity regions. Users are experiencing 502 Bad Gateway errors when trying to access the platform.

Updates will be provided in due course.
Posted Oct 29, 2025 - 16:26 UTC
This incident affected: Australia (OrbusInfinity), Canada (OrbusInfinity), Qatar (OrbusInfinity), South Africa (OrbusInfinity), United Arab Emirates (OrbusInfinity), United Kingdom (OrbusInfinity), Europe (OrbusInfinity), United States East (OrbusInfinity), United States West (OrbusInfinity), and OrbusInfinity Government (IRAP) (OrbusInfinity).