OrbusInfinity - Performance Degradation

Incident Report for Orbus Software

Postmortem

Incident Summary

On 9th October 2025, OrbusInfinity experienced significant service degradation affecting users in the EMEA and African regions. The incident began at 07:50 UTC and was fully resolved by 12:50 UTC and resulted in periodic HTTP errors or "Something Went Wrong" messages for impacted users. The root cause was traced to a critical failure in the underlying cloud provider’s global traffic management and content delivery service - a platform responsible for routing user requests to the most appropriate OrbusInfinity infrastructure node for optimal performance and reliability.

Impact

Regions Affected:

  • EMEA and Africa (including UK, Europe West, South Africa North, UAE North).

User Experience: 

  • Users in affected regions experienced timeouts, HTTP errors, and slow or failed logins.
  • Authentication and data access via Microsoft Entra ID and SharePoint were disrupted.

Root Cause Analysis

The outage was caused by a critical failure in the cloud provider’s global traffic management and content delivery service. This service is responsible for directing user requests to the nearest and healthiest OrbusInfinity infrastructure node, ensuring high availability and optimal performance.

  • A previously unknown software defect in the cloud provider's service control system led to the propagation of faulty configuration data.
  • This caused a significant portion of the cloud provider's traffic management infrastructure in EMEA and Africa to crash, resulting in approximately 30% of resources in these regions becoming unavailable.
  • The remaining healthy infrastructure was overloaded as it attempted to absorb redirected traffic, leading to further degradation.
  • The cloud provider’s engineering teams performed automated restarts, manual interventions, and traffic failover operations to restore service.

Service was restored as the cloud provider stabilized the infrastructure. OrbusInfinity operations returned to normal with no direct remediation required at the application layer.

It was not possible to mitigate the impact of the outage as it was a global incident for the cloud provider, affecting static content delivery and routing. Failover to an alternative availability zone would have had no impact in this instance. Alternatives to the OrbusInfinity routing and content delivery infrastructure are being investigated as fall backs for such scenarios, to reduce the impact of future global traffic management and content delivery service incidents.

Posted Nov 04, 2025 - 10:37 UTC

Resolved

After continuous monitoring and confirmation service levels are restored to normal levels in all regions, the incident has been fully resolved.

A postmortem report will be shared on completion of post-incident review activities.
Posted Oct 09, 2025 - 17:42 UTC

Monitoring

Service has been restored. HTTP errors should no longer be encountered although you may experience slight latency whilst the cloud provider continues to make adjustments and recover additional resources.

The incident remains open for monitoring.
Posted Oct 09, 2025 - 14:17 UTC

Update

The cloud provider continues to work towards full resolution on their side. All OrbusInfinity environments are stable at this time. The incident will remain open and active until confirmation of full service resolution is received from the cloud provider.
Posted Oct 09, 2025 - 13:15 UTC

Update

The cloud provider continues to restart the underlying and impacted instances. As a result, OrbusInfinity performance and stability is improving with most instances operating as normal at this stage.

Full mitigation is expected within the next 90 minutes.
Posted Oct 09, 2025 - 11:31 UTC

Update

It has been confirmed by the cloud provider, that a significant capacity loss of ~30% of impacted instances was recorded at 0740 UTC - predominantly across Europe and Africa. Investigations are ongoing.

All OrbusInfinity services across the North America & Australia regions are operational and have been throughout this incident.
Posted Oct 09, 2025 - 10:33 UTC

Identified

OrbusInfinity issues are being caused by an associated global incident being experienced by the underlying cloud provider. Work to address the incident is ongoing by the cloud provider.

Further updates will be provided as things progress.
Posted Oct 09, 2025 - 09:18 UTC

Investigating

We are currently reviewing issues relating to degraded performance. You may encounter slow loading of the OrbusInfinity platform and / or periodic HTTP errors.
Posted Oct 09, 2025 - 08:42 UTC
This incident affected: Australia (OrbusInfinity), Canada (OrbusInfinity), Qatar (OrbusInfinity), South Africa (OrbusInfinity), United Arab Emirates (OrbusInfinity), United Kingdom (OrbusInfinity), Europe (OrbusInfinity), United States East (OrbusInfinity), United States West (OrbusInfinity), and OrbusInfinity Government (IRAP) (OrbusInfinity).