Incident summary
Earlier today, an internal system responsible for processing appointment availability became overloaded and stopped responding, which temporarily interrupted availability updates across the platform. Service has now been fully restored and is operating normally. Below is a summary of what happened, the impact, and the steps we've taken.
Timeline
Time | Event |
|---|
~09:15 | The system processing availability became unresponsive and availability updates stopped. |
09:15 to 11:00 | Availability could not be processed. Slot information was not refreshed during this period. |
11:00 to 11:45 | Service restored and performance briefly degraded while the system caught up on outstanding work. |
11:45 | Normal service resumed and confirmed stable. |
Impact
While the system was affected, availability could not be processed, so slot information across the platform was not refreshed during the incident. In practice this meant that newly opened, changed, or filled slots may not have been surfacing, so availability views and booking outcomes could have been out of date for the affected period. Some booking attempts may have failed or needed to be retried while the system was unavailable.
The disruption was limited to availability processing. No patient or practice data was lost, and the rest of the platform remained accessible throughout. Once service was restored, all outstanding work was processed and slot information returned to an accurate, up-to-date state.
What caused it
We recently made a change to how the platform retrieves availability from connected systems. This change increased the everyday, background level of load on the underlying system that processes availability. Under normal daytime conditions this remained within safe limits, but it reduced the spare capacity available to absorb busier periods.
When the busiest period of the day arrived, the combination of this raised background load and the normal peak in demand exceeded what the system could handle at once. It reached full capacity and stopped responding, which halted availability processing until we intervened. The issue was the result of this specific change to how availability is retrieved, rather than a wider or ongoing limitation of the platform, which is why the preventive steps below focus on reducing that background load and giving ourselves earlier warning in future.
Corrective and preventive actions
Action | What it does |
|---|
Faster detection | We're implementing an improved alerting system so issues like this are flagged much sooner, before they become outages. |
Load protection | We're rolling out a caching layer to reduce load on the system and lower the risk of it becoming overwhelmed in the same way again. |
Safer rollout process | We're moving to a less aggressive, less continuous approach to retrieving availability, rather than syncing everything at once. |