The POP nodes died overnight. About a thousand users were affected for four hours because nobody answered the monitoring call at three in the morning.
That was the new entry in Peek's outage tracker.
I had not officially started at Peek yet. I was already reading through the technical operations material for the role I expected to take. The tracker was the fastest way to see what the service had been doing before I arrived.
Peek sold a dedicated mobile device for email and messaging. The device looked simple because the systems behind it handled the complexity. It connected to outside email providers, moved messages through Peek's service, and delivered them to something a customer carried.
When the POP nodes stopped, the person holding the device did not see a node failure. Email simply stopped arriving.
The tracker recorded the time, duration, affected users, cause, and proposed response. This entry was direct. Monitoring detected the failure and placed a call at three in the morning. Nobody answered. The outage continued for four hours and affected about a thousand people before service returned.
The incident did not need a long list of internal project names or system details to show the operating gap. Detection had worked. Response had not.
The next questions were practical. Who was covering the overnight hours? When did an unanswered alert escalate? Who owned the incident after the first call failed?
I knew the shape of that problem from Beatport. A service used across time zones did not stop when the office closed. Monitoring could identify a failure, but the on-call plan still needed somebody to answer.
The nodes were back. The four-hour response gap was now written down.