Three people. Three in the morning. Three and a half hours.
The site stayed up.
We were seeing twice the normal number of connections to the database replicas. Search servers were spiking, processes were consuming too much memory, and the machines could lock hard enough that a normal remote login was useless.
We had installed remote-management cards so we could force a reboot. That was better than driving to the data center. It was still slower than a service failing cleanly and recovering on its own.
The cause was not obvious. A search-indexing change might have increased the load. One API user might have been hammering the site. We blocked a suspicious address, watched connections drop, then lifted the block and saw them remain stable.
Useful clue. Not proof.
The team had built a small tool to inspect connections across the web servers. What we did not have was a good way to trend the heaviest API users over time.
The screenshots told the overnight story in pieces. Database lines climbed together. Search-load spikes appeared where the graph was normally flat. A before-and-after connection list showed what changed when we blocked one address. None of it supplied the answer by itself.
That became part of the follow-up. The DevOps group needed thresholds, history, and a way to distinguish a growing customer from abuse before the database told us at 3:00 a.m.
We scheduled a retrospective for later that morning.
Staying up was the immediate win. Understanding why we almost did not was the actual job.