Three people. Three in the morning. Three and a half hours.

The site stayed up.

We were seeing an abnormal number of database connections. The search layer was also spiking, consuming too much memory, and leaving machines difficult to reach normally.

We had a recovery path when a machine stopped responding. It was better than having no way in, but it still cost time while the site was under pressure.

The cause was not obvious. A search-indexing change might have increased the load. One API user might have been hammering the site. We blocked a suspicious address, watched connections drop, then lifted the block and saw them remain stable.

It was a useful clue, not proof.

The team had built a small tool that let us inspect current connection patterns. What we did not have was a useful history of the heaviest API consumers. We could see what was happening at that moment without being able to compare it cleanly to the hours or days before it.

The screenshots told the overnight story in pieces. Database activity climbed. Search-load spikes appeared where the graph was normally flat. A before-and-after connection list showed what changed when we blocked one address. None of it supplied the answer by itself.

I thought we would eventually need code to set sensible API thresholds and trend the largest consumers over time. That was something I intended to discuss with the DevOps group. It was an idea for the follow-up, not a conclusion from the night.

We scheduled a retrospective with the WebOps group and the soon-to-be DevOps group for later that morning.

For now, the important facts were limited. Three of us had been online from three to six-thirty. The site had not gone down again. The database and search behavior remained abnormal, and we still did not know whether the indexer, one API consumer, or something else was responsible.

The site stayed up. The explanation still needed work.

Archive