The data servers were not fine.

We had seen enough issues to cause recent outages. I sent a plain request to the person closest to that part of the system: can we talk today?

No incident manifesto. No giant distribution list. The problem needed two people with the right context in the same conversation.

Shutterstock was collecting and serving more than images. Contributor statistics, product data, operational metrics, and the systems behind them were growing too. That created new pressure in places that had been good enough at a smaller scale.

The dangerous phrase in infrastructure is “it usually works.”

Another dangerous phrase is “we should chat sometime.” I asked for that day because outages had already converted the topic from architecture into operations.

Usually is not a reliability target. It is a sign that the team has learned to live around a problem.

Recent outages meant the workaround period was over. We needed to understand which servers were failing, what the applications expected from them, and whether the system had enough separation to keep one bad component from spreading the damage.

The conversation also had to cross team boundaries. Systems could see the machines. Developers could see how the product used them. Neither view was complete by itself.

That was increasingly my job at Shutterstock: find the point where the technical problem stopped matching the org chart and pull the right people together.

Sometimes leadership was a strategy. Sometimes it was a two-line email that said we need to talk.

This one was the second kind.

Archive