I want a phone call any time the site goes down.

I joined Shutterstock in October, and I am still getting up to speed on the systems, the failure modes, and the way the team responds. A dashboard alert is not enough for me right now. If there is a critical failure, call my cell and pull me into it.

I need to know what is happening. I also want to talk through what failed and what fixed it so we can come back later and address the cause. The immediate response restores the site. The follow-up work should reduce the chance of the same incident happening again.

I also want to speak with the person who fixed it and give them some credit. Getting called into a broken site at night is difficult work. If somebody finds the problem and restores service, I want to acknowledge it directly.

We added me to the notification group. I tested my notifier and it called. Then I asked the team to trigger a failure from the site side so we could verify the full route. It is easy to test the phone in isolation and assume the system works. I want to know the production alert reaches the notifier, the notifier reaches me, and the person receiving it knows what to do next.

On Monday, the office firewall failed with extremely poor timing. We were already rebuilding the network during the sprint, but the old problem arrived before the new path was ready. The immediate response moved us onto new firewalls, and new switches are planned before the end of the year.

That should leave the office network in better shape, but other systems still need work.

There are data-center technologies and network systems we need to replace. The review will probably identify more work than we can complete in one sprint. I want the company to understand which problems already existed, which ones we have identified, and when we plan to address them.

The timing makes that harder to communicate. People experience the current outage, not the improved architecture we are trying to reach. They need to know what happened and what we are doing, without a technical explanation turning into an excuse.

I sent a note to the office after the firewall issue. I also asked for a broader state-of-the-union conversation when I return next week so the departments understand the rebuilds and can support the effort. Infrastructure work gets much easier when the company knows which risks are real, which changes are underway, and where a temporary disruption may appear while we replace something fragile.

Today is the last workday before the holiday for a lot of people. I told the systems team I do not want anybody leaving early until we talk as a group about the current issues, alerts, and on-call availability. They can choose the time. I am available whenever they call.

The note was short and blunt. Fewer people will be around during the holiday, so we need to confirm who is available and who owns each type of alert.

I saw similar problems at Beatport and Peek. Redundant firewalls still need correct configuration. Monitoring still needs a person to answer. If only one person understands a repair, we need to document it and share the knowledge with the rest of the team.

At Shutterstock, I am new enough that I need the calls. I need to hear the diagnosis, learn the system, meet the people who keep it running, and build the list of permanent work underneath the incidents.

Later, a strong on-call system should not depend on calling one director for every event. Right now, this is the fastest way for me to build context and make the responsibility explicit.

Before people leave for the holiday, we need a group call. We will review the current issues, confirm alerts, list on-call availability, and name the person responsible for each open risk. If the site fails, call me so I can help with the incident and the follow-up.

Archive