Part of the contributor base could not upload.

One upload server had failed. The storage array behind it needed a restore. While that happened, two people were looking for a way to move affected contributors onto another server.

That was the incident in three lines.

The customer impact made it urgent. Contributors supplied the images Shutterstock sold. If they could not upload, the marketplace stopped receiving new inventory even if the customer-facing site still looked normal.

This was the other side of availability. A homepage can return 200 while a business process is broken behind it.

For contributors, upload was the product. They had prepared files, keywords, and releases, then reached a server that could not take the work. A partial outage was still a complete stop for the people routed to the failed system.

We split the recovery work. One person owned the storage restore. Two others looked for a routing workaround. I handled the update and promised a clearer timeline when we had facts.

The first message did not guess at a restoration time. We did not yet know enough.

It also did not hide behind “degraded service.” I wrote what people could not do. That made the impact easy for the rest of the company to understand.

That restraint mattered. People wanted a number immediately, but a confident fiction would only create a second problem later.

The right sequence was simple: state the impact, name the work underway, assign ownership, and update when the recovery path became real.

It was not a sophisticated framework. It was what the situation needed.

Restore the array. Move whoever could move. Keep the update honest.

The contributors still needed their upload box back.

Archive