We had changed part of the Shutterstock network and needed to know how much traffic the new layout could handle. It was the weekend, which gave us room to test before normal customer traffic came back on Monday.

I asked the team to turn the usual workloads back on in a controlled order. That included FTP ingestion, a content-processing job we had paused, and the video publisher. My email said, “We wanna see what the new layout can handle.” That was really all there was to it.

I also asked people to call or email before adding load. If the network moved from 500 Mbps to 800 Mbps, we needed to know what had just changed. Turning everything on at once would give us a busy graph without much explanation.

The first results looked fine. Turning the video publisher back on added about 300 Mbps between the older network and the load balancers. Six video workers running in Amazon added another 400 Mbps across the firewall and core. The load balancer was moving about 1.4 Gbps through its trunk, and we were not seeing utilization alerts.

The team kept watching as traffic increased.

Closer to peak, the link between the firewall and the core routers reached 80 percent utilization. We reduced the Amazon workers from six to four, which brought the link down to a better level. Later, another 200 Mbps appeared between the older network and the database and backend network. By then we had reduced the video workers to three.

The site was still working. Nobody was looking at an outage. We had simply found a part of the network that would run out of room sooner than the rest of it.

The problem was the shared path between the firewall and the core routers. Several workloads could be fine on their own and still fill that link when they ran together. Reducing the video workers lowered traffic, but it also reduced the amount of processing we could do.

That was useful because the layout looked reasonable on paper. The test showed us how the shared links behaved when several real workloads were running together. We did not have to argue about whether the bottleneck might become a problem later. We had watched it reach 80 percent during the test.

On Monday morning, I wrote back that the current arrangement was not okay as a longer-term setup. We needed to separate some parts of the network earlier than we had planned.

This happened fairly often at Shutterstock. We would lay out a sensible order for infrastructure work, then traffic or a new product would grow faster than the schedule. The original plan was not necessarily wrong. The timing changed because we had better information.

There was also a staffing detail in the thread that I still like. On Saturday night, someone offered to turn the video publisher back on right away. I asked if it could wait until the next day so the person covering the network could get some sleep.

We still completed the test before Monday traffic. Waiting a few hours did not put the company at risk. It meant the person watching the network had a better chance of being useful when the load actually increased.

That matters in operations. It is easy to let the same few people carry every change because they know the system and are willing to stay up. It works until they are exhausted or unavailable. We were trying to build a team that could share the context, run a planned test, and make a decision from the numbers.

The test gave us enough to make that decision. One workload added about 300 Mbps without trouble. Six cloud workers added about 400 Mbps more. Together with normal traffic, that pushed a shared link to 80 percent. Cutting the worker count gave us temporary room, but it did not change the network design.

So we moved the separation work forward. The site stayed up, the team got a clear result, and we had a specific piece of infrastructure to fix next.

Archive