My whole email this morning was one sentence: “Let’s double the number of replicate workers :)”

That is not enough information to reconstruct the full change, and I am not going to pretend it is. It is enough to place the note inside a much larger stretch of storage work happening at Shutterstock in 2013.

We were moving and rebalancing an enormous image library while continuing to serve customers. Earlier in the year, the replication queue had been measured in hundreds of thousands of files. Hardware had to be added, data had to be copied, and the team had to watch whether the work itself was hurting the live system.

Replication sounds like background activity until it competes with customer traffic for disk, network, or CPU. More workers can move the queue faster. They can also create a new bottleneck or make an old one visible. The smiley face did not replace the monitoring.

I like this small artifact because it shows what a lot of infrastructure leadership actually looks like. It is not always a big architecture document or a conference presentation. Sometimes a team has already done the investigation, a safe lever is available, and the next decision fits on one line.

The hard part sits around the line: knowing the current queue, understanding the capacity of the storage nodes, making sure replicas remain healthy, and being willing to reverse course if latency moves the wrong way. A short instruction is only useful when it lands in a team with shared context.

There is also a tempo to this kind of work. Make a controlled change. Watch the graphs. Learn what the system can take. Repeat. Large migrations often finish through dozens of decisions that look too small to deserve their own story.

This one gets a story because the accumulated result mattered. We were trying to get data where it needed to go without treating the customer-facing system like a lab. On September 6, the next measured step was simple: give replication more workers and let the system tell us what happened.

Archive