The secondary firewall was supposed to fail over.

It did not.

The immediate result was losing VPN access to the data-center network. The separate console path was still reachable. That gave me a way back in.

I killed the secondary, forced traffic to the primary, and access returned.

The system was working again. The question was why the automatic part had not done its job.

I told the team I would spend that night and the next day digging through logs. The automatic failover had not worked, and I needed to find out why.

The recovery path held. The automatic failover had not. That was the part I needed to understand.

Those were two different states. Access had returned, so the immediate network problem was contained. The incident was not finished just because traffic was moving again. The secondary had been expected to take over automatically, and it had done the opposite of what we needed.

My note to the team separated the recovery from the investigation. I could describe exactly what I had done to restore access. I could not yet explain the failover. For that, I needed the logs, and I said I would spend that night and the following day going through them.

Then, halfway through the incident note, I realized I had left my laptop at happy hour.

“Going to get it :(”

The network was back. The log review was still ahead of me. The laptop, unfortunately, was somewhere else. First, I was going to get it.

Archive