The secondary firewall was supposed to fail over.
It did not.
The immediate result was losing VPN access to the data-center network. The console switch was still reachable because it sat outside the firewalls by design. That small bit of separation gave me a way back in.
I killed the secondary, forced traffic to the primary, and access returned.
The system was working again. The question was why the automatic part had not done its job.
I told the team I would spend that night and the next day digging through logs. A failover test is only useful if it tells the truth. “Redundant” on a diagram does not matter when the second box makes the wrong decision.
This was a normal kind of systems problem. The outage was brief. The recovery path held. Nobody got to declare victory because the failure had exposed a gap in the design or the configuration.
The console switch was the useful detail. We had kept one access path outside the devices most likely to trap us. When the normal route vanished, the boring backup route was exactly where we left it.
Then, halfway through the incident note, I realized I had left my laptop at happy hour.
“Going to get it :(”
That was also normal. Serious infrastructure work happened inside ordinary life. You could be tracing firewall logs one minute and retracing your steps to a bar the next.
The machine was back. The network was back. The investigation was not done.
At least I knew where both had gone.