Ze Random One
Member
- Joined
- 30 Apr 2011
- Messages
- 244
Working in IT, we (outside) do not know what has gone wrong. If, for example, their systems are subject to a ransomware attack, then BC/DR facilities are also potentially infected, as can be backups if the ransomware laid low for a while. If there was such a problem the last thing you do is give airtime to your attackers by publicly acknowledging the attack.
If, as is more likely, a key piece of kit has been damaged by the power cut, it may well have been a blind spot in the BC/DR arrangements.
It's very possible that BC/DR has not been tested with everyone actually hoofing it over to the DR ops centre during real operation hours, as enacting a planned full DR scenario has inherent risks of itself. It's more likely that DR was tested during quiet times (Christmas day perhaps, or in the early hours) when of course there will be few/no trains moving and far fewer staff needing access, so perhaps they have a situation where their DR systems are unreliable at full business load.
It's also possible that they've done many DR scenario tests, but the combination of failures they had yesterday was not envisioned.
As another example, I've encountered situations where we've purchased diverse routed telecoms/fibre from two separate companies, who failed to disclose that they were both connected to the same trunk, so it transpired the routing was only diverse as far as the telephone exchange, which left a single point of failure in place until we had an outage on both circuits, and we started asking pointed questions.
Finally, it's 10 days after Microsoft's "patch Tuesday" and is therefore a common weekend for production windows servers to get updated and rebooted (after having let your test systems try out the patches for a week or so, to check they don't break anything important). Perhaps they had a full DR in place, but critical servers in their DR datacentre were being patched at the time of the power cut, and are refusing to come back up until they are recovered from backups. Microsoft have, unfortunately, released Windows Updates over the last few months that have caused instability for some high availability database systems using older versions of Windows Server, and we've been forced to hold back some updates because of it - maybe something like this is compounding the issue.
If, as is more likely, a key piece of kit has been damaged by the power cut, it may well have been a blind spot in the BC/DR arrangements.
It's very possible that BC/DR has not been tested with everyone actually hoofing it over to the DR ops centre during real operation hours, as enacting a planned full DR scenario has inherent risks of itself. It's more likely that DR was tested during quiet times (Christmas day perhaps, or in the early hours) when of course there will be few/no trains moving and far fewer staff needing access, so perhaps they have a situation where their DR systems are unreliable at full business load.
It's also possible that they've done many DR scenario tests, but the combination of failures they had yesterday was not envisioned.
As another example, I've encountered situations where we've purchased diverse routed telecoms/fibre from two separate companies, who failed to disclose that they were both connected to the same trunk, so it transpired the routing was only diverse as far as the telephone exchange, which left a single point of failure in place until we had an outage on both circuits, and we started asking pointed questions.
Finally, it's 10 days after Microsoft's "patch Tuesday" and is therefore a common weekend for production windows servers to get updated and rebooted (after having let your test systems try out the patches for a week or so, to check they don't break anything important). Perhaps they had a full DR in place, but critical servers in their DR datacentre were being patched at the time of the power cut, and are refusing to come back up until they are recovered from backups. Microsoft have, unfortunately, released Windows Updates over the last few months that have caused instability for some high availability database systems using older versions of Windows Server, and we've been forced to hold back some updates because of it - maybe something like this is compounding the issue.