• Our new ticketing site is now live! Using either this or the original site (both powered by TrainSplit) helps support the running of the forum with every ticket purchase! Find out more and ask any questions/give us feedback in this thread!

Worldwide IT Outage (Crowdstrike update)

Status
Not open for further replies.

eoff

Member
Joined
15 Aug 2020
Messages
702
Location
East Lothian
I don't think there's anything wrong with the statement, it's accurate, a third party sent out a broken update and broke everything, are you saying it's their fault for using a trusted software provider used in almost every industry?
It is 100% their fault for designing business-critical infrastructure that has such a vulnerability. At one time companies tested and tested any updates, now they rely on online updates by third parties and are paying the price.
 
Sponsor Post - registered members do not see these adverts; click here to register, or click here to log in
R

RailUK Forums

dosxuk

Established Member
Joined
2 Jan 2011
Messages
2,435
It is 100% their fault for designing business-critical infrastructure that has such a vulnerability. At one time companies tested and tested any updates, now they rely on online updates by third parties and are paying the price.
With the danger, and fear, of ransomware attacks, companies pay for the speed of delivery. Nobody is going to be wanting to hold off on a patch to their security software so they can stage it through multiple levels of test system before deploying it to a wider set of systems, if you're doing that, you may as well wait for Microsoft to patch the underlying issue in the next month of updates.

Of course, everyone paying for this service will also be expecting the changes to have been tested to an extent by the provider, and that the software isn't written so badly that a "content update" can crash a system driver and make the system unbootable, or that the only way to fix such an issue is to manually remove the broken driver. The big question they'll be asking is not "why didn't we block the update and deploy it in a few days" but much more, "what on earth were they playing at, how did this update pass any tests".
 

jfollows

Established Member
Joined
26 Feb 2011
Messages
10,084
Location
Wilmslow
The big question they'll be asking is not "why didn't we block the update and deploy it in a few days" but much more, "what on earth were they playing at, how did this update pass any tests".
Exactly so. With some wisdom of hindsight it would appear that this update would have broken just about any Windows machine, so why wasn’t this caught? And, so, how can people who deliver this sort of stuff be trusted?
 

eoff

Member
Joined
15 Aug 2020
Messages
702
Location
East Lothian
With the danger, and fear, of ransomware attacks, companies pay for the speed of delivery. Nobody is going to be wanting to hold off on a patch to their security software so they can stage it through multiple levels of test system before deploying it to a wider set of systems, if you're doing that, you may as well wait for Microsoft to patch the underlying issue in the next month of updates.

Of course, everyone paying for this service will also be expecting the changes to have been tested to an extent by the provider, and that the software isn't written so badly that a "content update" can crash a system driver and make the system unbootable, or that the only way to fix such an issue is to manually remove the broken driver. The big question they'll be asking is not "why didn't we block the update and deploy it in a few days" but much more, "what on earth were they playing at, how did this update pass any tests".
There are choices on how systems are accessible, what networks they are on, if any, what those networks allow, how staff are trained, who has physical access etc. This is one of the aspects that plays into Availability / Business Continuity in any Information Security Management System, having security at the risk or no availability is not great. Someone made that decision and I don't think you should complain if something like this happens and you decided to do no testing. You may wish to relax some requirements for very high risk threats of course. I don't see how doing at least an install and restart test is going to take that long, presumably that can be automated in virtual machines pretty quickly.
 

Starmill

Veteran Member
Joined
18 May 2012
Messages
27,205
Location
Bolton
With the danger, and fear, of ransomware attacks, companies pay for the speed of delivery. Nobody is going to be wanting to hold off on a patch to their security software so they can stage it through multiple levels of test system before deploying it to a wider set of systems, if you're doing that, you may as well wait for Microsoft to patch the underlying issue in the next month of updates.

Of course, everyone paying for this service will also be expecting the changes to have been tested to an extent by the provider, and that the software isn't written so badly that a "content update" can crash a system driver and make the system unbootable, or that the only way to fix such an issue is to manually remove the broken driver. The big question they'll be asking is not "why didn't we block the update and deploy it in a few days" but much more, "what on earth were they playing at, how did this update pass any tests".
No doubt a good number of law and IT firms themselves have been affected, as they may well be customers given they're usually sensitive about security. As such I'd expect the contracts to be being pored over right now.
 

Busaholic

Veteran Member
Joined
7 Jun 2014
Messages
14,671
Someone in front of me in the supermarket had to wait 2-3 minutes for their contactless transaction to go through this morning. The next person used chip and pin and it went through immediately. I used cash, just in case :)
Wait for someone on here to be the first to call you a Luddite! Various supermarkets and other large shops are now admitting to problems with contactless payments earlier. Bet there were no problems in Moscow or Beijing!
 

Mcr Warrior

Veteran Member
Joined
8 Jan 2009
Messages
17,191
Simplistic question... Why were some businesses seemingly more susceptible than others, to today's disruptive outage?
 

74A

Member
Joined
27 Aug 2015
Messages
772
They
Simplistic question... Why were some businesses seemingly more susceptible than others, to today's disruptive outage?
They use different systems. Also most of the problems seem to be associated with anti virus company Crowdstrike. If you don't use them you won't be directly affected.
 

jfollows

Established Member
Joined
26 Feb 2011
Messages
10,084
Location
Wilmslow
Simplistic question... Why were some businesses seemingly more susceptible than others, to today's disruptive outage?
I am horribly biased, as a former IT professional, but I think the answer will be that although most businesses make use of Windows in some way, only a subset of them make Windows a key business-critical component of their IT. Those who didn't suffer may have made a choice to implement the key parts of their IT using Linux, Unix, z/OS or other platforms. I have experience with all of these and I know which one I'd chose - z/OS - which runs on IBM mainframes, but there are "cheaper" options with which I'm also familiar.

EDIT And even those using Windows didn't all use Crowdstrike, that's an important point too.
 
Last edited:

py_megapixel

Established Member
Joined
5 Nov 2018
Messages
7,454
Location
Northern England
That's why you boot into safe mode, that's fairly trivial to do on a consumer machine or a corporate machine that isn't locked down.
I think it still requires someone to physically be present because the point at which you would select safe mode is before the point at which any network hardware would start up. There might be some way of selecting "reboot to safe mode" from a system which is already up, though I can't think off the top of my head where that is. But if the system is bluescreening before it manages to get onto the network then of course that can't be done remotely either.

(Don't take any of this as gospel though - I don't work much with Windows so there may well be something I'm missing)
 

JamesT

Established Member
Joined
25 Feb 2015
Messages
4,836
I think it still requires someone to physically be present because the point at which you would select safe mode is before the point at which any network hardware would start up. There might be some way of selecting "reboot to safe mode" from a system which is already up, though I can't think off the top of my head where that is. But if the system is bluescreening before it manages to get onto the network then of course that can't be done remotely either.

(Don't take any of this as gospel though - I don't work much with Windows so there may well be something I'm missing)
You can have remote management down in the hardware before any operating system is involved. It’s more common in the server world, but some desktop systems also have it. It would need to be set up in advance though and even if the hardware is capable, it’s often disabled as a potential vector for hackers.
 

Egg Centric

Established Member
Joined
6 Oct 2018
Messages
2,840
Location
Land of the Prince Bishops
I did notice that in the office this morning in central London there were genuine concerns from those (several) who had parked at their local station and the parking machine rejected their credit card, that the railway would fine them an inordinate amount for not paying for parking. How on earth has the railway even come to this customer-unfriendly perception.

NCP (managing a TfL car park) did indeed try to penalty charge me in basically this scenario (I took it as far as POPLA which found against me, then dared them to take me to court, then nothing happened) so I think it's a perfectly sensible concern for the passengers to have!

== Doublepost prevention - post automatically merged: ==

I am horribly biased, as a former IT professional, but I think the answer will be that although most businesses make use of Windows in some way, only a subset of them make Windows a key business-critical component of their IT. Those who didn't suffer may have made a choice to implement the key parts of their IT using Linux, Unix, z/OS or other platforms. I have experience with all of these and I know which one I'd chose - z/OS - which runs on IBM mainframes, but there are "cheaper" options with which I'm also familiar.

EDIT And even those using Windows didn't all use Crowdstrike, that's an important point too.

Yup, we're not affected at all EXCEPT in some of our integrations with third party providers. And in at least one of those cases if you're paid fortnightly I'd be pretty concerned right now.
 

eoff

Member
Joined
15 Aug 2020
Messages
702
Location
East Lothian
Crowdstrike has updated their web page on this...


CrowdStrike is actively working with customers impacted by a defect found in a single content update for Windows hosts. Mac and Linux hosts are not impacted. This was not a cyberattack.


The issue has been identified, isolated and a fix has been deployed. We refer customers to the support portal for the latest updates and will continue to provide complete and continuous updates on our website.


We further recommend organizations ensure they’re communicating with CrowdStrike representatives through official channels.


Our team is fully mobilized to ensure the security and stability of CrowdStrike customers.


Updated 1:25pm ET, July 19, 2024:


We understand the gravity of the situation and are deeply sorry for the inconvenience and disruption. We are working with all impacted customers to ensure that systems are back up and they can deliver the services their customers are counting on.


We assure our customers that CrowdStrike is operating normally and this issue does not affect our Falcon platform systems. If your systems are operating normally, there is no impact to their protection if the Falcon Sensor is installed.


Below is the latest CrowdStrike Tech Alert with more information about the issue and workaround steps organizations can take. We will continue to provide updates to our community and the industry as they become available.

Summary​

  • CrowdStrike is aware of reports of crashes on Windows hosts related to the Falcon Sensor.

Details​

  • Symptoms include hosts experiencing a bugcheck\blue screen error related to the Falcon Sensor.
  • Windows hosts which have not been impacted do not require any action as the problematic channel file has been reverted.
  • Windows hosts which are brought online after 0527 UTC will also not be impacted
  • Hosts running Windows 7/2008 R2 are not impacted
  • This issue is not impacting Mac- or Linux-based hosts
  • Channel file "C-00000291*.sys" with timestamp of 0527 UTC or later is the reverted (good) version.
  • Channel file "C-00000291*.sys" with timestamp of 0409 UTC is the problematic version.
    • Note: It is normal for multiple "C-00000291*.sys files to be present in the CrowdStrike directory - as long as one of the files in the folder has a timestamp of 0527 UTC or later, that will be the active content.

Current Action​

  • CrowdStrike Engineering has identified a content deployment related to this issue and reverted those changes.
  • If hosts are still crashing and unable to stay online to receive the Channel File Changes, the workaround steps below can be used to address this issue.
  • We assure our customers that CrowdStrike is operating normally and this issue does not affect our Falcon platform systems. If your systems are operating normally, there is no impact to their protection if the Falcon Sensor is installed. Falcon Complete and Overwatch services are not disrupted by this incident.

Query to identify impacted hosts via Advanced event search​


Please see this KB article: How to identify hosts possibly impacted by Windows crashes.

Workaround Steps for individual hosts:​

  • Reboot the host to give it an opportunity to download the reverted channel file. If the host crashes again, then:
    • Boot Windows into Safe Mode or the Windows Recovery Environment
      • NOTE: Putting the host on a wired network (as opposed to WiFi) and using Safe Mode with Networking can help remediation.
    • Navigate to the %WINDIR%\System32\drivers\CrowdStrike directory
    • Locate the file matching “C-00000291*.sys”, and delete it.
    • Boot the host normally.
    • Note: Bitlocker-encrypted hosts may require a recovery key.

Workaround Steps for public cloud or similar environment including virtual:​

Option 1:​

  • Detach the operating system disk volume from the impacted virtual server
  • Create a snapshot or backup of the disk volume before proceeding further as a precaution against unintended changes
  • Attach/mount the volume to to a new virtual server
  • Navigate to the %WINDIR%\System32\drivers\CrowdStrike directory
  • Locate the file matching “C-00000291*.sys”, and delete it.
  • Detach the volume from the new virtual server
  • Reattach the fixed volume to the impacted virtual server

Option 2:​

  • Roll back to a snapshot before 0409 UTC.

AWS-specific documentation:​

Azure environments:​

User Access to Recovery Key in the Workspace ONE Portal​


When this setting is enabled, users can retrieve the BitLocker Recovery Key from the Workspace ONE portal without the need to contact the HelpDesk for assistance. To turn on the recovery key in the Workspace ONE portal, follow the next steps. Please see this Omnissa article for more information.


Bitlocker recovery-related KBs:​


 

birchesgreen

Established Member
Joined
18 Aug 2015
Messages
7,655
Location
Solihull
Wait for someone on here to be the first to call you a Luddite! Various supermarkets and other large shops are now admitting to problems with contactless payments earlier. Bet there were no problems in Moscow or Beijing!
What do you mean?
 

Essan

Member
Joined
22 Feb 2017
Messages
615
Location
Evesham / Lochailort
What do you mean?

The reality is China and Russia werent affected because they tend not to use western software. Although if they were affected they'd probably try and keep quiet about it. There were issues at Hong Kong Airport but otherwise no public reports of issues that I have seen.

Meanwhile, I had no issues either, along with millions of others in the west. I;d never even heard of crowdstrike before!
 

Energy

Established Member
Joined
29 Dec 2018
Messages
5,135
What do you mean?
I presume they are saying that if the Russians or Chinese were responsible for the attack then Moscow or Beijing respectively would not be affected. Of course, it wasn't an attack but a failing of Crowdstrike themselves.
Meanwhile, I had no issues either, along with millions of others in the west. I;d never even heard of crowdstrike before!
You wouldn't have heard of them unless you are involved in business IT. Cloudflare is another major company which most have not heard of but have devastating effects when an outage occurs.
 

jfollows

Established Member
Joined
26 Feb 2011
Messages
10,084
Location
Wilmslow
Actually, they’ve been crowing about it (https://www.theguardian.com/busines...june-figure-lowest-2019-horizon-business-live):

Russia, walled off by sanctions on tech, says it's doing fine amid global outage​

Blake Montgomery
Russian officials said Thursday that their country’s vital systems had not suffered outages anywhere near as bad as those in the UK, US, and much of the rest of the world.

Since Russia’s invasion of Ukraine began in early 2022, Microsoft and other software providers have drawn down their operations in the country.

In retaliation, the Kremlin has gone after them, speeding up the decreases in their operations, their departures, and the isolation of Russian digital systems.

Russia’s digital ministry issued a statement that read:

The situation once again highlights the significance of foreign software substitution.
Mikhail Klimarev of the the Internet Protection Society, a non-governmental organization, told Reuters:

CrowdStrike has not provided any services in Russia since February 2022.
 

johntea

Established Member
Joined
29 Dec 2010
Messages
3,030
The issue today is far from a one off, basically comes down to a dodgy definition update for a security product which has happened many times over the years with many of the vendors of such software

What has changed is companies buying into the whole 'cloud' business which at the end of the day mostly just means 'another company will run your services on their servers somewhere'

If the other company is running Crowdstrike on their servers you're buggered until they sort them out basically
 

miklcct

Established Member
Joined
2 May 2021
Messages
5,017
Location
Cricklewood
That's a variant on a very old expression, that on a date a girl should always have enough cash to suddenly get themselves a taxi to get home - if needs be!
That's a large amount of money. I normally have enough cash to get a bus home, but a taxi can easily cost you £50+!


I presume they are saying that if the Russians or Chinese were responsible for the attack then Moscow or Beijing respectively would not be affected. Of course, it wasn't an attack but a failing of Crowdstrike themselves.

You wouldn't have heard of them unless you are involved in business IT. Cloudflare is another major company which most have not heard of but have devastating effects when an outage occurs.
Supply chain attack is very real, but doesn't a normal redundancy setup consists of running independent nodes using completely different hardware and software located physically separately?
 

Ediswan

Established Member
Joined
15 Nov 2012
Messages
3,418
Location
Stevenage
The issue today is far from a one off, basically comes down to a dodgy definition update for a security product which has happened many times over the years with many of the vendors of such software
I'm not so sure. One of the error screens seen on TV showed HAL_INITIALIZATION_FAILED. I doubt a definition update would need to fiddle with the Hardware Abstraction Layer.
 

dosxuk

Established Member
Joined
2 Jan 2011
Messages
2,435
What has changed is companies buying into the whole 'cloud' business which at the end of the day mostly just means 'another company will run your services on their servers somewhere'
This isn't really a cloud issue - while this software has been installed on some servers, some of which are in the cloud, the vast majority of the affected devices are the ones out in the wild - in offices, in factories, in surgeries and behind tills. Nobody was paying Crowdstrike to provide a service on a server somewhere, they were paying them to protect all of their machines from bad actors.
 

Rail Quest

Member
Joined
8 Apr 2023
Messages
755
Location
Warrington
Will this make Crowdstrike go bankrupt?
Perhaps. The key thing for me that could seal their fate is the cataclysmic reputational damage this will do to the company - a cybersecurity supplier causing such a significant incident cannot be understated. Companies like LastPass have proven its not impossible for a cybersecurity-related firm to survive an attack or mistake of a large scale but they probably didn't have to pay anywhere near the compensation that CrowdStrike risk. Combine the short/mid-term financial loss of the compensation with such significant reputational costs is no doubt going to define the company's future, one way or another.
 

Freightmaster

Verified Rep
Joined
7 Jul 2009
Messages
4,428
I'm not so sure. One of the error screens seen on TV showed HAL_INITIALIZATION_FAILED. I doubt a definition update would need to fiddle with the Hardware Abstraction Layer.

Fiddling with HAL never ends well...

af029bcb8fc37037ed7dc00674441749.jpg

(image shows HAL from the movie '1999' with the quote "I'm sorry Dave, I'm afraid I can't do that")
 

Energy

Established Member
Joined
29 Dec 2018
Messages
5,135
Perhaps. The key thing for me that could seal their fate is the cataclysmic reputational damage this will do to the company - a cybersecurity supplier causing such a significant incident cannot be understated. Companies like LastPass have proven its not impossible for a cybersecurity-related firm to survive an attack or mistake of a large scale but they probably didn't have to pay anywhere near the compensation that CrowdStrike risk. Combine the short/mid-term financial loss of the compensation with such significant reputational costs is no doubt going to define the company's future, one way or another.
CrowdStrike will likely come off better as it wasn't a cyberattack (more a cyber-accident) and no vulnerable data was leaked. LastPass had passwords stolen.
 

Rail Quest

Member
Joined
8 Apr 2023
Messages
755
Location
Warrington
CrowdStrike will likely come off better as it wasn't a cyberattack (more a cyber-accident) and no vulnerable data was leaked. LastPass had passwords stolen.
Perhaps, though this may be dependent on how the accident came to be. If it was a single or small number of employees bypassing processes to release the update, then the fallout may not be as bad as if inherent development, testing or culture issues were to blame. If the latter is true and that became public knowledge via an investigation, I can imagine the headlines...;)
 

Pakenhamtrain

Established Member
Joined
26 Jan 2014
Messages
1,202
Location
Melbourne, Australia
Down here we had same as other things get affected.

It delayed entry into the football because the AFL and clubs use ticketmaster for thier member tickets

Metro trains lost thier backend systems but the safety critical stuff was working.
V/Line had to stop because they an issue with thier radio.
 
Status
Not open for further replies.

Top