All eyes are on Microsoft and cybersecurity giant Crowdstrike following a global IT outage. Paul here – co-owner and CISO at Valkyries’ InfoSec. Let’s break down exactly what happened on Friday, July 19th.
Who Are the Actors?
Let’s start with identifying the actors involved. The outage began when Crowdstrike deployed what should have been a normal security patch to their active customers for their Falcon Sensor software. Something important to note is that every one of the platforms affected allows for automatic patching of agents. This is a setting that can be controlled and/or turned off by the individual administrators of the account.
Because this patch affected Microsoft 365, Azure, and numerous other Microsoft platforms, it caused a cascading effect creating the “Blue Screen of Death.”
When Did It Start?
The outage began on July 19th, and affected businesses, banks, airlines, and even government agencies.
What Are the Facts?
What We Would Do…
One of the actions we take as an MSSP and MSP, while we provide support for IT teams, providers, and IT directors, is to help establish, and then enforce, proper policy and procedure. Following these procedures with you patch management is especially important, as it ensuring there is a methodical process in place to prevent faulty patches from ever being deployed into your environment.
Problematic patches have happened in the past and have impacted ecosystems as a whole. This isn’t the first and won’t be the last. In July of 2019, for example, Microsoft released Windows patches that systematically erased hard drives upon deployment.
What We Wouldn’t Do…
Saturday mornings are my typical day to do industry research. I grab a cup of coffee (or three) and spend time browsing Reddit, reviews, and looking for trends. Given this happened on Friday, I found articles all over the internet surrounding this event, plus forum posts on Reddit. A glaring issue became apparent very quickly – IT professionals asking “how do we prevent this from happening?”
Like any other business, I have some junior staff on my team. They’re early in their system administration career, and it’s common for them not to have practice or experience with situations like bad patches.
Because it’s common to trust your vendors.
It’s common to trust your tools implicitly.
While the military mantra is to “always trust your equipment,” you also want to listen to the other mantra of “trust but verify.”
In this case, with patch management, we would not deploy a patch we had not already tested and verified.
Trust But Verify
In many cases, testing and verifying a patch can be difficult and easier said than done. However, most, if not all IT firms have the ability to, and the staff capable of, replicating entire infrastructures in a virtual environment, using backups or virtual machines. Patches should be run through this test environment at minimum, once a month, shortly after patch Tuesday. That way, you can run all the patches from Microsoft, vet, and apply in your environment. Test them in the environment, test them on a sample size of your live environment, then deploy en masse when they check out. You should also use this process with any other vendors you have, and patches they put out.
This may be a redundant thing to say, but please also make sure your notifications are enabled to receive notice of new patches.
I can’t tell you how many live environments I’ve encountered where patches were missed because a staff member got irritated with the amount of notifications they were receiving.
United We…Fly?
Let’s narrow the scope down to United Airlines, a CrowdStrike customer. If United had tested these patches in a test environment, planes would not have been grounded nationwide because the error would have been caught in testing – not deployment. This means planes would have stayed in the air, generating income, rather than being grounded, costing revenue.
When you break down the release cycle from CrowdStrike, from midnight to 12:30 AM, the first patch was released/deployed. Roughly 35-40 minutes later, a second patch was released to fix the issue.
To me, it’s apparent there was a supply chain breakdown, and CrowdStrike, a cybersecurity firm known for its consistency, missed the mark. This patch may have been deployed through a manual process, deployed inadvertently to production, or erroneously put out without proper quality assurance. Regardless, the patch was faulty, creating a cascade effect of issues worldwide.
What are the takeaways?
- If you’re IT management or a C-suite executive reading this, I would professionally implore you, if you haven’t already, to provide your IT team the ability to have a warm test environment for your IT team.
- Once you have that environment in place, test your patches (and other security scenarios!) rigorously.
- Only deploy patches into your live environment after they are vetted, reviewed, and tested in your test environment.
- Make sure you have the IT policies and procedures in place to support this vetting process:
- Proper notification of a patch release;
- Reviewing and testing the patch code;
- Conducting a sample patch implementation in your testing environment;
- Doing additional testing in a sample size of your live environment;
- Then finally patch deployment to your entire platform.
This process ensures when your patch deployment goes live on your entire platform, you’ve conducted thorough due diligence and have worked out 90-98% of any potential bugs – making your life as an IT practitioner, and the next workday of your team, that much easier.
Stay alert and stay safe!