News
Cyber Resilience Expert: Backups Aren't a Recovery Strategy Until You Prove They Work
Having backups does not necessarily mean an enterprise can recover from a cyberattack. The more important question is whether those backups can be converted into a functioning, trustworthy business quickly enough when production systems, identities or even the recovery infrastructure itself have been compromised.
That was a central warning from technology author and longtime Microsoft MVP Brien Posey during "Building a Resilient Enterprise Before, During & After a Breach," the opening session of today's The Modern Enterprise's Cyberattack Survival Guide Summit, being made available for on-demand replay thanks to Rubrik, which also presented a session and shared these statistics:
[Click on image for larger view.] Statistics (source: Rubrik).
His broader point was that cyber recovery has to be designed around business survival rather than the simple existence of backup copies. That means deciding what matters most before an incident, proving recovery-time assumptions, isolating recovery systems from attackers and making sure the organization is not restoring compromised data straight back into production.
"A backup that you've never successfully restored is a theory, not necessarily a true recovery strategy."
Brien Posey, 22-Time Microsoft MVP
Recovery Starts With the Business
One of his first recommendations was to determine, before an attack happens, exactly what the organization would have to restore first if virtually everything went offline.
That answer might include critical business processes, applications and data, but it can also include identity infrastructure, communications systems, manufacturing systems, customer-facing services and third-party dependencies.
The key distinction, he said, is that recovery priorities should not simply reflect which technologies the IT department happens to consider important.
"You're really not recovering technology. You're recovering the business," he said.
That creates a role for both business and technical leadership. Executives and business owners should help determine which processes are essential to continued operations, while IT may sometimes need to override the expected sequence to contain an active compromise or restore an underlying dependency before a more visible application can safely return.
He framed the planning exercise around five questions:
- What must we recover?
- How quickly must we recover it?
- How much data can we afford to lose?
- Can we trust what we're restoring?
- Have we actually proven it works?
Those questions move recovery planning beyond a generic statement that backups exist. They force organizations to identify critical processes, establish realistic recovery time objectives and recovery point objectives, validate their recovery environment and actually test the process.
An RTO Is a Requirement, Not a Capability
Recovery time objectives can create a false sense of readiness when they exist only in documentation.
An organization may state that critical systems need to be back within four hours, for example, but he stressed the difference between defining that target and demonstrating that IT can actually meet it under realistic circumstances.
"An RTO is a business requirement. It doesn't actually become a capability until you've demonstrated it," he said.
That distinction matters because real recovery can involve far more than copying data back onto a server. Teams may need to rebuild indexes, restore application dependencies, deal with compromised credentials, validate backups, recover identity services and coordinate restoration across multiple systems.
He described one real-world case in which an attacker destroyed indexes associated with a backup system. The organization ultimately recovered, but rebuilding those indexes made the process take significantly longer than it normally would have.
RPOs require the same level of business scrutiny. If IT says the organization can tolerate four hours of data loss, he said business leadership should understand what those four hours might actually contain: financial transactions, customer orders, patient records, manufacturing data, contracts or engineering designs.
A technically acceptable RPO may therefore be unacceptable to the people responsible for the business process.
[Click on image for larger view.] Recovery Time Objectives vs. Reality (source: Brien Posey).
Attackers May Target the Recovery Path
He also warned against treating backup infrastructure as something separate from the attack surface.
A sophisticated attacker deploying ransomware has an obvious incentive to eliminate the victim's recovery options. That means backup servers, credentials, management consoles, replication platforms, disaster-recovery environments, cloud backup accounts and service accounts can themselves become high-value targets.
If those systems can be reached with the same identities and administrative paths used elsewhere in the environment, multiple copies of the data may provide less protection than they appear to.
He recommended considering immutable backups, offline or isolated copies, separate administrative credentials, geographic separation, multiple recovery locations and logical or physical isolation.
But Posey also cautioned against treating immutability itself as a complete answer. He told attendees that immutable backups are not foolproof and should be part of a broader protection strategy.
The objective is independence.
If an organization has five copies but an attacker can reach all five through the same compromised environment or credentials, those copies do not represent five truly independent recovery options.
[Click on image for larger view.] Immutable and Offline Backups (source: Brien Posey).
A Successful Restore Still Has To Be Trusted
Another potential failure point is assuming that a technically successful restoration is automatically a clean one.
Attackers sometimes remain inside an environment for an extended period before triggering ransomware or another visible disruption. If the intrusion began weeks or months before detection, backup sets created during that period may contain compromised systems, credentials, scripts or persistence mechanisms.
That means recovery teams have to establish when the environment was last known to be clean, determine which backups predate the compromise, decide how those backups will be validated and make sure the restored environment will not immediately become reinfected.
"Recovery isn't just restoring data; it's recovering trust," Posey said.
Identity is part of that calculation as well.
He noted that disaster-recovery plans often focus on servers, applications and databases while giving less attention to Active Directory, cloud identity providers, authentication systems and privileged accounts. Yet an organization can have healthy applications that nobody can securely access if the identity layer is gone.
He recounted an incident early in his IT career in which a major outage destroyed multiple production workloads along with the organization's identity system. Because there was no backup of that identity environment, the team had to reconstruct it manually from HR records, memory and calls to department heads about who should have access to which resources.
The experience lasted several days and illustrated why identity recovery has to be part of the broader continuity plan rather than an afterthought.
[Click on image for larger view.] Can You Trust What You're Restoring? (source: Brien Posey).
Test Under Conditions That Resemble a Real Failure
Testing ties all of those pieces together.
During the Q&A, Posey said the right exercise cadence depends on the size and structure of the organization, but for a larger enterprise he recommended ransomware-recovery exercises at least quarterly as a baseline.
A useful test should validate backups, establish criteria for determining that they are trustworthy and verify that recovery can happen within the allotted time.
Organizations with the resources can build a sandboxed environment resembling production and test recovery there rather than introducing risk into live systems. Even without a full simulation environment, he said teams can perform restoration tests and measure whether anything unexpected prevents them from meeting recovery objectives.
The larger lesson is that recovery planning begins well before the day an attacker succeeds.
It begins when the business decides which processes cannot be lost, how much downtime and data loss are actually acceptable, which systems and identities everything else depends on, and how those assumptions will be demonstrated rather than merely documented.
A backup can provide the raw material for recovery. Resilience depends on proving the organization can use it.
And More
An on-demand replay is useful, particularly for an event that took place today, but attending Virtualization & Cloud Review
and sister-site summits and webcasts live provides some advantages that a recording cannot fully reproduce. Live attendees can ask presenters questions about their own environments, see demonstrations in context as they happen and get direct guidance from subject-matter experts.
The Rubrik-sponsored summit also offered a $10 Starbucks gift card to the first 150 eligible attendees who stayed for the required portion of the live event. The session itself is now available through the on-demand replay.
With all that in mind, here are some upcoming summits and webcasts from Virtualization & Cloud Review: